Picture having mixed NAL unit type

The introduction of a flag to differentiate IRAP and non-IRAP subpictures in video coding systems addresses inefficiencies in handling mixed subpictures, improving coding efficiency and reducing latency in VR applications.

JP2025156376AActive Publication Date: 2025-10-14HUAWEI TECH CO LTD
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
JP2025119901
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Priority Date
2019-04-10
Filing Date
2025-07-16
Publication Date
2025-10-14
Estimated Expiration
2040-03-11

AI Technical Summary

Technical Problem

Existing video coding systems struggle with efficiently handling mixed Intra Random Access Point (IRAP) and non-IRAP subpictures within a single picture, leading to inefficiencies in bandwidth usage and decoding latency, particularly in applications like virtual reality (VR) where different subpictures may require different resolutions and access frequencies.

Method used

A mechanism is introduced to handle mixed NAL unit types within a picture by using a flag (mixed_nalu_types_in_pic_flag) to differentiate between IRAP and non-IRAP subpictures, allowing for dynamic resolution changes and efficient coding by treating subpictures differently during decoding.

Benefits of technology

This approach enhances coding efficiency by optimizing bandwidth allocation and reducing latency in decoding, ensuring seamless VR experiences without noticeable resolution changes, while minimizing resource usage.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2025156376000001_ABST
    Figure 2025156376000001_ABST
Patent Text Reader

Abstract

To disclose a video coding mechanism.SOLUTION: A mechanism contains a step of receiving a plurality of sub-pictures related to a picture and a bit stream containing a flag. Each sub-picture is contained in a plurality of video coding layer (VCL) network abstraction layer (NAL) unit. A value of the NAL unit type is similar to all of VCL NAL units related to the picture when the flag is set to a first value. When the flag is set to a second value, the value of a first NAL unit type related to the VCL NAL unit containing one or more of the sub-pictures of the picture is different from the value of a second NAL unit type related to the VCL NAL unit containing one or more of the sub-pictures of the picture. Each sub-picture is decoded on the basis of the value of the NAL unit type. Each sub-picture is transferred for displaying one part of a decoded video sequence.SELECTED DRAWING: Figure 7
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] CROSS-REFERENCE TO RELATED APPLICATIONS This patent application is a continuation of the "Support Of Mix" application filed by Ye-Kui Wang et al. on March 11, 2019. U.S. Provisional Patent Application No. 62, entitled "Encoded NAL Unit Types Within One Picture In Video Coding" / 816,749 and "Support Of Mixed U.S. Provisional Patent Application No. 62 / 8 entitled "NAL Unit Types Within One Picture In Video Coding" This application claims the benefit of provisional patent applications Nos. 32,132, the entire contents of which are incorporated herein by reference. To be incorporated.

[0002] TECHNICAL FIELD This disclosure relates generally to video coding, and more particularly to peak-based video coding. This relates to coding sub-pictures of a picture. [Background technology]

[0003] The amount of video data required to render even a relatively short video can be quite large. and that the data is streamed over a communications network having limited bandwidth capacity. even if it poses difficulties when the information should be mailed or otherwise communicated. Therefore, video data is generally transmitted over modern communication networks. Because memory resources may be limited, the video is compressed before it is stored. The size of the video can also be an issue when stored on a storage device. Video compression devices are used to code video data before transmission or storage. Use software and / or hardware at the source to This reduces the amount of data required to represent a single video image. The data is received at the destination by a video decompression device that decodes the video data. With limited network resources and ever-increasing demand for higher video quality, This results in improved compression, increasing the compression ratio with little or no sacrifice in image quality. and thawing techniques are preferred. Summary of the Invention [Means for solving the problem]

[0004] In an embodiment, the present disclosure provides a method implemented in a decoder, comprising: A bitstream including a plurality of associated sub-pictures and flags is transmitted to a decoder receiver. Thus, the sub-picture is received by the video coding layer (VCL). Network Abstraction Layer (NAL) Unit When the flag is set to a first value, the first NAL unit The process assumes that the type value is the same for all of the VCL NAL units associated with a picture. determining by the processor if the flag is set to a second value; of the first NAL unit type for VCL NAL units containing one or more of the subpictures The second NAL unit for a VCL NAL unit whose value contains one or more of the picture's subpictures. determining by a processor a value of the first NAL unit type different from the value of the second NAL unit type; one or more of the subpictures based on the value of the NAL unit type or the value of the second NAL unit type. and decoding by a processor.

[0005] A picture may be partitioned into multiple sub-pictures. Such sub-pictures may be separated into can be coded into sub-bitstreams, and then The bitstreams can be merged into a bitstream for transmission to a decoder. For example, subpictures may be used for virtual reality (VR) applications. In this example, the user may only see a portion of the VR picture at any one time. More bandwidth can be allocated to sub-pictures that are more likely to be displayed, Subpictures that are unlikely to be shown may be compressed to improve coding efficiency. As possible, different sub-pictures may be transmitted at different resolutions. VideoStream is an intra-random access point (IRAP). An IRAP picture may be coded using an intra-predicted (int) picture. It is coded exponentially and can be decoded without reference to other pictures. Non-IRAP pictures may be coded using inter prediction, and other pictures Non-IRAP pictures can be decoded by referencing the picture. However, IRAP pictures are decoded without reference to other pictures. Since the video sequence contains enough data to start decoding from the IRAP picture, IRAP pictures can be used within subpictures, This allows for dynamic resolution changes. for subpictures that are more likely to be seen (based on the user's current viewport) To transmit as many IRAP pictures as possible and further improve coding efficiency, Fewer IRAP pictures may be sent for less useful sub-pictures. The pictures are part of the same picture. Therefore, this method is similar to IRAP subpictures. Some video systems may result in pictures containing both IRAP and non-IRAP subpictures. does not provide for handling mixed pictures that have both IRAP and non-IRAP regions. Whether the picture is mixed and therefore contains both IRAP and non-IRAP components Based on this flag, the decoder determines whether the picture / subpicture Different sub-pictures are treated differently when decoding in order to properly decode and display the image. This flag may be stored in the PPS and is called mixed_nalu_types_in_pic_flag. Therefore, the disclosed mechanism allows for the implementation of additional functionality. Furthermore, the disclosed mechanism is useful when using subpicture bitstreams. It allows for dynamic resolution changes. Therefore, the disclosed mechanism Lower resolutions when streaming VR video without a noticeable loss of perspective This allows multiple sub-picture bitstreams to be transmitted. The mechanism used increases coding efficiency and therefore the encoder and decoder Use of network resources, memory resources, and / or processing resources in Reduce.

[0006] Optionally, in any of the above aspects, another implementation of the aspect comprises: This specifies that the system contains a Picture Parameter Set (PPS) that includes a flag.

[0007] Optionally, in any of the above aspects, another implementation of the aspect comprises: The value of the type indicates that the picture contains an Intra Random Access Point (IRAP) subpicture. The second NAL unit type value indicates that the picture contains a non-IRAP subpicture. It is stipulated that this indicates the following.

[0008] Optionally, in any of the above aspects, another implementation of the aspect comprises: The type value is a random access decodable leading picture. Instantaneous Decoding Refresh (IDR) with a leading picture fresh) (IDR_W_RADL), IDR without leading picture (IDR_N_LP), or clean It is defined as being equal to the clean random access (CRA) NAL unit type (CRA_NUT). Determine.

[0009] Optionally, in any of the above aspects, another implementation of the aspect comprises: The value of the type is the trailing picture NAL unit type (TRAIL_NU T), random-access decodable leading picture NAL unit type (RADL_NUT), or or random access skip reading pictures (RASL) The RFC 24482 specifies that the NAL unit type is equal to the RASL_NUT (apped leading picture) NAL unit type (RASL_NUT).

[0010] Optionally, in any of the above aspects, another implementation of the aspect is It is specified that the flag is nalu_types_in_pic_flag.

[0011] Optionally, in any of the above aspects, another implementation of the aspect refers to a PPS. A picture has two or more VCL NAL units, and the VCL NAL units are NAL unit types. When specifying that the numbers do not have the same value of nal_unit_type, If lag is equal to 1 and the picture that references the PPS has one or more of the VCL NAL units, and the VCL When specifying that NAL units have the same value of nal_unit_type, mixed_nalu_types_ Specify that in_pic_flag is equal to 0.

[0012] In an embodiment, the present disclosure provides a method implemented in an encoder, comprising: a step of determining by a processor whether the image contains multiple sub-pictures of different types; The video clip and subpictures of a picture are coded into multiple VCL NAL units in the bitstream. and converting all VCL NALs whose first NAL unit type value is associated with a picture. It is set to the first value when the subpictures of a picture are the same in terms of units. The first NAL unit type value for a VCL NAL unit containing one or more of the following is a picture subframe: a second NAL unit type value for a VCL NAL unit containing one or more of the pictures; The flags set to the second value when different are encoded into the bitstream by the processor. and a processor coupled to the decoder for generating a bitstream for transmission to the decoder. and storing the information by a memory.

[0013] A picture may be partitioned into multiple sub-pictures. Such sub-pictures may be separated into can be coded into sub-bitstreams, and then The bitstreams can be merged into a bitstream for transmission to a decoder. For example, subpictures may be used for virtual reality (VR) applications. In this example, the user may only see a portion of the VR picture at any one time. More bandwidth can be allocated to sub-pictures that are more likely to be displayed, Subpictures that are unlikely to be shown may be compressed to improve coding efficiency. As possible, different sub-pictures may be transmitted at different resolutions. Video streams are streamed using Intra Random Access Point (IRAP) pictures. An IRAP picture may be coded by intra prediction. Non-IRAP pictures can be decoded without reference to other pictures. It may be coded by prediction and restored by referencing other pictures. Non-IRAP pictures are significantly more condensed than IRAP pictures. contains enough data for a picture to be decoded without reference to other pictures Therefore, the video sequence must start decoding from the IRAP picture. The char can be used within a subpicture and allows for dynamic resolution changes. Therefore, the video system can determine the Send more IRAP pictures for sub-pictures that are more likely to be seen (even if the ,To further improve coding efficiency, for sub-pictures that are unlikely to be seen, Fewer IRAP pictures may be transmitted, but the sub-pictures must be part of the same picture. Therefore, this method can handle both IRAP and non-IRAP subpictures. Some video systems may result in a picture containing IRAP and non-IRAP regions. There is no provision to handle mixed pictures that have both. Therefore, it contains a flag indicating whether it contains both IRAP and non-IRAP components. Based on the flags, the decoder will decode and display the picture / subpicture appropriately. Therefore, different sub-pictures may be treated differently when decoding. This may be stored in the .mixed_nalu_types_in_pic_flag and may be called mixed_nalu_types_in_pic_flag. The mechanisms shown allow for the implementation of additional functionality. The system allows dynamic resolution changes when using subpicture bitstreams. Therefore, the disclosed mechanism does not significantly impair the user experience. Instead of using a lower resolution subpicture bitstream when streaming VR video, Therefore, the disclosed mechanism allows the codec improves the coding efficiency and therefore the network resources in the encoder and decoder , reducing memory resource and / or processing resource usage.

[0014] Optionally, in any of the above aspects, another implementation of the aspect is to convert the PPS into a bitstream. a step of encoding the flag into the PPS stream, wherein the flag is encoded into the PPS. This includes:

[0015] Optionally, in any of the above aspects, another implementation of the aspect comprises: The value of the second NAL unit type indicates that the picture contains an IRAP subpicture. Specifies that a value of 0 indicates that the picture contains non-IRAP sub-pictures.

[0016] Optionally, in any of the above aspects, another implementation of the aspect comprises: Specifies that the value of this parameter is equal to IDR_W_RADL, IDR_N_LP, or CRA_NUT.

[0017] Optionally, in any of the above aspects, another implementation of the aspect comprises: Specifies that the value of this type is equal to TRAIL_NUT, RADL_NUT, or RASL_NUT.

[0018] Optionally, in any of the above aspects, another implementation of the aspect is nalu_types_in_pic_flag.

[0019] Optionally, in any of the above aspects, another implementation of the aspect refers to a PPS. A picture has two or more VCL NAL units, and the VCL NAL units are of type nal_unit_type. When specifying that they do not have the same value, mixed_nalu_types_in_pic_flag is equal to 1 and PPS has one or more VCL NAL units, and the VCL NAL units are nal_u When specifying that they have the same value of nit_type, mixed_nalu_types_in_pic_flag is equal to 0. It is stipulated that

[0020] In an embodiment, the present disclosure provides a method for transmitting a signal to a processor, a receiver coupled to the processor, and a method for transmitting a signal to a processor. a memory coupled to the processor; and a transmitter coupled to the processor, a receiver, a memory, and a transmitter configured to perform the method of any of the above aspects. video coding devices.

[0021] In embodiments, the present disclosure provides a codec for use by a video coding device. A non-transitory computer-readable medium containing a computer program product, The program product, when executed by a processor, A method for performing the method of any of the above aspects, stored on a non-transitory computer-readable medium. It includes a non-transitory computer-readable medium containing computer-executable instructions.

[0022] In an embodiment, the present disclosure provides a method for generating a plurality of sub-pictures and flags associated with a picture. a receiving means for receiving a bitstream including a sub-picture, the sub-picture being a receiving means and a determining means, the receiving means including a flag set to a first value, When the first NAL unit type value is set to 0, all VCL NAL units associated with a picture are and if the flag is set to a second value, The first NAL unit type value for the VCL NAL unit containing one or more pictures is The second NAL unit for a VCL NAL unit containing one or more of the picture's subpictures. determining means for determining whether the first NAL unit type value or the second NAL unit type value is different from the first NAL unit type value; for decoding one or more of the subpictures based on the value of the second NAL unit type and a decoding means for decoding the received signal.

[0023] Optionally, in any of the above aspects, another implementation of the aspect is a decoder comprising: It is provided that the method is further configured to perform the method of any of the above aspects.

[0024] In an embodiment, the present disclosure provides a method for determining whether a picture contains multiple sub-pictures of different types. a determining means for determining whether a sub-picture of a picture is a picture; and an encoding means for encoding a sub-picture of a picture into a picture. The first NAL unit type value is It is set to the first value when it is the same for all VCL NAL units associated with a picture. The first NAL unit is for a VCL NAL unit that contains one or more of the picture's subpictures. The unit type value relates to a VCL NAL unit that contains one or more of the picture's subpictures. The flag set to the second value when it differs from the value of the second NAL unit type is bitstreamed. encoding means for encoding the bitstream into a bitstream for transmission to a decoder; and storage means for storing the program.

[0025] Optionally, in any of the above aspects, another implementation of the aspect is further characterized in that the encoder: It is provided that the device is further configured to perform the method of any of the above aspects.

[0026] For purposes of clarity, any one of the above-described embodiments may be considered a new embodiment within the scope of the present disclosure. may be combined with any one or more of the other above-mentioned embodiments to produce stomach.

[0027] These and other features are best understood from the following detailed description taken in conjunction with the accompanying drawings and claims. By understanding it will be understood more clearly.

[0028] For a more complete understanding of this disclosure, reference is made to the accompanying drawings, in which like reference numerals represent like parts, and in which: Reference is now made to the following brief description taken in conjunction with the detailed description. [Brief explanation of the drawings]

[0029] [Figure 1] 1 is a flow diagram of an exemplary method for coding a video signal. [Figure 2] 1 is a schematic diagram of an example coding and decoding (codec) system for video coding. [Figure 3] FIG. 1 is a schematic diagram illustrating an exemplary video encoder. [Figure 4] FIG. 1 is a schematic diagram illustrating an exemplary video decoder. [Figure 5] FIG. 2 is a schematic diagram illustrating an exemplary coded video sequence. [Figure 6] FIG. 1 is a schematic diagram illustrating multiple sub-picture video streams split from a virtual reality (VR) picture video stream. [Figure 7] FIG. 1 is a schematic diagram illustrating an exemplary bitstream containing pictures with mixed Network Abstraction Layer (NAL) unit types. [Figure 8] 1 is a schematic diagram of an exemplary video coding device. [Figure 9] 1 is a flow diagram of an example method for encoding a video sequence including pictures with mixed NAL unit types in a bitstream. [Figure 10] 1 is a flow diagram of an example method for decoding a video sequence including pictures with mixed NAL unit types from a bitstream. [Figure 11] 1 is a schematic diagram of an example system for coding a video sequence including pictures with mixed NAL unit types in a bitstream. DETAILED DESCRIPTION OF THE INVENTION

[0030] While exemplary implementations of one or more embodiments are provided below, the disclosed systems and / or or method may be any of a number of techniques whether currently known or in existence. It should be understood at the outset that the present disclosure may be implemented using techniques The exemplary implementations shown below, including the exemplary designs and implementations illustrated and described, are shown in Figures 1 and 2. The present invention should not be limited in any way to the aspects and techniques of the present invention, but should be considered in its entirety along with the full scope of equivalents of the appended claims. and may be modified within the scope of those claims.

[0031] The following acronyms are used: Coded Video Sequence (CVS), Decoded Picture Buffer (DPB), Instantaneous Decode Refresh (IDR), Intra Random Access Point Point (IRAP), Least Significant Bit (LSB), Most Significant Bit (MSB), Network Abstraction Layer (NAL ), Picture Order Count (POC), Raw Byte Sequence Payload ( RBSP: Raw Byte Sequence Payload), Sequence Parameter Set (SPS), and Working Draft The draft (WD) is used herein.

[0032] Many video compression techniques reduce the size of video files with minimal data loss. For example, video compression techniques can be used to reduce data redundancy in a video sequence. Spatial (e.g., intra-picture) prediction and temporal (e.g., inter-picture) prediction are used to reduce or eliminate Block-based video coding may include performing inter-picture (inter-picture) prediction. For example, a video slice (e.g., a video picture or a portion of a video picture) is The video blocks may be partitioned into video blocks, which may include tree blocks, coding blocks, and so on. Coding Tree Block (CTB), Coding Tree Unit (CTU), Coding Unit (CU) , and / or coding nodes. Video blocks in an (I) slice being processed are based on reference blocks in neighboring blocks in the same picture. The image is coded using spatial prediction relative to the reference sample. Video blocks within a unidirectionally predicted (P) or bidirectionally predicted (B) slice that are coded Spatial prediction or other reference samples relative to reference samples in neighboring blocks in the picture It is coded by using temporal prediction relative to a reference sample in the architecture. A picture may be referred to as a frame and / or an image, and a reference picture may be They may also be called reference frames and / or reference images. Spatial or temporal prediction is performed on an image block. The residual data is the difference between the original block and the predicted block. Therefore, an inter-coded block represents the pixel difference between the predicted block and the The motion vectors point to blocks of reference samples that form the block to be coded. The residual data indicates the difference between the predicted block and the block. The blocks to be coded are coded by the intra-coding mode and the residual data. For further compression, the residual data is transformed from the pixel domain to the transform domain. These result in residual transform coefficients, which may be quantized. First, the quantized transform coefficients may be arranged in a two-dimensional array. The coefficients may be scanned to generate a one-dimensional vector of transform coefficients. Video coding may be applied to achieve even greater compression. The compression techniques are discussed in detail below.

[0033] To ensure that the encoded video can be accurately decoded, the video is It is encoded and decoded according to a video coding standard. International Telecommunication Union (ITU) Standardization Sector (ITU-T) H.261, International Organization for Standardization / International Electrotechnical Commission (ISO / IEC) Moving Picture Experts Group (MPEG)-1 Part 2, ITU-T H.262 or ISO / IEC MPEG-2 Part Part 2, ITU-T H.263, ISO / IEC MPEG-4 Part 2, ITU-T H.264 or ISO / IEC MPEG-4 Part 10 Also known as Advanced Video Coding (AVC), ITU-T H.26 High Efficiency Video Coding (HEVC), also known as AVC Part 5 or MPEG-H Part 2. , Scalable Video Coding (SVC), multi-view video coding Multiview Video Coding (MVC) and Multiview Video Coding plus Depth (MVC+D) Multiview Video Coding plus Depth), and three-dimensional (3D) AVC (3D-AVC) extensions. HEVC includes Scalable HEVC (SHVC), Multi-View HEVC (MV-HEVC), and 3D HEVC (3D-HEVC). The Joint Video Experts Team (JVET) of ITU-T and ISO / IEC The erts team is developing a video coding technology called Versatile Video Coding (VVC). The development of the VVC decoding standard has begun. The VVC is a description of the algorithm, a VVC Working Draft (WD), The encoder side of the description and reference software provided by WD includes JVET-M1001-v6. It can be enjoyed.

[0034] Video coding systems use IRAP and non-IRAP pictures. The IRAP pictures are random pictures with respect to the video sequence. A picture coded by intra prediction that serves as an access point. In intra prediction, a block of a picture is predicted based on the number of other blocks in the same picture. This is the same as for non-IRAP pictures that use inter prediction. In contrast, in inter prediction, a block in the current picture is predicted by a block in the current picture. The IR is coded by reference to other blocks in a different reference picture. AP pictures are coded without reference to any other pictures, so they are not the first Other pictures can be decoded without decoding. Therefore, the decoder can decode any IRAP picture. In contrast, decoding of a video sequence can begin at a non-IRAP picture. A picture is coded with reference to other pictures, so that the decoder generally , decoding of a video sequence cannot start at a non-IRAP picture. AP pictures refresh the DPB because the IRAP picture is the start of the CVS and the CVS This is because pictures in S do not refer to pictures in the previous CVS. The decoder also ensures that coding errors associated with inter prediction are propagated through the IRAP picture. However, IRAP pictures can be used to prevent such errors. are significantly larger than non-IRAP pictures in terms of data size. The video sequence uses many non-IRAP It contains a 60-frame frame with fewer IRAP pictures interspersed among them. The CVS of a frame may contain one IRAP picture and 59 non-IRAP pictures.

[0035] In some cases, video coding systems are also referred to as 360-degree video. It may also be used to code virtual reality (VR) video, which allows users to It may also include a sphere of video content displayed as if the viewer were at the center of the sphere. Only a portion of the sphere, called the sphere, is visible to the user. A head-mounted display that selects and displays a spherical viewport based on the user's head movements. A head mounted display (HMD) may be used, which allows a person to physically exist in the virtual space depicted by the video. To achieve this result, each picture in the video sequence is , contains the entire sphere of video data for the corresponding instant. However, a small part of the picture (even Only a portion of the picture (for example, a single viewport) is displayed to the user. The rest of the picture is displayed in the The viewport is dynamically selected based on the user's head movements. Generally, the entire picture is transmitted so that it can be selected and displayed. This may lead to large video file sizes.

[0036] To improve coding efficiency, some systems divide pictures into subpictures. A subpicture is a defined spatial region of a picture. Each subpicture is , which contains the corresponding viewport of the picture. Video may be encoded in more than one resolution. Each resolution is encoded into a different sub-bitstream. When streaming, the coding system uses the current video being used by the user. Merging sub-bitstreams into a bitstream for transmission based on the In particular, the current viewport can be obtained from a high-resolution sub-bitstream. The unseen viewports are derived from the lower resolution bitstream. In this way, the highest quality video is displayed to the user, and lower quality videos are If the user selects a new viewport, the lower resolution video is presented to the user. The decoder determines that the new viewport is a higher resolution video. The encoder can then perform the merging process accordingly. When an IRAP picture is reached, the decoder will This approach improves the user's viewing experience. Significantly improves video compression without negative impact.

[0037] One concern with the above approach is that the amount of time required to change resolution is This is based on the length of time it takes to reach the RAP picture. The reader cannot start decoding a different video sequence at a non-IRAP picture. One way to reduce such latency is to use more IRAP pins. However, this increases the file size. To balance loading efficiency, different viewports / subpictures are It may include IRAP pictures at a frequency that is more likely to be seen, e.g., in viewports where the Some viewports may have more IRAP pictures than other viewports. In the context of soccer, a viewport that looks at the stands or ceiling is the view seen by the user. The basket and / or center are more likely to be exposed to such a viewport than to such a viewport. Viewports associated with the target may contain IRAP pictures more frequently.

[0038] This approach leads to other problems. In particular, the subpicture containing the viewport is simply Different sub-pictures have IRAP pictures at different frequencies. Some pictures contain both IRAP and non-IRAP subpictures. This is problematic because pictures are stored in the bitstream by using NAL units. A NAL unit is a parameter set or slice of a picture and the corresponding A storage unit that includes a slice header. An access unit is a unit that includes an entire picture. Thus, an access unit contains all of the NAL units associated with a picture. NAL units also contain a type that indicates the type of picture that contains the slice. In video systems, a single picture (e.g., contained in the same access unit) All related NAL units are required to be of the same type. The unit's storage mechanism allows pictures to be both IRAP and non-IRAP subpictures. When the

[0039] Disclosed herein is a method for encoding both IRAP and non-IRAP subpictures. It is a mechanism for adjusting the storage of NALs to support pictures containing This in turn allows for VR Rendering with different IRAP subpicture frequencies for different viewports. In a first example, the present disclosure provides a method for enabling mixed video. For example, the flag indicates whether the picture is an IRAP subpicture. This flag may indicate that the subpicture contains both IRAP and non-IRAP subpictures. The decoder must use different decoding methods to properly decode and display the picture / subpicture. This flag can be used in the Picture Parameter Set It may be stored in the PPS and may be called mixed_nalu_types_in_pic_flag.

[0040] In a second example, disclosed herein is a method for indicating whether a picture is mixed. A flag. For example, a flag indicates whether a picture is an IRAP subpicture or a non-IRAP subpicture. Furthermore, the flag may indicate that the mixed picture contains both the The picture is arranged to contain exactly two NAL unit types, one IRAP type and one non-IRAP type. For example, a picture is a random access decodable leading picture. Instantaneous Decoding Refresh (IDR) with leading picture (IDR_W_RADL), IDR without leading picture (IDR_N_LP), or clean random access (CRA) NAL unit type (CRA_NUT) Furthermore, a picture may contain an IRAP NAL unit containing only one of the following: Picture NAL unit type (TRAIL_NUT), random access decodable leading picture NAL unit type (RADL_NUT), or random access skip reading Non-IRAP NAL units containing only one of the picture (RASL) NAL unit types (RASL_NUT) Based on this flag, the decoder will select the picture / subpicture appropriately. Different sub-pictures may be treated differently when decoded in order to be decoded and displayed simultaneously. This flag may be stored in the PPS and is called mixed_nalu_types_in_pic_flag. Good too.

[0041] FIG. 1 is a flow diagram of an exemplary operational method 100 of coding a video signal. The video signal is encoded in an encoder. The encoding process uses various mechanisms. Compresses the video signal by reducing the video file size using The compact file size allows for the compression of compressed video while reducing the associated bandwidth overhead. The decoder then converts the compressed video to Decodes the video file and reconstructs the original video signal for display to the end user. In general, the decoding process is carried out in a way that allows the decoder to consistently reconstruct the video signal. It faithfully mimics the encoding process to achieve

[0042] In step 101, a video signal is input to an encoder. The signal may be an uncompressed video file stored in memory. Video files are captured by a video capture device such as a video camera and Video files may be encoded to support live streaming of video. may contain both audio and video components. A component is a sequence of image frames that, when viewed in sequence, give the visual impression of motion. A frame contains light and a sample, referred to herein as luma components (or luma samples). , and a color called a chroma component (or color sample). In some instances, the frames may also include depth values ​​to support three-dimensional viewing. stomach.

[0043] In step 103, the video is partitioned into blocks. This involves subdividing pixels into square and / or rectangular blocks for compression. For example, High Efficiency Video Coding (HEVC), also known as H.265 and MPEG-H Part 2 In the image, the frame is first cut to a predefined size (e.g., 64 pixels x 64 The CTU can be divided into coding tree blocks, which are blocks of lumped pixels. The CTU is divided into blocks, which then contain both chroma and chromatic samples. The code is then used to repeatedly subdivide the blocks until a structure that supports the coding is achieved. For example, the luma component of a frame may be represented as individual blocks: It may be subdivided to include values ​​for relatively uniform lighting. For example, The chroma components may be subdivided until each block contains a relatively uniform color value. Therefore, the segmentation mechanism varies depending on the content of the video frame.

[0044] In step 105, the image blocks divided in step 103 are compressed. Various compression mechanisms are used for this purpose, for example inter prediction and / or intra prediction. Inter prediction may be used. Inter prediction is the prediction that objects in a typical scene are detected in successive frames. It is designed to take advantage of the fact that the reference frame Blocks depicting objects in a frame do not need to be repeated in neighboring frames. In particular, objects such as tables may remain in a fixed position over multiple frames. Therefore, the table is shown once and adjacent frames are displayed by looking back at the reference frame. The pattern matching mechanism allows for multiple frames to be referenced. Furthermore, if a moving object is detected, e.g., It may be displayed across multiple frames due to object movement or camera movement. As an example of a specific example, a video may show a car moving across the screen over several frames. A motion vector may be used to indicate such movement. gives the offset from the coordinates of the object in the frame to the coordinates of the object in the reference frame It is a two-dimensional vector. Therefore, inter prediction is performed for image blocks in the current frame. as a set of motion vectors indicating the offsets from the corresponding blocks in the reference frame It can be encoded.

[0045] Intra prediction encodes blocks within a common frame. Take advantage of the fact that minute and chrominance components tend to clump together within a frame, e.g. ,Some green areas of a tree tend to be located in the vicinity of similar green areas. supports multiple directional prediction modes (e.g., 33 in HEVC), planar modes, and DC Use the (DC) mode. The directional mode is used to determine whether the current block is in the same direction as the neighboring blocks in the corresponding direction. Planar mode indicates a series of blocks along a row / column. This shows that blocks (e.g., planes) can be interpolated based on neighboring blocks at the ends of rows. In this case, the planar mode uses a relatively constant gradient of changing values ​​to separate the rows / columns. Shows smooth transitions of light / color. DC mode is used for boundary smoothing and blocks The average value related to the samples of all neighboring blocks related to the angular direction of the directional prediction mode is Therefore, the intra predicted block is the same as the image block. Instead of actual values, the blocks can be expressed as values ​​of various related prediction modes. The measurement block may represent the image block as a motion vector value instead of the actual value. In any case, the predicted block may not accurately represent the image block in some cases. All the differences are stored in the residual block. To further compress the file, A transformation may be applied to the block.

[0046] In step 107, various filtering techniques may be applied. In this case, the filter is applied by an in-loop filtering method. Clock-based prediction results in blocky images being generated at the decoder. Furthermore, block-based prediction schemes may encode blocks and then The blocks may be reconstructed for later use as reference blocks. The filtering method is noise suppression filter, deblocking filter, adaptive loop filter. , and apply a sample adaptive offset (SAO) filter iteratively to the block / frame These filters are used to ensure that the encoded file can be reconstructed accurately. These filters also reduce blocking artifacts. Further processing is performed in subsequent blocks where the object is coded based on the reconstructed reference block. in the reconstructed reference block so that it is less likely to introduce artifacts Reduces artifacts.

[0047] Once the video signal is segmented, compressed, and filtered, the resulting digital The data is encoded into a bitstream in step 109. The data discussed above and the and any signaling data desired for the purpose of the data, prediction data, residual blocks, and various The bitstream may contain flags for transmission to the decoder on request. The bitstream may also be broadcast to multiple decoders. The generation of the bitstream may be performed in an iterative process. Therefore, steps 101, 103, 105, 107, and 109 are performed for many frames. The steps may be performed sequentially and / or simultaneously across multiple blocks. The order given is presented for clarity and ease of review and is intended to be a guide for video coding processes. It is not intended to limit the processes to any particular order.

[0048] The decoder receives the bitstream and begins the decoding process in step 111. In particular, the decoder uses an entropy decoding scheme to convert the bitstream into the corresponding The decoder converts the bit stream into a syntax and video data in step 111. The syntax data from the stream is used to determine the partitions for the frame. The division should match the result of the block division in step 103. Step 1 The entropy coding / decoding used in 11 is described below. The imager selects a block from several possible choices based on the spatial location of values ​​in the input image. Many choices are made during the compression process, including the choice of block segmentation method. Signaling the selection may use multiple bins. When used, bins are binary values ​​that are treated as variables (e.g., bits that may change depending on the situation). Entropy coding is a technique that encoders can clearly do better in certain cases. It allows us to discard all unsatisfactory choices and leave a set of acceptable choices. Then, each allowable choice is assigned a codeword. The size is based on the number of allowable choices (e.g., one bin for two choices, one for four choices, etc.). (e.g., two bins for each choice). The encoder then generates a code for the selected choice. This method encodes the codeword as a potentially unique combination of all possible choices. A small subset of acceptable options as opposed to a unique choice from a larger set The size of the codeword is as large as desired to uniquely represent these choices. The decoder then determines the set of allowable choices in the same way as the encoder. The selection is decoded by determining the set of allowable choices. The decoder can read the codeword and determine the choice made by the encoder.

[0049] In step 113, the decoder performs the decoding of the block. In particular, the decoder: The decoder then uses the inverse transform to generate the residual block. The corresponding prediction block is used to reconstruct the image block according to the classification. The block is a combination of the intra-predicted block and the inter-predicted block generated by the encoder in step 105. The reconstructed image block may then be processed in step 11. 1. Position within the frame of the reconstructed video signal according to the segmentation data determined in 1. The syntax for step 113 is also the same as the entropy code discussed above. The information may be signaled in the bitstream by the

[0050] In step 115, the reconstructed video signal is processed in the same manner as in step 107 of the encoder. Filtering is performed on the frames of the signal. For example, noise suppression filters, digital filters, The blocking filter, adaptive loop filter, and SAO filter are used to may be applied to the frame to remove artifacts. The video signal is then displayed in step 117 for viewing by the end user. It can be output to play.

[0051] FIG. 2 illustrates an exemplary coding and decoding (codec) system for video coding. 1 is a schematic diagram of a codec system 200. In particular, the codec system 200 supports the implementation of the operating method 100. The codec system 200 provides the functionality for encoding and decoding. The codec system 200 is generalized to depict components used in both , receiving a video signal as discussed in connection with steps 101 and 103 of the method of operation 100. , resulting in a segmented video signal 201. System 200 may be configured as described in connection with steps 105, 107, and 109 of method 100. When acting as a coder, it converts the segmented video signal 201 into a coded bitstream. When acting as a decoder, the codec system 200 performs the operations of the method 100. From the bitstream, as discussed in connection with steps 111, 113, 115, and 117 The codec system 200 includes a general coder control component 211, which generates an output video signal. Transform, scaling and quantization component 213, intra-picture estimation component 215, A trap picture prediction component 217, a motion compensation component 219, a motion estimation component 221, a scaling component Transformation and Inversion Components 229, Filter Control Analysis Components 227, In-Loop Filter Configuration Components 225, the decoded picture buffer component 223, and the header format and context Context adaptive binary arithmetic coding (CABAC) structure 2. The black The solid lines indicate the movement of data being coded / decoded, while the dashed lines indicate other components. The components of the codec system 200 are The decoder may be a sub-series of the components of the codec system 200. For example, the decoder may include an intra-picture prediction component 217, Motion compensation component 219, scaling and inverse transform component 229, in-loop filter component These components may include a decoded picture buffer component 225, a decoded picture buffer component 223, and , which will be explained below.

[0052] The segmented video signal 201 is segmented into blocks of pixels by a coding tree. The coding tree is a representation of the various segments of a captured video sequence. Use divide mode to subdivide blocks of pixels into smaller blocks of pixels These blocks can then be subdivided into smaller blocks. They may also be called nodes in the coding tree. Larger parent nodes have smaller child nodes. The number of times a node is sub-divided depends on the depth of the node / coding tree. The divided blocks are sometimes included in coding units (CUs). For example, a CU may have a luma block, a red difference chroma (Cr) block, and the blue difference chroma (Cb) block in the corresponding syntax for CU. It may be a sub-portion of the CTU that contains the instructions. The partitioning mode is To partition a node into two, three, or four child nodes, each with a shape that changes depending on the node. This may include binary trees (BT), ternary trees (TT), and quad trees (QT) used in partitioned The video signal 201 passes through a general coder control component 211 for compression, conversion, scaling and and quantization component 213, intra-picture estimation component 215, filter control analysis component 22 7, as well as to the motion estimation component 221.

[0053] The general coder control component 211 converts the video sequence into a bitstream according to the application constraints. The system is configured to make decisions related to coding of the images of the sequence. The adaptive coder control component 211 optimizes the bit rate / bitstream size versus reconstruction quality. Such decisions depend on storage space / bandwidth availability and image resolution. The general coder control component 211 may also be configured to To mitigate underrun and overrun issues, the buffer size is adjusted based on the transmission rate. To manage these issues, the general coder control component 211 manages the segmentation, prediction, and filtering by other components of the The adaptive coder control component 211 adjusts the compression complexity to increase resolution and bandwidth usage. Dynamically increase or decrease the resolution and compression complexity to reduce bandwidth usage Therefore, the general coder control component 211 may dynamically reduce the playback speed of the video signal. The codec system 200's The general coder control component 211 controls the operation of the other components. The control data also generates parameters for decoding in the decoder. The header format is to be encoded into the bitstream to signal the and CABAC component 231.

[0054] The segmented video signal 201 is fed to a motion estimation component 221 and a motion vector estimation component 222 for inter prediction. Also transmitted to compensation component 219. Frames or slices of segmented video signal 201 The motion estimation component 221 and the motion compensation component 222 may be divided into multiple video blocks. The component 219 selects one or more blocks in one or more reference frames to provide temporal prediction. , and performs inter-predictive coding of the received video block for the The system 200 may, for example, select an appropriate coding mode for each block of video data. Multiple coding passes may be performed to select the code.

[0055] The motion estimation component 221 and the motion compensation component 219 may be highly integrated, but generally For illustrative purposes, they are shown separately. Motion estimation performed by motion estimation component 221 Motion vector estimation is the process of generating motion vectors that estimate motion for video blocks. The vector is, for example, the displacement of the object to be coded relative to the predicted block. A predicted block may refer to a block that is coded in terms of pixel differences. A predicted block is also called a reference block. Such pixel differences may be calculated using the sum of absolute differences (SAD), sum of squared differences (SSD), or may be determined by other difference metrics. HEVC uses CTU, coding tree Uses several coded objects, including blocks (CTBs) and CUs For example, a CTU can be split into CTBs, and then the CTBs are split into CUs. A CU can be divided into CBs for the purpose of predicting the CU. (PU) and / or a transform unit (TU) containing the transformed residual data. The motion estimation component 221 uses rate-distortion analysis as part of the rate-distortion optimization process. For example, motion estimation components are used to generate motion vectors, PUs, and TUs. The element 221 includes a plurality of reference blocks, a plurality of motion vectors, etc. for the current block / frame. The reference block, motion vector, etc., that has the best rate-distortion characteristics may be determined. The best rate-distortion performance is determined by the quality of the video reconstruction (e.g., compression Both the coding efficiency (e.g., the amount of data loss due to Balance the two.

[0056] In some examples, the codec system 200 stores the decoded picture in the decoded picture buffer component 223. It may also calculate values ​​for sub-integer pixel locations of stored reference pictures. , the video codec system 200 uses quarter pixel positions, eighth pixel positions, etc., of the reference picture. The values ​​of the pixel positions, or other fractional pixel positions, may be interpolated. Therefore, the motion estimation component 221 estimates the full pixel position and the fractional pixel position. May perform position-related motion searches and output motion vectors with fractional pixel accuracy The motion estimation component 221 compares the position of the PU with the position of the prediction block in the reference picture. The motion vectors of the video blocks in the slices inter-coded by The motion estimation component 221 calculates the calculated motion vectors for encoding. The motion is output as motion data to the header format and CABAC component 231, and the motion is motion compensated. Output to component 219.

[0057] The motion compensation performed by the motion compensation component 219 is determined by the motion estimation component 221. deriving or generating a prediction block based on the determined motion vector. Again, the motion estimation component 221 and the motion compensation component 219 may in some instances be receiving a motion vector for the PU of the current video block; Motion compensation component 219 may then find the predictive block to which the motion vector points. Then, the pixel values ​​of the predicted block are calculated from the pixel values ​​of the current video block being coded. A residual video block is formed by subtracting pixel values ​​to form pixel difference values. Generally, the motion estimation component 221 performs motion estimation related to the luma component and The compensation component 219 is calculated based on the luma component for both the chroma and luma components. The predicted block and the residual block are transformed, scaled, and and forwarded to the quantization component 213.

[0058] The segmented video signal 201 includes an intra-picture estimation component 215 and an intra-picture estimation component 216. The motion estimation component 221 and the motion compensation component 219 are also transmitted to the motion prediction component 217. Similarly, the intra-picture estimation component 215 and the intra-picture prediction component 217 are , may be highly integrated but are shown separately for conceptual purposes. The pixel estimation component 215 and the intra-picture prediction component 217 perform the inter-frame prediction as described above. Instead of the inter prediction performed by the motion estimation component 221 and the motion compensation component 219 Instead, the current block is intra-predicted relative to blocks in the current frame. The intra-picture estimation component 215 estimates the intra-picture data used to encode the current block. In some examples, the intra picture estimation component 215 determines the intra prediction mode. , select the appropriate intra-prediction mode for coding the current block from the multiple tested intra-prediction modes. An intra-prediction mode is selected. The selected intra-prediction mode is then used for encoding. The data is forwarded to the header format and CABAC component 231 for processing.

[0059] For example, the intra picture estimation component 215 may be implemented using various tested intra prediction models. Rate-distortion analysis was used to calculate the rate-distortion values ​​for the modes tested. The rate-distortion analysis is performed by selecting the intra prediction mode with the best rate-distortion characteristics. Generally, there are coded blocks and the blocks that were coded to generate the coded blocks. The amount of distortion (or error) between the original uncoded block and the coded block Determine the bit rate (e.g., number of bits) used to generate the image. The picture estimation component 215 determines which intra prediction mode is the best rate for the block. The distortion and level for the various coded blocks are then calculated to determine which of the blocks exhibits the highest distortion value. Furthermore, the intra-picture estimation component 215 calculates the rate-distortion ratio from the Depth Map Depth Blocks are extracted using Depth Modeling Mode (DMM) based on RDO optimization. It may be configured to code.

[0060] When implemented in an encoder, the intra-picture prediction component 217 The predicted block is generated based on the selected intra-prediction mode determined by the block estimation component 215. When implemented in a decoder, the residual block is generated from the bitstream. The residual block may be read from the predicted block, represented as a matrix. The residual block contains the difference between the original and the original block. The residual block is then transformed, scaled, and and quantization component 213. The picture prediction component 217 may operate on both the luma and chroma components.

[0061] The transform, scaling and quantization component 213 further compresses the residual block. The transform, scaling and quantization component 213 performs discrete computations on the residual block. Apply a transform such as the discrete sine transform (DCT), discrete sine transform (DST), or a similar transform, and then Generates video blocks containing differential transform coefficient values. A transform, or other types of transforms, may also be used. Transformations, scaling, and quantization can be performed using a transform domain such as the frequency domain. Component 213 may, for example, be configured to scale the transformed residual information based on frequency. Such scaling is further configured such that different frequency information is quantified with different granularity. This involves applying a scale factor to the residual information so that the reconstructed Transformation, scaling and quantization structures may affect the final visual quality of the video. The construction element 213 is further configured to quantize the transform coefficients to further reduce the bit rate. The quantization process reduces the bit depth associated with some or all of the coefficients. The degree of quantization may be modified by adjusting the quantization parameter. In some examples, the transform, scale and quantize component 213 then performs quantization. A scan of the matrix containing the quantized transform coefficients may be performed. The quantized transform coefficients are then bit-wise converted. Header format and CABAC component 231 for encoding into the data stream. do.

[0062] The scaling and inverse transform component 229 performs transformation and scaling to support motion estimation. Apply the inverse operation of the scaling and quantization component 213. Element 229 may be, for example, a reference block that may be a prediction block for another current block. Inverse scanning is performed to reconstruct the residual block in the pixel domain for later use as a block. Apply scaling, inverse transform, and / or inverse quantization. Or the motion compensation component 219 may use the motion estimation of a subsequent block / frame. Calculate the reference block by adding the residual block back to the predicted block corresponding to To reduce artifacts introduced during scaling, quantization, and transformation, To do this, a filter is applied to the reconstructed reference block. Artifacts result in inaccurate predictions when subsequent blocks are predicted (and and further artifacts).

[0063] The filter control analysis component 227 and the in-loop filter component 225 control the residual block and and / or apply filters to the reconstructed image blocks, e.g., scaling and the transformed residual block from the inverse transform component 229 to reconstruct the original image block. from the intra-picture prediction component 217 and / or the motion compensation component 219 to The reconstructed image block may then be combined with the corresponding predicted block. In some cases, a filter may be applied to the residual block instead. 2. As with the other components in FIG. 2, the filter control analysis component 227 and and in-loop filter component 225 may be highly integrated and implemented together, Shown separately for conceptual purposes: Filter applied to reconstructed reference block is applied to a particular spatial region and examines how such a filter is applied. The filter control analysis component 227 includes multiple parameters for adjusting such filters. analyzing the reconstructed reference block to determine where the filter should be applied; Set the corresponding parameters. Such data is used as filter control data for encoding. The data is forwarded to the header format and CABAC component 231. The composition element 225 applies such filters based on the filter control data. The deblocking filter, noise suppression filter, SAO filter, and adaptive loop filter are Such filters may include, depending on the example, (e.g., reconstructed in the spatial / pixel domain (for a given pixel block) or in the frequency domain. may be used.

[0064] When acting as an encoder, the filtered reconstructed image blocks , residual blocks, and / or predicted blocks are used in motion estimation, as discussed above. The decoded picture is stored in the decoded picture buffer component 223 for later use. When operating in this manner, the decoded picture buffer component 223 is reconstructed and filtered. The decoded blocks are stored and transmitted to the display as part of the output video signal. The picture buffer component 223 stores predicted blocks, residual blocks, and / or reconstructed blocks. The memory device may be any memory device capable of storing the generated image blocks.

[0065] The header format and CABAC components 231 are used to represent various configurations of the codec system 200. Coded to receive data from elements and transmit such data to a decoder In particular, the header format and CABAC components 231 for encoding control data such as general control data and filter control data. Generates various headers. In addition, prediction data including intra prediction and motion data The residual data in the form of quantized and unquantized transform coefficients is all coded into the bitstream. The final bitstream is then used to reconstruct the original segmented video signal 201. Contains all the information desired by the decoder. Such information includes intra-prediction modes index table (also called codeword mapping table), various block Definition of coding context for the block, indication of most likely intra prediction mode Such data may also include information about the engine, engine location, and indication of partition information. For example, the information may be encoded using tropy coding. Context adaptive variable length coding (CAVLC) g), CABAC, syntax-based context-adaptive binary arithmetic coding (SBAC) based context-adaptive binary arithmetic coding), Probability Interval Partition Entropy (PIPE: probability interval partitioning entropy) coding, or another entropy After entropy coding, the coded The encoded bitstream is then sent to another device (e.g., a video decoder). The information may be stored in a file or archived for later transmission or retrieval.

[0066] 3 is a block diagram illustrating an example video encoder 300. performs the encoding function of the codec system 200 and / or the steps of the operating method 100. may be used to perform steps 101, 103, 105, 107, and / or 109. The encoder 300 segments the input video signal and generates a signal substantially similar to the segmented video signal 201. The segmented video signal 301 is then The image is compressed and encoded into a bitstream by components of the encoder 300.

[0067] In particular, the segmented video signal 301 is subjected to an intra-picture prediction scheme for intra-prediction. The intra-picture prediction component 317 performs intra-picture estimation. The image prediction component 215 and the intra-picture prediction component 217 may be substantially similar to the image prediction component 215 and the intra-picture prediction component 217. The decoded video signal 301 is based on the reference blocks in the decoded picture buffer component 323. For inter prediction, the motion compensation component 321 also forwards the motion compensation The motion estimation component 221 and the motion compensation component 219 may be substantially similar to the intra-frame motion estimation component 221 and the motion compensation component 219. The prediction block and residual block from the picture prediction component 317 and the motion compensation component 321. The block is forwarded to the transform and quantization component 313 for transforming and quantizing the residual block. The transform and quantization component 313 is a transform, scaling and quantization component 213. The transformed and quantized residual block and the corresponding prediction The measurement block is then processed using an entropy coding structure for coding into the bitstream. The entropy coding configuration element 331 then forwards the entropy coding data (together with associated control data) to the entropy coding configuration element 331. Element 331 may be substantially similar to the header format and CABAC component 231. .

[0068] Also, the transformed and quantized residual block and / or the corresponding prediction block may be Transforms and quantities to reconstruct reference blocks for use by the motion compensation component 321. The inverse transform and quantization component 329 then transfers the inverse transform and quantization component 329 to the inverse transform and quantization component 329. Component 329 may be substantially similar to scaling and inverse transform component 229 . The in-loop filter of the in-loop filter component 325 may be configured to filter the residual block and / or the or is also applied to the reconstructed reference block. , substantially similar to the filter control analysis component 227 and the in-loop filter component 225. In-loop filter component 325 may be configured in conjunction with in-loop filter component 225. As discussed, multiple filters may be included. The filtered block is then The block is then added to the decoded picture buffer for use as a reference block by the motion compensation component 321. The decoded picture buffer component 323 stores the decoded picture buffer. The buffer component 223 may be substantially similar to the buffer component 223 .

[0069] 4 is a block diagram illustrating an example video decoder 400. Implementing the decoding function of the codec system 200 and / or step 1 of the operating method 100 11, 113, 115, and / or 117. Decoder 400 , for example, receives a bitstream from the encoder 300 and displays it to the end user. A reconstructed output video signal is generated based on the bitstream to

[0070] The bitstream is received by the entropy decoding component 433. The decoding component 433 may be CAVLC, CABAC, SBAC, PIPE coding, or other encoding. The decoder is configured to implement an entropy decoding scheme such as an entropy coding technique. For example, the entropy decoding component 433 may encode the data as codewords in the bitstream. Use header information to provide context for interpreting further data The decoded information may include general control data, filter control data, segment information, motion information, and the like. video data, such as the input data, prediction data, and quantized transform coefficients from the residual block The quantized transform coefficients contain any desired information for decoding the signal. , which is forwarded to the inverse transform and quantization component 429 for reconstruction into The component 429 may be similar to the inverse transform and quantize component 329 .

[0071] The reconstructed residual block and / or predicted block are based on an intra prediction operation. The image blocks are then forwarded to the intra-picture prediction component 417 for reconstruction into image blocks. The intra-picture prediction component 417 is a combination of the intra-picture estimation component 215 and the intra-picture In particular, the intra-picture prediction component may be substantially similar to the intra-picture prediction component 217. The component 417 uses the prediction mode to identify a reference block within the frame and performs intra prediction. Apply the residual block to the result to reconstruct the reconstructed image block. Intra-predicted image blocks and / or residual blocks and corresponding inter-predicted blocks The measured data is sent to the decoded picture buffer component 423 via the in-loop filter component 425. These are transferred to the decoded picture buffer component 223 and the in-loop filter component 224, respectively. In-loop filter component 425 may be substantially similar to in-loop filter component 225. Filtering the reconstructed image blocks, residual blocks, and / or predicted blocks Such information is stored in the decoded picture buffer component 423. The reconstructed image blocks from the buffer component 423 are motion compensated for inter prediction. The motion compensation component 421 is forwarded to the motion estimation component 221 and / or The motion compensation component 421 may be substantially similar to the motion compensation component 219. In particular, the motion compensation component 421 Generate a prediction block using the motion vector from the reference block to reconstruct the image block. The resulting reconstructed block is then , also forwarded to the decoded picture buffer component 423 via the in-loop filter component 425 The decoded picture buffer component 423 may store additional reconstructed image blocks. The reconstructed image blocks are then reconstructed into frames according to the partition information. Such frames may also be arranged in a sequence. The resulting signal is output to the display as a reconstructed output video signal.

[0072] FIG. 5 is a schematic diagram illustrating an exemplary CVS 500. For example, the CVS 500 may include a computer system according to the method 100. encoded by the decoder system 200 and / or an encoder such as encoder 300. Furthermore, the CVS 500 may be implemented in a manner similar to that of the codec system 200 and / or the decoder 400. The CVS 500 stores the coded pictures in decoding order 508. The decoding order 508 is the order in which the pictures are positioned in the bitstream. The pictures in CVS 500 are then output in presentation order 510. Presentation order 510 is the order in which the resulting This is the order in which pictures should be displayed by a decoder to properly display the video being For example, pictures in CVS 500 may generally be arranged in presentation order 510. However, certain pictures may be more closely related to other similar pictures, for example, to support inter-prediction. moved to different locations to increase coding efficiency by placing them closer together Shifting such pictures in this way results in a decoding order 508. In the example shown, the pictures are indexed in decoding order 508 from 0 to 4. In the presentation order 510, pictures with index 2 and index 3 are displayed before picture with index 0. It has been moved in front of Kucha.

[0073] The CVS 500 includes an IRAP picture 502. The IRAP picture 502 is a random access map for the CVS 500. A picture coded by intra prediction serves as an access point. In particular, blocks of the IRAP picture 502 are linked by reference to other blocks of the IRAP picture 502. An IRAP picture 502 is coded without reference to any other picture. Since it is mapped, it can be decoded without first decoding any other pictures. , the decoder can start decoding the CVS 500 at the IRAP picture 502. , the IRAP picture 502 may cause the DPB to be refreshed. The picture presented later is an IRAP picture 502 (e.g., picture There is no need to rely on the previous picture (index 0). Therefore, the picture buffer is The IRAP picture 502 may be refreshed as it is decoded. This is related to inter prediction. Since the coding error cannot propagate through the IRAP picture 502, This has the effect of preventing such errors. The IRAP picture 502 can be any type of picture. For example, an IRAP picture may be coded as IDR or CRA. IDR starts a new CVS 500 and executes an intraconverter that refreshes the picture buffer. CRA is not responsible for starting a new CVS 500 or for the Intra-coder that acts as a random access point without refreshing the buffer In this way, the leading picture5 related to the CRA is 04 may refer to a picture preceding the CRA, while the leading picture 50 associated with the IDR 4 does not need to refer to the picture before the IDR.

[0074] CVS 500 also includes various non-IRAP pictures. These include leading pictures 504 and The leading pictures 504 include the trailing pictures 506. The leading pictures 504 are IRAP pictures in decoding order 508. A picture positioned after the IRAP picture 502 in presentation order 510, but before the IRAP picture 502 in presentation order 510. The trailing picture 506 is an IRAP picture in both decoding order 508 and presentation order 510. The leading picture 504 and the trailing picture 505 are positioned after the leading picture 502. 6 are both coded using inter prediction. The code refers to the IRAP picture 502 or a picture positioned after the IRAP picture 502. Therefore, the trailing picture 506 is always decoded after the IRAP picture 502 is decoded. The leading picture 504 can be decoded by random access skip reading. may include Random Access Decodable Reading (RASL) and Random Access Decodable Reading (RADL) pictures. The RASL picture is coded with reference to the picture before the IRAP picture 502. However, it is coded at a position after the IRAP picture 502. Since it is picture-dependent, when the decoder starts decoding at IRAP picture 502, Therefore, the RASL picture cannot be decoded even if the IRAP picture 502 is a random access point. When used as a data entry, it is skipped and not decoded. However, RASL pictures are The coder randomly accesses the previous IRAP picture (before index 0, not shown). When used as a point, it is decoded and displayed. RADL pictures are IRAP pictures 50 2 and / or coded with reference to pictures following the IRAP picture 502, but presented The RADL picture is positioned before the IRAP picture 502 in the order 510. Since the IRAP picture 502 is a random access point, can be decoded and displayed.

[0075] Pictures from the CVS 500 may be stored in access units, respectively. A picture may be partitioned into slices, which are contained in NAL units. A NAL unit may be a parameter set or slice of a picture and the corresponding slice. A NAL unit is a storage unit that contains a data header. For example, from IRAP picture 502 The slice is an IDR (IDR_W_RADL) NAL unit with RADL, a leading picture, and a It may be included in an IDR (IDR_N_LP) NAL unit, a CRA NAL unit, etc. The L unit is an IDR picture where an IRAP picture 502 is associated with a RADL leading picture 504. The IDR_N_LP NAL unit indicates that the IRAP picture 502 is a leading CRA NAL unit 504. is a CRA picture that the IRAP picture 502 may be associated with a leading picture 504. Slices of non-IRAP pictures may also be placed in NAL units. For example, a slice of the trailing picture 506 is Trailing picture NAL unit tag indicating a predictively coded picture The slices of the leading picture 504 may be arranged in the corresponding The picture corresponds to a type of inter-predictive coded leading picture 504. RASL NAL unit type (RASL_NUT) and / or RADL NAL unit type (RASL_NUT) It may be included in the RADL_NUT type. It signals the slice of a picture within the corresponding NAL unit. Nulling allows the decoder to determine the appropriate decoding method to apply to each picture / slice. The mechanism can be easily determined.

[0076] FIG. 6 shows a VR picture video stream 600 divided into multiple sub-picture video streams. 6 is a schematic diagram showing streams 601, 602, and 603. For example, sub-picture video streams Each of the streams 601 to 603 and / or the VR picture video stream 600 is coded into the CVS 500. Therefore, the sub-picture video streams 601 to 603 and / or The VR picture video stream 600 may be generated by the codec system 200 according to the method 100 and / or may be encoded by an encoder such as the encoder 300. The video streams 601 to 603 and / or the VR picture video stream 600 are coded by It may be decoded by system 200 and / or a decoder such as decoder 400.

[0077] The VR picture video stream 600 includes multiple pictures presented over time. VR is a sphere of video content that can be displayed as if the user is at the center of the sphere. Each picture contains the entire sphere, while the viewport and Only a portion of the picture known as the A head-mounted display that selects and displays a spherical viewport based on the user's head movements. A head mounted display (HMD) may be used, which allows the user to physically enter the virtual space depicted by the video. To achieve this result, each picture in the video sequence contains the entire sphere of video data for the corresponding instant. However, a small part of the picture (e.g. Only a portion of the picture (e.g., a single viewport) is displayed to the user. The rest of the picture is displayed in the The viewports are dynamically changed depending on the user's head movements. Generally, the entire picture is transmitted so that it can be selected and displayed.

[0078] In the example shown, the pictures of the VR picture video stream 600 are Each picture and and corresponding subpictures are positioned at a temporal position (e.g., picture) as part of the temporal presentation. The sub-picture video streams 601 to 603 are arranged so that the subdivision is consistent over time. Such a consistent subdivision is generated when the subpicture video is applied as a Streams 601 to 603 are generated, and each stream has a predetermined size, shape, and VR picture. It contains a set of sub-pictures whose spatial locations are relative to the corresponding picture in the video stream 600. Furthermore, a set of sub-pictures in the sub-picture video streams 601-603 may be Therefore, the sub-picture video streams 601 to 603 are located at different time positions. The pictures may be aligned in the time domain based on their temporal positions. Sub-pictures from the intermediate sub-picture video streams 601-603 are Based on predefined spatial positions, the VR picture video stream 600 is reconstructed. In particular, the sub-picture video streams 601-603 can be merged in the spatial domain. Each of these sub-bitstreams can be coded into a separate sub-bitstream. When merged together, the frames result in a bitstream containing the entire set of pictures over time. The resulting bitstream is then rendered in the user's currently selected viewport. The image may be decoded based on the original image and sent to a decoder for display.

[0079] One of the problems with VR video is that all of the sub-picture video streams 601 to 603 are high-quality. This means that the decoder may transmit the image to the user in a high quality (e.g., high resolution) Dynamically selects the user's current viewport and provides the corresponding sub-picture video stream 60 It allows you to display subpictures from 1 to 603 in real time. For example, only a single viewport from the subpicture video stream 601 is seen. On the other hand, the sub-picture video streams 602-603 are discarded. Therefore, transmitting the sub-picture video streams 602-603 at high quality requires a significant amount of bandwidth. To improve coding efficiency, VR video is often split into multiple video streams. Each video stream 600 may be encoded into a different quality / resolution. In this way, the decoder encodes the current sub-picture video stream 601 In response, the encoder (or intermediate slicer) Ediate slicer or other content server) delivers higher quality video streams6 00 and a higher quality sub-picture video stream 601 from A lower quality sub-picture video stream 602-603 from stream 600 can be selected. The encoder then generates such sub-bitstreams for transmission to the decoder. can be merged into a complete coded bitstream. The reader receives a series of pictures, and the current viewport is of higher quality and the other The viewport is of lower quality. Furthermore, the highest quality subpicture is The lower quality subpicture is generally displayed to the user (when not in use), and the lower quality subpicture is generally discarded, this is a trade-off between functionality and coding efficiency.

[0080] When the user shifts his / her attention from the sub-picture video stream 601 to the sub-picture video stream 602, If the current sub-picture video stream 602 is shifted, the decoder will determine whether the new current sub-picture video stream 602 is of higher quality. The encoder then adjusts the merging mechanism accordingly. As mentioned above, the decoder can change the new CVS 500 only in the IRAP picture 502. Therefore, the sub-picture video stream 602 can start decoding the IRAP picture / sub-picture video stream 602. The IRAP picture is then displayed at a lower quality until it reaches the subpicture. A higher quality version of the video stream 602 is decoded. This technique allows for the reduction of video compression without adversely affecting the user's viewing experience. Significantly improves shrinkage.

[0081] One concern with the above approach is the amount of time required to change resolution. It is based on the length of time it takes to reach an IRAP picture in the video stream. indicates that the decoder is using different versions of the sub-picture video stream in a non-IRAP picture. This is because the decoding of the program 602 cannot be started. One way to do this is to include more IRAP pictures. To balance functionality and coding efficiency, different The viewport / subpicture video streams 601-603 are IRAP pictures at different frequencies. For example, viewport / subpicture video may be more likely to be seen. The video streams 601 to 603 are the other viewport / subpicture video streams 601 to 60 It is possible to have more than three IRAP pictures. For example, in the context of basketball, The viewport / subpicture video streams 601-603 are for viewing the stand or ceiling. Such viewport / subpicture video is less likely to be seen by Viewpoints related to the basket and / or center court rather than streams 601-603 Even if the IRAP / subpicture video streams 601 to 603 contain IRAP pictures more frequently, good.

[0082] This approach leads to further problems, especially when using sub-picture video streams that share POC. The sub-pictures from streams 601-603 are part of a single picture. Slices from a picture are included in NAL units based on the picture type. In a video coding system, all NAL units related to a single picture are , are constrained to contain the same NAL unit types. When the frames 601 to 603 have IRAP pictures at different frequencies, some of the pictures are IRAP sub-pictures. This includes both IRAP and non-IRAP subpictures, since each single picture is This violates the constraint that only NAL units of the same type should be used.

[0083] This disclosure provides a method for determining whether all NAL units for a slice in a picture have the same NAL unit type. This problem is addressed by removing the constraint that the picture must be By removing this restriction, access units can be An event may contain both IRAP and non-IRAP NAL unit types. Furthermore, pictures / access units are classified as IRAP and non-IRAP NAL unit types. A flag may be coded to indicate when the mix includes a mix with a The mixed NAL unit types in picture flag )(mixed_nalu_types_in_pic_flag). In addition, a single mixed picture / access unit A packet contains only one type of IRAP NAL unit and one type of non-IRAP NAL unit. This is because the constraint that unintended NAL units may be Prevents type mixing from occurring. If such mixing is allowed, the decoder The code must be designed to manage such a mix. Eliminates the complexity of the hardware required without adding any additional benefits to the process For example, a mixed picture can be selected from IDR_W_RADL, IDR_N_LP, or CRA_NUT. Additionally, a mixed picture may contain one of the types of IRAP NAL units specified in TRAIL Contains one type of non-IRAP NAL unit selected from _NUT, RADL_NUT, and RASL_NUT. An exemplary implementation of this scheme is discussed in more detail below.

[0084] FIG. 7 illustrates an example bitstream containing pictures with mixed NAL unit types. 7 is a schematic diagram illustrating a bitstream 700. For example, the bitstream 700 may be a codec according to the method 100. 400 for decoding by the codec system 200 and / or the decoder 400. The bitstream 700 may be generated by the encoder 300. VR picture merged from multiple sub-picture video streams 601 to 603 of different video resolutions. Each sub-picture video stream may include a sub-picture video stream 600, each sub-picture video stream having a different spatial resolution. Including CVS 500 in intermediate locations.

[0085] The bitstream 700 includes a sequence parameter set (SPS) 710, a number of picture parameters, and The SPS 71 includes a meter set (PPS) 711, a plurality of slice headers 715, and image data 720. 0 is a sequence common to all pictures in the video sequence contained in the bitstream 700. Such data includes picture sizing, bit depth, coding It may contain parameters for the coding tools, bitrate constraints, etc. Each picture in a video sequence contains parameters that apply to the entire image. , PPS 711. Each picture may reference a PPS 711, but in some cases, Note that a single PPS 711 may contain data for multiple pictures. For example, multiple similar pictures may be coded with similar parameters. In such cases, a single PPS 711 may contain data about such similar pictures. PPS 711 specifies the coding tools available for slices within the corresponding picture. The slice header 715 may indicate the size of the image within the picture, the quantization parameters, offsets, etc. It contains parameters specific to each slice in a video sequence. There may be one slice header 715 for each slice. The slice header 715 contains slice type information , Picture Order Count (POC), Reference Picture List, Prediction Weight, Tile Entry Point It may also contain tile entry points, deblocking parameters, etc. Note that da 715 may also be called a tile group header depending on the context. .

[0086] The image data 720 may be video data encoded by inter-prediction and / or intra-prediction. The video data includes the corresponding transformed and quantized residual data. A sequence includes multiple pictures 721 coded as image data 720. 721 is a single frame of a video sequence and therefore generally However, the subpicture 723 is displayed as a single unit when the sequence is displayed. The picture 721 may be displayed to implement certain technologies such as virtual reality. , PPS 711. Picture 721 may contain sub-pictures 723, tiles, and / or A sub-picture 723 may be divided into slices. A subpicture is a spatial region of a picture 721 that is consistently applied to all subpictures. The image 723 may be displayed by an HMD in a VR context. The sub-picture 723 is obtained from the sub-picture video streams 601 to 603 of the corresponding resolution. Subpictures 723 may reference SPS 710. In some systems, The slice 725 is called a tile group, which contains the tiles. The tile group of a tile references the slice header 715. A slice 725 is a single NAL unit. An integer number of complete tiles or an integer number within a tile of a picture 721 that are exclusively contained in the unit Thus, a slice 725 may be defined as a row of CTUs and The CTU / CTB is further divided into a coding tree based coding unit and / or a coding table. The coding blocks are then further divided into coding blocks. can be encoded / decoded by

[0087] Parameter sets and / or slices 725 are coded into NAL units. An AL unit is a syntactic structure that contains an indication of the type of data that follows, and and inserting emulation prevention bytes here and there as needed. It may be defined as the byte containing the data in the form of the entered RBSP. A NAL unit is a parameter set or slice 725 of a picture 721 and the corresponding slice In particular, a VCL NAL unit 740 is a storage unit that contains a picture 721 slice. 725 and the corresponding slice header 715. In addition, non-VCL NAL Unit 730 contains parameter sets such as SPS 710 and PPS 711. For example, SPS 710 and PPS 711 are both SPS NAL unit type (SPS_NUT) 731, which is a non-VCL NAL unit 730, and PPS NAL units They may each be included in type (PPS_NUT) 732.

[0088] As mentioned above, an IRAP picture, such as IRAP picture 502, is contained in an IRAP NAL unit 745. Non-IRAP pictures such as leading picture 504 and trailing picture 506 are , may be included in non-IRAP NAL units 749. In particular, IRAP NAL units 745 may be included in IRAP pictures or It is any NAL unit that contains a slice 725 obtained from a non-IRAP NAL unit or a sub-picture. The AL unit 749 can be used to capture any picture that is not an IRAP picture or a subpicture (e.g., Any slices 725 taken from the image (leading and trailing pictures) IRAP NAL unit 745 and non-IRAP NAL unit 749 are both Since both contain slice data, they are both VCL NAL units 740. In this case, IRAP NAL unit 745 is an IDR or RADL picture without a leading picture. The slice 725 from the IDR associated with the channel is converted to an IDR_N_LP NAL unit 741 or an IDR_w_RADL NAL In addition, the IRAP NAL unit 745 may be included in a CRA picture or a CRA unit. These slices 725 may be included in the CRA_NUT 743. In an exemplary embodiment, non-IRAP NAL Unit 749 receives a RASL picture, a RADL picture, or a SL from a trailing picture. The device 725 may be included in the RASL_NUT 746, the RADL_NUT 747, or the TRAIL_NUT 748, respectively. In an exemplary embodiment, the complete list of possible NAL units is listed in the NAL unit type The results are sorted by and shown below. [Table 1A] [Table 1B] [Table 1C]

[0089] As mentioned above, the VR video stream consists of sub-pictures with IRAP pictures at different frequencies. This may include a pixel 723, which may be more relevant for spatial areas that are less likely to be seen by the user. Fewer IRAP pictures are used, for spatial regions that users are likely to look at frequently. This allows more IRAP pictures to be used. Spatial regions that are likely to return frequently can be quickly adjusted to higher resolution. The method results in a picture 721 that contains both IRAP NAL units 745 and non-IRAP NAL units 749. When this happens, the picture 721 is called a mixed picture. This state is called an intra-picture mixed NAL unit. This can be signaled by the mixed_nalu_types_in_pic_flag 727. The mixed_nalu_types_in_pic_flag 727 may be set in the PPS 711. _types_in_pic_flag 727 indicates that each picture 721 that references a PPS 711 consists of two or more VCL NAL units. and the VCL NAL units 740 have the same value of NAL unit type (nal_unit_type). To specify that no mixed_nalu_type is used, it may be set equal to 1. s_in_pic_flag 727 indicates that each picture 721 that references a PPS 711 contains one or more VCL NAL units 740 and all VCL NAL units 740 of each picture 721 that references the PPS 711 are of type nal_unit_type may be set equal to 0 when both have the same value of

[0090] Additionally, when mixed_nalu_types_in_pic_flag 727 is set, the subpictures of picture 721 One or more VCL NAL units 740 in the structure 723 all have a first specification of the NAL unit type. and all other VCL NAL units 740 in picture 721 have a value of NAL unit type Constraints may be used that have different second specific values. For example, the constraint may be a mixture A picture 721 contains a single type of IRAP NAL unit 745 and a single type of non-IRAP NAL unit. For example, picture 721 may need to include one or more I DR_N_LP NAL unit 741, one or more IDR_w_RADL NAL units 742, or none Any combination of such IRAP NAL units 745 may contain one or more CRA_NUTs 743. Furthermore, a picture 721 may contain one or more RASL_NUTs 746, one or more or multiple RADL_NUTs 747, or one or more TRAIL_NUTs 748, It cannot contain any combination of IRAP NAL units 745 such as:

[0091] In the exemplary implementation, picture types are used to define the decoding process. Such a process may involve, for example, identifying pictures by picture order count (POC) Deriving information,marking the status of reference pictures in the Decoded Picture Buffer (DPB),D This includes the output of pictures from the PB. Pictures are all coded pictures. or its subparts may be identified by a type based on the NAL unit type that contains it. In some video coding systems, the picture type is instantaneous decoding refresh (I Other video coding systems may include IDR (Internal Dynamic Range) pictures and non-IDR pictures. In this example, the picture type is a trailing picture, a temporal sublayer access (TSA) temporal sub-layer) picture, step-wise temporal sub-layer access (STSA) sub-layer access) pictures, random access decodable reading (RADL) pictures, Random Access Skip Reading (RASL) Pictures, Broken Link Access (BLA : broken-link access) pictures, instant random access pictures, and clean run Such picture types may include dumb-access pictures. whether it is a sub-layer referenced picture or a sub-layer non-referenced picture. It may be further distinguished based on whether it is a sub-layer non-referenced picture. A BLA picture can be a BLA with a leading picture, a BLA with a RADL picture, or IDR pictures may be further distinguished as BLA pictures with no leading pictures. , further distinguished as IDRs with RADL pictures and IDRs without leading pictures. This may be done.

[0092] Such picture types are used to implement various video-related functions. For example, IDR, BLA, and / or CRA pictures may implement IRAP pictures. IRAP pictures may provide the following features / benefits: The presence of a picture may indicate that the decoding process can start from that picture. This function allows the decoding process to proceed as long as an IRAP picture is present at that position in the bitstream. Allows the implementation of random access features starting at a specified location. The presence of an IRAP picture does not necessarily mean that it is the beginning of a bitstream. A coded picture starting with an IRAP picture, excluding the IRAP picture, is placed before the IRAP picture. Redirect the decoding process so that the image is coded without any reference to the placed picture. Therefore, IRAP pictures located in the bitstream are refreshed. Therefore, coding errors placed before the IRAP picture are prevented from propagating. Decoding errors in the pictures that follow the IRAP picture in decoding order are reported through the IRAP picture. cannot be propagated to the picture.

[0093] IRAP pictures provide various features but come at a cost to compression efficiency. Therefore, the presence of IRAP pictures can cause a sudden increase in the bit rate. This rate penalty can have various causes. For example, IRAP pictures are much slower than non-IRAP pictures. are represented by significantly more bits than the inter-predicted pictures used as vectors. Furthermore, the existence of IRAP pictures is indicative of inter prediction. In particular, IRAP pictures do not use the previous reference picture from the DPB. Refresh the decoding process by removing the previous reference picture. It is the reference to be used for coding pictures that follow the IRAP picture in decoding order. This reduces the availability of pictures and therefore the efficiency of the process.

[0094] IDR pictures have a different signaling and derivation process than other IRAP picture types. For example, the signaling and derivation process associated with the IDR may be Instead of deriving the most significant bit (MSB) from the previous key picture, the MSB part of the POC is set to 0. Furthermore, the slice header of an IDR picture may contain a number of bits to assist in managing reference pictures. On the other hand, it may not contain information used for other purposes such as CRA, trailing, TSA, etc. The picture type is a reference picture that can be used to perform the reference picture marking process. Reference picture sets, such as a reference picture set (RPS) or a reference picture list The reference picture marking process may include reference picture information in the DPB. Whether the status is used for reference or not used for reference The presence of an IDR is the process by which the decryption process simply determines all references in the DPB. This indicates that the picture should be marked as unused for reference, so the IDR picture For architectures, such information does not need to be signaled.

[0095] In addition to the picture type, the picture identification information by POC is also used for reference in inter prediction. For managing the use of reference pictures, for outputting pictures from the DPB, the scaling of motion vectors is It is used for multiple purposes, such as for weighted prediction, for example, In some video coding systems, pictures in the DPB are used for short-term reference. used for long-term reference or not used for reference. A picture may be marked as not to be used for reference. Once such a picture is no longer available, it can no longer be used for prediction. Pictures can be deleted from the DPB when they are not needed for other video conferencing purposes. In a reading system, reference pictures are marked as short-term and long-term. A reference picture is a picture that is no longer needed for prediction reference. Between these statuses, The transition may be controlled by a marking process of the decoded reference pictures. Sliding window processes and / or explicit memory management control actions (MMCO) ) process may be used as a marking mechanism for decoded reference pictures. The sliding window process ensures that the number of reference frames in the SPS is max_num_ref_frame When the number of short-term reference pictures is equal to the specified maximum number, denoted as s, the short-term reference pictures are used for reference. The short-term reference picture is the most recently decoded short-term picture. The MMCO process may be stored in a first-in, first-out manner so that the data is held in the DPB. An MMCO command may contain multiple MMCO commands. An MMCO command may consist of one or more short-term or may mark long-term reference pictures as not being used for reference, All pictures may be marked as unused for reference or may be marking the reference picture or an existing short-term reference picture as long-term, and then The long-term reference picture may be assigned a long-term picture index.

[0096] In some video coding systems, reference picture marking operations and The process for outputting and removing pictures from the DPB after the picture is decoded is Other video coding systems use RPS for reference picture management. The most radical difference between the RPS mechanism and the MMCO / sliding window process is The main difference is that for each particular slice, the RPS is The goal is to provide a complete set of reference pictures to be used by subsequent pictures. Therefore, all pictures to be held in the DPB for use by current or future pictures The complete set of architectures is signaled in the RPS. This is relative to the DPB. Unlike the MMCO / sliding window method, where only changes are signaled, the RPS mechanism According to the algorithm, to maintain the correct status of reference pictures in the DPB, No information from the previous picture is required. The picture decoding order and DPB operation are It is used in some video coding systems to take advantage of the In some video coding systems, picture marking and Buffer operations, including both outputting and deleting decoded pictures from the DPB, are It may also be applied after the structure is decoded. First, the RPS is decoded from the slice header of the current picture, and then the RPS is decoded from the slice header of the picture. The marking and buffering operations of may be applied before decoding the current picture. .

[0097] In VVC, the management of reference pictures can be summarized as follows: Two reference picture lists, denoted as List 1 and List 2, are directly signaled and derived. They are based on the RPS or sliding window + MMCO process discussed above. The marking of reference pictures is done by dividing the active and inactive entries in the reference picture list. Based directly on the active entry and on the reference picture list 0 and 1, Only active entries are used as reference indices for inter prediction of CTUs. The information for the derivation of the two reference picture lists is the SPS, PPS, and SLPS. The syntax elements and structures in the device header signal the Predefined RPL structures are available for use by referencing them in slice headers. Two reference picture lists are signaled in a bidirectional inter-predicted (B) slice. All slices including unidirectional inter-predicted (P) slices and intra-predicted (I) slices The two reference picture lists are generated for the slices of the same frame. It may be constructed without the use of a synchronization process or a reference picture list modification process. The long-term reference picture (LTRP) is identified by the POC LSB. Delta POC MSB cycle (d elta POC MSB cycle) is signaled for LTRP as determined per picture. That's fine.

[0098] To code a video image, the image is first partitioned and the partitions are converted into a bitstream. Various picture partitioning schemes are available. The image is divided into regular slices, dependent slices, and tiles. , and / or Wavefront Parallel Processing (WPP). For simplicity, HEVC is When dividing slices into groups of CTBs for video coding, Only slices, subordinate slices, tiles, WPPs, and combinations of these may be used. Such partitioning is done by matching the maximum transmission unit (MTU) size. may be applied to support parallel processing, and reduced end-to-end delay. The MTU represents the maximum amount of data that can be sent in a single packet. If a payload exceeds the MTU, it is fragmented in a process called fragmentation. The data is split into two packets by the

[0099] Regular slicing, also known simply as slicing, is caused by loop filtering operations. with other regular slices in the same picture, despite some interdependencies. are segmented portions of the image that can be reconstructed independently. Each regular slice is , which are encapsulated in proprietary Network Abstraction Layer (NAL) units for transmission. Furthermore, intra-picture prediction across slice boundaries (intra-sample prediction, motion information prediction, The interdependence of coding mode prediction and entropy coding is due to the independent re- Such independent rebuilds may be disabled to support parallel For example, regular slice-based parallelism supports minimal inter-processor parallelism. Or use inter-core communication. However, since each normal slice is independent, Each slice is associated with a separate slice header. is due to the bit cost of the slice header for each slice and the slice boundaries. The lack of straddling prediction can result in significant coding overhead. ,Normal slices are used to support matching on MTU size requirements In particular, regular slices may be encapsulated in separate NAL units and When coded, each regular slice is divided into multiple packets. It should be smaller than the MTU of the MTU scheme to avoid fragmentation. For the purposes of parallelization and MTU size matching, the layout of slices within a picture is Conflicting requirements may be imposed on the outs.

[0100] A dependent slice is similar to a normal slice, but has an abbreviated slice header and It allows for the demarcation of image treeblock boundaries without impairing intra-block prediction. Therefore, a dependent slice is a normal slice fragmented into multiple NAL units. This allows for a portion of a normal slice to be encoded before the entire normal slice is completely coded. Reduced end-to-end delay by allowing packets to be sent before they finish results.

[0101] A picture may be divided into tile groups / slices and tiles. A tile is A sequence of CTUs that covers a rectangular area of ​​the picture. Tile Group / Slice contains several tiles of a picture. Raster scan tile group mode and long A rectangular tile group mode may be used to generate the tiles. In tile group mode, the tile group is a raster scan of the tiles in the picture. In rectangular tile group mode, the tile group A pixel contains several tiles of a picture that collectively form a rectangular region of the picture. The tiles within a rectangular tile group are in raster scan order for the tile group. For example, tiles have horizontal and vertical boundaries that create columns and rows of tiles. A tile may be a segmented portion of an image generated by a field. The scan order of the CTB may be coded in time sequence (right to left and top to bottom). Therefore, the CTB in the first tile must be replayed before proceeding to the CTB in the next tile. Like regular slices, tiles are coded in vertical scan order. This breaks the interdependence of prediction and entropy decoding. However, tiles are The tiles may not be contained in individual NAL units, and therefore the tiles may be sized to match the MTU. Each tile is processed by one processor / core. and can be used for intra-picture prediction between processing units that decode neighboring tiles. Inter-processor / inter-core communication used is shared (when neighboring tiles are in the same slice). Carrying the included slice header and the reconstructed sample and metadata It may be limited to performing group filtering related sharing. When a slice contains a The byte offset of the entry point for the file is signaled in the slice header. For each slice and tile, the following conditions may be satisfied: 1) Within a slice 1) all coded tree blocks of a tile belong to the same tile, and 2) all coded tree blocks of a tile belong to the same tile. All coded treeblocks in a slice belong to the same slice. At least one should be satisfied.

[0102] In WPP, the image is segmented into a single row of CTB. Entropy decoding and prediction mechanism The algorithm may use data from other rows of the CTB. This allows for parallel processing. For example, the current row may be decoded in parallel with the previous row. However, the decoding of the current row is delayed by 2CTB from the decoding process of the previous row. The extension is performed by the CTB above the current CTB in the current line and the CTB above it to the right before the current CTB is coded. This method ensures that relevant data for the CTB is available. When rendered, it appears as a wavefront. This offset onset is at most the same as the number of CTB rows the image contains. Enables parallelization up to the same number of processors / cores. Since intra-picture prediction between rows is allowed, the processor Inter-server / inter-core communication can be quite high. WPP division takes into account the size of the NAL unit. Therefore, WPP does not support MTU size matching. To enforce MTU size matching, normal slices are used for specific coding. It can be used in conjunction with WPP with some overhead. Finally, wavefront segments (wavefr A data segment may contain exactly one CTB row. When a slice starts within a CTB row, the slice should end within the same CTB row.

[0103] The tiles may also include a motion constrained tile set. A motion constrained tile set (MCTS) is a set of tiles whose associated motion vectors are located at full-sample locations within the MCTS. fractional samples, which require only full sample positions within the MCTS for interpolation. A tileset designed to be constrained to point to fractional samples. Furthermore, the motion vectors for temporal motion vector prediction derived from blocks outside the MCTS are In this way, each MCTS is not allowed to use tiles that are not included in the MCTS. Temporal MCTS Supplemental Enhancement Information (SEI) may be decoded independently without the need for a separate The "cement information" message indicates the presence of MCTS in the bitstream and synchronizes the MCTS. The MCTS SEI message may be used to signal the match for the MCTS set. (specified as part of the semantics of the SEI message) to generate a bitstream that It provides supplementary information that can be used in the extraction of the MCTS sub-bitstreams (specified by the MCTS sub-bitstream information). The information defines several sets of MCTSs, each of which is used for the MCTS sub-bitstream extraction process. Alternative Video Parameter Sets (VPS), Sequence Parameter Sets (SPS) used during and Picture Parameter Set (PPS) Raw Byte Sequence Payload (RBSP) bytes It contains several extracted information sets, including the MCTS sub-bitstream extraction process. When extracting the sub-bitstreams using the The slice header may be rewritten or replaced, Syntax related to slice addresses (including nt_in_pic_flag and slice_segment_address) One or all of the sub-bitstreams are different in the extracted sub-bitstream. The value may be used and therefore updated.

[0104] VR applications, also known as 360-degree video applications, are applications that display a portion of a complete sphere. and as a result may only display a subset of the entire picture. Dynamic Adaptive Streaming over H (DASH) The mechanism for 360 delivery relies on viewports via the Hypertext Transfer Protocol (HTP). Lower bitrates and support the delivery of 360-degree video via streaming mechanisms This mechanism may be used to, for example, Dividing the sphere / projected picture into multiple MCTSs by using a sphere / projected picture (bemap projection) Two or more bitstreams may be coded at different spatial resolutions or qualities. When delivering data to the decoder, the M from the higher resolution / quality bitstream is used. CTS is sent for the viewport that will be displayed (for example, the front viewport). MCTS from lower resolution / quality bitstreams are sent for other viewports. These MCTSs are packed in a specific way and then sent to the The viewport seen by the user creates a positive viewing experience. It is expected that the MCTS will be presented in high resolution / quality to allow the user to see the When you turn to look at a viewport (for example, the left or right viewport), the display The system fetches high resolution / quality MCTS for the new viewport. It comes from a lower resolution / quality viewport for a short period of time. When the user turns to look at the viewport, the viewport changes direction. There is a delay between when the higher resolution / quality representation of the port is seen. This delay is How quickly the system fetches higher resolution / quality MCTS for that viewport It depends on how fast you can fetch it, and how fast you can fetch it. The period of an IRAP depends on the period of the IRAP. The period of an IRAP is the interval between the occurrence of two IRAPs. This delay is Since the MCTS of the new viewport can only be decoded from the IRAP picture, the period of the IRAP Related to.

[0105] For example, if the period of the IRAP is coded as every 1 second, then the following applies: The best case scenario for delay is when the system starts fetching new segments / IRAP cycles. The network when the user turns their head to see a new viewport just before In this scenario, the system will Higher resolution / quality MCTS can be requested immediately, thus minimizing buffering. The matching delay can be set to almost zero, and the sensor delay is small and negligible. Assuming that the network is fully connected, the only delay is the network round trip delay. is the delay of the fetch request plus the transmission time of the requested MCTS. The round trip delay can be, for example, about 200 milliseconds. The worst case scenario for the delay is , the user's head moves to a new viewport after the system has already made a request for the next segment. The IRAP period when turning around to view + network round trip delay. The team will be more frequent with shorter IRAP periods to improve the worst case scenario above. The IRAP pictures can be coded using the IRAP picture scheme, which reduces the overall delay. However, this approach reduces compression efficiency and increases bandwidth requirements. become.

[0106] In an exemplary implementation, sub-pictures of the same coded picture are coded in different This mechanism is explained as follows: A picture may be divided into sub-pictures. A sub-picture is a picture with tile_ equal to 0. A rectangular set of tile groups / slice starting from the tile group with group_address Each sub-picture may refer to a corresponding PPS, and thus a separate tile The presence of sub-pictures may be indicated in the PPS. Subpictures are treated like pictures in the decoding process. In-loop filtering may always be disabled. The width and height of the subpicture are The position of the subpicture within the picture may be specified in units of luma CTU size. It may not be generalized, but may be derived using the following rules: A picture contains a CTU within a picture that is large enough to contain the subpicture within the picture's boundaries. To decode each subpicture, take the next unoccupied position in raster scan order. The reference picture is the current sub-picture and the corresponding reference picture in the decoded picture buffer. It is generated by extracting the area to be located. The extracted area is decoded. subpictures that are created and therefore have the same size and position within the picture. Inter prediction is performed between frames. In such cases, Allowing different nal_unit_type values ​​allows subpictures derived from random access pictures. and sub-pictures derived from non-random access pictures without much difficulty (e.g. can be merged into the same coded picture (e.g. without VCL level modifications) Such advantages also apply to MCTS-based coding.

[0107] Allowing different nal_unit_type values ​​within a coded picture is not supported by other systems. For example, a user may want to view a portion of 360-degree video content. Some areas may be viewed more frequently than others. Coding efficiency and average quality for 360° video delivery depend on the port. To make a better tradeoff between viewport switching latency and other More frequent IRAP pictures are coded for areas that are more commonly seen than for other areas. The same quality viewport switching latency can be achieved by switching from the first viewport to the second. When switching to a different viewport, the presentation quality of the second viewport is different from that of the first viewport. is the latency experienced by the user until a presentation quality equivalent to that of the

[0108] Another implementation is for mixed NALs within a picture, including derivation of POC and management of reference pictures. Use the following solution for unit type support: Mixed IRAP subpictures to specify whether there may be pictures with subpictures and non-IRAP subpictures. , flags referenced directly or indirectly by tile groups (sps_mixed_tile_gr oups_in_pic_flag) is present in the parameter set. For units, whether the POC MSB is reset when deriving the POC for a picture is specified. To specify this, a flag (poc_msb_reset_flag) must be present in the corresponding tile group header. A variable called PicRefreshFlag is defined and associated with the picture. Lag: The derivation of POC and the state of DPB should be refreshed when decoding a picture. The value of PicRefreshFlag is derived as follows: If the image group is included in the first access unit of the bitstream, PicRefreshF The lag is set equal to 1. Otherwise, the current tile group is set equal to the IDR tile. If it is a group, PicRefreshFlag is sps_mixed_tile_groups_in_pic_flag ? poc_ms b_reset_flag : Set to be equal to 1. Otherwise, the current tile group is reset to CR If it is an A tile group, the following applies: The current access unit is coded If this is the first access unit in the sequence being retrieved, PicRefreshFlag is equal to 1. The access unit is set to end of sequence NAL unit. The variable HandleCraAsFirstPicInCvsFlag immediately following or associated with the When set to , the current access unit is the first access unit. Otherwise, PicRefreshFlag is set equal to 0. (e.g., the current tile group is in the first access unit of the bitstream) Not part of the IRAP tile group isn't it).

[0109] When PicRefreshFlag is equal to 1, the value of POC MSB (PicOrderCntMsb) is Reset to 0 during derivation of POC. Reference Picture Set (RPS) or Reference Picture The information used for reference picture management, such as the reference picture list (RPL), is stored in the corresponding NAL unit. Signaled in the tile group / slice header regardless of the tile type. The feature list is built at the beginning of decoding each tile group, regardless of the NAL unit type. The reference picture list is RefPicList

[0000] and RefPicList

[0001] for the RPL method, RefPicList0[ ] and RefPicList1[ ] for the RPS method, or It may also contain a similar list containing reference pictures for center prediction operations. When ag is equal to 1, all reference pictures in the DPB are marked during the reference picture marking process. The structure is marked as unused for reference purposes.

[0110] Such implementations are associated with specific issues, e.g., nal_unit_t in a picture When mixing of type values ​​is not allowed and when determining whether a picture is an IRAP picture When the output and derivation of the variable NoRaslOutputFlag are described at the picture level, the decoder These derivations may be performed after receiving the first VCL NAL unit of any picture. However, due to the support for mixed NAL unit types within a picture, the decoder It must wait for the arrival of other VCL NAL units before performing the derivation. If so, the decoder must wait for the arrival of the last VCL NAL unit of the picture. Furthermore, such systems require that the POC MSB is reset during the derivation of the POC for a picture. A flag is added in the tile group header of the IDR NAL unit to specify whether or not This mechanism has the following problems: For IRAP and non-IRAP NAL unit types supported by this mechanism, Furthermore, this information is not shared within the tile group / slice header of the VCL NAL unit. Signaling means that an IRAP (IDR or CRA) NAL unit is interleaved with a non-IRAP NAL unit within a picture. When there is a change in the status of whether the bitstream is mixed with the requires that the value be changed during merging. Such a rewrite of the slice header is It occurs whenever a user requests video and therefore requires significant hardware resources. In addition, certain IRAP NAL unit types and certain non-IRAP NAL unit types require Some other mixture of different NAL unit types within a picture other than a mixture of bit types Such flexibility does not provide support for practical use cases. On the other hand, they complicate the codec design, which unnecessarily increases the complexity of the decoder. ,thus raising the associated implementation costs.

[0111] Generally, this disclosure relates to sub-picture or MCTS-based run-time coding in video coding. More particularly, this disclosure describes techniques for supporting subpicture access. The MCTS-based random access support is used to support the Describes an improved design for mixed NAL unit type support. The explanation is based on the VVC standard, but also applies to other video / media codec specifications.

[0112] To solve the above problem, the following exemplary implementation is disclosed. In one example, each picture is applied to the image at the time the picture is mixed. This is associated with an indication of whether the value of the specified nal_unit_type is included. This indication is signaled within the PPS. Marking the picture as unused for reference clears the POC MSB. Supports determining whether to reset and / or whether to reset the DPB When an indication is signaled within a PPS, a change in the value within the PPS is may be performed during a separate extraction. However, this does not affect the extraction of such a bitstream. or during a merger when the PPS is rewritten and replaced by other mechanisms. It is acceptable.

[0113] Alternatively, this indication may be signaled in the tile group header. may be required to be the same for all tile groups in a picture. In this case, the value is used during extraction of the sub-bitstream of the MCTS / Subpicture sequence. Alternatively, this indication may need to be changed to The tile group size is signaled in the tile header but is the same for all tile groups of a picture. However, in this case the value is the MCTS / subpicture sequence. The bitstream may need to be modified during extraction of the sub-bitstream. When used for a picture, the indication of all VCL NAL Define additional VCL NAL unit types whose units have the same NAL unit type value. However, in this case, the Nth bit of the VCL NAL unit The value of the AL unit type is used for extracting sub-bitstreams of MCTS / Subpicture sequences. This indication may need to be changed during the process. When used for a picture, all VCL NAL units of the picture have the same NAL unit type. signaling by defining additional IRAP VCL NAL unit types with the following values: However, in this case, the value of the NAL unit type of the VCL NAL unit is It must be modified during extraction of the sub-bitstreams of the MCTS / Subpicture sequence. Alternatively, at least one VCL N of any of the IRAP NAL unit types Each picture with an NAL unit contains a value of the NAL unit type with which the picture is mixed. The parameter may be associated with an indication of whether the

[0114] Furthermore, only mixed IRAP and non-IRAP NAL unit types are allowed. This allows for a limited mixing of nal_unit_type values ​​within a picture. Such constraints may be applied: For any particular picture, all VCL NAL units Either some VCL NAL units have the same NAL unit type, or some VCL NAL units have a specific IRAP NAL type. AL unit type and the rest have specific non-IRAP VCL NAL unit types. In other words, the VCL NAL units of any particular picture may be split into two or more IRAP NAL unit types and may have more than one non-IRAP NAL unit type. A picture does not contain a mixed picture nal_unit_type value and is not a VCL NAL A unit may be considered an IRAP picture only if it has an IRAP NAL unit type For all IRAP NAL units (including IDRs) that do not belong to an IRAP picture, the POC MSB is All IRAP NAL units (including IDRs) that do not belong to an IRAP picture may not be reset. For units, the DPB is not reset, and therefore all reference pictures are Marking the TemporalId as unused for the purpose of If at least one VCL NAL unit of a picture is an IRAP NAL unit, may be set equal to 0.

[0115] The following are specific implementations of one or more of the above aspects: If the value of u_types_in_pic_flag is equal to 0 and each VCL NAL unit has IDR_W_RADL and RSV_IRAP_ Codes with nal_unit_type in the range from IDR_W_RADL to RSV_IRAP_VCL13, including VCL13 An example PPS syntax and syntax The mantics are as follows: [Table 2] mixed_nalu_types_in_pic_flag indicates that each picture that references a PPS contains multiple VCL NAL units. and to specify that these NAL units do not have the same value of nal_unit_type The mixed_nalu_types_in_pic_flag is set equal to 0 for each picture that references a PPS. Equal to 0 to specify that the VCL NAL units of this It is set as follows.

[0116] An example tile group / slice header syntax is as follows: [Table 3A] [Table 3B]

[0117] The semantics of an exemplary NAL unit header are as follows: For VCL NAL units in a frame, one of the following two conditions is met: All VCL NAL units have the same value of nal_unit_type. , including specific IRAP NAL unit type values ​​(i.e., IDR_W_RADL and RSV_IRAP_VCL13) While all have a value of nal_unit_type in the range IDR_W_RADL to RSV_IRAP_VCL13 All other VCL NAL units are of a specific non-IRAP VCL NAL unit type (i.e., TRAIL_ Within the range from TRAIL_NUT to RSV_VCL_7, including NUT and RSV_VCL_7, or RSV_VCL14 and and RSV_VCL15, inclusive) nuh_temporal_id_plus1 minus 1 is the temporal identifier for the NAL unit. al identifier). The value of nuh_temporal_id_plus1 is not equal to 0.

[0118] The variable TemporalId is derived as follows: TemporalId = nuh_temporal_id_plus1 - 1 (7-1)

[0119] IDR_W_RADL and RSV_IRAP_VCL13 for VCL NAL units with nal_unit_type of picture If the other VCs of the picture are within the range from IDR_W_RADL to RSV_IRAP_VCL13, including L Regardless of the value of the nal_unit_type of the NAL unit, the TemporalId is used for all VCLs of the picture. For a NAL unit, the value of TemporalId is equal to 0. The same applies to the L unit. The value of emporalId is the VCL NAL unit of a coded picture or access unit. The TemporalId value of the

[0120] An exemplary decoding process for a coded picture is as follows: The decoding process operates as follows for the current picture CurrPic: NAL unit decoding The following decoding process is performed based on the tile group header Uses syntax elements from the layer and higher. The variables and functions related to the count are derived as detailed herein. This is only called for the first tile group / slice of the picture. At the beginning of the decoding process for a picture group / slice, The decoding process for this purpose uses reference picture list 0 (RefPicList

[0000] ) and reference picture list It is called to derive RefPicList

[0001] . The current picture is an IDR picture. If so, then the decoding process for constructing the reference picture list is It may be called to check conformance, but not for the current picture or the current picture in decoding order. It may not be necessary for decoding of pictures that follow a picture.

[0121] The decoding process for constructing the reference picture list is as follows: is called at the beginning of the decoding process for each tile group. The reference index is addressed by the reference picture list. When decoding a tile group, the reference picture list is an index into the tile group. It is not used in decoding P tile group data. Only reference picture list 0 (RefPicList

[0000] ) is used in decoding tile group data. When decoding a B tile group, reference picture list 0 and reference picture list Both RefPicList

[0001] and RefPicList

[0002] are used in decoding tile group data. At the beginning of the decoding process for a tile group, the reference picture list RefPicList

[0000] and The reference picture list is derived by marking the reference pictures. Used in decoding of IDR picture data or in decoding of tile group data. For all tile groups or I-tile groups of non-IDR pictures, RefPicList

[0000] and RefPicList

[0001] may be derived for the purpose of checking bitstream conformance. However, their derivation depends on the current picture or the picture that follows it in decoding order. For P tile groups, RefPicList

[0001] is the bitstream. It may be derived for the purpose of checking the conformance of the stream, but the derivation is based on the current picture or It is not required for the decoding of pictures that follow the current picture in decoding order.

[0122] 8 is a schematic diagram of an example video coding device 800. The device 800 is adapted to implement the disclosed examples / embodiments as described herein. The video coding device 800 is suitable for downstream ports 820, upstream ports 822, and downstream ports 824. Stream port 850 and / or upstream and / or downstream through the network a transceiver unit (including a transmitter and / or a receiver) for transmitting data in a stream; The video coding device 800 includes a logical unit for processing data. a processor 830 including a memory and / or a central processing unit (CPU), and a The video coding device 800 may be an electrical, optical, or wireless Upstream ports 850 and 851 for transmitting data over a wireless communication network and / or electrical, optical-electrical (OE) components, electrical- It may also include optical (EO) components and / or wireless communication components. The input device 800 includes an input and / or output port for communicating data to and from a user. The system may also include an input / output (I / O) device 860. The I / O device 860 displays video data. output devices such as a display for outputting the image data and a speaker for outputting the audio data. The I / O device 860 may include input devices such as a keyboard, a mouse, and a trackball. and / or corresponding interfaces for interacting with such output devices. It may also include an interface.

[0123] The processor 830 is implemented in hardware and software. 30 refers to one or more CPU chips, cores (e.g., as a multi-core processor), field Programmable gate arrays (FPGAs), application specific integrated circuits (ASICs), and digital signals The processor 830 may be implemented as a digital signal processor (DSP). 0, Tx / Rx 810, upstream port 850, and memory 832. 0 includes a coding module 814. The coding module 814 is The video stream 600 and / or the bitstream 700 may be used. Implementing the disclosed embodiments described herein, such as methods 100, 900, and 1000, The coding module 814 may use any other method / method described herein. Additionally, the coding module 814 may implement a codec system. 200, encoder 300, and / or decoder 400. For example, The encoding module 814 encodes the picture containing both IRAP NAL units and non-IRAP NAL units. Indicates when it contains a single type of IRAP NAL unit and a single type of non-IRAP NAL unit. A flag may be set in the PPS to constrain such pictures to contain only Therefore, the coding module 814 may use the following when coding the video data: providing video coding device 800 with additional functionality and / or coding efficiency; Thus, coding module 814 enhances the functionality of video coding device 800. It increases the functionality and addresses the problems inherent in video coding technology. Module 814 effects the transition of video coding device 800 to a different state. Generally, the coding module 814 is stored in a memory 832 and executed by a processor 830. as instructions to be executed (e.g., a computer program product stored on a non-transitory medium) It can be implemented as

[0124] Memory 832 can be a disk, tape drive, solid state drive, read-only Memory (ROM), Random Access Memory (RAM), Flash Memory, Ternary Content Addressable Memory (TCAM) ternary content-addressable memory (CRAM), static random access memory (SRAM), etc. The memory 832 includes one or more memory modules. The memory 832 is used when a program is selected for execution. for storing such programs in a memory and for storing instructions and and an overflow data storage device for storing data. It may also be used as a storage device.

[0125] FIG. 9 shows a video signal merged from multiple sub-picture video streams 601 to 603 of multiple video resolutions. bitstream 700 including the VR picture video stream 600 Video sequences such as CVS 500 that contain pictures with mixed NAL unit types 9 is a flow chart of an example method 900 for encoding a signal. When the codec system 200, the encoder 300, and / or the video coding device This may be used by an encoder such as the encoder 800.

[0126] The method 900 is for an encoder to generate a video sequence including multiple pictures, such as VR pictures. and encoding the video sequence into a bitstream based on, for example, user input. In step 901, the current picture is determined to be different from the current picture. The encoder determines whether the image contains multiple sub-pictures of different types. The type must be at least one slice of a picture that contains part of an IRAP subpicture and a non-IRA A P NAL sub-picture may contain at least one slice of a picture containing a portion of the NAL sub-picture. In step 903, the encoder converts the slices of the subpictures of the picture into a bitstream. Such a VCL NAL unit may be coded into one or more VCL NAL units within a frame. It may contain IRAP NAL units and one or more non-IRAP NAL units. The encoding step involves converting the sub-bitstreams of different resolutions into a single stream for transmission to the decoder. This may include merging the bitstreams into the original bitstream.

[0127] In step 905, an encoder encodes the PPS into a bitstream and sets the flag to As a specific example, encoding a PPS is For example, depending on the merging of sub-bitstreams, the already coded PPS may be modified to include the flag value. The flags may include changing the value of the NAL unit type associated with the picture. When the value is the same for all VCL NAL units in a Also, the flags relate to VCL NAL units that contain one or more of the picture's subpictures. The first NAL unit type value is a VCL NAL unit that contains one or more of the subpictures of a picture. If different from the value of the second NAL unit type for the unit, it may be set to a second value. For example, the first NAL unit type value indicates that the picture contains an IRAP subpicture. The second NAL unit type value may indicate that the picture also contains non-IRAP sub-pictures. Furthermore, the value of the first NAL unit type may indicate IDR_W_RADL, IDR_N_LP, or Additionally, the value of the second NAL unit type may be equal to TR_NUT or CRA_NUT. AIL_NUT, RADL_NUT, or RASL_NUT. The flag may be mixed_nalu_types_in_pic_flag. In a specific example, _types_in_pic_flag indicates that each picture that references the PPS containing the flag is a VCL NAL unit with two or more It may be set equal to 1 to specify that the flag The tag indicates that all VCL NAL units associated with the corresponding picture are of the NAL unit type (nal_uni t_type) does not have the same value. u_types_in_pic_flag indicates that each picture that references the PPS containing the flag is in one or more VCL NAL units. In addition, the flag ensures that all VCL NAL units of the corresponding picture have the same value of nal_unit_type Specify.

[0128] In step 907, the encoder generates a bitstream for transmission to the decoder. It may be memorized.

[0129] FIG. 10 shows a video signal merged from multiple sub-picture video streams 601 to 603 of multiple video resolutions. bitstream 700 including the VR picture video stream 600 A video sequence such as CVS 500 that contains pictures with mixed NAL unit types from different streams 10 is a flow diagram of an example method 1000 for decoding an instance of a digital signature. Sometimes the codec system 200, decoder 400, and / or video coding device This may be used by a decoder such as the processor 800.

[0130] The method 1000 may be implemented, for example, as a coded image representing a video sequence resulting from the method 900. It may start when the decoder begins to receive the bitstream of encoded data. In step 1001, a decoder receives a bitstream. In a particular example, the bitstream includes a plurality of sub-pictures and flags associated with the sub-pictures. A stream may contain a PPS containing a flag. Furthermore, a subpicture may be split into multiple VCL NAL units. For example, a slice associated with a subpicture is included in a VCL NAL unit. Included.

[0131] In step 1003, when the flag is set to a first value, the decoder The bit type value is determined to be the same for all VCL NAL units associated with a picture. Furthermore, when the flag is set to a second value, the decoder The first NAL unit type value for the VCL NAL unit containing one or more pictures is The second NAL unit for a VCL NAL unit containing one or more of the picture's subpictures. For example, the value of the first NAL unit type is determined to be different from the value of the picture The second NAL unit type value may indicate that the picture contains an IRAP subpicture. Furthermore, the first NAL unit type may indicate that the image also contains non-IRAP subpictures. The value may be equal to one of IDR_W_RADL, IDR_N_LP, or CRA_NUT. The value of NAL unit type 2 is equal to one of TRAIL_NUT, RADL_NUT, or RASL_NUT. As a specific example, the flag may be mixed_nalu_types_in_pic_flag mixed_nalu_types_in_pic_flag indicates that each picture that references a PPS contains two or more VCL NAL units. and the VCL NAL units do not have the same value of NAL unit type (nal_unit_type). It may be set equal to 1 when specifying that mixed_nalu_types_in_p ic_flag indicates that each picture that references a PPS has one or more VCL NAL units that reference the PPS. When the VCL NAL units of each picture in a may be set to

[0132] In step 1005, the decoder determines the subpicture based on the value of the NAL unit type. In step 1007, the decoder may decode one or more of the decoded Transfer one or more of the subpictures for display as part of the combined video sequence You may do so.

[0133] FIG. 11 shows a video signal merged from multiple sub-picture video streams 601 to 603 of multiple video resolutions. bitstream 700 including the VR picture video stream 600 Video sequences such as CVS 500 that contain pictures with mixed NAL unit types 11 is a schematic diagram of an example system 1100 for coding a Deck system 200, encoder 300, decoder 400, and / or video coding The system may be implemented by an encoder and a decoder, such as device 800. The system 1100 may be used when performing the methods 100, 900, and / or 1000.

[0134] The system 1100 includes a video encoder 1102. The video encoder 1102 is configured to encode a picture as a determination module 1101 for determining whether the image contains multiple sub-pictures of different types; The video encoder 1102 encodes sub-pictures of a picture into multiple sub-pictures in a bitstream. and a coding module 1103 for coding the VCL NAL unit of the encoded data into the VCL NAL unit of the encoded data. The decoding module 1103 decodes all VCL NAL units whose NAL unit type values ​​are related to pictures. It is set to the first value when the pixels are the same for all subpictures in the picture. The first NAL unit type value for a VCL NAL unit containing more than one subpicture of a picture The value of the second NAL unit type for the VCL NAL unit containing one or more of the and encoding into the bitstream a flag that is set to a second value when the The video encoder 1102 stores the bitstream for transmission to the decoder. The video encoder 1102 further includes a storage module 1105 for converting the bitstream into a video signal. The video encoder further includes a transmitting module 1107 for transmitting to the video decoder 1110. 1102 may be further configured to perform any of the steps of method 900.

[0135] The system 1100 also includes a video decoder 1110. The video decoder 1110 generates a video signal associated with a picture. receiving module for receiving a bitstream including a plurality of sub-pictures and a flag for A subpicture, including a video decoder module 1111, is contained in multiple VCL NAL units. When the flag is set to the first value, the value of the NAL unit type is set to picture. a determining module 1113 for determining that the VCL NAL units are the same for all related VCL NAL units; Furthermore, the determination module 1113 further includes, when the flag is set to a second value: The first NAL unit of a VCL NAL unit that contains one or more of the picture's subpictures. The value of this type is for a VCL NAL unit that contains one or more of the picture's subpictures. The video decoder 1110 determines that the value of the second NAL unit type is different from the value of the first NAL unit type. selects a decoder for decoding one or more of the subpictures based on the value of the NAL unit type. The video decoder 1110 further includes a signaling module 1115. a transfer module for transferring one or more of the subpictures for display as part of the The video decoder 1110 further includes a processor 1117. The video decoder 1110 performs any of the steps of the method 1000. It may be further configured as follows.

[0136] The first component may be connected to a line, trace, or other When there are no intervening components except for the medium, the first component is directly coupled to the second component. A component is an intermediary other than a line, trace, or another medium between a first component and a second component. When there is an intermediate component, it is indirectly connected to the second component. The term "about" and variations thereof include both directly and indirectly linked terms. The use of "+" means a range including ±10% of the number thereafter unless otherwise stated. do.

[0137] The steps of the exemplary methods described herein may not necessarily be performed in the order described. The order of steps in such methods is exemplary only and is not required to be performed in a specific order. It should also be understood that further steps may be taken to Various methods may be included, and specific steps may be included in methods consistent with various embodiments of the present disclosure. In some cases, they may be omitted or combined.

[0138] Although several embodiments have been given in this disclosure, the disclosed systems and methods , may be embodied in many other specific forms without departing from the spirit or scope of this disclosure. It will be understood that these examples are intended to be illustrative and not limiting. The intent should not be limited to the details given in this specification. For example, various elements or components may be combined or integrated into another system. Certain features may be omitted or not implemented.

[0139] Additionally, what is described in various embodiments as being separate or distinct may be The illustrated techniques, systems, subsystems, and methods may be modified without departing from the scope of the present disclosure. or be combined or integrated with any other system, component, technique, or methodology without Other examples of modifications, substitutions, and alterations may occur to those skilled in the art. may be made without departing from the spirit and scope of the disclosure herein. There is a possibility. [Explanation of symbols]

[0140] 100 How it works 200 Encoding and Decoding (Codec) Systems 201 segmented video signal 211 General Coda Control Components 213 Transformation, Scaling and Quantization Components 215 Intra-picture estimation components 217 Intra-picture prediction components 219 Motion Compensation Components 221 Motion Estimation Components 223 Decoded Picture Buffer Components 225 In-loop filter components 227 Filter Control Analysis Components 229 Scaling and Inverse Transformation Components 231 Header Format and Context-Adaptive Binary Arithmetic Coding (CABAC) Configuration element 300 Video Encoder 301 segmented video signal 313 Transformation and Quantization Components 317 Intra-picture prediction components 321 Motion Compensation Components 323 Decoded Picture Buffer Components 325 In-Loop Filter Components 329 Inverse Transform and Quantization Components 331 Entropy Coding Components 400 Video Decoder 417 Intra-picture prediction components 421 Motion Compensation Components 423 Decoded Picture Buffer Components 425 In-Loop Filter Components 429 Inverse Transform and Quantization Components 433 Entropy Decoding Components 500 CVS 502 IRAP Picture 504 Leading Picture 506 Trailing Picture 508 Decoding Order 510 Presentation order 600 VR picture video streams 601 Subpicture Video Stream 602 Subpicture Video Stream 603 Subpicture Video Stream 700 bitstream 710 Sequence Parameter Set (SPS) 711 Picture Parameter Set (PPS) 715 slice header 720 image data 721 Pictures 723 Subpicture 725 slices 727 Mixed NAL unit types in picture flag (mixed_nalu_types_in_pic_flag) 730 non-VCL NAL units 731 SPS NAL Unit Type (SPS_NUT) 732 PPS NAL unit type (PPS_NUT) 740 VCL NAL Unit 741 IDR_N_LP NAL unit 742 IDR_w_RADL NAL unit 743 CRA_NUT 745 IRAP NAL unit 746 RASL_NUT 747 RADL_NUT 748 TRAIL_NUT 749 Non-IRAP NAL Units 800 Video Coding Device 810 Transceiver Unit (Tx / Rx) 814 Coding Module 820 downstream ports 830 processor 832 memory 850 upstream ports 860 Input and / or Output (I / O) Devices 900 ways 1000 ways 1100 System 1101 Judgment Module 1102 Video Encoder 1103 Encoding Module 1105 Storage Module 1107 Transmitting Module 1110 Video Decoder 1111 Receiver Module 1113 Judgment Module 1115 Decryption Module 1117 Transfer Module

Claims

1. 1. A method implemented in a decoder, comprising: A bitstream including a plurality of sub-pictures and flags associated with a picture is stored in the data. receiving by a receiver of a video coder, said sub-picture being a step included in a Video Conferencing Layer (VCL) Network Abstraction Layer (NAL) unit; When the flag is set to a first value, a value of a first NAL unit type is set to the picture. the processor is configured to recognize that the VCL NAL units associated with the channel are the same for all of the channel. determining the When the flag is set to a second value, The value of the first NAL unit type for the VCL NAL unit includes one or more of the following: a second NAL unit for a VCL NAL unit containing one or more of the sub-pictures of the image; determining by said processor that the value is different from the bit type value; a value of the first NAL unit type or a value of the second NAL unit type; and decoding, by said processor, one or more of the sub-pictures. method.

2. The bitstream includes a picture parameter set (PPS) that includes the flag. The method according to claim 1.

3. The value of the first NAL unit type indicates that the picture is an intra random access point. and the value of the second NAL unit type indicates that the picture contains an Inter-Area Protocol (IRAP) subpicture.

3. The method of claim 1, wherein the subpicture indicates that the image contains non-IRAP subpictures.

4. The value of the first NAL unit type is a random-access decodable leading picture. Instantaneous Decoding Refresh (IDR) with leading picture (IDR_W_RADL), IDR without leading picture (IDR_N_LP), or the clean random access (CRA) NAL unit type (CRA_NUT). The method according to any one of claims 1 to 3.

5. The value of the second NAL unit type is the trailing picture NAL unit type (TRAIL_ NUT), Random Access Decodable Leading Picture NAL Unit Type (RADL_NUT), or Random Access Skip Leading Picture (RASL) NAL unit type 5. The method according to claim 1, wherein the parameter is equal to RASL_NUT.

6. 6. The method according to claim 1, wherein the flag is mixed_nalu_types_in_pic_flag. How to post.

7. The picture that references the PPS has two or more of the VCL NAL units, and the VCL N When specifying that AL units do not have the same value of NAL unit type (nal_unit_type) , mixed_nalu_types_in_pic_flag is equal to 1, and the picture that references the PPS is in the VC L NAL units, and the VCL NAL units have the same value of nal_unit_type. When specifying that the mixed_nalu_types_in_pic_flag is equal to 0, 10. The method according to any one of claims 1 to 9.

8. 1. A method implemented in an encoder, comprising: The processor determines whether a picture contains multiple subpictures of different types. determining the The sub-pictures of the picture are divided into multiple video coding levels in a bitstream. the processor encodes the VCL into Network Abstraction Layer (NAL) units. Steps and The first NAL unit type value is used for all VCL NAL units associated with the picture. and one of the sub-pictures of the picture is set to a first value when the sub-pictures are the same as the picture. the value of the first NAL unit type for the VCL NAL unit containing of the second NAL unit type for VCL NAL units containing one or more of the subpictures a flag set to a second value when the first value is different from the first value by the processor; encoding the data into a stream; a memory coupled to the processor for transmitting the bitstream to a decoder; and storing the data by the library.

9. encoding a picture parameter set (PPS) into the bitstream.

9. The method of claim 8, further comprising the step of: encoding the flag into the PPS. 。

10. The value of the first NAL unit type indicates that the picture is an intra random access point. and the value of the second NAL unit type indicates that the picture contains an Inter-Area Protocol (IRAP) subpicture.

10. The method of claim 8 or 9, wherein the sub-picture indicates that the image contains non-IRAP sub-pictures.

11. The value of the first NAL unit type is a random-access decodable leading picture. Instantaneous Decoding Refresh (IDR) with leading picture (IDR_W_RADL), IDR without leading picture (IDR_N_LP), or the clean random access (CRA) NAL unit type (CRA_NUT). The method according to any one of claims 8 to 10.

12. The value of the second NAL unit type is the trailing picture NAL unit type (TRAIL_ NUT), Random Access Decodable Leading Picture NAL Unit Type (RADL_NUT), or Random Access Skip Leading Picture (RASL) NAL unit type 12. The method according to claim 8, wherein the parameter is equal to RASL_NUT.

13. 13. The method according to claim 8, wherein the flag is mixed_nalu_types_in_pic_flag. The method described.

14. The picture that references the PPS has two or more of the VCL NAL units, and the VCL N When specifying that AL units do not have the same value of NAL unit type (nal_unit_type) , mixed_nalu_types_in_pic_flag is equal to 1, and the picture that references the PPS is in the VC L NAL units, and the VCL NAL units have the same value of nal_unit_type. When specifying that the mixed_nalu_types_in_pic_flag is equal to 0, 10. The method according to any one of claims 1 to 9.

15. a processor; a receiver coupled to the processor; and a memory coupled to the processor. a memory and a transmitter coupled to the processor, The memory and the transmitter perform the method of any one of claims 1 to 14. a video coding device configured to:

16. computer program products for use by video coding devices, a non-transitory computer-readable medium including a computer program product for executing a program on a processor; When executed by a processor, the video coding device is provided with any one of claims 1 to 14. A computer program stored on a non-transitory computer-readable medium that performs the method of any one of the preceding claims. A non-transitory computer-readable medium containing computer-executable instructions.

17. Receiving a bitstream containing multiple subpictures and flags associated with a picture a receiving means for receiving a sub-picture from a plurality of video coding layers (VCLs), ) Network Abstraction Layer (NAL) unit included in the receiving means; A determination means, When the flag is set to a first value, the value of the first NAL unit type is set to the first value. determining that the VCL NAL units are the same for all VCL NAL units associated with the image; When the flag is set to a second value, the sub-pictures of the picture the value of the first NAL unit type for a VCL NAL unit containing one or more of the above is a second NAL unit for a VCL NAL unit containing one or more of the sub-pictures of the a determination means for determining that the value is different from the value of the type; a value of the first NAL unit type or a value of the second NAL unit type; and decoding means for decoding one or more of the sub-pictures.

18. Claim 17: Further configured to perform the method of any one of claims 1 to 7.

2. A decoder according to claim 1 .

19. Method for determining whether a picture contains multiple sub-pictures of different types Step by step, An encoding means, The sub-pictures of the picture are then encoded into multiple video coding Layer (VCL) Network Abstraction Layer (NAL) units, The first NAL unit type value is set for all VCL NAL units associated with the picture. is set to a first value when the subpictures of the picture are the same, the value of the first NAL unit type for a VCL NAL unit containing one or more pictures before the picture a second NAL unit type for a VCL NAL unit that contains one or more of the subpictures; a flag set to a second value when different from the first value in the bitstream; encoding means for storage means for storing said bitstream for transmission to a decoder; Encoder.

20. Claim 1 further configured to perform the method of any one of claims 8 to 14.

9. The encoder according to claim 9.

Citation Information

Patent Citations

  • Video codec allowing sub-picture or region wise random access and concept for video composition using the same

    WO2020157287A1