Picture with Mixed NAL Unit Types
The introduction of a flag to differentiate IRAP and non-IRAP sub-pictures in video coding systems enhances coding efficiency and reduces resource usage by enabling dynamic resolution changes and efficient bandwidth allocation in virtual reality applications.
Patent Information
- Application Number
- JP2024180666
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Priority Date
- 2019-04-10
- Filing Date
- 2024-10-16
- Publication Date
- 2025-07-30
- Estimated Expiration
- 2040-03-11
AI Technical Summary
Existing video coding systems face challenges in efficiently compressing and transmitting video data, particularly in virtual reality applications, due to the need for mixed intra-random access point (IRAP) and non-IRAP sub-pictures within a single picture, which can lead to increased file size and decoding latency.
A mechanism is introduced to handle mixed NAL unit types within a picture by using a flag (mixed_nalu_types_in_pic_flag) to differentiate between IRAP and non-IRAP sub-pictures, allowing for dynamic resolution changes and efficient bandwidth allocation based on viewport likelihood, stored in the Picture Parameter Set (PPS).
This approach improves coding efficiency by allowing dynamic resolution changes without degrading user experience, reducing resource usage in encoders and decoders, and addressing the limitations of existing systems that require uniform NAL unit types within a single picture.
Smart Images

Figure 0007715905000007 
Figure 0007715905000008 
Figure 0007715905000009
Abstract
Description
Technical Field
[0002] The present disclosure generally relates to video coding, and more particularly to coding sub - pictures of a picture in video coding.
Background Art
[0003] The amount of video data required to depict even relatively short videos is quite large, which poses difficulties when it is to be streamed or otherwise transmitted over a communication network having a limited bandwidth capacity. Thus, video data is generally compressed before being transmitted over modern communication networks. Since memory resources may be limited, the size of the video can also be a problem when the video is stored on a storage device. In many cases, video compression devices use software and / or hardware at the transmitter to code the video data, thereby reducing the amount of data required to represent the digital video image. The compressed data is then received at the receiver by a video decompression device that decodes the video data. Due to limited network resources and the ever - increasing demand for higher video quality, improved compression and decompression techniques that increase the compression ratio without sacrificing much or any of the image quality are desirable.
Summary of the Invention
Means for Solving the Problems
[0004] In an embodiment, the present disclosure is a method implemented in a decoder, the method comprising receiving, by a receiver of the decoder, a bitstream including a plurality of sub-pictures and flags related to a picture, wherein the sub-pictures are included in video coding layer (VCL) network abstraction layer (NAL) units, and determining, by a processor, that a value of a first NAL unit type is the same for all VCL NAL units related to the picture when the flag is set to a first value, determining, by a processor, that a value of the first NAL unit type for a VCL NAL unit including one or more of the sub-pictures of the picture is different from a value of a second NAL unit type for a VCL NAL unit including one or more of the sub-pictures of the picture when the flag is set to a second value, and decoding, by a processor, one or more of the sub-pictures based on the value of the first NAL unit type or the value of the second NAL unit type. A picture may be divided into a plurality of sub-pictures. Such sub-pictures may be coded in separate sub-bitstreams, which may then be combined into a bitstream for transmission to a decoder. For example, the sub-pictures may be used for a virtual reality (VR) application. In a particular example, a user may always view only a portion of a VR picture. Thus, when the flag is set to the first value, the processor determines that the value of the first NAL unit type is the same for all VCL NAL units related to the picture; when the flag is set to the second value, the processor determines that the value of the first NAL unit type for a VCL NAL unit including one or more of the sub-pictures of the picture is different from the value of the second NAL unit type for a VCL NAL unit including one or more of the sub-pictures of the picture; and the processor decodes one or more of the sub-pictures based on the value of the first NAL unit type or the value of the second NAL unit type. The value of the first NAL unit type for a VCL NAL unit including one or more of the sub-pictures of the picture is different from the value of the second NAL unit type for a VCL NAL unit including one or more of the sub-pictures of the picture. The processor decodes one or more of the sub-pictures based on the value of the first NAL unit type or the value of the second NAL unit type. Based on the value of the first NAL unit type or the value of the second NAL unit type, the processor decodes one or more of the sub-pictures. One or more of the sub-pictures are decoded by the processor based on the value of the first NAL unit type or the value of the second NAL unit type. A picture may be divided into a plurality of sub-pictures. Such sub-pictures may be coded in separate sub-bitstreams, which may then be combined into a bitstream for transmission to a decoder. For example, the sub-pictures may be used for a virtual reality (VR) application. In a particular example, a user may always view only a portion of a VR picture. Thus,
[0005] a picture can be divided into a plurality of sub-pictures. Such sub-pictures can be coded in separate sub-bitstreams, which can then be combined into a bitstream for transmission to a decoder. For example, the sub-pictures can be used for a virtual reality (VR) application. In a specific example, a user can always view only a part of a VR picture. Therefore, when displayed For example, the sub-pictures may be used for a virtual reality (VR) application. In a particular example, a user may always view only a portion of a VR picture. Thus, when displayed For example, the sub-pictures may be used for a virtual reality (VR) application. In a particular example, a user may always view only a portion of a VR picture. Thus, when displayed a user may always view only a portion of a VR picture. Thus, when displayed More bandwidth can be allocated to sub-pictures that are likely to occur, as shown in the table Sub-pictures with a low likelihood of being shown can be compressed to increase coding efficiency so that different sub-pictures may be transmitted at different resolutions. Further, the video stream may be encoded by using intra-random access point (IRAP) pictures. IRAP pictures are encoded by intra-prediction and can be decoded without reference to other pictures . Non-IRAP pictures may be encoded by inter-prediction and can be decoded by referring to other pictures . Non-IRAP pictures are more highly condensed than IRAP pictures. However, since IRAP pictures contain enough data to be decoded without reference to other pictures, the video sequence must start decoding from an IRAP picture . IRAP pictures can be used within sub-pictures and can enable dynamic resolution changes. Thus, the video system may transmit more IRAP pictures for sub-pictures that are more likely to be viewed (e.g., based on the user's current viewport), and fewer IRAP pictures for sub-pictures that are less likely to be viewed, to further increase coding efficiency . However, sub-pictures are part of the same picture. Thus, this approach may result in pictures that include both IRAP and non-IRAP sub-pictures. Some video systems are not equipped to handle mixed pictures that have both IRAP and non-IRAP regions. This disclosure The picture is mixed and thus includes a flag indicating whether it includes both IRAP components and non-IRAP components. Based on this flag, the decoder can appropriately decode the picture / sub-picture and process different sub-pictures differently when decoding in order to display them. This flag may be stored in the PPS and may be called mixed_nalu_types_in_pic_flag. Therefore, the disclosed mechanism enables the implementation of additional functions. Furthermore, the disclosed mechanism enables dynamic resolution changes when using the sub-picture bitstream. Therefore, the disclosed mechanism enables a lower-resolution sub-picture bitstream to be transmitted when streaming VR video without significantly degrading the user experience. Therefore, the disclosed mechanism improves coding efficiency and thus reduces the use of network resources, memory resources, and / or processing resources in the encoder and decoder. Optionally, in any of the above aspects, another implementation of the aspect stipulates that the bitstream includes a picture parameter set (PPS) that includes a flag. Optionally, in any of the above aspects, another implementation of the aspect stipulates that the value of the first NAL unit type indicates that the picture includes an Intra Random Access Point (IRAP) sub-picture, and the value of the second NAL unit type indicates that the picture includes a non-IRAP sub-picture.
[0006]
[0007]
[0008] Optionally, in any of the above aspects, another implementation of the aspect comprises: The type value is a random access decodable leading picture. Instantaneous Decoding Refresh (IDR) with a leading picture fresh) (IDR_W_RADL), IDR without leading picture (IDR_N_LP), or clean It is defined as being equal to the clean random access (CRA) NAL unit type (CRA_NUT). Determine.
[0009] Optionally, in any of the above aspects, another implementation of the aspect comprises: The value of the type is the trailing picture NAL unit type (TRAIL_NU T), random-access decodable leading picture NAL unit type (RADL_NUT), or or random access skip reading pictures (RASL) The RFC 24482 specifies that the NAL unit type is equal to the RASL_NUT (apped leading picture) NAL unit type (RASL_NUT).
[0010] Optionally, in any of the above aspects, another implementation of the aspect is It is specified that the flag is nalu_types_in_pic_flag.
[0011] Optionally, in any of the above aspects, another implementation of the aspect refers to a PPS. A picture has two or more VCL NAL units, and the VCL NAL units are NAL unit types. When specifying that the same value of p(nal_unit_type) is not present, mixed_nalu_types_in_pic_f lag is equal to 1, the picture referring to the PPS has one or more of the VCL NAL units, and the VCL When specifying that the VCL NAL unit has the same value of nal_unit_type, mixed_nalu_types_ in_pic_flag is defined to be equal to 0.
[0012] In an embodiment, the present disclosure is a method implemented in an encoder, the method including steps of determining by a processor whether a picture includes a plurality of sub-pictures of different types, encoding the sub-pictures of the picture into a plurality of VCL NAL units in a bitstream, encoding, by the processor, into the bitstream a flag that is set to a first value when a value of a first NAL unit type is the same for all VCL NAL units related to the picture, and that is set to a second value when the value of the first NAL unit type for one or more of the VCL NAL units including one or more of the sub-pictures of the picture is different from a value of a second NAL unit type for one or more of the VCL NAL units including one or more of the sub-pictures of the picture, and storing, by memory coupled to the processor, the bitstream for transmission to a decoder. The picture may be partitioned into a plurality of sub-pictures. Such sub-pictures may be coded into separate sub-bitstreams, which may then be combined into a bitstream for transmission to a decoder. For example,
[0013] sub-bitstreams, and those sub-bitstreams may then be combined into a bitstream for transmission to a decoder. For example, sub-bitstreams, and those sub-bitstreams may then be combined into a bitstream for transmission to a decoder. For example, sub-bitstreams, and those sub-bitstreams may then be combined into a bitstream for transmission to a decoder. For example, For example, the sub-pictures may be used for virtual reality (VR) applications. In a specific example, the user may always view only a part of the VR picture. Therefore, it is possible to allocate more bandwidth to the sub-pictures that are more likely to be displayed, and the sub-pictures that are less likely to be displayed may be compressed to improve coding efficiency, so that different sub-pictures may be transmitted at different resolutions. Furthermore, the video stream may be encoded by using an intra-random access point (IRAP) picture. The IRAP picture is coded by intra prediction and can be decoded without reference to other pictures. Non-IRAP pictures may be coded by inter-prediction and can be decoded by referring to other pictures. Non-IRAP pictures are significantly more compressed than IRAP pictures. However, since the IRAP picture contains enough data to be decoded without referring to other pictures, the video sequence must start decoding from the IRAP picture. The IRAP picture can be used within a sub-picture and enables dynamic resolution change. Therefore, the video system may transmit more IRAP pictures for sub-pictures that are more likely to be viewed (e.g., based on the user's current viewport), and may transmit fewer IRAP pictures for sub-pictures that are less likely to be viewed to further improve coding efficiency. However, the sub-picture is a part of the same picture. Therefore, this method applies to both IRAP sub-pictures and non-IRAP sub-pictures. Some video systems may result in a picture containing IRAP and non-IRAP regions. There is no provision to handle mixed pictures that have both. Therefore, it contains a flag indicating whether it contains both IRAP and non-IRAP components. Based on the flags, the decoder will decode and display the picture / subpicture appropriately. Therefore, different sub-pictures may be treated differently when decoding. This may be stored in the .mixed_nalu_types_in_pic_flag and may be called mixed_nalu_types_in_pic_flag. The mechanisms shown allow for the implementation of additional functionality. The system allows dynamic resolution changes when using subpicture bitstreams. Therefore, the disclosed mechanism does not significantly impair the user experience. Instead of using a lower resolution subpicture bitstream when streaming VR video, Therefore, the disclosed mechanism allows the codec improves the coding efficiency and therefore the network resources in the encoder and decoder , reducing the use of memory resources and / or processing resources.
[0014] Optionally, in any of the above aspects, another implementation of the aspect is to convert the PPS into a bitstream. a step of encoding the flag into the PPS stream, wherein the flag is encoded into the PPS. This includes:
[0015] Optionally, in any of the above aspects, another implementation of the aspect comprises: The value of the type indicates that the picture contains an IRAP subpicture, and the second NAL unit type It is stipulated that the value of p indicates that the picture contains a non-IRAP sub-picture.
[0016] Optionally, in any of the above aspects, another implementation of the aspect is that the value of the first NAL unit type is stipulated to be equal to IDR_W_RADL, IDR_N_LP, or CRA_NUT.
[0017] Optionally, in any of the above aspects, another implementation of the aspect is that the value of the second NAL unit type is stipulated to be equal to TRAIL_NUT, RADL_NUT, or RASL_NUT.
[0018] Optionally, in any of the above aspects, another implementation of the aspect is that the flag is mixed_ nalu_types_in_pic_flag.
[0019] Optionally, in any of the above aspects, another implementation of the aspect is that the picture referring to the PPS has two or more of the VCL NAL units, and when it is specified that the VCL NAL units do not have the same value of nal_unit_type, mixed_nalu_types_in_pic_flag is equal to 1, and when the picture referring to the PPS has one or more of the VCL NAL units, and it is specified that the VCL NAL units have the same value of nal_u nit_type, mixed_nalu_types_in_pic_flag is stipulated to be equal to 0.
[0020] In an embodiment, the present disclosure includes a processor, a receiver coupled to the processor, a memory coupled to the processor, and a transmitter coupled to the processor, and the processor, the receiver, the memory, and the transmitter are configured to execute the method of any of the above aspects. including a video coding device.
[0021] In an embodiment, the present disclosure is a non-transitory computer-readable medium including a computer program product for use by a video coding device, wherein the computer program product includes computer-executable instructions stored on the non-transitory computer-readable medium that cause a video coding device to execute the method of any of the above aspects when executed by a processor. Including a non-transitory computer-readable medium containing computer-executable instructions.
[0022] In an embodiment, the present disclosure is a receiving means for receiving a bitstream including a plurality of sub-pictures and flags related to a picture, wherein the sub-pictures are included in a plurality of VCL NAL units, a receiving means, and a determining means, wherein when the flag is set to a first value, it is determined that the value of the first NAL unit type is the same for all VCL NAL units related to the picture, and when the flag is set to a second value, it is determined that the value of the first NAL unit type for VCL NAL units including one or more of the sub-pictures of the picture is different from the value of the second NAL unit type for VCL NAL units including one or more of the sub-pictures of the picture, a determining means for determining, and a decoding means for decoding one or more of the sub-pictures based on the value of the first NAL unit type or also the value of the second NAL unit type. Including a decoder.
[0023] Optionally, in any of the above aspects, another implementation of the aspect provides that the decoder is further configured to execute the method of any of the above aspects.
[0024] In an embodiment, the present disclosure provides a method for determining whether a picture contains multiple sub-pictures of different types. a determining means for determining whether a sub-picture of a picture is a picture; and an encoding means for encoding a sub-picture of a picture into a picture. The first NAL unit type value is It is set to the first value when it is the same for all VCL NAL units associated with a picture. The first NAL unit is for a VCL NAL unit that contains one or more of the picture's subpictures. The unit type value relates to a VCL NAL unit that contains one or more of the picture's subpictures. The flag set to the second value when it differs from the value of the second NAL unit type is bitstreamed. encoding means for encoding the bitstream into a bitstream for transmission to a decoder; and storage means for storing the program.
[0025] Optionally, in any of the above aspects, another implementation of the aspect is further characterized in that the encoder: It is provided that the device is further configured to perform the method of any of the above aspects.
[0026] For purposes of clarity, any one of the above-described embodiments may be considered a new embodiment within the scope of the present disclosure. may be combined with any one or more of the other above-mentioned embodiments to produce stomach.
[0027] These and other features are best understood from the following detailed description taken in conjunction with the accompanying drawings and claims. By understanding it will be understood more clearly.
[0028] For a more complete understanding of this disclosure, reference is made to the accompanying drawings, in which like reference numerals represent like parts, and in which: The following brief description, which is to be interpreted in connection with the detailed description, is hereby incorporated by reference.
Brief Description of the Drawings
[0029]
Figure 1
Figure 2
Figure 3
Figure 4
Figure 5
Figure 6
Figure 7
Figure 8
Figure 9
Figure 10
Figure 11
Modes for Carrying Out the Invention
[0030] Exemplary implementations of one or more embodiments are given below, but the disclosed systems and / It should first be understood that the method may be implemented using any number of techniques, whether currently known or existing. This disclosure is not to be limited in any way to the exemplary implementations, figures, and techniques shown below, including the exemplary designs and implementations illustrated and described herein, but may be modified within the full scope of the equivalents of the appended claims and within the scope of those claims. along with the scope of those claims.
[0031] The following acronyms, Coded Video Sequence (CVS), Decoded Picture Buffer (DPB), Instantaneous Decoding Refresh (IDR), Intra Random Access Point (IRAP), Least Significant Bit (LSB), Most Significant Bit (MSB), Network Abstraction Layer (NAL), Picture Order Count (POC), Raw Byte Sequence Payload (RBSP), Sequence Parameter Set (SPS), and Working Draft (WD) are used herein.
[0032] Many video compression techniques can be used to reduce the size of video files with minimal loss of data. For example, video compression techniques may include performing spatial (e.g., intra-picture) prediction and temporal (inter-picture) prediction to reduce or remove redundancy in the data of a video sequence. For block-based video coding, a video slice (e.g., a video picture or a portion of a video picture) may be partitioned into video blocks, and the video blocks may be tree blocks, coding blocks, etc. for example, a video slice (e.g., a video picture or a portion of a video picture) may be partitioned into video blocks, and the video blocks may be tree blocks, coding blocks, etc. GTS tree block (CTB), coding tree unit (CTU), coding unit (CU) and / or may also be referred to as a coding node. Video blocks within an intra-coded (I) slice are coded using spatial prediction related to reference samples in neighboring blocks within the same picture. Video blocks within an inter-coded unidirectional prediction (P) or bidirectional prediction (B) slice of a picture may be coded by using spatial prediction related to reference samples in neighboring blocks within the same picture or temporal prediction related to reference samples in other reference pictures. A picture may also be referred to as a frame and / or an image, and a reference picture may also be referred to as a reference frame and / or a reference image. Spatial or temporal prediction results in a prediction block representing the image block. Residual data represents the pixel difference between the original block and the prediction block. Thus, an inter-coded block is coded by a motion vector pointing to a block of reference samples forming the prediction block and residual data indicating the difference between the coded block and the prediction block. An intra-coded block is coded by an intra-coding mode and residual data . For further compression, the residual data may be transformed from the pixel domain to the transform domain. These result in residual transform coefficients, which may be quantized . First, the quantized transform coefficients may be arranged in a two-dimensional array. The quantized transform coefficients may be scanned to generate a one-dimensional vector of the transform coefficients. Entropy ... ... ... ... ... ... ... ... Coding may be applied to achieve further compression. Such video compression techniques are examined in detail below.
[0033] To ensure that the encoded video can be accurately decoded, the video is encoded and decoded in accordance with the corresponding video coding standard. Video coding standards include ITU-T H.261 of the International Telecommunication Union (ITU) Standardization Sector (ITU-T), ISO / IEC Moving Picture Experts Group (MPEG)-1 Part 2, ITU-T H.262 or ISO / IEC MPEG-2 Part 2, ITU-T H.263, ISO / IEC MPEG-4 Part 2, ITU-T H.264 or ISO / IEC MPEG-4 Part 10 also known as Advanced Video Coding (AVC), ITU-T H.26 5 or High Efficiency Video Coding (HEVC) also known as MPEG-H Part 2. AVC includes extensions such as Scalable Video Coding (SVC), Multiview Video Coding (MVC) and Multiview Video Coding plus Depth (MVC+D), as well as three-dimensional (3D) AVC (3D-AVC). HEVC includes extensions such as Scalable HEVC (SHVC), Multiview HEVC (MV-HEVC), and 3D HEVC (3D-HEVC). The joint video experts team (JVET) of ITU-T and ISO / IEC is developing a video coding standard called Versatile Video Coding (VVC). Scalable Video Coding (SVC), Multiview Video Coding (MVC), and Multiview Video Coding plus Depth (MVC+D), as well as extensions such as three-dimensional (3D) AVC (3D-AVC). HEVC includes extensions such as Scalable HEVC (SHVC), Multiview HEVC (MV-HEVC), and 3D HEVC (3D-HEVC). The joint video experts team (JVET) of ITU-T and ISO / IEC is developing a video coding standard called Versatile Video Coding (VVC). including extensions such as three-dimensional (3D) AVC (3D-AVC). HEVC includes extensions such as Scalable HEVC (SHVC), Multiview HEVC (MV-HEVC), and 3D HEVC (3D-HEVC). The joint video experts team (JVET) of ITU-T and ISO / IEC is developing a video coding standard called Versatile Video Coding (VVC). including extensions such as Scalable HEVC (SHVC), Multiview HEVC (MV-HEVC), and 3D HEVC (3D-HEVC). The joint video experts team (JVET) of ITU-T and ISO / IEC is developing a video coding standard called Versatile Video Coding (VVC). including extensions such as Scalable HEVC (SHVC), Multiview HEVC (MV-HEVC), and 3D HEVC (3D-HEVC). The joint video experts team (JVET) of ITU-T and ISO / IEC is developing a video coding standard called Versatile Video Coding (VVC). including extensions such as Scalable HEVC (SHVC), Multiview HEVC (MV-HEVC), and 3D HEVC (3D-HEVC). The joint video experts team (JVET) of ITU-T and ISO / IEC is developing a video coding standard called Versatile Video Coding (VVC). Started the development of the decoding standard. VVC includes the description of the algorithm, the encoder-side description of the VVC Working Draft (WD), and WD including JVET-M1001-v6 that provides reference software is included in the WD.
[0034] The video coding system may encode video by using IRAP pictures and non-IRAP pictures. An IRAP picture is a picture coded by intra prediction that serves as a random access point for the video sequence . In intra prediction, a block of a picture is coded by reference to other blocks within the same picture. This is in contrast to non-IRAP pictures that use inter prediction. In inter prediction, a block of the current picture is coded by reference to other blocks within a reference picture different from the current picture. Since an IRAP picture is coded without reference to other pictures, it can be decoded without first decoding any other pictures. Thus, the decoder can start decoding the video sequence at any IRAP picture. In contrast, non-IRAP pictures are coded with reference to other pictures and thus, generally, the decoder cannot start decoding the video sequence at a non-IRAP picture. Also, an IRAP picture refreshes the DPB. This is because the IRAP picture is the start point of the CVS and pictures within the CVS do not reference pictures within the previous CVS. Thus, an IRAP picture further prevents coding errors related to inter prediction from propagating through the IRAP picture Since it cannot be done, such an error can be stopped. However, the IRAP picture is significantly larger than the non-IRAP picture in terms of data size. Therefore, generally, the video sequence includes many non-IRAP pictures and a smaller number of IRAP pictures scattered among them to balance coding efficiency and functionality. For example, a 60-frame CVS may include one IRAP picture and 59 non-IRAP pictures.
[0035] In some cases, the video coding system may be used to code virtual reality (VR) video, sometimes also called 360-degree video. The VR video may include a sphere of video content that is displayed as if the user is at the center of the sphere. Only a portion of the sphere, called the viewport is displayed to the user. For example, the user may use a head-mounted display (HMD) to select and display the viewport of the sphere based on the movement of the user's head . This gives the impression of physically existing within the virtual space depicted by the video. To achieve this result, each picture of the video sequence includes the entire sphere of video data for the corresponding moment. However, only a small portion of the picture (e.g., just a single viewport) is displayed to the user. The rest of the picture is discarded without being rendered. Generally, the entire picture is transmitted so that different viewports can be dynamically selected and displayed according to the movement of the user's head. This approach may lead to a very large video file size. To improve coding efficiency, some systems divide pictures into sub-pictures . s lected and displayed, and generally, the entire picture is transmitted so that different viewports can be dynamically selected and displayed according to the movement of the user's head. This approach may lead to a very large video file size.
[0036] To improve coding efficiency, some systems divide pictures into sub-pictures is segmented. A sub-picture is a defined spatial region of a picture. Each sub-picture includes the corresponding viewport of the picture. The video can be encoded at two or more resolutions. Each resolution is encoded into a different sub-bitstream. When a user streams a VR video, the coding system can merge the sub-bitstreams into a bitstream for transmission based on the current viewport being used by the user. In particular, the current viewport is obtained from a high-resolution sub-bitstream, and the unseen viewports are obtained from a low-resolution bitstream. In this way, the highest quality video is presented to the user, and lower quality video is discarded. If the user selects a new viewport, lower resolution video is presented to the user. The decoder may request to receive higher resolution video for the new viewport. The encoder can then change the merge process accordingly. When an IRAP picture is reached, the decoder can start decoding the higher resolution video sequence for the new viewport. This approach significantly improves video compression without adversely affecting the user's
[0037] One concern with the above approach is that the length of time required to change resolutions is based on the length of time until an IRAP picture is reached. This is because, as described above, the decoder cannot start decoding different video sequences in non-IRAP pictures. One approach to reducing such latency is is to include artifacts. However, this leads to an increase in file size. To balance functionality and coding efficiency, different viewports / sub-pictures may include IRAP pictures at different frequencies. For example, a viewport that is more likely to be viewed may have more IRAP pictures than other viewports. For example, in the context of a basketball game, a viewport that views the stands or ceiling is less likely to be viewed by the user, so a viewport related to the basket and / or the court may include IRAP pictures more frequently. This approach leads to other problems. In particular, a sub-picture that includes a viewport is part of a single picture. When different sub-pictures have IRAP pictures at different frequencies, parts of the picture include both IRAP sub-pictures and non-IRAP sub-pictures. This is a problem because the picture is stored in the bitstream by using NAL units. An NAL unit is a storage unit that includes a picture parameter set or a slice and the corresponding slice header. An access unit is a unit that includes the entire picture. Therefore, an access unit includes all of the NAL units related to the picture. An NAL unit also includes a type that indicates the type of the picture that includes the slice. In some video systems, it is required that all NAL units related to a single picture (e.g., included in the same access unit) be of the same type. Therefore, the storage mechanism of the NAL unit is such that the picture includes both IRAP sub-pictures and non-IRAP sub-pictures. For example, a viewport that is more likely to be seen may have more IRAP pictures than other viewports. For example, in the context of a basketball game, a viewport that views the stands or ceiling is less likely to be seen by the user, so a viewport related to the basket and / or the court may include IRAP pictures more frequently. For example, in the context of a basketball game, a viewport that views the stands or ceiling is less likely to be seen by the user, so a viewport related to the basket and / or the court may include IRAP pictures more frequently. is less likely to be viewed by the user, so a viewport related to the basket and / or the court may include IRAP pictures more frequently than such a viewport. is less likely to be viewed by the user, so a viewport related to the basket and / or the court may include IRAP pictures more frequently than such a viewport. This approach leads to other problems. In particular, a sub-picture that includes a viewport is part of a single picture. When different sub-pictures have IRAP pictures at different frequencies, parts of the picture include both IRAP sub-pictures and non-IRAP sub-pictures. This is a problem because the picture is stored in the bitstream by using NAL units. An NAL unit is a storage unit that includes a picture parameter set or a slice and the corresponding slice header. An access unit is a unit that includes the entire picture. Therefore, an access unit includes all of the NAL units related to the picture. An NAL unit also includes a type that indicates the type of the picture that includes the slice. In some video systems, it is required that all NAL units related to a single picture (e.g., included in the same access unit) be of the same type. Therefore, the storage mechanism of the NAL unit is such that the picture includes both IRAP sub-pictures and non-IRAP sub-pictures.
[0038] This approach leads to other problems. In particular, a sub-picture that includes a viewport is part of a single picture. When different sub-pictures have IRAP pictures at different frequencies, parts of the picture include both IRAP sub-pictures and non-IRAP sub-pictures. This is a problem because the picture is stored in the bitstream by using NAL units. An NAL unit is a storage unit that includes a picture parameter set or a slice and the corresponding slice header. An access unit is a unit that includes the entire picture. Therefore, an access unit includes all of the NAL units related to the picture. An NAL unit also includes a type that indicates the type of the picture that includes the slice. In some video systems, it is required that all NAL units related to a single picture (e.g., included in the same access unit) be of the same type. Therefore, the storage mechanism of the NAL unit is such that the picture includes both IRAP sub-pictures and non-IRAP sub-pictures. This approach leads to other problems. In particular, a sub-picture that includes a viewport is part of a single picture. When different sub-pictures have IRAP pictures at different frequencies, parts of the picture include both IRAP sub-pictures and non-IRAP sub-pictures. This is a problem because the picture is stored in the bitstream by using NAL units. An NAL unit is a storage unit that includes a picture parameter set or a slice and the corresponding slice header. An access unit is a unit that includes the entire picture. Therefore, an access unit includes all of the NAL units related to the picture. An NAL unit also includes a type that indicates the type of the picture that includes the slice. In some video systems, it is required that all NAL units related to a single picture (e.g., included in the same access unit) be of the same type. Therefore, the storage mechanism of the NAL unit is such that the picture includes both IRAP sub-pictures and non-IRAP sub-pictures. This approach leads to other problems. In particular, a sub-picture that includes a viewport is part of a single picture. When different sub-pictures have IRAP pictures at different frequencies, parts of the picture include both IRAP sub-pictures and non-IRAP sub-pictures. This is a problem because the picture is stored in the bitstream by using NAL units. An NAL unit is a storage unit that includes a picture parameter set or a slice and the corresponding slice header. An access unit is a unit that includes the entire picture. Therefore, an access unit includes all of the NAL units related to the picture. An NAL unit also includes a type that indicates the type of the picture that includes the slice. In some video systems, it is required that all NAL units related to a single picture (e.g., included in the same access unit) be of the same type. Therefore, the storage mechanism of the NAL unit is such that the picture includes both IRAP sub-pictures and non-IRAP sub-pictures. This is a problem because the picture is stored in the bitstream by using NAL units. An NAL unit is a storage unit that includes a picture parameter set or a slice and the corresponding slice header. An access unit is a unit that includes the entire picture. Therefore, an access unit includes all of the NAL units related to the picture. An NAL unit also includes a type that indicates the type of the picture that includes the slice. In some video systems, it is required that all NAL units related to a single picture (e.g., included in the same access unit) be of the same type. Therefore, the storage mechanism of the NAL unit is such that the picture includes both IRAP sub-pictures and non-IRAP sub-pictures. This is a problem because the picture is stored in the bitstream by using NAL units. An NAL unit is a storage unit that includes a picture parameter set or a slice and the corresponding slice header. An access unit is a unit that includes the entire picture. Therefore, an access unit includes all of the NAL units related to the picture. An NAL unit also includes a type that indicates the type of the picture that includes the slice. In some video systems, it is required that all NAL units related to a single picture (e.g., included in the same access unit) be of the same type. Therefore, the storage mechanism of the NAL unit is such that the picture includes both IRAP sub-pictures and non-IRAP sub-pictures. An NAL unit is a storage unit that includes a picture parameter set or a slice and the corresponding slice header. An access unit is a unit that includes the entire picture. Therefore, an access unit includes all of the NAL units related to the picture. An NAL unit also includes a type that indicates the type of the picture that includes the slice. In some video systems, it is required that all NAL units related to a single picture (e.g., included in the same access unit) be of the same type. Therefore, the storage mechanism of the NAL unit is such that the picture includes both IRAP sub-pictures and non-IRAP sub-pictures. An access unit is a unit that includes the entire picture. Therefore, an access unit includes all of the NAL units related to the picture. An NAL unit also includes a type that indicates the type of the picture that includes the slice. In some video systems, it is required that all NAL units related to a single picture (e.g., included in the same access unit) be of the same type. Therefore, the storage mechanism of the NAL unit is such that the picture includes both IRAP sub-pictures and non-IRAP sub-pictures. An NAL unit also includes a type that indicates the type of the picture that includes the slice. In some video systems, it is required that all NAL units related to a single picture (e.g., included in the same access unit) be of the same type. Therefore, the storage mechanism of the NAL unit is such that the picture includes both IRAP sub-pictures and non-IRAP sub-pictures. In some video systems, it is required that all NAL units related to a single picture (e.g., included in the same access unit) be of the same type. Therefore, the storage mechanism of the NAL unit is such that the picture includes both IRAP sub-pictures and non-IRAP sub-pictures. In some video systems, it is required that all NAL units related to a single picture (e.g., included in the same access unit) be of the same type. Therefore, the storage mechanism of the NAL unit is such that the picture includes both IRAP sub-pictures and non-IRAP sub-pictures. In some video systems, it is required that all NAL units related to a single picture (e.g., included in the same access unit) be of the same type. Therefore, the storage mechanism of the NAL unit is such that the picture includes both IRAP sub-pictures and non-IRAP sub-pictures. It may malfunction when a square is included.
[0039] Disclosed herein is a mechanism for adjusting the storage mode of NAL to support pictures including both IRAP sub - pictures and non - IRAP sub - pictures. This enables VR video that includes different frequencies of IRAP sub - pictures for different viewports. In a first example, disclosed herein is a flag indicating whether a picture is mixed. For example, the flag may indicate that the picture includes both IRAP sub - pictures and non - IRAP sub - pictures. Based on this flag, the decoder may process different types of sub - pictures differently when decoding to properly decode and display the picture / sub - picture. This flag may be stored in the Picture Parameter Set (PPS) and may be called mixed_nalu_types_in_pic_flag. In a second example, disclosed herein is a flag indicating whether a picture is mixed. For example, the flag may indicate that the picture includes both IRAP sub - pictures and non - IRAP sub - pictures. Further, the flag constrains the picture such that the mixed picture includes exactly two NAL unit types, one IRAP type and one non - IRAP type. For example, the picture may be an Instantaneous Decoding Refresh (IDR) with Random Access Decodable Leading Picture (IDR_W_RADL), an IDR without a leading picture (IDR_N_LP), or a Clean Random Access (CRA) NAL unit type (CRA_NUT). In a first example, disclosed herein is a flag indicating whether a picture is mixed. For example, the flag may indicate that the picture includes both IRAP sub - pictures and non - IRAP sub - pictures. Based on this flag, the decoder may process different types of sub - pictures differently when decoding to properly decode and display the picture / sub - picture. This flag may be stored in the Picture Parameter Set (PPS) and may be called mixed_nalu_types_in_pic_flag. In a first example, disclosed herein is a flag indicating whether a picture is mixed. For example, the flag may indicate that the picture includes both IRAP sub - pictures and non - IRAP sub - pictures. Based on this flag, the decoder may process different types of sub - pictures differently when decoding to properly decode and display the picture / sub - picture. This flag may be stored in the Picture Parameter Set (PPS) and may be called mixed_nalu_types_in_pic_flag. In a first example, disclosed herein is a flag indicating whether a picture is mixed. For example, the flag may indicate that the picture includes both IRAP sub - pictures and non - IRAP sub - pictures. Based on this flag, the decoder may process different types of sub - pictures differently when decoding to properly decode and display the picture / sub - picture. This flag may be stored in the Picture Parameter Set (PPS) and may be called mixed_nalu_types_in_pic_flag. In a first example, disclosed herein is a flag indicating whether a picture is mixed. For example, the flag may indicate that the picture includes both IRAP sub - pictures and non - IRAP sub - pictures. Based on this flag, the decoder may process different types of sub - pictures differently when decoding to properly decode and display the picture / sub - picture. This flag may be stored in the Picture Parameter Set (PPS) and may be called mixed_nalu_types_in_pic_flag. In a first example, disclosed herein is a flag indicating whether a picture is mixed. For example, the flag may indicate that the picture includes both IRAP sub - pictures and non - IRAP sub - pictures. Based on this flag, the decoder may process different types of sub - pictures differently when decoding to properly decode and display the picture / sub - picture. This flag may be stored in the Picture Parameter Set (PPS) and may be called mixed_nalu_types_in_pic_flag. In a first example, disclosed herein is a flag indicating whether a picture is mixed. For example, the flag may indicate that the picture includes both IRAP sub - pictures and non - IRAP sub - pictures. Based on this flag, the decoder may process different types of sub - pictures differently when decoding to properly decode and display the picture / sub - picture. This flag may be stored in the Picture Parameter Set (PPS) and may be called mixed_nalu_types_in_pic_flag.
[0040] In a second example, disclosed herein is a flag indicating whether a picture is mixed. For example, the flag may indicate that the picture includes both IRAP sub - pictures and non - IRAP sub - pictures. In a second example, disclosed herein is a flag indicating whether a picture is mixed. For example, the flag may indicate that the picture includes both IRAP sub - pictures and non - IRAP sub - pictures. In a second example, disclosed herein is a flag indicating whether a picture is mixed. For example, the flag may indicate that the picture includes both IRAP sub - pictures and non - IRAP sub - pictures. In a second example, disclosed herein is a flag indicating whether a picture is mixed. For example, the flag may indicate that the picture includes both IRAP sub - pictures and non - IRAP sub - pictures. Further, the flag constrains the picture such that the mixed picture includes exactly two NAL unit types, one IRAP type and one non - IRAP type. In a second example, disclosed herein is a flag indicating whether a picture is mixed. For example, the flag may indicate that the picture includes both IRAP sub - pictures and non - IRAP sub - pictures. Further, the flag constrains the picture such that the mixed picture includes exactly two NAL unit types, one IRAP type and one non - IRAP type. In a second example, disclosed herein is a flag indicating whether a picture is mixed. For example, the flag may indicate that the picture includes both IRAP sub - pictures and non - IRAP sub - pictures. Further, the flag constrains the picture such that the mixed picture includes exactly two NAL unit types, one IRAP type and one non - IRAP type. In a second example, disclosed herein is a flag indicating whether a picture is mixed. For example, the flag may indicate that the picture includes both IRAP sub - pictures and non - IRAP sub - pictures. Further, the flag constrains the picture such that the mixed picture includes exactly two NAL unit types, one IRAP type and one non - IRAP type. It may include an IRAP NAL unit that contains only one of them. Further, the picture may be a trailing Picture NAL unit type (TRAIL_NUT), a random access decodable leading picture NAL unit type (RADL_NUT), or a non-IRAP NAL unit that contains only one of the random access skip leading picture (RASL) NAL unit types (RASL_NUT). Based on this flag, the decoder may decode and display the picture / sub-picture appropriately, and may process different sub-pictures differently when decoding. This flag may be stored in the PPS and may be called mixed_nalu_types_in_pic_flag.
[0041] FIG. 1 is a flowchart of an exemplary operating method 100 for coding a video signal. In particular, the video signal is encoded in an encoder. The encoding process compresses the video signal by using various mechanisms to reduce the video file size. A smaller file size enables the compressed video file to be transmitted to the user while reducing the associated bandwidth overhead. Then, the decoder decodes the compressed video file and reconstructs the original video signal for display to the end user. Generally, the decoding process faithfully mimics the encoding process to enable the decoder to reconstruct the video signal without contradictions.
[0042] In step 101, the video signal is input to the encoder. For example, the video signal may be an uncompressed video file stored in memory. As another example, the video The off-file is captured by a video capture device such as a video camera and may be encoded to support live streaming of video. The video file may include both an audio component and a video component. The video component includes a series of image frames that give a visual impression of movement when viewed in sequence. The frames include pixels represented by light, referred to herein as luma components (or luma samples), and color, referred to as chroma components (or color samples). In some examples, the frame may also include depth values to support three-dimensional viewing.
[0043] In step 103, the video is segmented into blocks. Segmentation involves sub-dividing the pixels of each frame into square and / or rectangular blocks for compression. For example, in High Efficiency Video Coding (HEVC), also known as H.265 and MPEG-H Part 2, the frame may first be divided into Coding Tree Blocks that are blocks of a predefined size (e.g., 64 pixels × 64 pixels). The CTU includes both luma samples and chroma samples. The CTU is divided into blocks, and a coding tree may be used to repeatedly sub-divide the blocks until a configuration that supports further encoding is achieved. For example, the luma component of the frame may be sub-divided until the individual blocks include relatively uniform lighting values. For example, the chroma component of the frame may be sub-divided until the individual blocks include relatively uniform color values. Thus, the segmentation mechanism varies depending on the content of the video frame.
[0044] In step 105, various compression mechanisms are used to compress the image blocks segmented in step 103. For example, inter prediction and / or intra prediction may be used. Inter prediction is designed to utilize the fact that objects in a normal scene tend to appear in consecutive frames. Therefore, blocks depicting an object in a reference frame need not be repeatedly shown in neighboring frames. In particular, an object such as a table may remain in a fixed position over multiple frames. Thus, the table is shown once, and adjacent frames can refer back to the reference frame. A pattern matching mechanism may be used to match an object over multiple frames. Further, a moving object may be represented over multiple frames, for example, due to the movement of the object or the movement of the camera. As a specific example, a video may show a car moving across the screen over multiple frames. A motion vector may be used to indicate such movement. A motion vector is a two-dimensional vector that gives the offset from the coordinates of an object in a frame to the coordinates of the object in a reference frame. Therefore, inter prediction can encode an image block in the current frame as a set of motion vectors indicating the offsets from the corresponding blocks in the reference frame. For example, inter prediction and / or intra prediction may be used. Inter prediction is designed to utilize the fact that objects in a normal scene tend to appear in consecutive frames. Therefore, blocks depicting an object in a reference frame need not be repeatedly shown in neighboring frames. In particular, an object such as a table may remain in a fixed position over multiple frames. Thus, the table is shown once, and adjacent frames can refer back to the reference frame. A pattern matching mechanism may be used to match an object over multiple frames. Further, a moving object may be represented over multiple frames, for example, due to the movement of the object or the movement of the camera. As a specific example, a video may show a car moving across the screen over multiple frames. A motion vector may be used to indicate such movement. A motion vector is a two-dimensional vector that gives the offset from the coordinates of an object in a frame to the coordinates of the object in a reference frame. Therefore, inter prediction can encode an image block in the current frame as a set of motion vectors indicating the offsets from the corresponding blocks in the reference frame. Inter prediction is designed to utilize the fact that objects in a normal scene tend to appear in consecutive frames. Therefore, blocks depicting an object in a reference frame need not be repeatedly shown in neighboring frames. In particular, an object such as a table may remain in a fixed position over multiple frames. Thus, the table is shown once, and adjacent frames can refer back to the reference frame. A pattern matching mechanism may be used to match an object over multiple frames. Further, a moving object may be represented over multiple frames, for example, due to the movement of the object or the movement of the camera. As a specific example, a video may show a car moving across the screen over multiple frames. A motion vector may be used to indicate such movement. A motion vector is a two-dimensional vector that gives the offset from the coordinates of an object in a frame to the coordinates of the object in a reference frame. Therefore, inter prediction can encode an image block in the current frame as a set of motion vectors indicating the offsets from the corresponding blocks in the reference frame. Therefore, blocks depicting an object in a reference frame need not be repeatedly shown in neighboring frames. In particular, an object such as a table may remain in a fixed position over multiple frames. Thus, the table is shown once, and adjacent frames can refer back to the reference frame. A pattern matching mechanism may be used to match an object over multiple frames. Further, a moving object may be represented over multiple frames, for example, due to the movement of the object or the movement of the camera. As a specific example, a video may show a car moving across the screen over multiple frames. A motion vector may be used to indicate such movement. A motion vector is a two-dimensional vector that gives the offset from the coordinates of an object in a frame to the coordinates of the object in a reference frame. Therefore, inter prediction can encode an image block in the current frame as a set of motion vectors indicating the offsets from the corresponding blocks in the reference frame. In particular, an object such as a table may remain in a fixed position over multiple frames. Thus, the table is shown once, and adjacent frames can refer back to the reference frame. A pattern matching mechanism may be used to match an object over multiple frames. Further, a moving object may be represented over multiple frames, for example, due to the movement of the object or the movement of the camera. As a specific example, a video may show a car moving across the screen over multiple frames. A motion vector may be used to indicate such movement. A motion vector is a two-dimensional vector that gives the offset from the coordinates of an object in a frame to the coordinates of the object in a reference frame. Therefore, inter prediction can encode an image block in the current frame as a set of motion vectors indicating the offsets from the corresponding blocks in the reference frame. Thus, the table is shown once, and adjacent frames can refer back to the reference frame. A pattern matching mechanism may be used to match an object over multiple frames. Further, a moving object may be represented over multiple frames, for example, due to the movement of the object or the movement of the camera. As a specific example, a video may show a car moving across the screen over multiple frames. A motion vector may be used to indicate such movement. A motion vector is a two-dimensional vector that gives the offset from the coordinates of an object in a frame to the coordinates of the object in a reference frame. Therefore, inter prediction can encode an image block in the current frame as a set of motion vectors indicating the offsets from the corresponding blocks in the reference frame. A pattern matching mechanism may be used to match an object over multiple frames. Further, a moving object may be represented over multiple frames, for example, due to the movement of the object or the movement of the camera. As a specific example, a video may show a car moving across the screen over multiple frames. A motion vector may be used to indicate such movement. A motion vector is a two-dimensional vector that gives the offset from the coordinates of an object in a frame to the coordinates of the object in a reference frame. Therefore, inter prediction can encode an image block in the current frame as a set of motion vectors indicating the offsets from the corresponding blocks in the reference frame. Further, a moving object may be represented over multiple frames, for example, due to the movement of the object or the movement of the camera. As a specific example, a video may show a car moving across the screen over multiple frames. A motion vector may be used to indicate such movement. A motion vector is a two-dimensional vector that gives the offset from the coordinates of an object in a frame to the coordinates of the object in a reference frame. Therefore, inter prediction can encode an image block in the current frame as a set of motion vectors indicating the offsets from the corresponding blocks in the reference frame. As a specific example, a video may show a car moving across the screen over multiple frames. A motion vector may be used to indicate such movement. A motion vector is a two-dimensional vector that gives the offset from the coordinates of an object in a frame to the coordinates of the object in a reference frame. Therefore, inter prediction can encode an image block in the current frame as a set of motion vectors indicating the offsets from the corresponding blocks in the reference frame. A motion vector may be used to indicate such movement. A motion vector is a two-dimensional vector that gives the offset from the coordinates of an object in a frame to the coordinates of the object in a reference frame. Therefore, inter prediction can encode an image block in the current frame as a set of motion vectors indicating the offsets from the corresponding blocks in the reference frame. A motion vector is a two-dimensional vector that gives the offset from the coordinates of an object in a frame to the coordinates of the object in a reference frame. Therefore, inter prediction can encode an image block in the current frame as a set of motion vectors indicating the offsets from the corresponding blocks in the reference frame. Therefore, inter prediction can encode an image block in the current frame as a set of motion vectors indicating the offsets from the corresponding blocks in the reference frame. A motion vector is a two-dimensional vector that gives the offset from the coordinates of an object in a frame to the coordinates of the object in a reference frame. Therefore, inter prediction can encode an image block in the current frame as a set of motion vectors indicating the offsets from the corresponding blocks in the reference frame. Therefore, inter prediction can encode an image block in the current frame as a set of motion vectors indicating the offsets from the corresponding blocks in the reference frame. as a set of motion vectors indicating the offsets from the corresponding blocks in the reference frame. encoded.
[0045] Intra prediction encodes blocks within a common frame. Intra prediction utilizes the fact that the luma components and chroma components tend to form blocks within a frame. For example, the luma components and chroma components tend to form blocks within a frame. For example, , The green regions of a part of the wood tend to be located adjacent to similar green regions. Intra prediction uses a plurality of directional prediction modes (e.g., 33 in HEVC), a planar mode, and a direct current (DC) mode. The directional mode indicates that the current block is the same as / similar to the samples of neighboring blocks in the corresponding direction. The planar mode indicates that a series of blocks (e.g., a plane) along a row / column can be interpolated based on the neighboring blocks near the end of the row. In fact, the planar mode shows a smooth transition of light / color between rows / columns by using a relatively constant gradient of changing values. The DC mode is used for boundary smoothing and indicates that the block is the same as / similar to the average value related to the samples of all neighboring blocks related to the angular direction of the directional prediction mode. Therefore, an intra prediction block can represent an image block as values of various relationship prediction modes instead of actual values. Furthermore, an inter prediction block can represent an image block as values of motion vectors instead of actual values. In either case, the prediction block may not accurately represent the image block in some cases. All differences are stored in the residual block. To further compress the file, a transform may be applied to the residual block. In step 107, various filtering techniques may be applied. In HEVC, the filter is applied by an in-loop filtering method. The block-based prediction discussed above may result in the generation of a block-noisy image at the decoder. Furthermore, the block-based prediction method encodes the block and then In step 107, various filtering techniques may be applied. In HEVC, the filter is applied by an in-loop filtering method. The block-based prediction discussed above may result in the generation of a block-noisy image at the decoder. Furthermore, the block-based prediction method encodes the block and then All differences are stored in the residual block. To further compress the file, a transform may be applied to the residual block. In step 107, various filtering techniques may be applied. In HEVC, the filter is applied by an in-loop filtering method. The block-based prediction discussed above may result in the generation of a block-noisy image at the decoder. Furthermore, the block-based prediction method encodes the block and then [[ID=1)7]] In step 107, various filtering techniques may be applied. In HEVC, the filter is applied by an in-loop filtering method. The block-based prediction discussed above may result in the generation of a block-noisy image at the decoder. Furthermore, the block-based prediction method encodes the block and then [[ID=!20]]In step 107, various filtering techniques may be applied. In HEVC, the filter is applied by an in-loop filtering method. The block-based prediction discussed above may result in the generation of a block-noisy image at the decoder. Furthermore, the block-based prediction method encodes the block and then In step 107, various filtering techniques may be applied. In HEVC, the filter is applied by an in-loop filtering method. The block-based prediction discussed above may result in the generation of a block-noisy image at the decoder. Furthermore, the block-based prediction method encodes the block and then In step 107, various filtering techniques may be applied. In HEVC, the filter is applied by an in-loop filtering method. The block-based prediction discussed above may result in the generation of a block-noisy image at the decoder. Furthermore, the block-based prediction method encodes the block and then All differences are stored in the residual block. To further compress the file, a transform may be applied to the residual block.
[0046] In step 107, various filtering techniques may be applied. In HEVC, the filter is applied by an in-loop filtering method. The block-based prediction discussed above may result in the generation of a block-noisy image at the decoder. Furthermore, the block-based prediction method encodes the block and then In HEVC, the filter is applied by an in-loop filtering method. The block-based prediction discussed above may result in the generation of a block-noisy image at the decoder. Furthermore, the block-based prediction method encodes the block and then In step 107, various filtering techniques may be applied. In HEVC, the filter is applied by an in-loop filtering method. The block-based prediction discussed above may result in the generation of a block-noisy image at the decoder. Furthermore, the block-based prediction method encodes the block and then In step 107, various filtering techniques may be applied. In HEVC, the filter is applied by an in-loop filtering method. The block-based prediction discussed above may result in the generation of a block-noisy image at the decoder. Furthermore, the block-based prediction method encodes the block and then It may be reconstructed to use the resulting block as a reference block later. In-loop fil The loop filtering method repeatedly applies a noise reduction filter, a deblocking filter, an adaptive loop filter , and a sample adaptive offset (SAO) filter to blocks / frames. These filters reduce such blocking artifacts so that the encoded file can be accurately reconstructed. Further, these filters reduce artifacts within the reconstructed reference block so that there is a lower likelihood of additional artifacts occurring in subsequent blocks that are encoded based on the reconstructed reference block.
[0047] When the video signal is segmented, compressed, and filtered, the resulting data is encoded into a bitstream at step 109. The bitstream includes the data considered above and any signaling data desirable to support the proper reconstruction of the video signal at the decoder. For example, such data may include partition data, prediction data, residual blocks, and various flags that give coding instructions to the decoder. The bitstream may be stored in memory for transmission to the decoder upon request. The bitstream may also be broadcast and / or multicast to multiple decoders. The generation of the bitstream is an iterative process. Thus, steps 101, 103, 105, 107, and 109 may be continuously and / or simultaneously performed over many frames and blocks. Shown in FIG. 1 and blocks. Shown in FIG. 1 The order is presented for clarity and ease of consideration and is not intended to limit the video coding process to a specific order.
[0048] The decoder receives the bitstream and starts the decoding process at step 111 . In particular, the decoder uses an entropy decoding method to convert the bitstream into the corresponding syntax and video data. At step 111, the decoder determines the partitions for the frames using the syntax data from the bitstream. The partitioning should match the result of the block partitioning at step 103. The entropy coding / decoding used at step 111 will be described later. The encoder makes many selections during the compression process, such as selecting a block partitioning method from several possible options based on the spatial positioning of the values in the input image. Signaling the exact same selection may use a large number of bins. As used herein, a bin is a binary value treated as a variable (e.g., a bit value that may change depending on the situation). Entropy coding enables the encoder to discard all options that are clearly not good in a particular case and leave a set of acceptable options. Next, each acceptable option is assigned a codeword. The length of the codeword is based on the number of acceptable options (e.g., one bin for two options, two bins for four options, etc.). Then, the encoder encodes the codeword for the selected option. This method ensures that the codeword potentially represents all possible options . The length is based on the number of acceptable options (e.g., one bin for two options, two bins for four options, etc.). Then, the encoder encodes the codeword for the selected option. This method ensures that the codeword potentially represents all possible options . represents all possible options In contrast to uniquely indicating a selection from a large set, it is of a size only as large as desirable to uniquely indicate a selection from a small subset of acceptable options, so the size of the codeword is reduced. The decoder then decodes the selection by determining a set of acceptable options in the same manner as the encoder. By determining the set of acceptable options, the decoder can read the codeword and determine the selection made by the encoder. In step 113, the decoder performs decoding of the block. In particular, the decoder generates a residual block using an inverse transform. The decoder then reconstructs the image block according to the partition using the residual block and the corresponding
[0049] prediction block. The prediction block may include both an intra prediction block and an inter prediction block generated by the encoder in step 105. The reconstructed image block is then positioned within the frame of the reconstructed video signal according to the partition data determined in step 111. The syntax for step 113 may also be signaled in the bitstream by the entropy coding discussed above. In step 115, filtering is performed on the frame of the reconstructed video signal in the same manner as step 107 of the encoder. For example, a noise suppression filter, a de blocking filter, an adaptive loop filter, and an SAO filter may be applied to the frame to remove blocking artifacts. When the frame is filtered, the video signal is displayed in step 117 for viewing by the end user.
[0050] In step 115, in the same manner as step 107 of the encoder, filtering is performed on the frame of the reconstructed video signal. For example, a noise suppression filter, a de blocking filter, an adaptive loop filter, and an SAO filter may be applied to the frame to remove blocking artifacts. When the frame is filtered, the video signal is displayed in step 117 for viewing by the end user. When the frame is filtered, the video signal is Can be output for play.
[0051] FIG. 2 is a schematic diagram of an exemplary coding and decoding (codec) system 200 for video coding. In particular, the codec system 200 provides functions to support the implementation of the operation method 100. The codec system 200 is generalized to depict components used in both the encoder and the decoder. The codec system 200 receives, as considered in connection with steps 101 and 103 of the operation method 100, a video signal, partitions it, resulting in a partitioned video signal 201. Next, when acting as an encoder, the codec system 200 compresses the partitioned video signal 201 into a coded bitstream, as considered in connection with steps 105, 107, and 109 of method 100. When acting as a decoder, the codec system 200 generates an output video signal from the bitstream, as considered in connection with steps 111, 113, 115, and 117 of the operation method 100. The codec system 200 includes general codec control components 211, transform / scaling and quantization components 213, intra-picture estimation components 215, intra-picture prediction components 217, motion compensation components 219, motion estimation components 221, scaling and inverse transform components 229, filter control analysis components 227, in-loop filter components 225, decoded picture buffer components 223, as well as header format and context adaptive binary arithmetic coding (CABAC) components. It includes component 231. Such components are combined as shown. In FIG. 2, the black solid lines indicate the movement of the data to be encoded / decoded, while the dashed lines indicate the movement of the control data that controls the operation of the other components. The components of the codec system 200 may all be present in the encoder. The decoder may include a subset of the components of the codec system 200. For example, the decoder may include an intra-picture prediction component 217, a motion compensation component 219, a scaling and inverse transform component 229, an in-loop filter component 225, and a decoded picture buffer component 223. These components will be described later. The segmented video signal 201 is a captured video sequence segmented into blocks of pixels by a coding tree. The coding tree uses various partitioning modes to sub-divide blocks of pixels into smaller blocks of pixels. These blocks can then be sub-divided into even smaller blocks. The blocks may be referred to as nodes of the coding tree. Larger parent nodes are divided into smaller child nodes. The number of times a node is sub-divided is called the depth of the node / coding tree. The sub-divided blocks may optionally be included in a coding unit (CU). For example, a CU may be a lower part of a CTU that includes a luma block, a red difference chroma (Cr) block, and a blue difference chroma (Cb) block along with corresponding syntax commands related to the CU. The partitioning mode depends on the partitioning mode used. a motion compensation component 219, a scaling and inverse transform component 229, an in-loop filter component 225, and a decoded picture buffer component 223. These components will be described later. The segmented video signal 201 is a captured video sequence segmented into blocks of pixels by a coding tree. The coding tree uses various partitioning modes to sub-divide blocks of pixels into smaller blocks of pixels. These blocks can then be sub-divided into even smaller blocks. The blocks may be referred to as nodes of the coding tree. Larger parent nodes are divided into smaller child nodes. The number of times a node is sub-divided is called the depth of the node / coding tree. The sub-divided blocks may optionally be included in a coding unit (CU). For example, a CU may be a lower part of a CTU that includes a luma block, a red difference chroma (Cr) block, and a blue difference chroma (Cb) block along with corresponding syntax commands related to the CU. The partitioning mode depends on the partitioning mode used.
[0052] The segmented video signal 201 is a captured video sequence segmented into blocks of pixels by a coding tree. The coding tree uses various partitioning modes to sub-divide blocks of pixels into smaller blocks of pixels. These blocks can then be sub-divided into even smaller blocks. The blocks may be referred to as nodes of the coding tree. Larger parent nodes are divided into smaller child nodes. The number of times a node is sub-divided is called the depth of the node / coding tree. The sub-divided blocks may optionally be included in a coding unit (CU). For example, a CU may be a lower part of a CTU that includes a luma block, a red difference chroma (Cr) block, and a blue difference chroma (Cb) block along with corresponding syntax commands related to the CU. The partitioning mode depends on the partitioning mode used. The segmented video signal 201 is a captured video sequence segmented into blocks of pixels by a coding tree. The coding tree uses various partitioning modes to sub-divide blocks of pixels into smaller blocks of pixels. These blocks can then be sub-divided into even smaller blocks. The blocks may be referred to as nodes of the coding tree. Larger parent nodes are divided into smaller child nodes. The number of times a node is sub-divided is called the depth of the node / coding tree. The sub-divided blocks may optionally be included in a coding unit (CU). For example, a CU may be a lower part of a CTU that includes a luma block, a red difference chroma (Cr) block, and a blue difference chroma (Cb) block along with corresponding syntax commands related to the CU. The partitioning mode depends on the partitioning mode used. The segmented video signal 201 is a captured video sequence segmented into blocks of pixels by a coding tree. The coding tree uses various partitioning modes to sub-divide blocks of pixels into smaller blocks of pixels. These blocks can then be sub-divided into even smaller blocks. The blocks may be referred to as nodes of the coding tree. Larger parent nodes are divided into smaller child nodes. The number of times a node is sub-divided is called the depth of the node / coding tree. The sub-divided blocks may optionally be included in a coding unit (CU). For example, a CU may be a lower part of a CTU that includes a luma block, a red difference chroma (Cr) block, and a blue difference chroma (Cb) block along with corresponding syntax commands related to the CU. The partitioning mode depends on the partitioning mode used. The sub-divided blocks may optionally be included in a coding unit (CU). For example, a CU may be a lower part of a CTU that includes a luma block, a red difference chroma (Cr) block, and a blue difference chroma (Cb) block along with corresponding syntax commands related to the CU. The partitioning mode depends on the partitioning mode used. and a blue difference chroma (Cb) block along with corresponding syntax commands related to the CU. The partitioning mode depends on the partitioning mode used. and a blue difference chroma (Cb) block along with corresponding syntax commands related to the CU. The partitioning mode depends on the partitioning mode used. to respectively partition a node into two, three, or four child nodes with varying shapes may include a binary tree (BT), a ternary tree (TT), and a quaternary tree (QT) used for the partitioning. The partitioned video signal 201 is transferred to a general coder control component 211, a transform / scaling and quantization component 213, an intra-picture estimation component 215, a filter control analysis component 22 7, and a motion estimation component 221 for compression.
[0053] The general coder control component 211 is configured to make decisions related to the coding of the video sequence of images into a bitstream according to the constraints of the application. For example, the general coder control component 211 manages the optimization of the bitrate / bitstream size versus the quality of reconstruction. Such decisions may be made based on the availability of storage space / bandwidth and the requirements of the image resolution . Also, the general coder control component 211 manages the use of the buffer in light of the transmission speed to mitigate the problems of buffer underrun and overrun. To manage these problems, the general coder control component 211 manages the partitioning, prediction, and filtering by its other components. For example, the general coder control component 211 may dynamically increase the complexity of compression to increase the resolution and increase the use of bandwidth, or may dynamically decrease the complexity of compression to decrease the resolution and decrease the use of bandwidth . Therefore, the general coder control component 211 controls the other components of the codec system 200 to balance the quality of reconstruction of the video signal and the concern of the bitrate . The general coder control component 211 controls the operation of the other components . Generate control data for controlling. The control data is also header formatted so as to be encoded into the bit stream for signaling the parameters for decoding in the decoder and transferred to the CABAC component 231. The segmented video signal 201 is also sent to the motion estimation component 221 and the motion
[0054] compensation component 219 for inter prediction. The frame or slice of the segmented video signal 201 may be divided into a plurality of video blocks. The motion estimation component 221 and the motion compensation component 219 perform inter prediction coding of the received video block with respect to one or more blocks in one or more reference frames to provide temporal prediction. The codec system 200 may perform a plurality of coding passes, for example, to select an appropriate coding mode for each block of video data. The motion estimation component 221 and the motion compensation component 219 may be highly integrated, but are shown separately for conceptual purposes. The motion estimation
[0055] performed by the component 221 is a process of generating a motion vector for estimating the motion for the video block. The motion vector may indicate, for example, the displacement of the coded object with respect to the prediction block. The prediction block is a block that is known to closely match the block to be coded in terms of the difference between pixels. The prediction block may also be referred to as a reference block. Such a difference between pixels may be determined by the sum of absolute differences (SAD), the sum of squared differences (SSD), or other difference measurement criteria. HEVC is a CTU, a coding tree is known to closely match the block to be coded in terms of the difference between pixels. The prediction block may also be referred to as a reference block. Such a difference between pixels may be determined by the sum of absolute differences (SAD), the sum of squared differences (SSD), or is also determined by other difference measurement criteria. HEVC is CTU, coding tree Using several coded objects including a block (CTB) and a CU . For example, a CTU can be split into CTBs, and then the CTBs can be split into CUs for inclusion in a CB. A CU can be coded as a prediction unit (PU) containing prediction data for the CU and / or a transform unit (TU) containing transformed residual data . The motion estimation component 221 generates motion vectors, PUs, and TUs by using rate-distortion analysis as part of the rate-distortion optimization process. For example, the motion estimation component 221 may determine multiple reference blocks, multiple motion vectors, etc. for the current block / frame, and select a reference block, a motion vector, etc. having the best rate-distortion characteristics . The best rate-distortion characteristics balance both the quality of video reconstruction (e.g., the amount of data loss due to compression) and the coding efficiency (e.g., the final coded size) . In some examples, the codec system 200 may calculate values at pixel positions finer than the integers of the reference pictures stored in the decoded picture buffer component 223. For example, the video codec system 200 may interpolate values at quarter-pixel positions, eighth-pixel positions, or other fractional pixel positions of the reference picture. Accordingly, the motion estimation component 221 may perform motion search related to full pixel positions and fractional pixel positions and output motion vectors with fractional pixel accuracy . The motion estimation component 221 compares the position of the PU with the position of the predicted block of the reference picture . For example, the motion estimation component 221 may determine multiple reference blocks, multiple motion vectors, etc. for the current block / frame, and select a reference block, a motion vector, etc. having the best rate-distortion characteristics . The best rate-distortion characteristics balance both the quality of video reconstruction (e.g., the amount of data loss due to compression) and the coding efficiency (e.g., the final coded size) . In some examples, the codec system 200 may calculate values at pixel positions finer than the integers of the reference pictures stored in the decoded picture buffer component 223. For example, the video codec system 200 may interpolate values at quarter-pixel positions, eighth-pixel positions, or other fractional pixel positions of the reference picture. Accordingly, the motion estimation component 221 may perform motion search related to full pixel positions and fractional pixel positions and output motion vectors with fractional pixel accuracy . The motion estimation component 221 compares the position of the PU with the position of the predicted block of the reference picture . Using several coded objects including a block (CTB) and a CU
[0056] . In some examples, the codec system 200 may calculate values at pixel positions finer than the integers of the reference pictures stored in the decoded picture buffer component 223. For example, the video codec system 200 may interpolate values at quarter-pixel positions, eighth-pixel positions, or other fractional pixel positions of the reference picture. Accordingly, the motion estimation component 221 may perform motion search related to full pixel positions and fractional pixel positions and output motion vectors with fractional pixel accuracy . For example, the video codec system 200 may interpolate values at quarter-pixel positions, eighth-pixel positions, or other fractional pixel positions of the reference picture. Accordingly, the motion estimation component 221 may perform motion search related to full pixel positions and fractional pixel positions and output motion vectors with fractional pixel accuracy . The motion estimation component 221 compares the position of the PU with the position of the predicted block of the reference picture . For example, the motion estimation component 221 may determine multiple reference blocks, multiple motion vectors, etc. for the current block / frame, and select a reference block, a motion vector, etc. having the best rate-distortion characteristics . The best rate-distortion characteristics balance both the quality of video reconstruction (e.g., the amount of data loss due to compression) and the coding efficiency (e.g., the final coded size) . In some examples, the codec system 200 may calculate values at pixel positions finer than the integers of the reference pictures stored in the decoded picture buffer component 223. For example, the video codec system 200 may interpolate values at quarter-pixel positions, eighth-pixel positions, or other fractional pixel positions of the reference picture. Accordingly, the motion estimation component 221 may perform motion search related to full pixel positions and fractional pixel positions and output motion vectors with fractional pixel accuracy . The motion estimation component 221 compares the position of the PU with the position of the predicted block of the reference picture Calculate the motion vector for the PU of the video block within the slice inter-coded thereby. The motion estimation component 221 outputs the calculated motion vector as motion data to the header format and the CABAC component 231 for encoding, and outputs the motion to the motion compensation component 219.
[0057] The motion compensation executed by the motion compensation component 219 may include retrieving or generating a prediction block based on the motion vector determined by the motion estimation component 221. Also, the motion estimation component 221 and the motion compensation component 219 may be functionally integrated in some examples. Upon receiving the motion vector for the current video block's PU, the motion compensation component 219 may find the prediction block pointed to by the motion vector. Then, by subtracting the pixel values of the prediction block from the pixel values of the currently coded video block, a residual video block is formed by forming a pixel difference value. Generally, the motion estimation component 221 performs motion estimation related to the luma component, and the motion compensation component 219 uses the motion vector calculated based on the luma component for both the chroma component and the luma component. The prediction block and the residual block are transferred to the transform / scaling and quantization component 213.
[0058] The segmented video signal 201 is also transmitted to the intra-picture estimation component 215 and the intra-picture prediction component 217. Similar to the motion estimation component 221 and the motion compensation component 219, the intra-picture estimation component 215 and the intra-picture prediction component 217 , may be highly integrated, but are shown separately for conceptual purposes. Intra-picture Estimation component 215 and intra-picture prediction component 217 perform intra-prediction on the current block instead of inter-prediction executed by motion estimation component 221 and motion compensation component 219 between frames as described above. In particular , the intra-picture estimation component 215 determines the intra-prediction mode to be used for encoding the current block. In some examples, the intra-picture estimation component 215 selects an appropriate intra-prediction mode for encoding the current block from a plurality of tested intra-prediction modes. The selected intra-prediction mode is then transferred to the header format and CABAC component 231 for encoding. , for example, the intra-picture estimation component 215 calculates rate-distortion values using rate-distortion analysis for various tested intra-prediction modes and selects the intra-prediction mode with the best rate-distortion characteristics among the tested modes. Rate-distortion analysis generally determines the amount of distortion (or error) between the encoded block and the original unencoded block used to generate the encoded block, and the bit rate (e.g., number of bits) used to generate the encoded block. The intra-picture estimation component 215 calculates a ratio from the distortion and rate for various encoded blocks to determine which intra-prediction mode exhibits the best rate-distortion value for the block. Further, the intra-picture estimation component 215 is the rate-distortion best
[0059] For example, the intra-picture estimation component 215 calculates rate-distortion values using rate-distortion analysis for various tested intra-prediction modes and selects the intra-prediction mode with the best rate-distortion characteristics among the tested modes. Rate-distortion analysis generally determines the amount of distortion (or error) between the encoded block and the original unencoded block used to generate the encoded block, and the bit rate (e.g., number of bits) used to generate the encoded block. The intra-picture estimation component 215 calculates a ratio from the distortion and rate for various encoded blocks to determine which intra-prediction mode exhibits the best rate-distortion value for the block. Rate-distortion analysis generally determines the amount of distortion (or error) between the encoded block and the original unencoded block used to generate the encoded block, and the bit rate (e.g., number of bits) used to generate the encoded block. The intra-picture estimation component 215 calculates a ratio from the distortion and rate for various encoded blocks to determine which intra-prediction mode exhibits the best rate-distortion value for the block. Rate-distortion analysis generally determines the amount of distortion (or error) between the encoded block and the original unencoded block used to generate the encoded block, and the bit rate (e.g., number of bits) used to generate the encoded block. The intra-picture estimation component 215 calculates a ratio from the distortion and rate for various encoded blocks to determine which intra-prediction mode exhibits the best rate-distortion value for the block. Rate-distortion analysis generally determines the amount of distortion (or error) between the encoded block and the original unencoded block used to generate the encoded block, and the bit rate (e.g., number of bits) used to generate the encoded block. The intra-picture estimation component 215 calculates a ratio from the distortion and rate for various encoded blocks to determine which intra-prediction mode exhibits the best rate-distortion value for the block. Further, the intra-picture estimation component 215 is the rate-distortion best The depth blocks of the depth map may be configured to be coded using a depth modeling mode (DMM) based on rate-distortion optimization (RDO). It may be configured to code.
[0060] When implemented in an encoder, the intra-picture prediction component 217 generates a residual block from a prediction block based on a selected intra prediction mode determined by the intra-picture estimation component 215, or when implemented in a decoder, may read the residual block from a bitstream. The residual block contains the difference in values between the prediction block and the original block, represented as a matrix. The residual block is then transferred to the transform-scaling and quantization component 213. The intra-picture estimation component 215 and the intra-picture prediction component 217 may operate on both the luma component and the chroma component. Based on the selected intra prediction mode determined by the intra-picture estimation component 215, generate a residual block from the prediction block, or when implemented in the decoder, read the residual block from the bitstream. When implemented in the decoder, it may read the residual block from the bitstream. The residual block contains the difference in values between the prediction block and the original block, represented as a matrix. The residual block is then transferred to the transform-scaling and quantization component 213. The intra-picture estimation component 215 and the intra-picture prediction component 217 may operate on both the luma component and the chroma component. It may read the residual block from the bitstream. The residual block contains the difference in values between the prediction block and the original block, represented as a matrix. The residual block is then transferred to the transform-scaling and quantization component 213. The intra-picture estimation component 215 and the intra-picture prediction component 217 may operate on both the luma component and the chroma component. The difference between the prediction block and the original block, and then the residual block is transferred to the transform-scaling and quantization component 213. The intra-picture estimation component 215 and the intra-picture prediction component 217 may operate on both the luma component and the chroma component. The intra-picture estimation component 215 and the intra-picture prediction component 217 may operate on both the luma component and the chroma component. It may operate on both the luma component and the chroma component.
[0061] The transform-scaling and quantization component 213 is configured to further compress the residual block. The transform-scaling and quantization component 213 applies a transform such as a discrete cosine transform (DCT), a discrete sine transform (DST), or a transform of a similar concept to the residual block to generate a video block containing residual transform coefficient values. A wavelet transform, an integer transform, a subband transform, or other types of transforms may also be used. The transform may transform the residual information from the pixel value domain to a transform domain such as the frequency domain. The transform-scaling and quantization component 213 is further configured to scale the residual information, for example, based on frequency. Such scaling includes applying a scale factor to the residual information so that different frequency information is quantized at different granularities, which is for the reconstructed The transform-scaling and quantization component 213 is configured to further compress the residual block. The transform-scaling and quantization component 213 applies a transform such as a discrete cosine transform (DCT), a discrete sine transform (DST), or a transform of a similar concept to the residual block to generate a video block containing residual transform coefficient values. A wavelet transform, an integer transform, a subband transform, or other types of transforms may also be used. The transform may transform the residual information from the pixel value domain to a transform domain such as the frequency domain. The transform-scaling and quantization component 213 is further configured to scale the residual information, for example, based on frequency. Such scaling includes applying a scale factor to the residual information so that different frequency information is quantized at different granularities, which is for the reconstructed The transform-scaling and quantization component 213 applies a transform such as a discrete cosine transform (DCT), a discrete sine transform (DST), or a transform of a similar concept to the residual block to generate a video block containing residual transform coefficient values. A wavelet transform, an integer transform, a subband transform, or other types of transforms may also be used. The transform may transform the residual information from the pixel value domain to a transform domain such as the frequency domain. The transform-scaling and quantization component 213 is further configured to scale the residual information, for example, based on frequency. Such scaling includes applying a scale factor to the residual information so that different frequency information is quantized at different granularities, which is for the reconstructed The transform-scaling and quantization component 213 applies a transform such as a discrete cosine transform (DCT), a discrete sine transform (DST), or a transform of a similar concept to the residual block to generate a video block containing residual transform coefficient values. A wavelet transform, an integer transform, a subband transform, or other types of transforms may also be used. The transform may transform the residual information from the pixel value domain to a transform domain such as the frequency domain. The transform-scaling and quantization component 213 is further configured to scale the residual information, for example, based on frequency. Such scaling includes applying a scale factor to the residual information so that different frequency information is quantized at different granularities, which is for the reconstructed The transform-scaling and quantization component 213 applies a transform such as a discrete cosine transform (DCT), a discrete sine transform (DST), or a transform of a similar concept to the residual block to generate a video block containing residual transform coefficient values. A wavelet transform, an integer transform, a subband transform, or other types of transforms may also be used. The transform may transform the residual information from the pixel value domain to a transform domain such as the frequency domain. The transform-scaling and quantization component 213 is further configured to scale the residual information, for example, based on frequency. Such scaling includes applying a scale factor to the residual information so that different frequency information is quantized at different granularities, which is for the reconstructed The transform-scaling and quantization component 213 applies a transform such as a discrete cosine transform (DCT), a discrete sine transform (DST), or a transform of a similar concept to the residual block to generate a video block containing residual transform coefficient values. A wavelet transform, an integer transform, a subband transform, or other types of transforms may also be used. The transform may transform the residual information from the pixel value domain to a transform domain such as the frequency domain. The transform-scaling and quantization component 213 is further configured to scale the residual information, for example, based on frequency. Such scaling includes applying a scale factor to the residual information so that different frequency information is quantized at different granularities, which is for the reconstructed The transform-scaling and quantization component 213 is further configured to scale the residual information, for example, based on frequency. Such scaling includes applying a scale factor to the residual information so that different frequency information is quantized at different granularities, which is for the reconstructed The transform-scaling and quantization component will scale the residual information, for example, based on frequency. Such scaling includes applying a scale factor to the residual information so that different frequency information is quantized at different granularities, which is for the reconstructed The transform-scaling and quantization component 213 is further configured to scale the residual information, for example, based on frequency. Such scaling includes applying a scale factor to the residual information so that different frequency information is quantized at different granularities, which is for the reconstructed Transformation, scaling and quantization structures may affect the final visual quality of the video. The construction element 213 is further configured to quantize the transform coefficients to further reduce the bit rate. The quantization process reduces the bit depth associated with some or all of the coefficients. The degree of quantization may be modified by adjusting the quantization parameter. In some examples, the transform, scale and quantize component 213 then performs quantization. A scan of the matrix containing the quantized transform coefficients may be performed. The quantized transform coefficients are then bit-wise converted. Header format and CABAC component 231 for encoding into the data stream. do.
[0062] The scaling and inverse transform component 229 performs transformation and scaling to support motion estimation. Apply the inverse operation of the scaling and quantization component 213. Element 229 may be, for example, a reference block that may be a prediction block for another current block. Inverse scanning is performed to reconstruct the residual block in the pixel domain for later use as a block. Apply scaling, inverse transform, and / or inverse quantization. Or the motion compensation component 219 may use the motion estimation of a subsequent block / frame. Calculate the reference block by adding the residual block back to the predicted block corresponding to To reduce artifacts introduced during scaling, quantization, and transformation, To do this, a filter is applied to the reconstructed reference block. Artifacts result in inaccurate predictions when subsequent blocks are predicted (and and further artifacts).
[0063] The filter control analysis component 227 and the in-loop filter component 225 apply a filter to the residual block and / or the reconstructed image block. For example, the transformed residual block from the scaling and inverse transform component 229 may be combined with the corresponding prediction block from the intra-picture prediction component 217 and / or the motion compensation component 219 to reconstruct the original image block. Then, a filter may be applied to the reconstructed image block. In some examples, the filter may instead be applied to the residual block. Similar to the other components in FIG. 2, the filter control analysis component 227 and the in-loop filter component 225 are highly integrated and may be implemented together, but are shown separately for conceptual purposes. The filter applied to the reconstructed reference block is applied to a specific spatial region and includes a plurality of parameters for adjusting how such a filter is applied. The filter control analysis component 227 analyzes the reconstructed reference block to determine where such a filter should be applied and sets the corresponding parameters. Such data is transferred as filter control data to the header format and the CABAC component 231 for encoding. The in-loop filter component 225 applies such a filter based on the filter control data. The filter may include a deblocking filter, a noise suppression filter, an SAO filter, and an adaptive loop filter. Such a filter may be applied in the spatial / pixel domain (e.g., to the reconstructed pixel block) or in the frequency domain, depending on the example. and / or the reconstructed image block. For example, the transformed residual block from the scaling and inverse transform component 229 may be combined with the corresponding prediction block from the intra-picture prediction component 217 and / or the motion compensation component 219 to reconstruct the original image block. to reconstruct the original image block. Then, a filter may be applied to the reconstructed image block. In some examples, the filter may instead be applied to the residual block. Similar to the other components in FIG. 2, the filter control analysis component 227 and the in-loop filter component 225 are highly integrated and may be implemented together, but are shown separately for conceptual purposes. The filter applied to the reconstructed reference block is applied to a specific spatial region and includes a plurality of parameters for adjusting how such a filter is applied. The filter control analysis component 227 analyzes the reconstructed reference block to determine where such a filter should be applied and sets the corresponding parameters. Such data is transferred as filter control data to the header format and the CABAC component 231 for encoding. The in-loop filter component 225 applies such a filter based on the filter control data. The filter may include a deblocking filter, a noise suppression filter, an SAO filter, and an adaptive loop filter. Such a filter may be applied in the spatial / pixel domain (e.g., to the reconstructed pixel block) or in the frequency domain, depending on the example. to reconstruct the original image block. Then, a filter may be applied to the reconstructed image block. In some examples, the filter may instead be applied to the residual block. Similar to the other components in FIG. 2, the filter control analysis component 227 and the in-loop filter component 225 are highly integrated and may be implemented together, but are shown separately for conceptual purposes. The filter applied to the reconstructed reference block is applied to a specific spatial region and includes a plurality of parameters for adjusting how such a filter is applied. The filter control analysis component 227 analyzes the reconstructed reference block to determine where such a filter should be applied and sets the corresponding parameters. Such data is transferred as filter control data to the header format and the CABAC component 231 for encoding. The in-loop filter component 225 applies such a filter based on the filter control data. The filter may include a deblocking filter, a noise suppression filter, an SAO filter, and an adaptive loop filter. Such a filter may be applied in the spatial / pixel domain (e.g., to the reconstructed pixel block) or in the frequency domain, depending on the example. to the reconstructed image block. In some examples, the filter may instead be applied to the residual block. Similar to the other components in FIG. 2, the filter control analysis component 227 and the in-loop filter component 225 are highly integrated and may be implemented together, but are shown separately for conceptual purposes. The filter applied to the reconstructed reference block is applied to a specific spatial region and includes a plurality of parameters for adjusting how such a filter is applied. The filter control analysis component 227 analyzes the reconstructed reference block to determine where such a filter should be applied and sets the corresponding parameters. Such data is transferred as filter control data to the header format and the CABAC component 231 for encoding. The in-loop filter component 225 applies such a filter based on the filter control data. The filter may include a deblocking filter, a noise suppression filter, an SAO filter, and an adaptive loop filter. Such a filter may be applied in the spatial / pixel domain (e.g., to the reconstructed pixel block) or in the frequency domain, depending on the example. to the residual block. Similar to the other components in FIG. 2, the filter control analysis component 227 and the in-loop filter component 225 are highly integrated and may be implemented together, but are shown separately for conceptual purposes. The filter applied to the reconstructed reference block is applied to a specific spatial region and includes a plurality of parameters for adjusting how such a filter is applied. The filter control analysis component 227 analyzes the reconstructed reference block to determine where such a filter should be applied and sets the corresponding parameters. Such data is transferred as filter control data to the header format and the CABAC component 231 for encoding. The in-loop filter component 225 applies such a filter based on the filter control data. The filter may include a deblocking filter, a noise suppression filter, an SAO filter, and an adaptive loop filter. Such a filter may be applied in the spatial / pixel domain (e.g., to the reconstructed pixel block) or in the frequency domain, depending on the example. and the in-loop filter component 225 are highly integrated and may be implemented together, but are shown separately for conceptual purposes. The filter applied to the reconstructed reference block is applied to a specific spatial region and includes a plurality of parameters for adjusting how such a filter is applied. The filter control analysis component 227 analyzes the reconstructed reference block to determine where such a filter should be applied and sets the corresponding parameters. Such data is transferred as filter control data to the header format and the CABAC component 231 for encoding. The in-loop filter component 225 applies such a filter based on the filter control data. The filter may include a deblocking filter, a noise suppression filter, an SAO filter, and an adaptive loop filter. Such a filter may be applied in the spatial / pixel domain (e.g., to the reconstructed pixel block) or in the frequency domain, depending on the example. The filter applied to the reconstructed reference block is applied to a specific spatial region and includes a plurality of parameters for adjusting how such a filter is applied. The filter control analysis component 227 analyzes the reconstructed reference block to determine where such a filter should be applied and sets the corresponding parameters. Such data is transferred as filter control data to the header format and the CABAC component 231 for encoding. The in-loop filter component 225 applies such a filter based on the filter control data. The filter may include a deblocking filter, a noise suppression filter, an SAO filter, and an adaptive loop filter. Such a filter may be applied in the spatial / pixel domain (e.g., to the reconstructed pixel block) or in the frequency domain, depending on the example. is applied to a specific spatial region and includes a plurality of parameters for adjusting how such a filter is applied. The filter control analysis component 227 analyzes the reconstructed reference block to determine where such a filter should be applied and sets the corresponding parameters. Such data is transferred as filter control data to the header format and the CABAC component 231 for encoding. The in-loop filter component 225 applies such a filter based on the filter control data. The filter may include a deblocking filter, a noise suppression filter, an SAO filter, and an adaptive loop filter. Such a filter may be applied in the spatial / pixel domain (e.g., to the reconstructed pixel block) or in the frequency domain, depending on the example. is applied to a specific spatial region and includes a plurality of parameters for adjusting how such a filter is applied. The filter control analysis component 227 analyzes the reconstructed reference block to determine where such a filter should be applied and sets the corresponding parameters. Such data is transferred as filter control data to the header format and the CABAC component 231 for encoding. The in-loop filter component 225 applies such a filter based on the filter control data. The filter may include a deblocking filter, a noise suppression filter, an SAO filter, and an adaptive loop filter. Such a filter may be applied in the spatial / pixel domain (e.g., to the reconstructed pixel block) or in the frequency domain, depending on the example. The filter control analysis component 227 analyzes the reconstructed reference block to determine where such a filter should be applied and sets the corresponding parameters. Such data is transferred as filter control data to the header format and the CABAC component 231 for encoding. The in-loop filter component 225 applies such a filter based on the filter control data. The filter may include a deblocking filter, a noise suppression filter, an SAO filter, and an adaptive loop filter. Such a filter may be applied in the spatial / pixel domain (e.g., to the reconstructed pixel block) or in the frequency domain, depending on the example. The filter control analysis component 227 analyzes the reconstructed reference block to determine where such a filter should be applied and sets the corresponding parameters. Such data is transferred as filter control data to the header format and the CABAC component 231 for encoding. The in-loop filter component 225 applies such a filter based on the filter control data. The filter may include a deblocking filter, a noise suppression filter, an SAO filter, and an adaptive loop filter. Such a filter may be applied in the spatial / pixel domain (e.g., to the reconstructed pixel block) or in the frequency domain, depending on the example. The filter control analysis component 227 analyzes the reconstructed reference block to determine where such a filter should be applied and sets the corresponding parameters. Such data is transferred as filter control data to the header format and the CABAC component 231 for encoding. The in-loop filter component 225 applies such a filter based on the filter control data. The filter may include a deblocking filter, a noise suppression filter, an SAO filter, and an adaptive loop filter. Such a filter may be applied in the spatial / pixel domain (e.g., to the reconstructed pixel block) or in the frequency domain, depending on the example. The in-loop filter component 225 applies such a filter based on the filter control data. The filter may include a deblocking filter, a noise suppression filter, an SAO filter, and an adaptive loop filter. Such a filter may be applied in the spatial / pixel domain (e.g., to the reconstructed pixel block) or in the frequency domain, depending on the example. The filter may include a deblocking filter, a noise suppression filter, an SAO filter, and an adaptive loop filter. Such a filter may be applied in the spatial / pixel domain (e.g., to the reconstructed pixel block) or in the frequency domain, depending on the example. The filter may include a deblocking filter, a noise suppression filter, an SAO filter, and an adaptive loop filter. Such a filter may be applied in the spatial / pixel domain (e.g., to the reconstructed pixel block) or in the frequency domain, depending on the example. Such a filter may be applied in the spatial / pixel domain (e.g., to the reconstructed pixel block) or in the frequency domain, depending on the example. Such a filter may be applied in the spatial / pixel domain (e.g., to the reconstructed pixel block) or in the frequency domain, depending on the example.
[0064] When operating as an encoder, the filtered and reconstructed image blocks , residual blocks, and / or prediction blocks are stored in the decoded picture buffer component 223 for later use in motion estimation as discussed above. When operating as a decoder, the decoded picture buffer component 223 stores the reconstructed and filtered blocks and transfers them to the display as part of the output video signal. The decoded picture buffer component 223 may be any memory device capable of storing prediction blocks, residual blocks, and / or reconstructed image blocks. When operating as an encoder, the filtered and reconstructed image blocks , residual blocks, and / or prediction blocks are stored in the decoded picture buffer component 223 for later use in motion estimation as discussed above. When operating as a decoder, the decoded picture buffer component 223 stores the reconstructed and filtered blocks and transfers them to the display as part of the output video signal. The decoded picture buffer component 223 may be any memory device capable of storing prediction blocks, residual blocks, and / or reconstructed image blocks. When operating as an encoder, the filtered and reconstructed image blocks , residual blocks, and / or prediction blocks are stored in the decoded picture buffer component 223 for later use in motion estimation as discussed above. When operating as a decoder, the decoded picture buffer component 223 stores the reconstructed and filtered blocks and transfers them to the display as part of the output video signal. The decoded picture buffer component 223 may be any memory device capable of storing prediction blocks, residual blocks, and / or reconstructed image blocks. When operating as an encoder, the filtered and reconstructed image blocks
[0065] The header format and CABAC component 231 receive data from various components of the codec system 200 and encode such data into a coded bitstream for transmission to the decoder. In particular, the header format and CABAC component 231 generate various headers for encoding control data such as general control data and filter control data. Further, prediction data including intra prediction and motion data as well as residual data in the form of quantized transform coefficients are all encoded into the bitstream. The final bitstream contains all the information desired by the decoder to reconstruct the original segmented video signal 201. Such information includes an index table (also called a codeword mapping table) for the intra prediction mode, the definition of the coding context for various blocks, an indication of the most likely intra prediction mode, The header format and CABAC component 231 receive data from various components of the codec system 200 and encode such data into a coded bitstream for transmission to the decoder. In particular, the header format and CABAC component 231 generate various headers for encoding control data such as general control data and filter control data. Further, prediction data including intra prediction and motion data as well as residual data in the form of quantized transform coefficients are all encoded into the bitstream. The final bitstream contains all the information desired by the decoder to reconstruct the original segmented video signal 201. Such information includes an index table (also called a codeword mapping table) for the intra prediction mode, the definition of the coding context for various blocks, an indication of the most likely intra prediction mode, The header format and CABAC component 231 receive data from various components of the codec system 200 and encode such data into a coded bitstream for transmission to the decoder. In particular, the header format and CABAC component 231 generate various headers for encoding control data such as general control data and filter control data. Further, prediction data including intra prediction and motion data as well as residual data in the form of quantized transform coefficients are all encoded into the bitstream. The final bitstream contains all the information desired by the decoder to reconstruct the original segmented video signal 201. Such information includes an index table (also called a codeword mapping table) for the intra prediction mode, the definition of the coding context for various blocks, an indication of the most likely intra prediction mode, The header format and CABAC component 231 receive data from various components of the codec system 200 and encode such data into a coded bitstream for transmission to the decoder. In particular, the header format and CABAC component 231 generate various headers for encoding control data such as general control data and filter control data. Further, prediction data including intra prediction and motion data as well as residual data in the form of quantized transform coefficients are all encoded into the bitstream. The final bitstream contains all the information desired by the decoder to reconstruct the original segmented video signal 201. Such information includes an index table (also called a codeword mapping table) for the intra prediction mode, the definition of the coding context for various blocks, an indication of the most likely intra prediction mode, The header format and CABAC component 231 receive data from various components of the codec system 200 and encode such data into a coded bitstream for transmission to the decoder. In particular, the header format and CABAC component 231 generate various headers for encoding control data such as general control data and filter control data. Further, prediction data including intra prediction and motion data as well as residual data in the form of quantized transform coefficients are all encoded into the bitstream. The final bitstream contains all the information desired by the decoder to reconstruct the original segmented video signal 201. Such information includes an index table (also called a codeword mapping table) for the intra prediction mode, the definition of the coding context for various blocks, an indication of the most likely intra prediction mode, The header format and CABAC component 231 receive data from various components of the codec system 200 and encode such data into a coded bitstream for transmission to the decoder. In particular, the header format and CABAC component 231 generate various headers for encoding control data such as general control data and filter control data. Further, prediction data including intra prediction and motion data as well as residual data in the form of quantized transform coefficients are all encoded into the bitstream. The final bitstream contains all the information desired by the decoder to reconstruct the original segmented video signal 201. Such information includes an index table (also called a codeword mapping table) for the intra prediction mode, the definition of the coding context for various blocks, an indication of the most likely intra prediction mode, The header format and CABAC component 231 receive data from various components of the codec system 200 and encode such data into a coded bitstream for transmission to the decoder. In particular, the header format and CABAC component 231 generate various headers for encoding control data such as general control data and filter control data. Further, prediction data including intra prediction and motion data as well as residual data in the form of quantized transform coefficients are all encoded into the bitstream. The final bitstream contains all the information desired by the decoder to reconstruct the original segmented video signal 201. Such information includes an index table (also called a codeword mapping table) for the intra prediction mode, the definition of the coding context for various blocks, an indication of the most likely intra prediction mode, The header format and CABAC component 231 receive data from various components of the codec system 200 and encode such data into a coded bitstream for transmission to the decoder. In particular, the header format and CABAC component 231 generate various headers for encoding control data such as general control data and filter control data. Further, prediction data including intra prediction and motion data as well as residual data in the form of quantized transform coefficients are all encoded into the bitstream. The final bitstream contains all the information desired by the decoder to reconstruct the original segmented video signal 201. Such information includes an index table (also called a codeword mapping table) for the intra prediction mode, the definition of the coding context for various blocks, an indication of the most likely intra prediction mode, The header format and CABAC component 231 receive data from various components of the codec system 200 and encode such data into a coded bitstream for transmission to the decoder. In particular, the header format and CABAC component 231 generate various headers for encoding control data such as general control data and filter control data. Further, prediction data including intra prediction and motion data as well as residual data in the form of quantized transform coefficients are all encoded into the bitstream. The final bitstream contains all the information desired by the decoder to reconstruct the original segmented video signal 201. Such information includes an index table (also called a codeword mapping table) for the intra prediction mode, the definition of the coding context for various blocks, an indication of the most likely intra prediction mode, The header format and CABAC component 231 receive data from various components of the codec system 200 and encode such data into a coded bitstream for transmission to the decoder. In particular, the header format and CABAC component 231 generate various headers for encoding control data such as general control data and filter control data. Further, prediction data including intra prediction and motion data as well as residual data in the form of quantized transform coefficients are all encoded into the bitstream. The final bitstream contains all the information desired by the decoder to reconstruct the original segmented video signal 201. Such information includes an index table (also called a codeword mapping table) for the intra prediction mode, the definition of the coding context for various blocks, an indication of the most likely intra prediction mode, -tion, indication of partition information, etc. may also be included. Such data may be encoded by using entropy coding. For example, the information may be encoded by context adaptive variable length coding (CAVLC), CABAC, syntax-based context-adaptive binary arithmetic coding (SBAC), probability interval partitioning entropy (PIPE) coding, or another entropy coding technique. After entropy coding, the encoded bitstream may be transmitted to another device (e.g., a video decoder), or may be archived for later transmission or retrieval.
[0066] FIG. 3 is a block diagram showing an exemplary video encoder 300. The video encoder 300 may implement the encoding function of the codec system 200 and / or be used to implement steps 101, 103, 105, 107, and / or 109 of the operation method 100. The encoder 300 partitions the input video signal to produce a partitioned video signal 301 that is substantially the same as the partitioned video signal 201. The partitioned video signal 301 is then compressed by the
[0067] components of the encoder 300 and encoded into a bitstream. It is transferred to the intra-picture prediction component 317. The intra-picture prediction component 317 may be substantially the same as the intra-picture estimation component It may be substantially the same as the components 215 and the intra-picture prediction component 217. The segmented The video signal 301 is also transferred to the motion compensation component 321 for inter-prediction based on the reference blocks in the decoded picture buffer component 323. The motion compensation component 321 may be substantially the same as the motion estimation component 221 and the motion compensation component 219. The prediction blocks and residual blocks from the intra picture prediction component 317 and the motion compensation component 321 are transferred to the transform and quantization component 313 for the transformation and quantization of the residual blocks The transform and quantization component 313 may be substantially the same as the transform, scaling, and quantization component 213 The transformed and quantized residual blocks and the corresponding prediction blocks are transferred to the entropy coding component 331 (together with the relevant control data) for coding into the bitstream The entropy coding component 331 may be substantially the same as the header format and the CABAC component 231 The transformed and quantized residual blocks and the corresponding prediction blocks are transferred to the inverse transform and quantization component 329 from the transform and quantization component 313 for reconstruction into reference blocks for use by the motion compensation component 321 The inverse transform and quantization component 329 may be substantially the same as the scaling and inverse transform component 229 The in-loop filter of the in-loop filter component 325, depending on the example, for the residual blocks and / It may be substantially the same as the header format and the CABAC component 231 .
[0068] Also, the transformed and quantized residual blocks and / or the corresponding prediction blocks are Transferred to the inverse transform and quantization component 329 from the transform and quantization component 313 for reconstruction into reference blocks for use by the motion compensation component 321 The inverse transform and quantization component 329 may be substantially the same as the scaling and inverse transform component 229 The in-loop filter of the in-loop filter component 325, depending on the example, for the residual blocks and / [[ID=3�]]The in-loop filter of the in-loop filter component 325, depending on the example, for the residual blocks and / Or is also applied to the reconstructed reference block. The in-loop filter component 325 is substantially the same as the filter control analysis component 227 and the in-loop filter component 225 and may be. The in-loop filter component 325 may include multiple filters as considered in relation to the in-loop filter component 225 . Then, the filtered block is stored in the decoded picture buffer component 323 for use as a reference block by the motion compensation component 321. The decoded picture buffer component 323 may be substantially the same as the decoded picture buffer component 223 .
[0069] FIG. 4 is a block diagram showing an exemplary video decoder 400. The video decoder 400 implements the decoding function of the codec system 200 and / or may be used to implement steps 1 11, 113, 115, and / or 117 of the operation method 100. The decoder 400 receives, for example, a bitstream from the encoder 300 and generates an output video signal reconstructed based on the bitstream for display to an end user .
[0070] The bitstream is received by the entropy decoding component 433. The entropy decoding component 433 is configured to implement an entropy decoding method such as CAVLC, CABAC, SBAC, PIPE coding, or other entropy coding techniques. For example, the entropy decoding component 433 uses the header information to provide a context for interpreting further data encoded as codewords in the bitstream . For example, the entropy decoding component 433 This may be the case. The decoded information includes any desired information for decoding a video signal, such as general control data, filter control data, section information, motion data, prediction data, and quantized transform coefficients from residual blocks. The quantized transform coefficients are transferred to the inverse transform and quantization component 429 for reconstruction of the residual blocks. The inverse transform and quantization component 429 may be the same as the inverse transform and quantization component 329. The reconstructed residual blocks and / or prediction blocks are transferred to the intra-picture prediction component 417 for reconstruction into image blocks based on the intra prediction operation. The intra-picture prediction component 417 may be substantially the same as the intra-picture estimation component 215 and the intra-picture prediction component 217. In particular, the intra-picture prediction component 417 identifies reference blocks within the frame using the prediction mode and applies the residual blocks to the result to reconstruct the intra-predicted image blocks. The reconstructed intra-predicted image blocks and / or residual blocks and the corresponding inter-prediction data are transferred to the decoded picture buffer component 423 via the in-loop filter component 425, which may be substantially the same as the decoded picture buffer component 223 and the in-loop filter component 225, respectively. The in-loop filter component 425 filters the reconstructed image blocks, residual blocks, and / or prediction blocks, and such information is stored in the decoded picture buffer component 423. The reconstructed image blocks from the decoded picture buffer component 423 are motion compensated for inter prediction. including any desired information for decoding a video signal, such as general control data, filter control data, section information, motion data, prediction data, and quantized transform coefficients from residual blocks. The quantized transform coefficients are transferred to the inverse transform and quantization component 429 for reconstruction of the residual blocks. The inverse transform and quantization component 429 may be the same as the inverse transform and quantization component 329. to the inverse transform and quantization component 429 for reconstruction of the residual blocks. The inverse transform and quantization component 429 may be the same as the inverse transform and quantization component 329. The inverse transform and quantization component 429 may be the same as the inverse transform and quantization component 329.
[0071] The reconstructed residual blocks and / or prediction blocks are transferred to the intra-picture prediction component 417 for reconstruction into image blocks based on the intra prediction operation. The intra-picture prediction component 417 may be substantially the same as the intra-picture estimation component 215 and the intra-picture prediction component 217. In particular, the intra-picture prediction component 417 identifies reference blocks within the frame using the prediction mode and applies the residual blocks to the result to reconstruct the intra-predicted image blocks. The reconstructed intra-predicted image blocks and / or residual blocks and the corresponding inter-prediction data are transferred to the decoded picture buffer component 423 via the in-loop filter component 425, which may be substantially the same as the decoded picture buffer component 223 and the in-loop filter component 225, respectively. In particular, the intra-picture prediction component 417 identifies reference blocks within the frame using the prediction mode and applies the residual blocks to the result to reconstruct the intra-predicted image blocks. The reconstructed intra-predicted image blocks and / or residual blocks and the corresponding inter-prediction data are transferred to the decoded picture buffer component 423 via the in-loop filter component 425, which may be substantially the same as the decoded picture buffer component 223 and the in-loop filter component 225, respectively. The reconstructed intra-predicted image blocks and / or residual blocks and the corresponding inter-prediction data are transferred to the decoded picture buffer component 423 via the in-loop filter component 425, which may be substantially the same as the decoded picture buffer component 223 and the in-loop filter component 225, respectively. The reconstructed intra-predicted image blocks and / or residual blocks and the corresponding inter-prediction data are transferred to the decoded picture buffer component 423 via the in-loop filter component 425, which may be substantially the same as the decoded picture buffer component 223 and the in-loop filter component 225, respectively. The reconstructed intra-predicted image blocks and / or residual blocks and the corresponding inter-prediction data are transferred to the decoded picture buffer component 423 via the in-loop filter component 425, which may be substantially the same as the decoded picture buffer component 223 and the in-loop filter component 225, respectively. The in-loop filter component 425 filters the reconstructed image blocks, residual blocks, and / or prediction blocks, and such information is stored in the decoded picture buffer component 423. The in-loop filter component 425 filters the reconstructed image blocks, residual blocks, and / or prediction blocks, and such information is stored in the decoded picture buffer component 423. The in-loop filter component 425 filters the reconstructed image blocks, residual blocks, and / or prediction blocks, and such information is stored in the decoded picture buffer component 423. The reconstructed image blocks from the decoded picture buffer component 423 are motion compensated for inter prediction. It is transferred to the component 421. The motion compensation component 421 may be substantially the same as the motion estimation component 221 and / or the motion compensation component 219. In particular, the motion compensation component 421 uses the motion vector from the reference block to generate a prediction block and applies the residual block to the result to reconstruct the image block. The resulting reconstructed block is also transferred to the decoded picture buffer component 423 via the in-loop filter component 425. The decoded picture buffer component 423 continues to store the further reconstructed image blocks, and those reconstructed image blocks can be reconstructed into frames by the partition information. Also, such frames may be arranged in a sequence. The sequence is output to the display as the reconstructed output video signal. Figure 5 is a schematic diagram showing an exemplary CVS 500. For example, the CVS 500 may be encoded by an encoder such as the codec system 200 and / or the encoder 300 according to the method 100. Furthermore, the CVS 500 may be decoded by a decoder such as the codec system 200 and / or the decoder 400. The CVS 500 includes pictures coded in the decoding order 508. The decoding order 508 is the order in which the pictures are positioned within the bitstream. Next, the pictures of the CVS 500 are output in the presentation order 510. The presentation order 510 is the order in which the pictures should be displayed by the decoder to properly display the resulting video.
[0072] For example, the pictures of the CVS 500 may generally be positioned in the presentation order 510. However, the pictures of the CVS 500 may also be decoded by a decoder such as the codec system 200 and / or the decoder 400. The CVS 500 includes pictures coded in the decoding order 508. The decoding order 508 is the order in which the pictures are positioned within the bitstream. Next, the pictures of the CVS 500 are output in the presentation order 510. The presentation order 510 is the order in which the pictures should be displayed by the decoder to properly display the resulting video. For example, the pictures of the CVS 500 may generally be positioned in the presentation order 510. However, Moreover, specific pictures may be moved to different positions in order to increase coding efficiency, for example, by arranging similar pictures closer to support inter prediction. Moving such pictures in this way results in a decoding order 508. In the example shown, pictures are indexed in decoding order 508 from 0 to 4. In presentation order 510, the pictures of index 2 and index 3 are moved before the picture of index 0. Moving such pictures in this way results in a decoding order 508. In the example shown, pictures are indexed in decoding order 508 from 0 to 4. In presentation order 510, the pictures of index 2 and index 3 are moved before the picture of index 0.
[0073] CVS 500 includes an IRAP picture 502. The IRAP picture 502 is a picture coded by intra prediction that serves as a random access point for CVS 500. In particular, the blocks of the IRAP picture 502 are coded by reference to other blocks of the IRAP picture 502. Since the IRAP picture 502 is coded without reference to other pictures, it can be decoded without first decoding any other pictures. Therefore, the decoder can start decoding CVS 500 at the IRAP picture 502. Furthermore, the IRAP picture 502 may refresh the DPB. For example, pictures presented after the IRAP picture 502 may not rely on pictures before the IRAP picture 502 (e.g., picture index 0) for inter prediction. Therefore, the picture buffer can be refreshed when the IRAP picture 502 is decoded. This has the effect of stopping all such errors because coding errors related to inter prediction cannot propagate through the IRAP picture 502. The IRAP picture 502 can be various types of pictures. For example, pictures presented after the IRAP picture 502 may not rely on pictures before the IRAP picture 502 (e.g., picture index 0) for inter prediction. Therefore, the picture buffer can be refreshed when the IRAP picture 502 is decoded. This has the effect of stopping all such errors because coding errors related to inter prediction cannot propagate through the IRAP picture 502. The IRAP picture 502 can be various types of pictures. This has the effect of stopping all such errors because coding errors related to inter prediction cannot propagate through the IRAP picture 502. The IRAP picture 502 can be various types of pictures. may be included. For example, an IRAP picture may be coded as an IDR or a CRA. An IDR is an intracoded picture that starts a new CVS 500 and refreshes the picture buffer. A CRA is an intracoded picture that acts as a random access point without starting a new CVS 500 or refreshing the picture buffer. In this way, the leading picture 504 related to the CRA may refer to the picture before the CRA, while the leading picture 504 related to the IDR need not refer to the picture before the IDR. CVS 500 also includes various non-IRAP pictures. These include the leading picture 504 and the trailing picture 506. The leading picture 504 is positioned after the IRAP picture 502 in decoding order 508, but before the IRAP picture 502 in presentation order 510.
[0074] The trailing picture 506 is positioned after the IRAP picture 502 in both decoding order 508 and presentation order 510. Both the leading picture 504 and the trailing picture 50 6 are coded by inter prediction. The trailing picture 506 is coded by referring to the IRAP picture 502 or a picture positioned after the IRAP picture 502. Therefore, the trailing picture 506 can always be decoded when the IRAP picture 502 is decoded. The leading picture 504 may include random access skip leading (RASL) and random access decodable leading (RADL) pictures. Yes. The RASL picture is coded by reference to the picture before the IRAP picture 502 but is coded at a position after the IRAP picture 502. Since the RASL picture relies on the previous picture, it cannot be decoded when the decoder starts decoding at the IRAP picture 502. Therefore, the RASL picture is skipped and not decoded when the IRAP picture 502 is used as a random access point. However, the RASL picture is decoded and displayed when the encoder uses a previous IRAP picture (before index 0 not shown) as a random access point. The RADL picture is coded with reference to the IRAP picture 502 and / or the pictures following the IRAP picture 502, but is positioned before the IRAP picture 502 in the presentation order 510. Since the RADL picture does not rely on the picture before the IRAP picture 502, it can
[0075] be decoded and displayed when the IRAP picture 502 is a random access point. Pictures from the CVS 500 may each be stored in an access unit. Furthermore, a picture may be segmented into slices, and the slices may be included in NAL units. The NAL unit is a storage unit including a picture parameter set or a slice and the corresponding slice header. The NAL unit is assigned a type to indicate to the decoder the type of data It may be included in an IDR (IDR_N_LP) NAL unit, a CRA NAL unit, etc. IDR_W_RADL NA The L unit indicates that the IRAP picture 502 is an IDR picture associated with the RADL reading picture 504. The IDR_N_LP NAL unit indicates that the IRAP picture 502 is an IDR picture not associated with any reading picture 504. The CRA NAL unit indicates that the IRAP picture 502 may be a CRA picture associated with the reading picture 504. Slices of non-IRAP pictures may also be placed in the NAL unit. For example, the slice of the trailing picture 506 may be placed in the trailing picture NAL unit type (TRAIL_NUT) that indicates that the trailing picture 506 is an inter-predicted picture. The slice of the reading picture 504 may be included in the RASL NAL unit type (RASL_NUT) and / or the RADL NAL unit type (RADL_NUT) that indicate that the corresponding picture is the corresponding type of inter-predicted reading picture 504. By signaling the slice of the picture within the corresponding NAL unit, the decoder can easily determine the appropriate decoding mechanism to apply to each picture / slice. For example, the slice of the trailing picture 506 may be placed in the trailing picture NAL unit type (TRAIL_NUT) that indicates that the trailing picture 506 is an inter-predicted picture. The slice of the reading picture 504 may be included in the RASL NAL unit type (RASL_NUT) and / or the RADL NAL unit type (RADL_NUT) that indicate that the corresponding picture is the corresponding type of inter-predicted reading picture 504. By signaling the slice of the picture within the corresponding NAL unit, the decoder can easily determine the appropriate decoding mechanism to apply to each picture / slice. For example, the slice of the trailing picture 506 may be placed in the trailing picture NAL unit type (TRAIL_NUT) that indicates that the trailing picture 506 is an inter-predicted picture. The slice of the reading picture 504 may be included in the RASL NAL unit type (RASL_NUT) and / or the RADL NAL unit type (RADL_NUT) that indicate that the corresponding picture is the corresponding type of inter-predicted reading picture 504. By signaling the slice of the picture within the corresponding NAL unit, the decoder can easily determine the appropriate decoding mechanism to apply to each picture / slice. For example, the slice of the trailing picture 506 may be placed in the trailing picture NAL unit type (TRAIL_NUT) that indicates that the trailing picture 506 is an inter-predicted picture. The slice of the reading picture 504 may be included in the RASL NAL unit type (RASL_NUT) and / or the RADL NAL unit type (RADL_NUT) that indicate that the corresponding picture is the corresponding type of inter-predicted reading picture 504. By signaling the slice of the picture within the corresponding NAL unit, the decoder can easily determine the appropriate decoding mechanism to apply to each picture / slice. By signaling the slice of the picture within the corresponding NAL unit, the decoder can easily determine the appropriate decoding mechanism to apply to each picture / slice. By signaling the slice of the picture within the corresponding NAL unit, the decoder can easily determine the appropriate decoding mechanism to apply to each picture / slice.
[0076] Figure 6 is a schematic diagram showing a plurality of sub-picture video streams 601, 602, and 603 split from the VR picture video stream 600. For example, each of the sub-picture video streams 601 to 603 and / or the VR picture video stream 600 is coded into the CVS 500. For example, each of the sub-picture video streams 601 to 603 and / or the VR picture video stream 600 is coded into the CVS 500. For example, each of the sub-picture video streams 601 to 603 and / or the VR picture video stream 600 is coded into the CVS 500. It may be inged. Therefore, the sub-picture video streams 601 to 603 and / or The VR picture video stream 600 may be encoded by an encoder such as the codec system 200 according to the method 100 and / or It may be encoded by an encoder such as the encoder 300. Furthermore, the sub-picture Video streams 601 to 603 and / or the VR picture video stream 600 may be decoded by a decoder such as the codec System 200 and / or decoder 400.
[0077] The VR picture video stream 600 includes a plurality of pictures presented over time. In particular , VR operates by coding a sphere of video content that can be displayed as if the user were at the center of the sphere. Each picture includes the entire sphere. On the other hand, only a part of the picture known as the viewport Is displayed to the user. For example, the user A head-mounted display (HMD) that selects and displays the viewport of the sphere based on the movement of the user's head may be used. This gives the impression of physically Existing within the virtual space depicted by the video. To achieve this result, each picture of the video sequence Includes the entire sphere of video data at the corresponding moment. However, only a small part of the picture (for Example, a single viewport) is displayed to the user. The rest of the picture is Discarded without being rendered. Generally, the entire picture is transmitted so that different viewports can be dynamically Selected and displayed according to the movement of the user's head. In the example shown, the pictures of the VR picture video stream 600 are available views
[0078] In the example shown, the pictures of the VR picture video stream 600 are available views It can be sub-divided into sub-pictures respectively based on the ports. Thus, each picture and the corresponding sub-pictures include a temporal position (e.g., the order of the pictures) as part of the temporal presentation. The sub-picture video streams 601-603 are generated when the sub-division is applied consistently over time. Such a consistent sub-division generates the sub-picture video streams 601-603, and each stream includes a set of sub-pictures with a predetermined size, shape, and spatial position relative to the corresponding picture in the VR picture video stream 600. Further, the set of sub-pictures within the sub-picture video streams 601-603 have different temporal positions during presentation. Thus, the sub-pictures of the sub-picture video streams 601-603 can be aligned in the temporal domain based on the temporal position. Then, the sub-pictures from the sub-picture video streams 601-603 at each temporal position can be merged in the spatial domain based on a predetermined spatial position to reconstruct the VR picture video stream 600 for display. In particular, the sub-picture video streams 601-603 can be encoded into separate sub-bitstreams respectively. Such sub-bitstreams, when merged together, result in a bitstream that includes the entire set of pictures over time. The resulting bitstream can be decoded based on the currently selected viewport of the user and sent to the decoder for display. One of the problems with VR video is that all of the sub-picture video streams 601-603 may be sent
[0079] to the user at high quality (e.g., high resolution). This is because the decoder may Dynamically select the user's current viewport and the corresponding sub-picture video stream 60 Enable real-time display of sub-pictures from 1 to 603. However, the user may, for example, only view a single viewport from the sub-picture video stream 601 while the sub-picture video streams 602 to 603 are discarded. Therefore, transmitting high-quality sub-picture video streams 602 to 603 may waste a large amount of bandwidth width. To improve coding efficiency, VR videos may be encoded into multiple video streams 600, with each video stream 600 encoded at a different quality / resolution In this way, the decoder can send a request for the current sub-picture video stream 601 Accordingly, the encoder (or intermediate slicer or other content server) can select a higher-quality sub-picture video stream 601 from a higher-quality video stream 600 and a lower-quality sub-picture video stream 602 to 603 from a lower-quality video stream 600 The encoder can then merge such sub-bitstreams into a fully encoded bitstream for transmission to the decoder In this way, the decoder receives a series of pictures, with the current viewport being of higher quality and the other viewports being of lower quality. Furthermore, the highest-quality sub-pictures are generally displayed to the user (when there is no head movement), and the lower-quality sub-pictures are generally discarded, which strikes a balance between functionality and coding efficiency. discarded, which strikes a balance between functionality and coding efficiency. viewports being of lower quality. Furthermore, the highest-quality sub-pictures are generally displayed to the user (when there is no head movement), and the lower-quality sub-pictures are generally discarded, which strikes a balance between functionality and coding efficiency. When there is no head movement), generally displayed to the user, and the lower-quality sub-pictures are generally discarded, which strikes a balance between functionality and coding efficiency.
[0080] When a user moves their eyes from sub - picture video stream 601 to sub - picture video stream 602, the decoder requests that the new current sub - picture video stream 602 be transmitted at a higher quality. The encoder can then change the merging mechanism accordingly. As described above, the decoder can start decoding a new CVS 500 only at IRAP picture 502. Therefore, the sub - picture video stream 602 is displayed at a lower quality until an IRAP picture / sub - picture is reached. Then, the IRAP picture can be decoded at a higher quality to start decoding a higher - quality version of the sub - picture video stream 602. This approach significantly improves video compression without adversely affecting the user's viewing experience.
[0081] One concern with the above approach is that the length of time required to change the resolution is based on the length of time until an IRAP picture is reached within the video stream. This is because the decoder cannot start decoding different versions of the sub - picture video stream 602 at non - IRAP pictures. One way to reduce such latency is to include more IRAP pictures. However, this leads to an increase in file size. To balance functionality and coding efficiency, different viewports / sub - picture video streams 601 - 603 may include IRAP pictures at different frequencies. For example, viewports / sub - picture video streams 601 - 603 that are more likely to be viewed may include IRAP pictures more frequently than other viewports / sub - picture video streams 601 - 60 It may have more IRAP pictures than 3. For example, in the context of basketball, views ports / sub-picture video streams 601-603 that look at the stands or ceiling are less likely to be seen by the user, so views related to the basket and / or center court ports / sub-picture video streams 601-603 may contain IRAP pictures more frequently than such viewports / sub-picture video streams 601-603.
[0082] This approach leads to further problems. In particular, sub-pictures that share a POC from sub-picture video streams 601-603 are part of a single picture. As described above, slices from a picture are included in NAL units based on the picture type. In some video decoding systems, all NAL units related to a single picture are constrained to include the same NAL unit type. When different sub- picture video streams 601-603 have IRAP pictures at different frequencies, parts of the picture include both IRAP sub-pictures and non-IRAP sub-pictures. This violates the constraint that each single picture should use only NAL units of the same type.
[0083] The present disclosure addresses this problem by removing the constraint that all NAL units related to slices within a picture use the same NAL unit type. For example, a picture is included in an access unit. By removing this Furthermore, pictures / access units are classified as IRAP and non-IRAP NAL unit types. A flag may be coded to indicate when the mix includes a mix with a The mixed NAL unit types in picture flag )(mixed_nalu_types_in_pic_flag). In addition, a single mixed picture / access unit A packet contains only one type of IRAP NAL unit and one type of non-IRAP NAL unit. This is because the constraint that unintended NAL units may be Prevents type mixing from occurring. If such mixing is allowed, the decoder The code must be designed to manage such a mix. Eliminates the complexity of the hardware required without adding any additional benefits to the process For example, a mixed picture can be selected from IDR_W_RADL, IDR_N_LP, or CRA_NUT. Additionally, a mixed picture may contain one of the types of IRAP NAL units specified in TRAIL Contains one type of non-IRAP NAL unit selected from _NUT, RADL_NUT, and RASL_NUT. An exemplary implementation of this scheme is discussed in more detail below.
[0084] FIG. 7 illustrates an example bitstream containing pictures with mixed NAL unit types. 7 is a schematic diagram illustrating a bitstream 700. For example, the bitstream 700 may be a codec according to the method 100. 400 for decoding by the codec system 200 and / or the decoder 400. It may be generated by the encoder 300 and / or. Further, the bitstream 700 may include a VR picture video stream 600 merged from a plurality of sub-picture video streams 601 to 603 of a plurality of video resolutions, and each sub-picture video stream includes a CVS 500 at different spatial positions. The bitstream 700 may include a sequence parameter set (SPS) 710, a plurality of picture parameter sets (PPS) 711, a plurality of slice headers 715, and picture data 720. The SPS 710 includes sequence data common to all pictures in the video sequence included in the bitstream 700. Such data may include picture size settings, bit depth, coding tool parameters, bitrate constraints, etc. The PPS 711 includes parameters applied to the entire picture. Therefore, each picture in the video sequence may refer to the PPS 711. Each picture refers to the PPS 711, but it should be noted that in some examples, a single PPS 711 may include data regarding a plurality of pictures. For example, a plurality of similar pictures may be coded with similar parameters. In such a case, a single PPS 711 may include data regarding such similar pictures. The PPS 711 may indicate coding tools, quantization parameters, offsets, etc. available for slices within the corresponding picture. The slice header 715 includes parameters specific to each slice within the picture. Therefore, there may be one slice header 715 for each slice in the video sequence. The slice header 715 includes slice type information. It may be included, and each sub-picture video stream includes a CVS 500 at different spatial positions. Including the CVS 500 at different spatial positions.
[0085] The bitstream 700 includes a sequence parameter set (SPS) 710, a plurality of picture parameter sets (PPS) 711, a plurality of slice headers 715, and picture data 720. The SPS 710 includes sequence data common to all pictures in the video sequence included in the bitstream 700. Such data may include picture size settings, bit depth, coding tool parameters, bitrate constraints, etc. The PPS 711 includes parameters applied to the entire picture. Therefore, each picture in the video sequence may refer to the PPS 711. The SPS 710 includes sequence data common to all pictures in the video sequence included in the bitstream 700. Such data may include picture size settings, bit depth, coding tool parameters, bitrate constraints, etc. The SPS 710 includes sequence data common to all pictures in the video sequence included in the bitstream 700. Such data may include picture size settings, bit depth, coding tool parameters, bitrate constraints, etc. The PPS 711 includes parameters applied to the entire picture. Therefore, each picture in the video sequence may refer to the PPS 711. The PPS 711 includes parameters applied to the entire picture. Therefore, each picture in the video sequence may refer to the PPS 711. Each picture may refer to the PPS 711, but in some examples, it should be noted that a single PPS 711 may include data regarding a plurality of pictures. Each picture may refer to the PPS 711, but in some examples, it should be noted that a single PPS 711 may include data regarding a plurality of pictures. For example, a plurality of similar pictures may be coded with similar parameters. In such a case, a single PPS 711 may include data regarding such similar pictures. The PPS 711 may indicate coding tools, quantization parameters, offsets, etc. available for slices within the corresponding picture. The slice header 715 includes parameters specific to each slice within the picture. Therefore, there may be one slice header 715 for each slice in the video sequence. The slice header 715 includes parameters specific to each slice within the picture. Therefore, there may be one slice header 715 for each slice in the video sequence. , picture order count (POC), reference picture list, prediction weight, tile entry point nt (tile entry point), deblocking parameters, etc. may be included. Note that slice he ader 715 may also be called a tile group header depending on the context. .
[0086] Image data 720 includes video data encoded by inter prediction and / or intra prediction as well as corresponding transformed and quantized residual data. For example, a video sequence includes a plurality of pictures 721 coded as image data 720. A picture 721 is a single frame of a video sequence and thus is generally displayed as a single unit when displaying the video sequence. However, sub-pictures 723 may be displayed to implement certain technologies such as virtual reality. Pictures 721 each reference a PPS 711. A picture 721 may be divided into sub-pictures 723, tiles, and / or slices. A sub-picture 723 is a spatial region of a picture 721 that is consistently applied to a coded video sequence. Thus, a sub-picture 723 may be displayed by an HMD in the context of VR. Further, a sub-picture 723 having a specified POC may be obtained from sub-picture video streams 601 - 603 of corresponding resolutions. A sub-picture 723 may reference an SPS 710. In some systems, a slice 725 is called a tile group that includes tiles. A slice 725 and / or a tile group of tiles reference a slice header 715. A slice 725 is a single NAL unit An integer number of complete tiles of picture 721 exclusively included in the knit or an integer number within the tile may be defined as consecutive complete rows of CTUs. Thus, slice 725 is further divided into CTUs and / or CTBs. The CTU / CTB is further divided into coding blocks based on a coding tree. Then, the coding blocks can be encoded / decoded by a prediction mechanism. As described above, and / or CTB. The CTU / CTB is further divided into coding blocks based on a coding tree. Then, the coding blocks can be encoded / decoded by a prediction mechanism. As described above, and / or CTB. The CTU / CTB is further divided into coding blocks based on a coding tree. Then, the coding blocks can be encoded / decoded by a prediction mechanism.
[0087] The parameter set and / or slice 725 is coded into an NAL unit. The NAL unit may be defined as a byte including a syntax structure including an indication of the type of subsequent data, and / or CTB. The CTU / CTB is further divided into coding blocks based on a coding tree. Then, the coding blocks can be encoded / decoded by a prediction mechanism. / or CTB. The CTU / CTB is further divided into coding blocks based on a coding tree. Then, the coding blocks can be encoded / decoded by a prediction mechanism. / or CTB. The CTU / CTB is further divided into coding blocks based on a coding tree. Then, the coding blocks can be encoded / decoded by a prediction mechanism. / or CTB. The CTU / CTB is further divided into coding blocks based on a coding tree. Then, the coding blocks can be encoded / decoded by a prediction mechanism. / or CTB. The CTU / CTB is further divided into coding blocks based on a coding tree. Then, the coding blocks can be encoded / decoded by a prediction mechanism. / or CTB. The CTU / CTB is further divided into coding blocks based on a coding tree. Then, the coding blocks can be encoded / decoded by a prediction mechanism. / or CTB. The CTU / CTB is further divided into coding blocks based on a coding tree. Then, the coding blocks can be encoded / decoded by a prediction mechanism. / or CTB. The CTU / CTB is further divided into coding blocks based on a coding tree. Then, the coding blocks can be encoded / decoded by a prediction mechanism. / or CTB. The CTU / CTB is further divided into coding blocks based on a coding tree. Then, the coding blocks can be encoded / decoded by a prediction mechanism. / or CTB. The CTU / CTB is further divided into coding blocks based on a coding tree. Then, the coding blocks can be encoded / decoded by a prediction mechanism.
[0088] As described above, an IRAP picture such as IRAP picture 502 is included in an IRAP NAL unit 745. can be obtained. Non-IRAP pictures such as the leading picture 504 and the trailing picture 506 are , and can be included in the non-IRAP NAL unit 749. In particular, the IRAP NAL unit 745 is any NAL unit that includes a slice 725 obtained from an IRAP picture or a sub-picture. The non-IRAP N AL unit 749 is any NAL unit that includes a slice 725 obtained from a picture that is not an IRAP picture or a sub-picture (e.g., a leading picture and a trailing picture). The IRAP NAL unit 745 and the non-IRAP NAL unit 749 are both VCL NAL units 740 because both contain slice data. In an exemplary embodiment, the IRAP NAL unit 745 may include the slice 725 from an IDR picture without a leading picture or an IDR related to a RADL picture in the IDR_N_LP NAL unit 741 or the IDR_w_RADL NAL unit 742, respectively. Further, the IRAP NAL unit 745 may include the slice 725 from a CRA picture in the CRA_NUT 743. In an exemplary embodiment, the non-IRAP NAL unit 749 may include the slice 725 from a RASL picture, a RADL picture, or a trailing picture in the RASL_NUT 746, the RADL_NUT 747, or the TRAIL_NUT 748, respectively . In an exemplary embodiment, a complete list of possible NAL units is sorted by NAL unit type and shown below.
Table 1A
Table 1B
Table 1C
[0089] As described above, the VR video stream may include sub-pictures having IRAP pictures at different frequencies. This allows for fewer IRAP pictures to be used for spatial regions that the user is less likely to view, and more IRAP pictures to be used for spatial regions that the user is more likely to view frequently. In this way, spatial regions that the user is likely to return to frequently can be quickly adjusted to a higher resolution. When this method results in a picture that includes both IRAP NAL units and non-IRAP NAL units , the picture is called a mixed picture. This state may be signaled by the mixed_nalu_types_in_pic_flag 727. The mixed_nalu_types_in_pic_flag 727 may be set in the PPS 711. Further, the mixed_nalu_types_in_pic_flag 727 may be set to be equal to 1 when each picture that refers to the PPS 711 has two or more VCL NAL units and the VCL NAL units of each picture that refers to the PPS 711 do not have the same value of the nal_unit_type. Further, the mixed_nalu_types_in_pic_flag 727 may be set to be equal to 1 when each picture that refers to the PPS 711 has one or more VCL NAL units and all of the VCL NAL units of each picture that refers to the PPS 711 have the nal_unit_type for spatial regions that the user is less likely to view, and more IRAP pictures to be used for spatial regions that the user is more likely to view frequently. In this way, spatial regions that the user is likely to return to frequently can be quickly adjusted to a higher resolution. When this method results in a picture that includes both IRAP NAL units and non-IRAP NAL units , the picture is called a mixed picture. This state may be signaled by the mixed_nalu_types_in_pic_flag 727. The mixed_nalu_types_in_pic_flag 727 may be set in the PPS 711. Further, the mixed_nalu_types_in_pic_flag 727 may be set to be equal to 1 when each picture that refers to the PPS 711 has two or more VCL NAL units and the VCL NAL units of each picture that refers to the PPS 711 do not have the same value of the nal_unit_type. Further, the mixed_nalu_types_in_pic_flag 727 may be set to be equal to 1 when each picture that refers to the PPS 711 has one or more VCL NAL units and all of the VCL NAL units of each picture that refers to the PPS 711 have the nal_unit_type for spatial regions that the user is less likely to view, and more IRAP pictures to be used for spatial regions that the user is more likely to view frequently. In this way, spatial regions that the user is likely to return to frequently can be quickly adjusted to a higher resolution. When this method results in a picture that includes both IRAP NAL units and non-IRAP NAL units , the picture is called a mixed picture. This state may be signaled by the mixed_nalu_types_in_pic_flag 727. The mixed_nalu_types_in_pic_flag 727 may be set in the PPS 711. Further, the mixed_nalu_types_in_pic_flag 727 may be set to be equal to 1 when each picture that refers to the PPS 711 has two or more VCL NAL units and the VCL NAL units of each picture that refers to the PPS 711 do not have the same value of the nal_unit_type. Further, the mixed_nalu_types_in_pic_flag 727 may be set to be equal to 1 when each picture that refers to the PPS 711 has one or more VCL NAL units and all of the VCL NAL units of each picture that refers to the PPS 711 have the nal_unit_type for spatial regions that the user is less likely to view, and more IRAP pictures to be used for spatial regions that the user is more likely to view frequently. In this way, spatial regions that the user is likely to return to frequently can be quickly adjusted to a higher resolution. When this method results in a picture that includes both IRAP NAL units and non-IRAP NAL units , the picture is called a mixed picture. This state may be signaled by the mixed_nalu_types_in_pic_flag 727. The mixed_nalu_types_in_pic_flag 727 may be set in the PPS 711. Further, the mixed_nalu_types_in_pic_flag 727 may be set to be equal to 1 when each picture that refers to the PPS 711 has two or more VCL NAL units and the VCL NAL units of each picture that refers to the PPS 711 do not have the same value of the nal_unit_type. Further, the mixed_nalu_types_in_pic_flag 727 may be set to be equal to 1 when each picture that refers to the PPS 711 has one or more VCL NAL units and all of the VCL NAL units of each picture that refers to the PPS 711 have the nal_unit_type for spatial regions that the user is less likely to view, and more IRAP pictures to be used for spatial regions that the user is more likely to view frequently. In this way, spatial regions that the user is likely to return to frequently can be quickly adjusted to a higher resolution. When this method results in a picture that includes both IRAP NAL units <00012When having the same value as, it may be set to be equal to 0.
[0090] Furthermore, when the mixed_nalu_types_in_pic_flag 727 is set, one or more VCL NAL units 740 among the sub-pictures 723 of the picture 721 all have the value of the first specific type of NAL unit, and constraints may be used such that all other VCL NAL units 740 within the picture 721 have a different second specific value of the NAL unit type. For example, the constraint may require that the mixed picture 721 contains a single type of IRAP NAL unit 745 and a single type of non-IRAP NAL unit 749. For example, the picture 721 may include one or more I DR_N_LP NAL units 741, one or more IDR_w_RADL NAL units 742, or one or more CRA_NUTs 743, but may not include any combination of such IRAP NAL units 745. Furthermore, the picture 721 may include one or more RASL_NUTs 746, one or more RADL_NUTs 747, or one or more TRAIL_NUTs 748, but may not include any combination of such IRAP NAL units 745.
[0091] In an exemplary implementation, the picture type is used to define the decoding process. Such a process may include, for example, deriving picture identification information based on the picture order count (POC), marking the status of reference pictures in the decoded picture buffer (DPB), outputting pictures from the D PB, etc. A picture is all of the coded pictures from the DPB, etc. A picture is all of the coded pictures It can be identified by a type based on the NAL unit type including it or its lower part. One In some video coding systems, the picture type may include instantaneous decoding refresh (I DR) pictures and non-IDR pictures. In other video coding systems the picture type may include trailing pictures, temporal sub-layer access (TSA: te mporal sub-layer) pictures, step-wise temporal sub-layer access (STSA: step-wise temporal sub-layer access) pictures, random access decodable leading (RADL) pictures, random access skip leading (RASL) pictures, broken-link access (BLA : broken-link access) pictures, instantaneous random access pictures, and clean ran dom access pictures. Such picture types may be further distinguished based on whether the picture is a sub-layer referenced picture or a sub-layer non-referenced picture sub-layer non-referenced picture). BLA pictures may be further distinguished as BLA with a leading picture, BLA with a RADL picture, and BLA without a leading picture. IDR pictures may be further distinguished as IDR with a RADL picture and IDR without a leading picture and may be further distinguished. Such picture types may be used to implement functions related to various videos. For example, IDR, BLA, and / or CRA pictures implement IRAP pictures
[0092] and may be used to implement functions related to various videos. For example, IDR, BLA, and / or CRA pictures implement IRAP pictures and may be used to implement functions related to various videos. For example, IDR, BLA, and / or CRA pictures implement IRAP pictures It may be used for. The IRAP picture may provide the following functions / advantages. The IRAP pi The presence of the picture may indicate that the decoding process can be started from that picture. This function enables the implementation of the random access feature where the decoding process starts at a specified position within the bitstream as long as the IRAP picture is present at that position. Such a position is not necessarily at the beginning of the bitstream. Also, the presence of the IRAP picture refreshes the decoding process so that coded pictures starting with the IRAP picture, except for RASL pictures, are coded without referring to any pictures positioned before the IRAP picture at all. Therefore, the IRAP picture positioned within the bitstream stops the propagation of decoding errors. Therefore, decoding errors of coded pictures positioned before the IRAP picture cannot be propagated through the IRAP picture to pictures following the IRAP picture in decoding order.
[0093] The IRAP picture provides various functions but causes a disadvantage to the compression efficiency. Therefore, the presence of the IRAP picture may cause a sharp increase in the bitrate. This disadvantage to the compression efficiency has various causes. For example, the IRAP picture is an intra-predicted picture represented by significantly more bits than an inter-predicted picture used as a non-IRAP picture. Furthermore, the presence of the IRAP picture impairs the temporal prediction used in inter prediction. In particular, the IRAP picture refreshes the decoding process by removing the previous reference picture from the DPB. Removing the previous reference picture References for use in coding pictures that follow an IRAP picture in decoding order reduce the availability of pictures and thus reduce the efficiency of this process.
[0094] IDR pictures may use different signaling and derivation processes than other IRAP picture types. For example, the signaling and derivation processes related to IDR may set the MSB part of the POC to 0 instead of deriving the most significant bit (MSB) from the previous key picture. Furthermore, the slice header of an IDR picture may not contain information used to assist in the management of reference pictures. On the other hand, other picture types such as CRA, trailing, TSA, etc. may contain reference picture information such as a reference picture set (RPS: reference picture set) or a reference picture list used to perform the reference picture marking process. The reference picture marking process is a process that determines whether the status of a reference picture in the DPB is used for reference or not used for reference. The presence of an IDR indicates that the decoding process simply marks all reference pictures in the DPB as not used for reference, so such information may not be signaled for IDR pictures. picture information such as a reference picture set (RPS: reference picture set) or a reference picture list used to perform the reference picture marking process. The reference picture marking process is a process that determines whether the status of a reference picture in the DPB is used for reference or not used for reference. The presence of an IDR indicates that the decoding process simply marks all reference pictures in the DPB as not used for reference, so such information may not be signaled for IDR pictures. picture information such as a reference picture set (RPS: reference picture set) or a reference picture list used to perform the reference picture marking process. The reference picture marking process is a process that determines whether the status of a reference picture in the DPB is used for reference or not used for reference. The presence of an IDR indicates that the decoding process simply marks all reference pictures in the DPB as not used for reference, so such information may not be signaled for IDR pictures. The reference picture marking process is a process that determines whether the status of a reference picture in the DPB is used for reference or not used for reference. The presence of an IDR indicates that the decoding process simply marks all reference pictures in the DPB as not used for reference, so such information may not be signaled for IDR pictures. The reference picture marking process is a process that determines whether the status of a reference picture in the DPB is used for reference or not used for reference. The presence of an IDR indicates that the decoding process simply marks all reference pictures in the DPB as not used for reference, so such information may not be signaled for IDR pictures. The presence of an IDR indicates that the decoding process simply marks all reference pictures in the DPB as not used for reference, so such information may not be signaled for IDR pictures. picture indicates that the decoding process simply marks all reference pictures in the DPB as not used for reference, so such information may not be signaled for IDR pictures.
[0095] In addition to the picture type, the identification information of a picture by POC is also used for multiple purposes such as the management of the use of reference pictures in inter prediction, the output of pictures from the DPB, the scaling of motion vectors, and weighted prediction. For example, one In a video coding system of a portion, pictures in a DPB are marked as being used for short-term reference, being used for long-term reference, or not being used for reference. If a picture is marked as not being used for reference, the picture can no longer be used for prediction. When such a picture is no longer needed for output, the picture can be deleted from the DPB. In other video coding systems, reference pictures may be marked as short-term and long-term. A reference picture may be marked as not being used for reference when the picture is no longer needed for predictive reference. The transition between these statuses may be controlled by a marking process of decoded reference pictures. An implicit sliding window process and / or an explicit memory management control operation (MMCO) process may be used as a marking mechanism for decoded reference pictures. The sliding window process marks short-term reference pictures as not being used for reference when the number of reference frames is equal to the specified maximum number denoted as max_num_ref_frames in the SPS. Short-term reference pictures may be stored in a first-in-first-out manner such that the most recently decoded short-term picture is held in the DPB. The explicit MMCO process may include a plurality of MMCO commands. An MMCO command may mark one or more short-term or long-term reference pictures as not being used for reference, may mark all pictures as not being used for reference, or may mark the current picture as not being used for reference. The transition between these statuses may be controlled by a marking process of decoded reference pictures. An implicit sliding window process and / or an explicit memory management control operation (MMCO) process may be used as a marking mechanism for decoded reference pictures. The implicit sliding window process and / or the explicit memory management control operation (MMCO) process may be used as a marking mechanism for decoded reference pictures. The sliding window process marks short-term reference pictures as not being used for reference when the number of reference frames is equal to the specified maximum number denoted as max_num_ref_frames in the SPS. Short-term reference pictures may be stored in a first-in-first-out manner such that the most recently decoded short-term picture is held in the DPB. The explicit MMCO process may include a plurality of MMCO commands. An MMCO command may mark one or more short-term or long-term reference pictures as not being used for reference, may mark all pictures as not being used for reference, or may mark the current picture as not being used for reference. Short-term reference pictures may be stored in a first-in-first-out manner such that the most recently decoded short-term picture is held in the DPB. The explicit MMCO process may include a plurality of MMCO commands. An MMCO command may mark one or more short-term or long-term reference pictures as not being used for reference, may mark all pictures as not being used for reference, or may mark the current picture as not being used for reference. Mark the reference picture or the existing short-term reference picture as long-term, and then, A long-term picture index may be assigned to the long-term reference picture.
[0096] In some video coding systems, the marking operation of the reference picture as well as the process for output and deletion of pictures from the DPB is performed after the pictures are decoded. Other video coding systems use the RPS for reference picture management. The most fundamental difference between the RPS mechanism and the MMCO / sliding window process is that, for each specific slice, the RPS provides the complete set of reference pictures used by the current picture or any subsequent picture. Therefore, the complete set of all pictures to be held in the DPB for use by the current or future pictures is signaled in the RPS. This is different from the MMCO / sliding window scheme where only relative changes to the DPB are signaled. According to the RPS mechanism, in order to maintain the correct status of the reference pictures in the DPB, information from the previous pictures in decoding order is not required. The order of picture decoding and the operation of the DPB are modified in some video coding systems to take advantage of the RPS and enhance error resilience. the complete set of all pictures to be held in the DPB for use by the current or future pictures is signaled in the RPS. This is different from the MMCO / sliding window scheme where only relative changes to the DPB are signaled. According to the RPS mechanism, in order to maintain the correct status of the reference pictures in the DPB, information from the previous pictures in decoding order is not required. The order of picture decoding and the operation of the DPB are modified in some video coding systems to take advantage of the RPS and enhance error resilience. the complete set of all pictures to be held in the DPB for use by the current or future pictures is signaled in the RPS. This is different from the MMCO / sliding window scheme where only relative changes to the DPB are signaled. According to the RPS mechanism, in order to maintain the correct status of the reference pictures in the DPB, information from the previous pictures in decoding order is not required. The order of picture decoding and the operation of the DPB are modified in some video coding systems to take advantage of the RPS and enhance error resilience. the complete set of all pictures to be held in the DPB for use by the current or future pictures is signaled in the RPS. This is different from the MMCO / sliding window scheme where only relative changes to the DPB are signaled. According to the RPS mechanism, in order to maintain the correct status of the reference pictures in the DPB, information from the previous pictures in decoding order is not required. The order of picture decoding and the operation of the DPB are modified in some video coding systems to take advantage of the RPS and enhance error resilience. the complete set of all pictures to be held in the DPB for use by the current or future pictures is signaled in the RPS. This is different from the MMCO / sliding window scheme where only relative changes to the DPB are signaled. According to the RPS mechanism, in order to maintain the correct status of the reference pictures in the DPB, information from the previous pictures in decoding order is not required. The order of picture decoding and the operation of the DPB are modified in some video coding systems to take advantage of the RPS and enhance error resilience. the complete set of all pictures to be held in the DPB for use by the current or future pictures is signaled in the RPS. This is different from the MMCO / sliding window scheme where only relative changes to the DPB are signaled. According to the RPS mechanism, in order to maintain the correct status of the reference pictures in the DPB, information from the previous pictures in decoding order is not required. The order of picture decoding and the operation of the DPB are modified in some video coding systems to take advantage of the RPS and enhance error resilience. the complete set of all pictures to be held in the DPB for use by the current or future pictures is signaled in the RPS. This is different from the MMCO / sliding window scheme where only relative changes to the DPB are signaled. According to the RPS mechanism, in order to maintain the correct status of the reference pictures in the DPB, information from the previous pictures in decoding order is not required. The order of picture decoding and the operation of the DPB are modified in some video coding systems to take advantage of the RPS and enhance error resilience. the complete set of all pictures to be held in the DPB for use by the current or future pictures is signaled in the RPS. This is different from the MMCO / sliding window scheme where only relative changes to the DPB are signaled. According to the RPS mechanism, in order to maintain the correct status of the reference pictures in the DPB, information from the previous pictures in decoding order is not required. The order of picture decoding and the operation of the DPB are modified in some video coding systems to take advantage of the RPS and enhance error resilience. In some video coding systems, the operation of the buffer including both the marking of pictures and the output and deletion of the decoded pictures from the DPB may be applied after the current picture is decoded. In other video coding systems, first, the RPS is decoded from the slice header of the current picture, and then, the picture In some video coding systems, the operation of the buffer including both the marking of pictures and the output and deletion of the decoded pictures from the DPB may be applied after the current picture is decoded. In other video coding systems, first, the RPS is decoded from the slice header of the current picture, and then, the picture The marking and buffer operations may be applied before decoding the current picture. 。
[0097] In VVC, the reference picture management method may be summarized as follows. Two reference picture lists, denoted as List 0 and List 1, are directly signaled and derived. They may be based on the RPS or the sliding window + MMCO process discussed above. The marking of reference pictures is directly based on reference picture lists 0 and 1 using both active and non-active entries of the reference picture lists, but only active entries may be used as reference indexes in the inter prediction of CTUs. The information for deriving the two reference picture lists is signaled by syntax elements and syntax structures in the SPS, PPS, and slice header. A predefined RPL structure is signaled in the SPS for use by referring to it in the slice header. The two reference picture lists are generated for all types of slices, including bi-directional inter prediction (B) slices, uni-directional inter prediction (P) slices, and intra prediction (I) slices. The two reference picture lists may be constructed without using the reference picture list initialization process or the reference picture list modification process. The long-term reference picture (LTRP) is identified by the POC LSB. A delta POC MSB cycle may be signaled for the LTRP as determined for each picture.
[0098] To code a video image, the image is first segmented, and the segments are coded into a bitstream Various picture segmentation methods are available. For example, an image can be segmented into regular slices, dependent slices, tiles, and / or wavefront parallel processing (WPP). For simplicity, HEVC constrains the encoder such that only regular slices, dependent slices, tiles, WPP, and combinations thereof are used when segmenting slices into groups of CTBs for video coding. Such segmentation can be applied to support maximum transmission unit (MTU) size matching, parallel processing, and reduced end-to-end delay. The MTU represents the maximum amount of data that can be sent in a single packet. If the payload of a packet exceeds the MTU, the payload is split into two packets by a process called fragmentation. A regular slice, also simply called a slice, is a segmented part of an image that can be reconstructed independently of other regular slices within the same picture, despite some dependencies due to loop filtering operations. Each regular slice is encapsulated into its own network abstraction layer (NAL) unit for transmission. Additionally, the interdependencies of intra-picture prediction (intra-sample prediction, motion information prediction, coding mode prediction) and entropy coding across slice boundaries may be made unavailable to support independent reconstruction. Such independent reconstruction enables parallel processing. processing. processing. A regular slice, also simply called a slice, is a segmented part of an image that can be reconstructed independently of other regular slices within the same picture, despite some dependencies due to loop filtering operations. Each regular slice is encapsulated into its own network abstraction layer (NAL) unit for transmission. Additionally, the interdependencies of intra-picture prediction (intra-sample prediction, motion information prediction, coding mode prediction) and entropy coding across slice boundaries may be made unavailable to support independent reconstruction. Such independent reconstruction enables parallel processing. processing.
[0099] A regular slice, also simply called a slice, is a segmented part of an image that can be reconstructed independently of other regular slices within the same picture, despite some dependencies due to loop filtering operations. Each regular slice is encapsulated into its own network abstraction layer (NAL) unit for transmission. Additionally, the interdependencies of intra-picture prediction (intra-sample prediction, motion information prediction, coding mode prediction) and entropy coding across slice boundaries may be made unavailable to support independent reconstruction. Such independent reconstruction enables parallel processing. processing. processing. Furthermore, the interdependencies of intra-picture prediction (intra-sample prediction, motion information prediction, coding mode prediction) and entropy coding across slice boundaries may be made unavailable to support independent reconstruction. Such independent reconstruction enables parallel processing. processing. Supports parallelization. For example, parallelization based on normal slices uses minimal inter-processor or inter-core communication. However, since each normal slice is independent, each slice is associated with a separate slice header. The use of normal slices can result in significant coding overhead due to the bit cost of the slice header for each slice and the lack of prediction to straddle the slice boundaries. Furthermore , normal slices may be used to support matching regarding the MTU size requirements . In particular, when normal slices are encapsulated into separate NAL units and can be coded independently , each normal slice should be smaller than the MTU of the MTU method to avoid splitting the slice into multiple packets . Thus, the purpose of parallelization and the purpose of MTU size matching may impose conflicting requirements on the layout of slices within a picture .
[0100] Dependent slices are similar to normal slices but have a shortened slice header and allow for the partitioning of the boundaries of picture tree blocks without sacrificing intra-picture prediction . Thus, dependent slices allow for fragmentation of normal slices into multiple NAL units, which results in reduced end-to-end delay by enabling some parts of a normal slice to be sent before the encoding of the entire normal slice is complete . .
[0101] A picture may be divided into tile groups / slices and tiles. A tile is a sequence of CTUs that covers a rectangular region of the picture . A tile group / slice includes some tiles of a picture. The raster scan tile group mode and the long rectangular tile group mode may be used to generate the tiles. In the raster scan tile group mode, the tile group includes a sequence of tiles of the raster scan of the tiles of the picture. In the rectangular tile group mode, the tile group includes some tiles of the picture that collectively form a rectangular area of the picture. The tiles within the rectangular tile group are in the raster scan order of the tile group. For example, a tile may be a segmented part of an image generated by horizontal and vertical boundaries that generate columns and rows of tiles. The tiles may be coded in raster scan order (from right to left and from top to bottom). The scan order of the CTB is limited within a tile. Thus, the CTB within the first tile is coded in raster scan order before proceeding to the CTB within the next tile. Similar to a normal slice, a tile does not impair the prediction interdependencies and entropy decoding interdependencies within a picture. However, a tile may not be included in individual NAL units and thus may not be used for MTU size matching. Each tile can be processed by one processor / core, and the inter-processor / core communication used for intra-picture prediction between processing units decoding neighboring tiles may be limited to carrying a shared slice header (when neighboring tiles are within the same slice) and performing sharing related to loop filtering of reconstructed samples and metadata. For two or more tiles are in the raster scan order of the tile group. For example, a tile may be a segmented part of an image generated by horizontal and vertical boundaries that generate columns and rows of tiles. The tiles may be coded in raster scan order (from right to left and from top to bottom). The scan order of the CTB is limited within a tile. Thus, the CTB within the first tile is coded in raster scan order before proceeding to the CTB within the next tile. Similar to a normal slice, a tile does not impair the prediction interdependencies and entropy decoding interdependencies within a picture. However, a tile may not be included in individual NAL units and thus may not be used for MTU size matching. Each tile can be processed by one processor / core, and the inter-processor / core communication used for intra-picture prediction between processing units decoding neighboring tiles may be limited to carrying a shared slice header (when neighboring tiles are within the same slice) and performing sharing related to loop filtering of reconstructed samples and metadata. For two or more tiles are in the raster scan order of the tile group. For example, a tile may be a segmented part of an image generated by horizontal and vertical boundaries that generate columns and rows of tiles. The tiles may be coded in raster scan order (from right to left and from top to bottom). The scan order of the CTB is limited within a tile. Thus, the CTB within the first tile is coded in raster scan order before proceeding to the CTB within the next tile. Similar to a normal slice, a tile does not impair the prediction interdependencies and entropy decoding interdependencies within a picture. However, a tile may not be included in individual NAL units and thus may not be used for MTU size matching. Each tile can be processed by one processor / core, and the inter-processor / core communication used for intra-picture prediction between processing units decoding neighboring tiles may be limited to carrying a shared slice header (when neighboring tiles are within the same slice) and performing sharing related to loop filtering of reconstructed samples and metadata. For two or more tiles are in the raster scan order of the tile group. For example, a tile may be a segmented part of an image generated by horizontal and vertical boundaries that generate columns and rows of tiles. The tiles may be coded in raster scan order (from right to left and from top to bottom). The scan order of the CTB is limited within a tile. When a slice contains a The byte offset of the entry point for the file is signaled in the slice header. For each slice and tile, the following conditions may be satisfied: 1) Within a slice 1) all coded tree blocks of a tile belong to the same tile, and 2) all coded tree blocks of a tile belong to the same tile. All coded treeblocks in a slice belong to the same slice. At least one should be satisfied.
[0102] In WPP, the image is segmented into a single row of CTB. Entropy decoding and prediction mechanism The algorithm may use data from other rows of the CTB. This allows for parallel processing. For example, the current row may be decoded in parallel with the previous row. However, the decoding of the current row is delayed by 2CTB from the decoding process of the previous row. The extension is performed by the CTB above the current CTB in the current line and the CTB above it to the right before the current CTB is coded. This method ensures that relevant data for the CTB is available. When rendered, it appears as a wavefront. This offset onset is at most the same as the number of CTB rows the image contains. Enables parallelization up to the same number of processors / cores. Since intra-picture prediction between rows is allowed, the processor Inter-server / inter-core communication can be quite high. WPP division takes into account the size of the NAL unit. Therefore, WPP does not support MTU size matching. To enforce MTU size matching, normal slices are used for specific coding. It can be used in conjunction with the WPP while incurring overhead. Finally, a wavefront segment may contain exactly one row of CTBs. Further, when using the WPP and when a slice starts within a row of CTBs, the slice should end within the same row of CTBs.
[0103] A tile may also include a motion constrained tile set . A motion constrained tile set (MCTS) is a tile set designed such that the associated motion vectors point to full-sample positions within the MCTS and fractional-sample points only to full-sample positions within the MCTS for interpolation. . Further, the use of motion vector candidates for temporal motion vector prediction derived from blocks outside the MCTS is not allowed. In this way, each MCTS may be decoded independently without the presence of tiles not included in the MCTS. A temporal MCTS supplemental enhancement information (SEI) message indicates the presence of the MCTS within the bitstream and may be used to signal the MCTS. The MCTS SEI message provides supplemental information that can be used in MCTS sub-bitstream extraction to generate a bitstream that conforms to the set of MCTSs (as specified as part of the semantics of the SEI message). The information defines sets of several MCTSs each, and alternative video parameter sets (VPSs), sequence parameter sets (SPSs), that are used during the MCTS sub-bitstream extraction process and the bytes of the raw byte sequence payload (RBSP) of the picture parameter set (PPS) including several sets of extracted information. When extracting the sub-bitstream by the MCTS sub-bitstream extraction process the parameter sets (VPS, SPS, and PPS) may be rewritten or replaced, and the slice header may be updated (first_slice_segme nt_in_pic_flag and slice_segment_address) because one or all of the syntax elements related to the slice address may use different values in the extracted sub-bitstream and may thus be updated
[0104] A VR application, also called a 360-degree video application, may display only a subset of the entire picture, which is a part of a complete sphere and the result. A mechanism for 360-degree delivery that depends on the viewport via dynamic adaptive streaming over hypertext transfer protocol (DASH) may be used to lower the bitrate and support the delivery of 360-degree video by the streaming mechanism. This mechanism divides the sphere / projected picture into multiple MCTSs, for example, by using cube map projection (CMP) Two or more bitstreams may be encoded at different spatial resolutions or qualities. When delivering data to the decoder, MCTSs from the higher-resolution / quality bitstream are sent for the viewport to be displayed (e.g., the front viewport) MCTS from a low-resolution / quality bitstream is sent for other viewports. These MCTS are packed in a specific way and then sent to the receiver for decoding. The viewport seen by the user is expected to be represented by high-resolution / quality MCTS to create a positive viewing experience. When the user turns their head to look at another viewport (e.g., the left or right viewport), the content being displayed comes from a lower-resolution / quality viewport for a short period while the system fetches the high-resolution / quality MCTS for the new viewport. There is a delay between when the user turns their head to look at another viewport and when the higher-resolution / quality representation of the viewport is seen. This delay depends on how quickly the system can fetch the high-resolution / quality MCTS for the new viewport, and further, how quickly it can be fetched depends on the IR AP period. The IRAP period is the interval between the occurrences of two IRAPs. This delay is related to the IRAP period because the MCTS for the new viewport may only be decodable from the IRAP picture. For example, if the IRAP period is coded at 1 second intervals, then the following applies. The best-case scenario for the delay is the network round-trip delay when the user turns their head to look at a new viewport just before the system starts fetching the new segment / IRAP period. In this scenario, the system fetches the high-resolution / quality MCTS for the new viewport. fetch the high-resolution / quality MCTS for the new viewport, and further, how quickly it can be fetched depends on the IR AP period. The IRAP period is the interval between the occurrences of two IRAPs. This delay is related to the IRAP period because the MCTS for the new viewport may only be decodable from the IRAP picture. For example, if the IRAP period is coded at 1 second intervals, then the following applies. The best-case scenario for the delay is the network round-trip delay when the user turns their head to look at a new viewport just before the system starts fetching the new segment / IRAP period. In this scenario, the system fetches the high-resolution / quality MCTS for the new viewport. related to the IRAP period.
[0105] For example, if the IRAP period is coded at 1 second intervals, then the following applies. The best-case scenario for the delay is the network round-trip delay when the user turns their head to look at a new viewport just before the system starts fetching the new segment / IRAP period. In this scenario, the system fetches the high-resolution / quality MCTS for the new viewport. In this scenario, the system fetches the high-resolution / quality MCTS for the new viewport. Higher resolution / quality MCTS can be requested immediately, thus minimizing the buffer ring delay can be set to almost zero, the sensor delay is small and can be ignored Assuming that it can be ignored, the only delay is the network round-trip delay, which is the network round-trip delay is the sum of the transmission time of the MCTS requested for the fetch request delay. The network round-trip delay can be, for example, about 200 milliseconds. The worst-case scenario for delay is when the system has already made a request for the next segment and the user's head turns to view a new viewport is the IRAP period + network round-trip delay. The bitstream can be encoded using more frequent IRAP pictures so that the IRAP period is shorter to improve the worst-case scenario above, which is because this reduces the overall delay. However, this method increases the bandwidth requirement because the compression efficiency decreases becomes large.
[0106] In an exemplary implementation, sub-pictures of the same coded picture are allowed to contain different nal_unit_type values. This mechanism is described as follows A picture may be divided into sub-pictures. A sub-picture is a rectangular set of tile groups / slices starting from a tile group with a tile_ group_address equal to 0 Each sub-picture may refer to the corresponding PPS and thus may have a separate tile segmentation. The existence of sub-pictures may be indicated within the PPS. Each sub-pic ture is treated like a picture in the decoding process. Across the boundaries of sub-pictures Loop filtering may always be disabled. The width and height of the sub-picture may be specified in units of the luma CTU size. The position of the sub-picture within the picture need not be signaled, but may be derived using the following rules. The sub-picture takes the next unoccupied position in the raster scan order of the CTUs within the picture, being large enough to contain the sub-picture within the picture boundaries. The reference picture for decoding each sub-picture is generated by extracting the area located with the current sub-picture from the reference pictures in the decoded picture buffer. The area extracted is the decoded sub-picture, and thus inter prediction is performed between sub-pictures of the same size and at the same position within the picture. In such a case, allowing different values of nal_unit_type within the coded picture enables sub-pictures derived from random access pictures and sub-pictures derived from non-random access pictures to be merged into the same coded picture without significant difficulty (e.g., without modification at the VCL level). Such an advantage also applies to coding based on MCTS. Allowing different values of nal_unit_type within the coded picture may also be beneficial in other scenarios. For example, a user may view some areas of 360-degree video content more frequently than other areas. To yield a better trade-off between coding efficiency and average equivalent quality viewport switching latency in 360-degree video delivery depending on MCTS / sub-picture based viewports, other...
[0107] For areas that are better visible than the area, more frequent IRAP pictures can be coded. The viewport switching latency of equal quality is the latency experienced by the user until the presentation
[0108] quality of the second viewport reaches the presentation quality equal to that of the first viewport when switching from the first viewport to the second viewport. Another implementation uses the following solution for supporting mixed NAL unit types in pictures, including the derivation of POC and the management of reference pictures. A flag (sps_mixed_tile_groups_in_pic_flag) is present in the parameter set to specify whether pictures with mixed IRAP sub-pictures and non-IRAP sub-pictures may exist. For NAL units including IDR tile groups, a flag (poc_msb_reset_flag) is present in the corresponding tile group header to specify If it is a group, PicRefreshFlag is set to be equal to poc_ms b_reset_flag : 1. Otherwise, if the current tile group is CR A tile group, the following applies. If the current access unit is the first access unit of the coded sequence, PicRefreshFlag is set to be equal to 1 . If the access unit follows immediately after an end of sequence NAL unit or the associated variable HandleCraAsFirstPicInCvsFlag is set to be equal to 1, the current access unit is the first access unit of the coded sequence. Otherwise, PicRefreshFlag is set to be equal to 0 (for example, the current tile group does not belong to the first access unit of the bitstream and is not an IRAP tile group ).
[0109] When PicRefreshFlag is equal to 1, the value of POC MSB (PicOrderCntMsb) is reset to be equal to 0 during the derivation of the POC for the picture. Information used for reference picture management such as the reference picture set (RPS) or reference
[0109]
[0109] picture list (RPL) is signaled within the tile group / slice header regardless of the corresponding NAL unit type. The reference picture list is constructed at the start of decoding of each tile group regardless of the NAL unit type . It can be. The reference picture list may include RefPicList
[0000] and RefPicList
[0001] for the RPL method, RefPicList0[ ] and RefPicList1[ ] for the RPS method, or a similar list including reference pictures for the inter-prediction operation related to the picture. When PicRefreshFl ag is equal to 1, during the marking process of the reference pictures, all reference pi ctures in the DPB are marked as not being used for reference.
[0110] Such an implementation is associated with specific problems. For example, when the mixing of nal_unit_t ype values within a picture is not allowed, and when the derivation of whether a picture is an IRAP picture and the derivation of the variable NoRaslOutputFlag are described at the picture level, the decoder can perform these derivations after receiving the first VCL NAL unit of any picture. However, due to the support for mixed NAL unit types within a picture, the decoder has to wait for the arrival of other VCL NAL units before performing the above derivations. In the worst case, the decoder has to wait for the arrival of the last VCL NAL unit of the picture. Furthermore, such a system may signal a flag in the tile group header of the IDR NAL unit to specify whether the POC MSB is reset during the derivation of the POC for a picture. This mechanism has the following problems. In the case of mixed CRA NAL unit types and non-IRAP NAL unit types, this mechanism is not supported. Additionally, this information can be sig naled within the tile group / slice header of the VCL NAL unit. Moreover, such a system may signal a flag in the tile group header of the IDR NAL unit to specify whether the POC MSB is reset during the derivation of the POC for a picture. This mechanism has the following problems. In the case of mixed CRA NAL unit types and non-IRAP NAL unit types, this mechanism is not supported. Additionally, this information can be sig naled within the tile group / slice header of the VCL NAL unit. This mechanism has the following problems. For mixed CRA NAL unit types and non-IRAP NAL unit types, this mechanism is not supported. Furthermore, this information can be sig naled within the tile group / slice header of the VCL NAL unit. Gnaling requires that when there is a change in the status of whether an IRAP (IDR or CRA) NAL unit is mixed with non-IRAP NAL units within a picture, the value be changed during bitstream extraction or merging. Such rewriting of the slice header occurs whenever the user requests a video and thus requires significant hardware resources. Further, some other mixtures of different NAL unit types within a picture other than the mixing of specific IRAP NAL unit types and specific non-IRAP NAL unit types are allowed. Such flexibility does not provide support for practical use cases, on the one hand, while they complicate codec design, which unnecessarily increases decoder complexity and thus raises the cost of related implementations. When there is a change in the status of whether an IRAP (IDR or CRA) NAL unit is mixed with non-IRAP NAL units within a picture, it is necessary for the value to be changed during bitstream extraction or merging. Such rewriting of the slice header occurs whenever the user requests a video and thus requires significant hardware resources. When there is a change in the status of whether an IRAP (IDR or CRA) NAL unit is mixed with non-IRAP NAL units within a picture, it is necessary for the value to be changed during bitstream extraction or merging. Such rewriting of the slice header occurs whenever the user requests a video and thus requires significant hardware resources. Such rewriting of the slice header occurs whenever the user requests a video and thus requires significant hardware resources. In addition to the mixing of specific IRAP NAL unit types and specific non-IRAP NAL unit types, some other mixtures of different NAL unit types within a picture are allowed. In addition to the mixing of specific IRAP NAL unit types and specific non-IRAP NAL unit types, some other mixtures of different NAL unit types within a picture are allowed. Such flexibility does not support practical use cases, on the one hand, while they complicate codec design, which unnecessarily increases decoder complexity and thus raises the cost of related implementations. Such flexibility does not support practical use cases, on the one hand, while they complicate codec design, which unnecessarily increases decoder complexity and thus raises the cost of related implementations.
[0111] Generally, the present disclosure describes techniques for supporting sub-picture or MCTS-based random access in video coding. More specifically, the present disclosure describes an improved design for supporting the support of mixed NAL unit types within a picture that is used to support sub-picture or MCTS-based random access. The description of the technology is based on the VVC standard but is also applicable to other video / media codec specifications. Generally, the present disclosure describes techniques for supporting sub-picture or MCTS-based random access in video coding. More specifically, the present disclosure describes an improved design for supporting the support of mixed NAL unit types within a picture that is used to support sub-picture or MCTS-based random access. More specifically, the present disclosure describes an improved design for supporting the support of mixed NAL unit types within a picture that is used to support sub-picture or MCTS-based random access. The description of the technology is based on the VVC standard but is also applicable to other video / media codec specifications.
[0112] To solve the above problems, the following exemplary implementations are disclosed. Such implementations can be applied individually or in combination. In one example, each picture is associated with an indication of whether the picture contains a value of mixed nal_unit_type. This indication To solve the above problems, the following exemplary implementations are disclosed. Such implementations can be applied individually or in combination. In one example, each picture is associated with an indication of whether the picture contains a value of mixed nal_unit_type. The indication is signaled within the PPS. This indication supports the determination of whether the POC MSB should be reset and / or whether the DPB should be reset by marking all references pictures that are not used for reference purposes. When the indication is signaled within the PPS, the change of the value within the PPS may be done during merging or separate extraction. However, this is acceptable when the PPS is rewritten and replaced by some other mechanism during the extraction or merging of such bitstreams. Alternatively, this indication is signaled within the tile group header, but may be required to be the same for all tile groups of a picture. However, in this case, the value may need to be changed during the extraction of the sub-bitstream of the MCTS / sub-picture sequence. Alternatively, this indication is signaled within the NAL unit header, but may be required to be the same for all tile groups of a picture. However, in this case, the value may need to be changed during the extraction of the sub-bitstream of the MCTS / sub-picture sequence. Alternatively, this indication may be signaled by defining an additional VCL NAL unit type such that all VCL NAL units of a picture have the same value of the NAL unit type when used for the picture. However, in this case, the value of the NAL unit type of the VCL NAL unit may need to be changed during the extraction of the sub-bitstream of the MCTS / sub-picture sequence.
[0113] header, but may be required to be the same for all tile groups of a picture. However, in this case, the value may need to be changed during the extraction of the sub-bitstream of the MCTS / sub-picture sequence. Alternatively, this indication may be signaled by defining an additional VCL NAL unit type such that all VCL NAL units of a picture have the same value of the NAL unit type when used for the picture. However, in this case, the value of the NAL unit type of the VCL NAL unit may need to be changed during the extraction of the sub-bitstream of the MCTS / sub-picture sequence. header, but may be required to be the same for all tile groups of a picture. However, in this case, the value may need to be changed during the extraction of the sub-bitstream of the MCTS / sub-picture sequence. Alternatively, this indication may be signaled by defining an additional VCL NAL unit type such that all VCL NAL units of a picture have the same value of the NAL unit type when used for the picture. However, in this case, the value of the NAL unit type of the VCL NAL unit may need to be changed during the extraction of the sub-bitstream of the MCTS / sub-picture sequence. header, but may be required to be the same for all tile groups of a picture. However, in this case, the value may need to be changed during the extraction of the sub-bitstream of the MCTS / sub-picture sequence. Alternatively, this indication may be signaled by defining an additional VCL NAL unit type such that all VCL NAL units of a picture have the same value of the NAL unit type when used for the picture. However, in this case, the value of the NAL unit type of the VCL NAL unit may need to be changed during the extraction of the sub-bitstream of the MCTS / sub-picture sequence. header, but may be required to be the same for all tile groups of a picture. However, in this case, the value may need to be changed during the extraction of the sub-bitstream of the MCTS / sub-picture sequence. Alternatively, this indication may be signaled by defining an additional VCL NAL unit type such that all VCL NAL units of a picture have the same value of the NAL unit type when used for the picture. However, in this case, the value of the NAL unit type of the VCL NAL unit may need to be changed during the extraction of the sub-bitstream of the MCTS / sub-picture sequence. header, but may be required to be the same for all tile groups of a picture. However, in this case, the value may need to be changed during the extraction of the sub-bitstream of the MCTS / sub-picture sequence. Alternatively, this indication may be signaled by defining an additional VCL NAL unit type such that all VCL NAL units of a picture have the same value of the NAL unit type when used for the picture. However, in this case, the value of the NAL unit type of the VCL NAL unit may need to be changed during the extraction of the sub-bitstream of the MCTS / sub-picture sequence. header, but may be required to be the same for all tile groups of a picture. However, in this case, the value may need to be changed during the extraction of the sub-bitstream of the MCTS / sub-picture sequence. There may be a need to be changed in the middle. Alternatively, this indication is used for pictures when all VCL NAL units of the picture have the same NAL unit type value for pictures by defining additional IRAP VCL NAL unit types. It may be signaled However, in this case, the value of the NAL unit type of the VCL NAL unit needs to be changed during the extraction of the sub-bitstream of the MCTS / sub-picture sequence There may be a need. Alternatively, each picture having at least one VCL N AL unit of any of the IRAP NAL unit types may be associated with an indication of whether the picture contains a value of a mixed NAL unit type.
[0114] Furthermore, restrictions may be applied to allow mixing of the values of nal_unit_type within a picture in a limited way by allowing only mixed IRAP NAL unit types and non-IRAP NAL unit types. For any particular picture, all VCL NAL units either have the same NAL unit type, or some VCL NAL units have a particular IRAP N AL unit type and the rest have a particular non-IRAP VCL NAL unit type. In other words, the VCL NAL units of any particular picture cannot have two or more IRAP NAL unit types and cannot have two or more non-IRAP NAL unit types. A picture may be considered an IRAP picture only if it does not contain a value of a mixed nal_unit_type and the VCL NAL units have an IRAP NAL unit type. units do not contain a value of a mixed nal_unit_type and the VCL NAL units have an IRAP NAL unit type. For all IRAP NAL units that do not belong to the IRAP picture (including IDR), the POC MSB may not be reset. For all IRAP NAL units that do not belong to the IRAP picture (including IDR), the DPB is not reset, and thus, marking all reference pictures as not used for reference is not performed. The TemporalId may be set to 0 for a picture if at least one VCL NAL unit of the picture is an IRAP NAL unit.
[0115] The following are one or more specific implementations of the above aspects. An IRAP picture may be defined as a coded picture where the value of mixed_nalu_types_in_pic_flag is equal to 0 and each VCL NAL unit has a nal_unit_type within the range from IDR_W_RADL to RSV_IRAP_VCL13, including IDR_W_RADL and RSV_IRAP_ VCL13. Exemplary PPS syntax and semantics are as follows. [Table 2] mixed_nalu_types_in_pic_flag is set to 0 to specify that each picture referring to the PPS has multiple VCL NAL units and these NAL units do not have the same value of nal_unit_type. mixed_nalu_types_in_pic_flag is set to 0 to specify that the VCL NAL units of each picture referring to the PPS have the same value of nal_unit_type.
[0116] The syntax of an exemplary tile group / slice header is as follows.
Table 3A
Table 3B
[0117] The semantics of an exemplary NAL unit header are as follows. For any particular pic ture's VCL NAL units, one of the following two conditions is satisfied. Namely, All VCL NAL units have the same value of nal_unit_type. Some of the VCL NAL units have a value of a particular IRAP NAL unit type (i.e., values of nal_unit_type in the range from IDR_W_RADL to RSV_IRAP_VCL13 including IDR_W_RADL and RSV_IRAP_VCL13), while all other VCL NAL units have a value of a particular non-IRAP VCL NAL unit type (i.e., values of nal_unit_type in the range from TRAIL_NUT to RSV_VCL_7 including TRAIL_NUT and RSV_VCL_7 or in the range from RSV_VCL14 to RSV_VCL15 including RSV_VCL14 and RSV_VCL15). The value obtained by subtracting 1 from nuh_temporal_id_plus1 specifies the temporal identifier for the NAL unit. The value of nuh_temporal_id_plus1 is not equal to 0.
[0118] The variable TemporalId is derived as follows. TemporalId = nuh_temporal_id_plus1 - 1 (7-1)
[0119] When nal_unit_type is within the range from IDR_W_RADL to RSV_IRAP_VCL13 for the VCL NAL unit of a picture and includes IDR_W_RADL and RSV_IRAP_VCL13, regardless of the value of nal_unit_type of other VCL NAL units of the picture, TemporalId is equal to 0 for all VCL NAL units of the picture. The value of TemporalId is the same for all VCL NA L units of the access unit. The value of TemporalId of a coded picture or access unit is the value of TemporalId of the VCL NAL unit of the coded picture or access unit.
[0120] <� An exemplary decoding process for a coded picture is as follows. The decoding process operates as follows for the current picture CurrPic. The decoding of the NAL unit is shown in detail herein. The following decoding process uses the syntax elements of the layer of the tile group header and the layers above it. Variables and functions related to the picture order count are derived as shown in detail herein. This is called only for the first tile group / slice of the picture. At the beginning of the decoding process for each tile group / slice, the decoding process for constructing the reference picture list is called for the derivation of reference picture list 0 (RefPicList
[0000] ) and reference picture list 1 (RefPicList
[0001] ). When the current picture is an IDR picture The decoding process operates as follows for the current picture CurrPic. The decoding of the NAL unit is shown in detail herein. The following decoding process uses the syntax elements of the layer of the tile group header and the layers above it. Variables and functions related to the picture order count are derived as shown in detail herein. This is called only for the first tile group / slice of the picture. At the beginning of the decoding process for each tile group / slice, the decoding process for constructing the reference picture list is called for the derivation of reference picture list 0 (RefPicList
[0000] ) and reference picture list 1 (RefPicList
[0001] ). When the current picture is an IDR picture is shown in detail herein. The following decoding process uses the syntax elements of the layer of the tile group header and the layers above it. Variables and functions related to the picture order count are derived as shown in detail herein. This is called only for the first tile group / slice of the picture. At the beginning of the decoding process for each tile group / slice, the decoding process for constructing the reference picture list is called for the derivation of reference picture list 0 (RefPicList
[0000] ) and reference picture list 1 (RefPicList
[0001] ). When the current picture is an IDR picture and functions are derived as shown in detail herein. This is called only for the first tile group / slice of the picture. At the beginning of the decoding process for each tile group / slice, the decoding process for constructing the reference picture list is called for the derivation of reference picture list 0 (RefPicList
[0000] ) and reference picture list 1 (RefPicList
[0001] ). When the current picture is an IDR picture At the beginning of the decoding process for each tile group / slice, the decoding process for constructing the reference picture list is called for the derivation of reference picture list 0 (RefPicList
[0000] ) and reference picture list 1 (RefPicList
[0001] ). When the current picture is an IDR picture is called for the derivation of reference picture list 0 (RefPicList
[0000] ) and reference picture list 1 (RefPicList
[0001] ). When the current picture is an IDR picture It should be noted that there seems to be an incorrect "�" in the original text at line 18 which has been translated as "�" in the above translation. Please check and correct if necessary.In this case, a decoding process for constructing a reference picture list may then be called for the purpose of checking the compliance of the bitstream but may not be necessary for decoding the current picture or a picture following the current picture in decoding order.
[0121] The decoding process for constructing a reference picture list is as follows. This process is called at the beginning of the decoding process for each tile group. The reference pictures are addressed by a reference index. The reference index is an index into the reference picture list. When decoding an I tile group, the reference picture list is not used in decoding the tile group data. When decoding a P tile group, only the reference picture list 0 (RefPicList
[0000] ) is used in decoding the tile group data. When decoding a B tile group, both the reference picture list 0 and the reference picture list 1 (RefPicList
[0001] ) are used in decoding the tile group data. At the beginning of the decoding process for each tile group, the reference picture lists RefPicList
[0000] and RefPicList
[0001] are derived. The reference picture lists are used in the marking of reference pictures or in decoding the tile group data. For all tile groups of an IDR picture or for I tile groups of a non-IDR picture, RefPicList
[0000] and RefPicList
[0001] may be derived for the purpose of checking the compliance of the bitstream but their derivation is not necessary for decoding the current picture or a picture following the current picture in decoding order. It is not necessary for decoding. For the P tile group, RefPicList
[0001] may be derived for the purpose of checking the conformity of the picture group, but the derivation is not necessary for decoding the current picture or the picture following the current picture in the decoding order. It may be derived for the purpose of checking the conformity of the picture group, but the derivation is not necessary for decoding the current picture or the picture following the current picture in the decoding order. It is not necessary for decoding the current picture or the picture following the current picture in the decoding order.
[0122] FIG. 8 is a schematic diagram of an exemplary video coding device 800. The video coding device 800 is suitable for implementing the examples / embodiments disclosed as described herein. The video coding device 800 includes a transceiver unit (Tx / Rx) 810 including a downstream port 820, an upstream port 850, and / or a transmitter and / or receiver for transmitting data upstream and / or downstream via a network. The video coding device 800 further includes a processor 830 including a logic unit and / or a central processing unit (CPU) for processing data, and a memory 832 for storing data. The video coding device 800 may also include electrical, optical, or wireless communication components coupled to the upstream port 850 and / or the downstream port 820 for communication of data via an electrical, optical-electrical (OE) network, an electro-optical (EO) network, and / or a wireless communication network. The video coding device 800 may also include an input and / or output (I / O) device 860 for transmitting data to and from a user. The I / O device 860 may include output devices such as a display for displaying video data, a speaker for outputting audio data, etc. It is suitable for implementing the examples / embodiments disclosed as described herein. The video coding device 800 includes a transceiver unit (Tx / Rx) 810 including a downstream port 820, an upstream port 850, and / or a transmitter and / or receiver for transmitting data upstream and / or downstream via a network. The video coding device 800 includes a transceiver unit (Tx / Rx) 810 including a downstream port 820, an upstream port 850, and / or a transmitter and / or receiver for transmitting data upstream and / or downstream via a network. The video coding device 800 includes a transceiver unit (Tx / Rx) 810 including a downstream port 820, an upstream port 850, and / or a transmitter and / or receiver for transmitting data upstream and / or downstream via a network. The video coding device 800 further includes a processor 830 including a logic unit and / or a central processing unit (CPU) for processing data, and a memory 832 for storing data. The video coding device 800 further includes a processor 830 including a logic unit and / or a central processing unit (CPU) for processing data, and a memory 832 for storing data. The video coding device 800 may also include electrical, optical, or wireless communication components coupled to the upstream port 850 and / or the downstream port 820 for communication of data via an electrical, optical-electrical (OE) network, an electro-optical (EO) network, and / or a wireless communication network. The video coding device 800 may also include electrical, optical, or wireless communication components coupled to the upstream port 850 and / or the downstream port 820 for communication of data via an electrical, optical-electrical (OE) network, an electro-optical (EO) network, and / or a wireless communication network. The video coding device 800 may also include electrical, optical, or wireless communication components coupled to the upstream port 850 and / or the downstream port 820 for communication of data via an electrical, optical-electrical (OE) network, an electro-optical (EO) network, and / or a wireless communication network. The video coding device 800 may also include electrical, optical, or wireless communication components coupled to the upstream port 850 and / or the downstream port 820 for communication of data via an electrical, optical-electrical (OE) network, an electro-optical (EO) network, and / or a wireless communication network. The video coding device 800 may also include an input and / or output (I / O) device 860 for transmitting data to and from a user. The video coding device 800 may also include an input and / or output (I / O) device 860 for transmitting data to and from a user. The I / O device 860 may include output devices such as a display for displaying video data, a speaker for outputting audio data, etc. It may be included. The I / O device 860 includes input devices such as a keyboard, a mouse, a trackball, etc., and / or corresponding interfaces for interacting with such output devices may also be included.
[0123] The processor 830 is implemented by hardware and software. The processor 8 30 may be implemented as one or more CPU chips, cores (e.g., as a multi-core processor), field programmable gate arrays (FPGAs), application-specific integrated circuits (ASICs), and digital signal processors (DSPs). The processor 830 communicates with the downstream port 82 0, Tx / Rx 810, the upstream port 850, and the memory 832. The processor 83 0 includes a coding module 814. The coding module 814 may use the CVS 500, VR pixel video stream 600, and / or the bitstream 700 to implement the disclosed embodiments described herein, such as methods 100, 900, and 1000. The coding module 814 may implement any other method / mechanism described herein. Furthermore, the coding module 814 may implement the codec system 200, the encoder 300, and / or the decoder 400. For example, the coding module 814 may indicate when a picture includes both IRAP NAL units and non-IRAP NAL units and set a flag in the PPS to constrain such a picture to include only a single type of IRAP NAL unit and a single type of non-IRAP NAL unit. <I Therefore, when coding the video data, the coding module 814 provides additional functions and / or coding efficiency to the video coding device 800. Therefore, the coding module 814 enhances the capabilities of the video coding device 800 and addresses issues specific to video coding technology. Furthermore, the coding module 814 brings about a transformation of the video coding device 800 to different states. Alternatively, the coding module 814 can be implemented as instructions stored in the memory 832 and executed by the processor 830 (e.g., as a computer program product stored in a non-transitory medium). The memory 832 includes any one or more of a disk, a tape drive, a solid state drive, a read-only memory (ROM), a random access memory (RAM), a flash memory, a ternary content-addressable memory (TCAM), a static random access memory (SRAM), etc. The memory 832 stores such a program when selected for program execution and also serves as an overflow data storage device for storing instructions and data read during program execution. Figure 9 shows a bitstream 700, such as a bitstream including a VR picture video stream 600 merged from a plurality of sub-picture video streams 601 - 603 of multiple video resolutions.
[0124]
[0125] A video sequence such as CVS 500 including pictures having NAL unit types mixed in the slice is an exemplary flowchart of method 900 for encoding. Method 900 is being executed by method 100 when the codec system 200, encoder 300, and / or video coding device such as encoder 800 may be used.
[0126] Method 900 may begin when the encoder determines to receive a video sequence including a plurality of pictures such as VR pictures and encode the video sequence into a bitstream, for example, based on user input. In step 901, the encoder determines whether the current picture includes a plurality of sub-pictures of different types. Such types may include at least one slice of a picture including a part of an IRAP sub-picture and at least one slice of a picture including a part of a non-IRAP NAL sub-picture. In step 903, the encoder encodes the slices of the sub-pictures of the picture into a plurality of VCL NAL units within the bitstream. Such VCL NAL units may include one or more IRAP NAL units and one or more non-IRAP NAL units. For example, the encoding step may include merging sub-bitstreams of different resolutions into a single bitstream for transmission to the decoder. In step 905, the encoder encodes the PPS into the bitstream and encodes a flag into the PPS within the bitstream. As a specific example, encoding the PPS may, for example In step 903, the encoder encodes the slices of the sub-pictures of the picture into a plurality of VCL NAL units within the bitstream. Such VCL NAL units may include one or more IRAP NAL units and one or more non-IRAP NAL units. For example, the encoding step may include merging sub-bitstreams of different resolutions into a single bitstream for transmission to the decoder. In step 905, the encoder encodes the PPS into the bitstream and encodes a flag into the PPS within the bitstream. As a specific example, encoding the PPS may, for example include merging sub-bitstreams of different resolutions into a single bitstream for transmission to the decoder.
[0127] In step 905, the encoder encodes the PPS into the bitstream and encodes a flag into the PPS within the bitstream. As a specific example, encoding the PPS may, for example include merging sub-bitstreams of different resolutions into a single bitstream for transmission to the decoder. If so, it may include changing the already encoded PPS to include the value of the flag. The flag may be set to a first value when the value of the NAL unit type is the same for all VCL NAL units related to the picture. Also, the flag may be set to a second value when the value of the first NAL unit type for a VCL NAL unit including one or more of the sub-pictures of the picture is different from the value of the second NAL unit type for a VCL NAL unit including one or more of the sub-pictures of the picture. For example, the value of the first NAL unit type may indicate that the picture includes an IRAP sub-picture, and the value of the second NAL unit type may indicate that the picture also includes a non-IRAP sub-picture. Further, the value of the first NAL unit type may be equal to one of IDR_W_RADL, IDR_N_LP, or CRA_NUT. Further, the value of the second NAL unit type may be equal to one of TRAIL_NUT, RADL_NUT, or RASL_NUT. As a specific example, the flag may be mixed_nalu_types_in_pic_flag. In a specific example, mixed_nalu_types_in_pic_flag may be set to 1 to specify that each picture referring to the PPS including the flag has two or more VCL NAL units. Furthermore, the flag specifies that not all VCL NAL units related to the corresponding picture have the same value of the NAL unit type (nal_unit_type). In another specific example, mixed_nal It may be included to change it as follows including the value of the flag. The flag may be set to the first value when the value of the NAL unit type is related to the picture for all VCL NAL units. Also, when the value of the first NAL unit type for a VCL NAL unit including one or more of the sub-pictures of the picture is different from the value of the second NAL unit type for a VCL NAL unit including one or more of the sub-pictures of the picture among the sub-pictures of the picture, the flag may be set to the second value. For example, the value of the first NAL unit type may indicate that the picture includes an IRAP sub-picture and the value of the second NAL unit type may indicate that the picture also includes one or more non-IRAP sub-pictures among the sub-pictures of the picture for a VCL NAL unit including one or more of the sub-pictures of the picture where the value of the second NAL unit type for a VCL NAL unit including one or more of the sub-pictures of the picture is different from the value of the first NAL unit type for a VCL NAL unit including one or more of the sub-pictures of the picture, it may be set to the second value . For example, the value of the first NAL unit type may indicate that the picture includes an IRAP sub-picture and the value of the second NAL unit type may indicate that the picture also includes a non-IRAP sub-picture . Further, the value of the first NAL unit type may be equal to one of IDR_W_RADL, IDR_N_LP, or CRA_NUT. Further, the value of the second NAL unit type may be equal to one of TRAIL_NUT, RADL_NUT, or RASL_NUT. As a specific example, the flag may be mixed_nalu_types_in_pic_flag. In a specific example, mixed_nalu _types_in_pic_flag may be set to 1 to specify that each picture referring to the PPS including the flag has two or more VCL NAL units . Further, the flag may specify that not all VCL NAL units related to the corresponding picture have the same value of the NAL unit type (nal_unit_type). In another specific example, mixed_nal It may be set to be equal to 1 to specify that each picture referring to the PPS including the flag has two or more VCL NAL units. Furthermore, the flag specifies that not all VCL NAL units related to the corresponding picture have the same value of the NAL unit type (nal_unit_type). In another specific example, mixed_nal It may be specified that not all VCL NAL units related to the corresponding picture have the same value of the NAL unit type (nal_unit_type). In another specific example, mixed_nal The u_types_in_pic_flag may be set to be equal to 0 to specify that each picture referring to the PPS containing the flag has one or more VCL NAL units. Further, the flag specifies that all VCL NAL units of the corresponding picture have the same value of nal_unit_type.
[0128] In step 907, the encoder may store the bitstream for transmission to the decoder.
[0129] FIG. 10 is a flowchart of an exemplary method 1000 for decoding a video sequence, such as a CVS 500, that includes pictures having a mixed NAL unit type from a bitstream, such as a bitstream 700, that includes a VR picture video stream 600 combined from a plurality of sub-picture video streams 601-603 of a plurality of video resolutions. Method 1000 may be used by a decoder, such as codec system 200, decoder 400, and / or video coding device 800, when executing method 100. Method 1000 may begin, for example, when the decoder begins to receive a bitstream of coded data representing a video sequence as a result of method 900. In step 1001, the decoder receives the bitstream. The bitstream includes a plurality of sub-pictures and flags related to the picture. In a particular example, the bitstream may include a PPS containing a flag. Further, the sub-pictures are included in a plurality of VCL NAL units. For example, a slice related to the sub-picture is in the VCL NAL unit
[0130] Method 1000 may begin, for example, when the decoder begins to receive a bitstream of coded data representing a video sequence as a result of method 900. In step 1001, the decoder receives the bitstream. The bitstream includes a plurality of sub-pictures and flags related to the picture. In a particular example, the bitstream may include a PPS containing a flag. Further, the sub-pictures are included in a plurality of VCL NAL units. For example, a slice related to the sub-picture is in the VCL NAL unit is included.
[0131] In step 1003, when the flag is set to the first value, the decoder determines that the value of the NAL unit type is the same for all VCL NAL units related to the picture. Furthermore, when the flag is set to the second value, the decoder determines that the value of the first NAL unit type for VCL NAL units including one or more of the sub-pictures of the picture is different from the value of the second NAL unit type for VCL NAL units including one or more of the sub-pictures of the picture. For example, the value of the first NAL unit type may indicate that the picture includes an IRAP sub-picture, and the value of the second NAL unit type may indicate that the picture also includes a non-IRAP sub-picture. Furthermore, the value of the first NAL unit type may be equal to one of IDR_W_RADL, IDR_N_LP, or CRA_NUT. Furthermore, the value of the second NAL unit type may be equal to one of TRAIL_NUT, RADL_NUT, or RASL_NUT. As a specific example, the flag may be mixed_nalu_types_in_pic_flag. The mixed_nalu_types_in_pic_flag may be set to 1 when each picture referring to the PPS has two or more VCL NAL units and the VCL NAL units do not have the same value of the NAL unit type (nal_unit_type). Also, the mixed_nalu_types_in_pic_flag may be set to 1 when each picture referring to the PPS has one or more VCL NAL units and the PPS refers to one or more VCL NAL units that include one or more of the sub-pictures of the picture and the value of the first NAL unit type for the VCL NAL units including one or more of the sub-pictures of the picture is different from the value of the second NAL unit type for the VCL NAL units including one or more of the sub-pictures of the picture. For example, the value of the first NAL unit type may indicate that the picture includes an IRAP sub-picture, and the value of the second NAL unit type may indicate that the picture also includes a non-IRAP sub-picture. Furthermore, the value of the first NAL unit type may be equal to one of IDR_W_RADL, IDR_N_LP, or CRA_NUT. Furthermore, the value of the second NAL unit type may be equal to one of TRAIL_NUT, RADL_NUT, or RASL_NUT. As a specific example, the flag may be mixed_nalu_types_in_pic_flag. The mixed_nalu_types_in_pic_flag may be set to 1 when each picture referring to the PPS has two or more VCL NAL units and the VCL NAL units do not have the same value of the NAL unit type (nal_unit_type). Also, the mixed_nalu_types_in_pic_flag may be set to 1 when each picture referring to the PPS has one or more VCL NAL units and the PPS refers to one or more VCL NAL units that include one or more of the sub-pictures of the picture and the value of the first NAL unit type for the VCL NAL units including one or more of the sub-pictures of the picture is different from the value of the second NAL unit type for the VCL NAL units including one or more of the sub-pictures of the picture. For example, the value of the first NAL unit type may indicate that the picture includes an IRAP sub-picture, and the value of the second NAL unit type may indicate that the picture also includes a non-IRAP sub-picture. Furthermore, the value of the first NAL unit type may be equal to one of IDR_W_RADL, IDR_N_LP, or CRA_NUT. Furthermore, the value of the second NAL unit type may be equal to one of TRAIL_NUT, RADL_NUT, or RASL_NUT. As a specific example, the flag may be mixed_nalu_types_in_pic_flag. The mixed_nalu_types_in_pic_flag may be set to 1 when each picture referring to the PPS has two or more VCL NAL units and the VCL NAL units do not have the same value of the NAL unit type (nal_unit_type). Also, the mixed_nalu_types_in_pic_flag may be set to 1 when each picture referring to the PPS has one or more VCL NAL units and the PPS refers to one or more VCL NAL units that include one or more of the sub-pictures of the picture and the value of the first NAL unit type for the VCL NAL units including one or more of the sub-pictures of the picture is different from the value of the second NAL unit type for the VCL NAL units including one or more of the sub-pictures of the picture. For example, the value of the first NAL unit type may indicate that the picture includes an IRAP sub-picture, and the value of the second NAL unit type may indicate that the picture also includes a non-IRAP sub-picture. Furthermore, the value of the first NAL unit type may be equal to one of IDR_W_RADL, IDR_N_LP, or CRA_NUT. Furthermore, the value of the second NAL unit type may be equal to one of TRAIL_NUT, RADL_NUT, or RASL_NUT. As a specific example, the flag may be mixed_nalu_types_in_pic_flag. The mixed_nalu_types_in_pic_flag may be set to 1 when each picture referring to the PPS has two or more VCL NAL units and the VCL NAL units do not have the same value of the NAL unit type (nal_unit_type). Also, the mixed_nalu_types_in_pic_flag may be set to 1 when each picture referring to the PPS has one or more VCL NAL units and the PPS refers to one or more VCL NAL units that include one or more of the sub-pictures of the picture and the value of the first NAL unit type for the VCL NAL units including one or more of the sub-pictures of the picture is different from the value of the second NAL unit type for the VCL NAL units including one or more of the sub-pictures of the picture. For example, the value of the first NAL unit type may indicate that the picture includes an IRAP sub-picture, and the value of the second NAL unit type may indicate that the picture also includes a non-IRAP sub-picture. Furthermore, the value of the first NAL unit type may be equal to one of IDR_W_RADL, IDR_N_LP, or CRA_NUT. Furthermore, the value of the second NAL unit type may be equal to one of TRAIL_NUT, RADL_NUT, or RASL_NUT. As a specific example, the flag may be mixed_nalu_types_in_pic_flag. The mixed_nalu_types_in_pic_flag may be set to 1 when each picture referring to the PPS has two or more VCL NAL units and the VCL NAL units do not have the same value of the NAL unit type (nal_unit_type). Also, the mixed_nalu_types_in_pic_flag may be set to 1 when each picture referring to the PPS has one or more VCL NAL units and the PPS refers to When the VCL NAL units of each picture are equal to the same value of nal_unit_type, it may be set to 0. It may be set to.
[0132] In step 1005, the decoder may decode one or more of the subpictures based on the value of the NAL unit type. Also, in step 1007, the decoder may transfer one or more of the subpictures for display as part of the decoded video sequence. It may be decoded. It may be transferred. It may be.
[0133] FIG. 11 is a schematic diagram of an exemplary system 1100 for coding a video sequence such as a CVS 500 that includes pictures having NAL unit types mixed in a bitstream such as a bitstream 700 that includes a VR picture video stream 600 merged from a plurality of subpicture video streams 601-603 of a plurality of video resolutions. The system 1100 may be implemented by an encoder and a decoder such as an encoder-decoder system 200, an encoder 300, a decoder 400, and / or a video coding device 800. Further, the system 1100 may be used when implementing methods 100, 900, and / or 1000. It may be merged. It may be mixed. It is a schematic diagram. It may be implemented by. Further, the system 1100 may be used. It may be used when implementing.
[0134] The system 1100 includes a video encoder 1102. The video encoder 1102 includes a determination module 1101 for determining whether a picture includes a plurality of different types of subpictures. The video encoder 1102 further includes an encoding module 1103 for encoding the subpictures of the picture into a plurality of VCL NAL units in the bitstream. Further, the symbol It may be determined. It may be included. It may be encoded. The quantization module 1103 is for encoding a flag set to a first value when the values of the NAL unit type are the same for all VCL NAL units related to a picture, and set to a second value when the value of the first NAL unit type for a VCL NAL unit including one or more of the sub-pictures of the picture is different from the value of the second NAL unit type for a VCL NAL unit including one or more of the sub-pictures of the picture into a bitstream. The video encoder 1102 further includes a storage module 1105 for storing the bitstream for transmission to the decoder. The video encoder 1102 further includes a transmission module 1107 for transmitting the bitstream to the video decoder 1110. The video encoder 1102 may be further configured to execute any of the steps of method 900. when the values of the NAL unit type are the same for all VCL NAL units related to a picture, and set to a second value when the value of the first NAL unit type for a VCL NAL unit including one or more of the sub-pictures of the picture is different from the value of the second NAL unit type for a VCL NAL unit including one or more of the sub-pictures of the picture into a bitstream. The video encoder 1102 further includes a storage module 1105 for storing the bitstream for transmission to the decoder. The video encoder 1102 further includes a transmission module 1107 for transmitting the bitstream to the video decoder 1110. The video encoder 1102 may be further configured to execute any of the steps of method 900. when the value of the first NAL unit type for a VCL NAL unit including one or more of the sub-pictures of the picture is different from the value of the second NAL unit type for a VCL NAL unit including one or more of the sub-pictures of the picture into a bitstream. The video encoder 1102 further includes a storage module 1105 for storing the bitstream for transmission to the decoder. The video encoder 1102 further includes a transmission module 1107 for transmitting the bitstream to the video decoder 1110. The video encoder 1102 may be further configured to execute any of the steps of method 900. from the value of the second NAL unit type for a VCL NAL unit including one or more of the sub-pictures of the picture into a bitstream. The video encoder 1102 further includes a storage module 1105 for storing the bitstream for transmission to the decoder. The video encoder 1102 further includes a transmission module 1107 for transmitting the bitstream to the video decoder 1110. The video encoder 1102 may be further configured to execute any of the steps of method 900. from the value of the second NAL unit type for a VCL NAL unit including one or more of the sub-pictures of the picture into a bitstream. The video encoder 1102 further includes a storage module 1105 for storing the bitstream for transmission to the decoder. The video encoder 1102 further includes a transmission module 1107 for transmitting the bitstream to the video decoder 1110. The video encoder 1102 may be further configured to execute any of the steps of method 900. The video encoder 1102 further includes a storage module 1105 for storing the bitstream for transmission to the decoder. The video encoder 1102 further includes a transmission module 1107 for transmitting the bitstream to the video decoder 1110. The video encoder 1102 may be further configured to execute any of the steps of method 900. The video encoder 1102 further includes a transmission module 1107 for transmitting the bitstream to the video decoder 1110. The video encoder 1102 may be further configured to execute any of the steps of method 900. The video encoder 1102 further includes a transmission module 1107 for transmitting the bitstream to the video decoder 1110. The video encoder 1102 may be further configured to execute any of the steps of method 900. The video encoder 1102 may be further configured to execute any of the steps of method 900.
[0135] The system 1100 also includes a video decoder 1110. The video decoder 1110 includes a receiving module 1111 for receiving a bitstream including a plurality of sub-pictures and flags related to a picture. The sub-pictures are included in a plurality of VCL NAL units. The video decoder 1110 further includes a determination module 1113 for determining that the values of the NAL unit type are the same for all VCL NAL units related to a picture when the flag is set to the first value. The video decoder 1110 includes a receiving module 1111 for receiving a bitstream including a plurality of sub-pictures and flags related to a picture. The sub-pictures are included in a plurality of VCL NAL units. The video decoder 1110 further includes a determination module 1113 for determining that the values of the NAL unit type are the same for all VCL NAL units related to a picture when the flag is set to the first value. The sub-pictures are included in a plurality of VCL NAL units. The video decoder 1110 further includes a determination module 1113 for determining that the values of the NAL unit type are the same for all VCL NAL units related to a picture when the flag is set to the first value. The video decoder 1110 further includes a determination module 1113 for determining that the values of the NAL unit type are the same for all VCL NAL units related to a picture when the flag is set to the first value. The video decoder 1110 further includes a determination module 1113 for determining that the values of the NAL unit type are the same for all VCL NAL units related to a picture when the flag is set to the first value. Further, the determination module 1113, when the flag is set to the second value, for a VCL NAL unit including one or more of the sub-pictures of the picture, the value of the first NAL unit type is for a VCL NAL unit including one or more of the sub-pictures of the picture, the value of the first NAL unit type is It is for determining that the value is different from the value of the second NAL unit type. Video decoder 1110 further includes a decoding module 1115 for decoding one or more of the sub-pictures based on the value of the NAL unit type. The video decoder 1110 further includes a transfer module 1117 for transferring one or more of the sub-pictures for display as part of the decoded video sequence. The video decoder 1110 may be further configured to perform any of the steps of method 1000.
[0136] The first component is directly coupled to the second component when there is no intermediate component except for the line, trace, or other medium between the first component and the second component. The first component is indirectly coupled to the second component when there is an intermediate component other than the line, trace, or other medium between the first component and the second component. The term "coupled" and its variations include both being directly coupled and being indirectly coupled. The use of the term "about" means a range including ±10% of the subsequent number unless stated otherwise.
[0137] The steps of the exemplary methods described herein are not necessarily required to be performed in the order described, and it should be understood that the order of such steps is merely exemplary. Similarly, additional steps may be included in such methods, and specific steps may be omitted or combined in methods consistent with various embodiments of the present disclosure.
[0138] Although several embodiments have been provided in this disclosure, the disclosed systems and methods may be embodied in many other specific forms without departing from the spirit or scope of this disclosure. It will be understood that these examples are to be considered illustrative and not restrictive, and the intention is not to be limited to the details provided herein. For example, various elements or components may be combined or integrated into another system, or certain features may be omitted or not implemented. It should be considered illustrative and not restrictive, and the intention should not be limited to the details provided herein. For example, various elements or components may be combined or integrated into another system, or certain features may be omitted or not implemented. For example, various elements or components may be combined or integrated into another system, or certain features may be omitted or not implemented. For example, various elements or components may be combined or integrated into another system, or certain features may be omitted or not implemented.
[0139] In addition, the technologies, systems, subsystems, and methods described as separate or distinct in various embodiments may be combined or integrated with other systems, components, technologies, or methods without departing from the scope of this disclosure. In addition, the technologies, systems, subsystems, and methods described as separate or distinct in various embodiments may be combined or integrated with other systems, components, technologies, or methods without departing from the scope of this disclosure. In addition, the technologies, systems, subsystems, and methods described as separate or distinct in various embodiments may be combined or integrated with other systems, components, technologies, or methods without departing from the scope of this disclosure. Other examples of changes, replacements, and modifications may be found by those skilled in the art and made without departing from the spirit and scope disclosed herein. Other examples of changes, replacements, and modifications may be found by those skilled in the art and made without departing from the spirit and scope disclosed herein. Other examples of changes, replacements, and modifications may be found by those skilled in the art and made without departing from the spirit and scope disclosed herein.
Description of Reference Numerals
[0140] 100 Operating Method 200 Encoding and Decoding (Codec) System 201 Segmented Video Signal 211 General Coder Control Component 213 Transformation, Scaling, and Quantization Component 215 Intra Picture Estimation Component 217 Intra Picture Prediction Component 219 Motion Compensation Component 221 Motion Estimation Component 223 Decoded Picture Buffer Component 225 In-Loop Filter Component 227 Filter control analysis component 229 Scaling and inverse transformation component 231 Header format and context adaptive binary arithmetic coding (CABAC) component Element 300 Video encoder 301 Segmented video signal 313 Transformation and quantization component 317 Intra-picture prediction component 321 Motion compensation component 323 Decoded picture buffer component 325 In-loop filter component 329 Inverse transformation and quantization component 331 Entropy coding component 400 Video decoder 417 Intra-picture prediction component 421 Motion compensation component 423 Decoded picture buffer component 425 In-loop filter component 429 Inverse transformation and quantization component 433 Entropy decoding component 500 CVS 502 IRAP picture 504 Leading picture 506 Trailing picture 508 Decoding order 510 Presentation order 600 VR picture video stream 601 Sub-picture video stream 602 Sub-picture video stream 603 Sub-picture video stream 700 Bitstream 710 Sequence parameter set (SPS) 711 Picture parameter set (PPS) 715 Slice header 720 Image data 721 Picture 723 Sub - picture 725 Slice 727 Mixed NAL unit types in picture flag (mixed_nalu_types_in_pic_flag) 730 Non - VCL NAL unit 731 SPS NAL unit type (SPS_NUT) 732 PPS NAL unit type (PPS_NUT) 740 VCL NAL unit 741 IDR_N_LP NAL unit 742 IDR_w_RADL NAL unit 743 CRA_NUT 745 IRAP NAL unit 746 RASL_NUT 747 RADL_NUT 748 TRAIL_NUT 749 Non - IRAP NAL unit 800 Video coding device 810 Transceiver unit (Tx / Rx) 814 Coding module 820 Down - stream port 830 Processor 832 Memory 850 Up - stream port 860 Input and / or output (I / O) device 900 Method 1000 Method 1100 System 1101 Decision module 1102 Video encoder 1103 Encoding module 1105 Memory module 1107 Transmission module 1110 Video decoder 1111 Reception module 1113 Decision module 1115 Decryption Module 1117 Transfer Module
Claims
Claim 1 A method implemented in a decoder, comprising: receiving a bitstream including a sequence parameter set (SPS), one or more picture parameter sets (PPS), and encoded data of a plurality of pictures, wherein the encoded data of the plurality of pictures is included in a plurality of video coding layer (VCL) network abstraction layer (NAL) units, the SPS is included in a non-VCL NAL unit having an SPS NAL unit type (SPS_NUT), the PPS is included in a non-VCL NAL unit having a PPS NAL unit type (PPS_NUT), and the PPS includes a flag; when the flag is set to a first value, determining that all of the VCL NAL units related to a picture have the same NAL unit type; when the flag is set to a second value, determining that one or more VCL NAL units of a picture have a value of a first NAL unit type and all other VCL NAL units of the picture have a value of a second NAL unit type, wherein the value of the first NAL unit type is equal to an instantaneous decoding refresh (IDR) with a random access decodable leading picture (IDR_W_RADL), an IDR without a leading picture (IDR_N_LP), or a clean random access (CRA) NAL unit type (CRA_NUT); decoding one or more of the pictures based on the flag. Claim 2. The method according to claim 1, wherein the value of the second NAL unit type is equal to a trailing picture NAL unit type (TRAIL_NUT). Claim 3 The method according to claim 1 or 2, wherein the flag is a mixed_nalu_types_in_pic_flag. Claim 4. The method according to any one of claims 1 to 3, wherein the first value is 0 and the second value is 1. Claim 5 A method implemented in an encoder, comprising: Encoding a sequence parameter set (SPS) and one or more picture parameter sets (PPS) into a bitstream, and encoding a plurality of pictures into the bitstream, wherein the encoded data of the plurality of pictures is included in a plurality of video coding layer (VCL) network abstraction layer (NAL) units, the SPS is included in a non-VCL NAL unit having an SPS NAL unit type (SPS_NUT), the PPS is included in a non-VCL NAL unit having a PPS NAL unit type (PPS_NUT), and the PPS includes a flag, the flag is set to a first value when all of the VCL NAL units related to a picture have the same NAL unit type, the flag is set to a second value when one or more VCL NAL units of a picture have a value of a first NAL unit type and all other VCL NAL units of the picture have a value of a second NAL unit type, wherein the value of the first NAL unit type is equal to an instantaneous decoding refresh (IDR) with a random access decodable leading picture (IDR_W_RADL), an IDR without a leading picture (IDR_N_LP), or a clean random access (CRA) NAL unit type (CRA_NUT). **Claim 6**: The method according to claim 5, wherein the value of the second NAL unit type is equal to a trailing picture NAL unit type (TRAIL_NUT). **Claim 7** The method according to claim 5 or 6, wherein the flag is mixed_nalu_types_in_pic_flag. **Claim 8**: The method according to any one of claims 5 to 7, wherein the first value is 0 and the second value is 1. **Claim 9** A video coding device, comprising a processor, a receiver coupled to the processor, a memory coupled to the processor, and a transmitter coupled to the processor, wherein the processor is configured to execute the method according to any one of claims 1 to 4. **Claim 10**: A video coding device, comprising a processor, a receiver coupled to the processor, a memory coupled to the processor, and a transmitter coupled to the processor, wherein the processor is configured to execute the method according to any one of claims 5 to 8. **Claim 11** A non-transitory computer-readable medium including a computer program product for use by a video coding device, the computer program product including computer-executable instructions stored on the non-transitory computer-readable medium that, when executed by a processor, cause the video coding device to execute the method according to any one of claims 1 to 4 or any one of claims 5 to 8. **Claim 12**: Receiving means for receiving a bitstream including a sequence parameter set (SPS), one or more picture parameter sets (PPS) and encoded data of a plurality of pictures, wherein the encoded data of the plurality of pictures are included in a plurality of video coding layer (VCL) network abstraction layer (NAL) units, the SPS is included in a non-VCL NAL unit having an SPS NAL unit type (SPS_NUT), the PPS is included in a non-VCL NAL unit having a PPS NAL unit type (PPS_NUT), and the PPS includes a flag; and Determining means, when the flag is set to a first value, determining that all of the VCL NAL units associated with the picture have the same NAL unit type; when the flag is set to a second value, determining that one or more VCL NAL units of the picture have a value of a first NAL unit type and all other VCL NAL units of the picture have a value of a second NAL unit type, and the value of the first NAL unit type is equal to an instantaneous decoding refresh (IDR) having a random access decodable leading picture (IDR_W_RADL), an IDR having no leading picture (IDR_N_LP), or a clean random access (CRA) NAL unit type (CRA_NUT); determining means A decoder including decoding means for decoding one or more of the pictures based on the flag.
13. The decoder according to claim 12, further configured to execute the method according to any one of claims 2 to 4.
14. An encoder, including encoding means for encoding a sequence parameter set (SPS) and one or more picture parameter sets (PPS) into a bitstream and encoding a plurality of pictures into the bitstream, wherein the encoded data of the plurality of pictures is included in a plurality of video coding layer (VCL) network abstraction layer (NAL) units, the SPS is included in a non-VCL NAL unit having an SPS NAL unit type (SPS_NUT), the PPS is included in a non-VCL NAL unit having a PPS NAL unit type (PPS_NUT), and the PPS includes a flag, wherein the flag is set to a first value when all of the VCL NAL units related to the picture have the same NAL unit type, wherein the flag is set to a second value when one or more VCL NAL units of the picture have a value of a first NAL unit type and all other VCL NAL units of the picture have a value of a second NAL unit type, wherein the value of the first NAL unit type is equal to an instantaneous decoding refresh (IDR) having a random access decodable leading picture (IDR_W_RADL), an IDR having no leading picture (IDR_N_LP), or a clean random access (CRA) NAL unit type (CRA_NUT), the encoder.
15. The encoder according to claim 14, further configured to execute the method according to any one of claims 6 to 8.
16. A device for storing a bitstream, comprising at least one storage medium and at least one communication interface, wherein the at least one communication interface is configured to receive or transmit the bitstream, wherein the at least one storage medium is configured to store the bitstream, The bitstream includes a sequence parameter set (SPS), one or more picture parameter sets (PPS), and encoded data of a plurality of pictures, the encoded data of the plurality of pictures is included in a plurality of video coding layer (VCL) network abstraction layer (NAL) units, the SPS is included in a non-VCL NAL unit having an SPS NAL unit type (SPS_NUT), the PPS is included in a non-VCL NAL unit having a PPS NAL unit type (PPS_NUT), and the PPS includes a flag. When all of the VCL NAL units related to a picture have the same NAL unit type, the flag is set to a first value. When one or more VCL NAL units of a picture have a value of a first NAL unit type and all other VCL NAL units of the picture have a value of a second NAL unit type, the flag is set to a second value. The value of the first NAL unit type is equal to an instant refresh with decode access (IDR) having a random access decodable leading picture (IDR_W_RADL), an IDR having no leading picture (IDR_N_LP), or a clean random access (CRA) NAL unit type (CRA_NUT). **Claim 17**: A method for storing a bitstream, comprising: receiving or transmitting the bitstream via a communication interface; storing the bitstream in at least one storage medium, the bitstream including a sequence parameter set (SPS), one or more picture parameter sets (PPS), and encoded data of a plurality of pictures, the encoded data of the plurality of pictures being included in a plurality of video coding layer (VCL) network abstraction layer (NAL) units, the SPS being included in a non-VCL NAL unit having an SPS NAL unit type (SPS_NUT), the PPS being included in a non-VCL NAL unit having a PPS NAL unit type (PPS_NUT), and the PPS including a flag. When all of the VCL NAL units associated with a picture have the same NAL unit type, the flag is set to a first value. When one or more VCL NAL units of a picture have a value of a first NAL unit type and all other VCL NAL units of the picture have a value of a second NAL unit type, the flag is set to a second value. The method includes steps where the value of the first NAL unit type is equal to an Instantaneous Decoding Refresh (IDR) with a Random Access Decodable Leading Picture (IDR_W_RADL), an IDR without a Leading Picture (IDR_N_LP), or a Clean Random Access (CRA) NAL unit type (CRA_NUT).
18. A device for transmitting a bitstream, at least one storage medium configured to store at least one bitstream; at least one processor configured to obtain one or more bitstreams from the at least one storage medium; a transmitter configured to transmit the one or more bitstreams, and the at least one bitstream includes a Sequence Parameter Set (SPS), one or more Picture Parameter Sets (PPS), and encoded data of a plurality of pictures, the encoded data of the plurality of pictures is included in a plurality of Video Coding Layer (VCL) Network Abstraction Layer (NAL) units, the SPS is included in a non-VCL NAL unit having an SPS NAL unit type (SPS_NUT), the PPS is included in a non-VCL NAL unit having a PPS NAL unit type (PPS_NUT), the PPS includes a flag, when all of the VCL NAL units associated with a picture have the same NAL unit type, the flag is set to a first value; when one or more VCL NAL units of a picture have a value of a first NAL unit type and all other VCL NAL units of the picture have a value of a second NAL unit type, the flag is set to a second value. A device in which the value of the first NAL unit type is equal to an Instantaneous Decoding Refresh (IDR) with a random access decodable leading picture (IDR_W_RADL), an IDR without a leading picture (IDR_N_LP), or a Clean Random Access (CRA) NAL unit type (CRA_NUT). [
19. ] A method for transmitting a bitstream, comprising: storing at least one bitstream in one or more storage media; obtaining one or more bitstreams from the one or more storage media; transmitting the one or more bitstreams, wherein the at least one bitstream includes a Sequence Parameter Set (SPS), one or more Picture Parameter Sets (PPS), and encoded data of a plurality of pictures, the encoded data of the plurality of pictures is included in a plurality of Video Coding Layer (VCL) Network Abstraction Layer (NAL) units, the SPS is included in a non-VCL NAL unit having an SPS NAL unit type (SPS_NUT), the PPS is included in a non-VCL NAL unit having a PPS NAL unit type (PPS_NUT), the PPS includes a flag, the flag is set to a first value when all of the VCL NAL units related to a picture have the same NAL unit type; the flag is set to a second value when one or more VCL NAL units of a picture have a value of a first NAL unit type and all other VCL NAL units of the picture have a value of a second NAL unit type; a method in which the value of the first NAL unit type is equal to an Instantaneous Decoding Refresh (IDR) with a random access decodable leading picture (IDR_W_RADL), an IDR without a leading picture (IDR_N_LP), or a Clean Random Access (CRA) NAL unit type (CRA_NUT). [
20. ] A system for processing a bitstream, comprising a server, a source device, one or more storage devices, and a destination device, wherein the source device is configured to obtain a video source from the server The source device is further configured to encode the video source to obtain one or more bitstreams, the source device is configured to store the one or more bitstreams in the one or more storage devices, and / or the source device is configured to transmit the one or more bitstreams to the destination device via a communication interface, the destination device is configured to decode the one or more bitstreams to obtain video data, at least one of the one or more bitstreams includes a sequence parameter set (SPS), one or more picture parameter sets (PPS), and encoded data of a plurality of pictures, the encoded data of the plurality of pictures is included in a plurality of video coding layer (VCL) network abstraction layer (NAL) units, the SPS is included in a non-VCL NAL unit having an SPS NAL unit type (SPS_NUT), the PPS is included in a non-VCL NAL unit having a PPS NAL unit type (PPS_NUT), and the PPS includes a flag, when all of the VCL NAL units related to a picture have the same NAL unit type, the flag is set to a first value, when one or more VCL NAL units of a picture have a value of a first NAL unit type and all other VCL NAL units of the picture have a value of a second NAL unit type, the flag is set to a second value, a system in which the value of the first NAL unit type is equal to an instantaneous decoding refresh (IDR) with a random access decodable leading picture (IDR_W_RADL), an IDR without a leading picture (IDR_N_LP), or a clean random access (CRA) NAL unit type (CRA_NUT).
Citation Information
Patent Citations
Video codec allowing sub-picture or region wise random access and concept for video composition using the same
WO2020157287A1