Picture with mixed NAL unit types
By dividing video pictures into sub-pictures and employing flags to distinguish IRAP and non-IRAP types, the method optimizes video coding for efficient resource use and user experience in limited bandwidth scenarios.
Patent Information
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- HUAWEI TECH CO LTD
- Filing Date
- 2025-07-16
- Publication Date
- 2026-04-27
AI Technical Summary
Existing video coding technologies face challenges in efficiently compressing video data for streaming and storage, particularly in limited bandwidth scenarios, without compromising image quality.
The method involves dividing a picture into multiple sub-pictures, encoding them into separate sub-bitstreams, and using flags to differentiate between intra-random access point (IRAP) and non-IRAP sub-pictures, allowing for dynamic resolution adjustment based on viewport likelihood, thereby improving coding efficiency and resource utilization.
This approach enhances coding efficiency by allocating resources more effectively, reducing network, memory, and processing requirements while maintaining user experience in applications like virtual reality, without significant degradation.
Smart Images

Figure 0007852128000007 
Figure 0007852128000008 
Figure 0007852128000009
Abstract
Description
Technical Field
[0002] The present disclosure generally relates to video coding, and more particularly to coding sub - pictures of a picture in video coding.
Background Art
[0003] The amount of video data required to depict even relatively short videos is quite large, which poses difficulties when it is to be streamed over a communication network having a limited bandwidth capacity or transmitted in some other way. Thus, video data is generally compressed before being transmitted over modern communication networks. Since memory resources may be limited, the size of the video can also be a problem when the video is stored on a storage device. In many cases, video compression devices use software and / or hardware at the transmitter to code the video data before transmission or storage, thereby reducing the amount of data required to represent the digital video image. The compressed data is then received at the receiver by a video decompression device that decodes the video data. [[ID=�2]] Due to limited network resources and the ever - increasing demand for higher video quality, improved compression and decompression techniques that increase the compression ratio without sacrificing much or any of the image quality are desirable.
Summary of the Invention
Means for Solving the Problems
[0004] In embodiments, the present disclosure describes a method implemented in a decoder, wherein a picture The bitstream containing multiple related subpictures and flags is sent to the decoder's receiver. Therefore, in the receiving step, the subpicture is the video coding layer (VCL: vi Network Abstraction Layer (NAL) Unit The step included in the first NAL unit is set to the first value when the flag is set to the first value. The process assumes that the value of type is the same for all VCL NAL units related to the picture. The step of determining by the sasser, and when the flag is set to the second value, the picture A first NAL unit type relating to a VCL NAL unit containing one or more subpictures The second NAL unit for a VCL NAL unit whose value contains one or more of the subpictures of the picture. A step in which the value of the nit type is determined by a different processor, and the first NAL unit Based on the value of Ip or the value of the second NAL unit type, one or more of the subpictures A method including the step of decoding by a processor.
[0005] A picture can be divided into multiple subpictures. Such subpictures can be separated. It is possible to code into sub-bitstreams of and then those sub The bitstream can be merged into a bitstream for transmission to the decoder. For example, subpictures may be used for virtual reality (VR) applications. In this example, the user may always view only a portion of the VR picture. Therefore, the displayed More bandwidth can be allocated to subpictures that are more likely to be displayed. Subpictures that are less likely to be shown are compressed to improve coding efficiency. As possible, different sub-pictures may be transmitted at different resolutions. DeoStream is an intra-random access point (IRAP). It may be encoded by using an IRAP picture. IRAP pictures are intra-prepared It is coded by measurement and can be decoded without referencing other pictures. Non-IRAP pictures may be coded by interpretation, and other pictures It can be decrypted by referencing the chat. Non-IRAP pictures are significantly more complex than IRAP pictures. It is condensed into this. However, the IRAP picture is decoded without referencing other pictures. Since it contains enough data, the video sequence begins decoding from the IRAP picture. It must be done. IRAP pictures can be used within subpictures. It may be possible to change the resolution dynamically. Therefore, the video system can (for example, you Regarding subpictures that are more likely to be seen (based on the current viewport) Send more IRAP pictures and see them to further improve coding efficiency. You may send fewer IRAP pictures for sub-pictures of lower quality. However, sub The picture is part of the same picture. Therefore, this method is IRAP subpicture and It may result in a picture that includes both non-IRAP subpictures and some video systems. This does not provide the necessary documentation for handling mixed pictures that have both IRAP and non-IRAP regions. The picture is mixed and thus includes a flag indicating whether it includes both IRAP components and non-IRAP components. Based on this flag, the decoder can appropriately decode the picture / sub-picture and process different sub-pictures differently when decoding to display them. This flag may be stored in the PPS and may be called mixed_nalu_types_in_pic_flag. Therefore, the disclosed mechanism enables the implementation of additional functions. Further, the disclosed mechanism enables dynamic resolution change when using the sub-picture bitstream. Therefore, the disclosed mechanism enables a lower-resolution sub-picture bitstream to be transmitted when streaming VR video without significantly degrading the user experience. Therefore, the disclosed mechanism improves coding efficiency and thus reduces the use of network resources, memory resources, and / or processing resources in the encoder and decoder. Optionally, in any of the above aspects, another implementation of the aspect defines that the bitstream includes a picture parameter set (PPS) that includes a flag. Optionally, in any of the above aspects, another implementation of the aspect defines that the value of the first NAL unit type indicates that the picture includes an intra-random access point (IRAP) sub-picture, and the value of the second NAL unit type indicates that the picture includes a non-IRAP sub-picture.
[0006]
[0007]
[0008] Optionally, in any of the above aspects, another implementation of the aspect is that the value of the first NAL unit type is defined to be equal to an instantaneous decoding refresh (IDR: Instantaneous Decoding Refresh) (IDR_W_RADL) with a random access decodable leading picture, an IDR without a leading picture (IDR_N_LP), or a clean random access (CRA: clean random access) NAL unit type (CRA_NUT).
[0009] Optionally, in any of the above aspects, another implementation of the aspect is that the value of the second NAL unit type is defined to be equal to a trailing picture NAL unit type (TRAIL_NUT), a random access decodable leading picture NAL unit type (RADL_NUT), or a random access skipped leading picture (RASL: random access skipped leading picture) NAL unit type (RASL_NUT).
[0010] Optionally, in any of the above aspects, another implementation of the aspect is that the flag is defined as mixed_nalu_types_in_pic_flag.
[0011] Optionally, in any of the above aspects, another implementation of the aspect is that the picture referring to the PPS has two or more of the VCL NAL units, and the VCL NAL units are of the NAL unit type When specifying that there should not be the same value for `P(nal_unit_type)`, use `mixed_nalu_types_in_pic_f`. If lag is equal to 1, and the picture referencing PPS has one or more VCL NAL units, When specifying that NAL units have the same value for nal_unit_type, mixed_nalu_types_ This defines in_pic_flag as equal to 0.
[0012] In embodiments, the present disclosure describes a method implemented in an encoder, which involves picture The processor determines whether it contains multiple sub-pictures of different types. The picture and its subpictures are coded into multiple VCL NAL units within the bitstream. The steps to convert and the values of the first NAL unit type to all VCL NAL related to the picture It is set to the first value when it is the same with respect to the unit, and among the subpictures of the picture The first NAL unit type value for a VCL NAL unit containing one or more of the following is a sub-picture The second NAL unit type value for a VCL NAL unit that contains one or more pictures and The processor codes the bitstream with a flag that is set to a second value when it is different. The process involves converting the bitstream and transmitting it to the decoder, then combining the bitstream into the processor. A method including the step of storing in a memory.
[0013] A picture can be divided into multiple subpictures. Such subpictures can be separated. It is possible to code into sub-bitstreams of and then those sub The bitstream can be merged into a bitstream for transmission to the decoder. For example, subpictures may be used for virtual reality (VR) applications. In this example, the user may always view only a portion of the VR picture. Therefore, the displayed More bandwidth can be allocated to subpictures that are more likely to be displayed. Subpictures that are less likely to be shown are compressed to improve coding efficiency. As possible, different sub-pictures may be transmitted at different resolutions. DeoStream will use intra-random access point (IRAP) pictures. Therefore, it may be encoded. IRAP pictures are coded by intra prediction. It is possible to decrypt without referencing other pictures. Non-IRAP pictures are interactive. - May be coded by prediction, and can be restored by referencing other pictures. It can be numbered. Non-IRAP pictures are significantly more condensed than IRAP pictures. However, IRAP pictures The kucha contains enough data to be decrypted without referencing other pictures. Therefore, the video sequence must start decoding from the IRAP picture. The 'cha' can be used within a sub-picture and allows for dynamic resolution changes. Therefore, the video system is based on (for example, the user's current viewport) Send more IRAP pictures for sub-pictures that are more likely to be viewed. To further improve coding efficiency, regarding subpictures that are less likely to be seen You may send fewer IRAP pictures. However, subpictures are one of the same picture. This is the part. Therefore, this method handles both IRAP subpictures and non-IRAP subpictures. It may bring about pictures that include. Some video systems have an IRAP area and a non-IRAP area. There is no provision for handling mixed pictures that have both. This disclosure is for when the picture is mixed, however This includes a flag indicating whether it includes both IRAP components and non-IRAP components. Based on the flags, the decoder properly decodes and displays the picture / subpicture. This allows different subpictures to be treated differently during decryption. This flag is for PPS It may be stored in and may be called mixed_nalu_types_in_pic_flag. Therefore, open The revealed mechanism enables the implementation of additional functions. Furthermore, the disclosed mechanism This allows for dynamic resolution changes when using the bitstream of a subpicture. Therefore, the disclosed mechanism will not significantly impair the user experience. However, when streaming VR video, the bitstack of lower-resolution sub-pictures is reduced. This enables the transmission of the Ream. Therefore, the disclosed mechanism is Cody To improve network efficiency, and therefore network resources in encoders and decoders This reduces the use of memory resources and / or processing resources.
[0014] Optionally, in any of the above embodiments, another implementation of the embodiment uses the PPS bit. A step in which the flag is encoded in PPS. It also includes.
[0015] Optionally, in any of the above embodiments, another implementation of the embodiment is the first NAL unit The type value indicates that the picture contains an IRAP subpicture, and the second NAL unit type The value of 'P' indicates that the picture contains a non-IRAP subpicture.
[0016] Optionally, in any of the above embodiments, another implementation of the embodiment is the first NAL unit Specify that the value of type is equal to IDR_W_RADL, IDR_N_LP, or CRA_NUT.
[0017] Optionally, in any of the above embodiments, another implementation of the embodiment is a second NAL unit Specifies that the value of type is equal to TRAIL_NUT, RADL_NUT, or RASL_NUT.
[0018] Optionally, in any of the above embodiments, another implementation of the embodiment is if the flag is mixed_ It is defined as nalu_types_in_pic_flag.
[0019] Optionally, in any of the above embodiments, another implementation of the embodiment refers to PPS. The picture has two or more VCL NAL units, and the VCL NAL units are of type nal_unit_type When specifying that they do not have the same value, mixed_nalu_types_in_pic_flag is equal to 1, PPS The picture that references has one or more VCL NAL units, and the VCL NAL unit is nal_u When specifying that nit_type has the same value, mixed_nalu_types_in_pic_flag is equal to 0. It is defined as such.
[0020] In embodiments, the present disclosure includes a processor, a receiver coupled to the processor, and a processor It includes memory coupled to a separator and a transmitter coupled to a processor, and the processor, receiver The QR, memory, and transmitter are configured to perform any of the methods described above. This includes video coding devices.
[0021] In embodiments, the present disclosure provides a code for use by a video coding device. Non-temporary computer-readable media including computer program products, computer When the program product is executed by the processor, it is used by the video coding device. A method of any of the above embodiments, stored in a non-temporary computer-readable medium Includes non-temporary computer-readable media containing computer-executable instructions.
[0022] In embodiments, the Disclosure relates to a plurality of subpictures and flags associated with a picture. A receiving means for receiving a bitstream containing subpictures, wherein the subpictures are multiple VCLs The NAL unit includes a receiving means and a determination means, wherein the flag is set to a first value. When the value of the first NAL unit type is associated with all VCL NAL units of the picture, If it is determined that they are the same and the flag is set to the second value, the sub-picture The value of the first NAL unit type for a VCL NAL unit that contains one or more pictures A second NAL unit relating to a VCL NAL unit that includes one or more subpictures of a picture. A determination means for determining that the value of the type is different from the value of the first NAL unit type and This decodes one or more of the subpictures based on the value of the second NAL unit type. It includes a decoder that includes a decoding means.
[0023] Optionally, in any of the above embodiments, another implementation of the embodiment is that the decoder is above It is further specified that it is configured to perform any of the methods described above.
[0024] In embodiments, the disclosure may include multiple sub-pictures of different types. A determination means for determining whether or not, and an encoding means for encoding the subpicture of the picture. Encode the value of the first NAL unit type into multiple VCL NAL units within the stream. Set to the first value when it is the same for all VCL NAL units associated with the picture. The first NAL unit relating to a VCL NAL unit that includes one or more subpictures of a picture. The value of the knit type relates to a VCL NAL unit that contains one or more subpictures of the picture. When the value of the second NAL unit type is different, the flag set to the second value is bit An encoding means for encoding into a stream, and a bitstream for transmitting to a decoder. Includes an encoder that includes a storage means for storing data.
[0025] Optionally, in any of the above embodiments, another implementation of the embodiment is that the encoder is It is further specified that it is configured to perform any of the methods described above.
[0026] For the purpose of clarity, any one of the embodiments described above may be a new embodiment within the scope of this disclosure. It may be combined with any one or more of the other embodiments described above to generate stomach.
[0027] These and other features are described in detail below in conjunction with the attached drawings and claims. Understanding will lead to a clearer understanding.
[0028] To better understand this disclosure, see the attached drawings where similar reference numbers represent similar parts. The following brief explanation, which will be interpreted in relation to the more detailed explanation, is referenced here. [Brief explanation of the drawing]
[0029] [Figure 1] This is a flowchart illustrating an exemplary method for coding a video signal. [Figure 2] This is a schematic diagram illustrating an exemplary coding and decoding (codec) system for video coding. [Figure 3] This is a schematic diagram illustrating an example video encoder. [Figure 4] This is a schematic diagram illustrating an exemplary video decoder. [Figure 5] This is a schematic diagram illustrating an example of coded video sequences. [Figure 6] This is a schematic diagram showing multiple sub-picture video streams separated from a virtual reality (VR) picture video stream. [Figure 7] This is a schematic diagram showing an exemplary bitstream containing a picture with mixed network abstraction layer (NAL) unit types. [Figure 8] This is a schematic diagram of an exemplary video coding device. [Figure 9] This is a flowchart illustrating an exemplary method for encoding a video sequence containing pictures with mixed NAL unit types in a bitstream. [Figure 10] This is a flowchart illustrating an exemplary method for decoding a video sequence containing pictures with mixed NAL unit types from a bitstream. [Figure 11] This is a schematic diagram of an exemplary system for coding a video sequence containing pictures with mixed NAL unit types in a bitstream. [Modes for carrying out the invention]
[0030] One or more exemplary implementations of the embodiments are given below, but the disclosed systems and / Alternatively, the method is any number of techniques, whether currently known or existing. It should be understood from the outset that this may be implemented using techniques. This disclosure is used herein The following exemplary implementations, including the exemplary designs and implementations illustrated and explained, are shown in the diagrams below. Not to be limited in any way to surfaces and technology, but together with the entire scope of equivalents of the attached claims They may be modified within the scope of those claims.
[0031] The following acronyms stand for Coded Video Sequence (CVS): Decoded Picture Buffer (DPB), Instantaneous Decoded Refresh (IDR), Intra Random Access Port Integer (IRAP), least significant bit (LSB), most significant bit (MSB), Network Abstraction Layer (NAL) ), Picture Order Count (POC), Raw byte sequence payload ( RBSP (Raw Byte Sequence Payload), Sequence Parameter Set (SPS), and Working Grass The draft (WD) is used herein.
[0032] Many video compression techniques reduce the size of video files with minimal data loss. It can be used for purposes such as reducing the redundancy of data in a video sequence. For example, video compression techniques can reduce the redundancy of data in a video sequence. To reduce or eliminate spatial (e.g., intrapicture) prediction and temporal (i This may include performing predictions (inter-picture). Block-based video coding. Therefore, a video slice (for example, a video picture or part of a video picture) It can also be divided into video blocks, and video blocks are tree blocks, coding blocks Coding tree block (CTB), coding tree unit (CTU), coding unit (CU) , and / or may also be called coding nodes. Picture intracoding The video block within the slice being sliced (I) is referenced in neighboring blocks within the same picture. Coding is performed using spatial predictions related to the light sample. Picture intercoding The video blocks within the unidirectional prediction (P) or bidirectional prediction (B) slices being processed are the same Spatial prediction or other reference particles related to reference samples in neighboring blocks within a picture Coded by using time predictions related to reference samples within the kucha. That's fine too. A picture may also be called a frame and / or image, and a reference picture is Reference frame and / or reference image. Spatial or temporal predictions are made in the image frame. This brings up a predicted block representing the lock. The residual data is the difference between the original block and the predicted block. It represents the difference between pixels. Therefore, the intercoded block is the predicted block. A motion vector pointing to a block of reference samples that form a lock, and the coded block Encoded by residual data showing the difference between the lock and the predicted block. The blocks being coded are coded by intracoding mode and residual data. It is converted to a number. For further compression, the residual data is converted from the pixel region to the conversion region. These may be used. These result in residual transformation coefficients, which may be quantized. First, the quantized transformation coefficients may be arranged in a two-dimensional array. The coefficients may be scanned to generate a one-dimensional vector of transformation coefficients. Entropy - Coding may be applied to achieve even greater compression. The compression technique will be discussed in detail below.
[0033] To ensure that the encoded video can be decoded accurately, the video is... It is encoded and decoded according to the video coding standard. The video coding standard is, International Telecommunication Union (ITU) Standardization Sector (ITU-T) H.261, International Organization for Standardization / International Electrotechnical Commission (ISO / IEC) Video Expert Group (MPEG)-1 Part 2, ITU-T H.262 or ISO / IEC MPEG-2 Part 2 ITU-T H.263, ISO / IEC MPEG-4 Part 2, ITU-T H.264, or ISO / IEC MPEG-4 Part 10 Also known as Advanced Video Coding (AVC), ITU-T H.26 AVC includes High Efficiency Video Coding (HEVC), also known as AVC 5 or MPEG-H Part 2. Scalable Video Coding (SVC), multi-view video coding Multiview Video Coding (MVC) and multiview video coding plus depth (MVC+D): Includes extensions such as Multiview Video Coding plus Depth, and 3D AVC (3D-AVC). Yes. HEVC includes extensible HEVC (SHVC), multi-view HEVC (MV-HEVC), and 3D HEVC (3D-HEVC). Includes any extensions. ITU-T and ISO / IEC Joint Video Expert Team (JVET: Joint Video Expansion) The erts team is developing a video coding technology called Versatile Video Coding (VVC). Development of the decoding standard has begun. VVC includes an algorithm description and a VVC Working Draft (WD). The WD includes JVET-M1001-v6, which provides a description of the encoder side and reference software. Born.
[0034] The video coding system uses IRAP and non-IRAP pictures. Therefore, video may be encoded. IRAP picture is random about the video sequence. It is a picture coded by intra-prediction that acts as an access point. In intra prediction, a block of picture is compared with other blocks within the same picture. Coding by reference. This is a non-IRAP picture that uses interpretation and This is in contrast. In inter prediction, the current picture block is the current picture Coding is done by references to other blocks within a different reference picture. AP pictures are coded without any reference to other pictures, so first any Other pictures can also be decrypted without decryption. Therefore, the decoder can decrypt any IRAP picture. In the chat, decoding of the video sequence can be initiated. In contrast, non-IRAP picture The image is coded by referencing other pictures, and therefore, generally speaking, the decoder is Furthermore, it is not possible to start decoding a video sequence in a non-IRAP picture. The AP picture refreshes the DPB. This is because the IRAP picture is the starting point for CVS, and CVS This is because the picture in S does not reference the picture in the previous CVS. Therefore, IRAP pict Furthermore, coding errors related to interpretation propagate through IRAP pictures. Since it is not possible to do so, such errors can be stopped. However, IRAP picture In terms of data size, it is significantly larger than non-IRAP pictures. Therefore, generally speaking, Deosequencing is a common non-IRAP approach to balance coding efficiency and functionality. Includes pictures and a smaller number of IRAP pictures scattered between them. For example, 60 frames. The CVS of the room may contain one IRAP picture and 59 non-IRAP pictures.
[0035] In some cases, video coding systems are also called 360-degree video. It may be used to code virtual reality (VR) videos. VR videos allow users to... The video content may include a sphere that appears as if the viewer is at the center of the sphere. Only a portion of the sphere, called the sphere, is displayed to the user. For example, if the user is Based on the user's head movements, the viewport of the sphere is selected and the head-mounted display shown... You may use a Ray (HMD). This allows you to physically exist within the virtual space depicted by the video. It gives the impression of being present. To achieve this result, each picture in the video sequence is , including the entire sphere of video data for the corresponding moment. However, a small portion of the picture (even Only a single viewport is displayed to the user. The rest of the picture is rendered. It is discarded without dulling. Different viewports are dynamically selected according to the user's head movements. Generally, the entire picture is sent so that it can be selected and displayed. This technique is very This may result in a large video file size.
[0036] To improve coding efficiency, some systems divide pictures into subpictures. To divide. A subpicture is a defined spatial area of a picture. Each subpicture is , including the corresponding viewport for the picture. Video can be encoded in two or more resolutions. Each resolution is encoded into a different sub-bitstream. When reaming, the coding system uses the current bit used by the user. Based on the newport, the sub-bitstreams are merged into the bitstream for transmission. This is possible. In particular, the current viewport is derived from a high-resolution sub-bitstream. The viewport that is not being viewed is obtained from a low-resolution bitstream. In this way, the highest quality video is displayed to the user, and lower quality videos are displayed. It will be discarded. If the user selects a new viewport, the lower resolution video will be discarded. However, this is presented to the user. The decoder determines that the new viewport is a higher resolution view. It may request to receive the deo. The encoder then processes the merge process accordingly. It can be changed. When it reaches an IRAP picture, the decoder will use the new viewport's higher resolution. Decoding of a video sequence can be initiated. This method enhances the user's viewing experience. Significantly improves video compression without causing any adverse effects.
[0037] One concern with the above method is the length of time required to change the resolution. This is based on the length of time it takes to reach the RAP picture. The driver is unable to start decoding different video sequences in non-IRAP pictures. Therefore, one way to reduce such latency is to increase the number of IRAP pins. This involves including kucha. However, this leads to an increase in file size. Functionality and To balance loading efficiency, different viewports / subpictures are different. IRAP pictures may be included at a certain frequency. For example, in viewports where they are more likely to be seen. A viewport may have more IRAP pictures than other viewports. For example, a bus In the context of kettleball, a viewport that shows the stands or ceiling is seen by the user. Because the possibility is lower, baskets and / or sensors are better than such viewports. Viewports related to turquoise may more frequently include IRAP pictures.
[0038] This method leads to other problems. In particular, subpictures that include viewports are single It is part of one picture. Different sub-pictures have IRAP pictures at different frequencies. Some parts of the picture include both IRAP subpictures and non-IRAP subpictures. The picture is stored in the bitstream by using the NAL unit, so the question is... The topic is: The NAL unit is a parameter set or slice of a picture and its corresponding A slice is a storage unit that includes the slice header. An access unit is a unit that includes the entire picture. Yes. Therefore, the access unit includes all NAL units related to the picture. Hmm. The NAL unit also includes types that indicate the type of picture, including slices. Some video In the system, a single picture (for example, one included in the same access unit) All related NAL units are required to be of the same type. Therefore, NAL The unit's memory mechanism allows pictures to be stored as both IRAP subpictures and non-IRAP subpictures. When this is included, it may not function correctly.
[0039] Disclosed herein are both IRAP subpictures and non-IRAP subpictures. This is a mechanism for adjusting the NAL's memory scheme to support the inclusion of pictures. This, in turn, involves VR, including different IRAP subpicture frequencies for different viewports. Enables video. In the first example disclosed herein, pictures are mixed This is a flag that indicates whether or not it exists. For example, the flag indicates that the picture is an IRAP subpicture and It may also indicate that it includes both non-IRAP subpictures. Based on this flag, decorate The software uses different decoding methods to properly decode and display the picture / subpicture. Subpictures of different types can be handled differently. This flag is used in the picture parameters. It may be stored in the PPS and may be called mixed_nalu_types_in_pic_flag.
[0040] In the second example, what is disclosed herein is whether the picture is a mixture. It is a flag. For example, the flag indicates whether the picture is an IRAP subpicture or a non-IRAP subpicture. It may also indicate that it includes both. Furthermore, the flag indicates that the mixed picture is one IRAP Picchi includes just two NAL unit types, including one IRP and one non-IRAP type. This restricts the access. For example, a picture is a random access decryptable reading picture. Instantaneous Decoded Refresh (IDR) (IDR_W_RADL) with a reading picture, IDR without a reading picture (IDR_N_LP), or Clean Random Access (CRA) NAL Unit Type (CRA_NUT) It may also include an IRAP NAL unit containing only one of the following. Furthermore, the picture may include a trailing Picture NAL unit type (TRAIL_NUT), random access decryptable reading picture NAL unit type (RADL_NUT), or Random Access Skip Reading • Non-IRAP NAL unit containing only one of the following: Picture (RASL) NAL unit type (RASL_NUT) It may include a sub-picture. Based on this flag, the decoder will appropriately process the picture / sub-picture. To decrypt and display it, different sub-pictures may be treated differently during decryption. This flag may be stored in the PPS and is called mixed_nalu_types_in_pic_flag. That's good too.
[0041] Figure 1 is a flowchart of an exemplary operation method 100 for coding a video signal. In particular, The audio signal is encoded in the encoder. The encoding process involves various mechanisms. It compresses the video signal by using it to reduce the video file size. The file size is reduced while the associated bandwidth overhead is reduced, and the compressed video This allows the file to be sent to the user. The decoder then processes the compressed video. The audio file is decrypted, and the original video signal is reconstructed for display to the end user. Generally, the decoding process allows the decoder to reconstruct the video signal without inconsistencies. To achieve this, we faithfully mimic the encoding process.
[0042] In step 101, the video signal is input to the encoder. For example, the video signal The file may be an uncompressed video file stored in memory. Another example is a video The files are captured by video capture devices such as video cameras. Video files may be encoded to support live streaming of audio. This may include both audio and video components. The component is a series of image frames that, when viewed in sequence, give a visual impression of movement. It includes light and It includes pixels represented by a color called the chroma component (or color sample). In some cases, the frame may also include depth values to support three-dimensional viewing. stomach.
[0043] In step 103, the video is divided into blocks. The division is by each frame. This includes subdividing pixels into square and / or rectangular blocks for compression. For example, High Efficiency Video Coding (HEVC) (also known as H.265 and MPEG-H Part 2) In this context, the frame is first defined by a predetermined size (for example, 64 pixels x 64 pixels). It can be divided into coding tree blocks, which are blocks of pixels. CTU is Lumasa Includes both sample and chroma sample. Divide the CTU into blocks, and then further code The code is used to repeatedly subdivide blocks until a configuration that supports numbering is achieved. A framing tree may be used. For example, the framing components are that the individual blocks are It may be subdivided until it contains relatively uniform lighting values. For example, the frame The chroma component may be subdivided until each individual block contains relatively uniform color values. Therefore, the segmentation mechanism changes depending on the content of the video frame.
[0044] In step 105, the image blocks separated in step 103 are compressed. Various compression mechanisms are used for this purpose. For example, interprediction and / or inter R-prediction may be used. Interpretation is used when objects in a normal scene are in a continuous frame. It is designed to take advantage of the fact that it tends to appear in the 'm'. Therefore, the reference frame In a frame, the blocks used to draw objects do not need to be shown repeatedly in neighboring frames. In particular, objects such as tables may remain in the same position across multiple frames. Therefore, the table is shown once, and adjacent frames refer back to the reference frame. This can be referenced. The pattern matching mechanism spans multiple frames. It may be used to match objects. Furthermore, if the moving object is, for example, An object may be represented across multiple frames due to its movement or the camera's movement. As a typical example, the video shows a car moving across the screen over multiple frames. It may be done. A motion vector can be used to indicate such movement. This gives an offset from the coordinates of the object in the frame to the coordinates of the object in the reference frame. It is a two-dimensional vector. Therefore, the interpretation is the image block within the current frame. This represents a set of motion vectors indicating the offset from the corresponding block within the reference frame. It can be encoded.
[0045] Intra prediction encodes blocks within a common frame. Intra prediction encodes blocks. This takes advantage of the fact that minute and chroma components tend to cluster together within a frame. Some green areas of a tree tend to be located near similar green areas. Intra-prediction This includes multiple directional prediction modes (e.g., 33 in HEVC), planar modes, and DC. Use (DC) mode. Directional mode is when the current block is in the direction of neighboring blocks. This indicates that it is similar to / the same as the sample. Planar mode is a series of blocks along rows / columns. This demonstrates that a row (for example, a plane) can be interpolated based on neighboring blocks at the edges of the row. In this case, the planar mode uses a relatively constant gradient of changing values between rows / columns. It shows a smooth transition of light / color. DC mode is used for boundary smoothing, and blocks Planar sample of all neighboring blocks related to the angular direction of the direction prediction mode This indicates that it is similar to / the same as the average. Therefore, the intra prediction block is the image block. The value of 'k' can be represented as a value of various relation prediction modes instead of an actual value. Furthermore, the interpretation A measurement block can represent an image block as a motion vector value instead of an actual value. In any case, the predicted block may not accurately represent the image block in some cases. All differences are stored in the residual block. To further compress the file, the residuals The conversion may be applied to the block.
[0046] In step 107, various filtering techniques may be applied. Then, the filter is applied by an intra-loop filtering scheme. Lock-based predictions result in the generation of blocky images in the decoder. Alternatively, the block-based prediction method encodes the blocks, and then encodes The created block may be reconstructed to be used later as a reference block. The rethreading method includes noise suppression filters, deblocking filters, and adaptive loop filters. , and iteratively apply a sample adaptive offset (SAO) filter to the block / frame. These filters ensure that the encoded file can be reconstructed accurately. Reduces blocking artifacts. Furthermore, these filters reduce artifacts. In subsequent blocks that are encoded based on the reconstructed reference block, Within the reconstructed reference block, the likelihood of generating artifacts is reduced. Reduces the artifact cost.
[0047] When a video signal is segmented, compressed, and filtered, the resulting data The data is encoded into a bitstream in step 109. The bitstream is Based on the data discussed above, and to support the appropriate reconstruction of the video signal in the decoder... This includes any signaling data that is desirable for partitioning. For example, such data could be partitioned. Various methods for giving coding instructions to the data, prediction data, residual blocks, and decoder. It may include flags. The bitstream is sent to the decoder upon request. It may be stored in memory. The bitstream can also be broadcast to multiple decoders. Bitstream generation may be via and / or multicast. This is a process. Therefore, steps 101, 103, 105, 107, and 109 are many steps. This may be done continuously and / or simultaneously across the m and block. As shown in Figure 1. The order presented is for clarity and ease of consideration, and is for video coding professionals. It is not intended to restrict the order of the Seths.
[0048] The decoder receives the bitstream and starts the decoding process in step 111. In particular, the decoder uses an entropy decoding scheme to convert the bitstream into corresponding Convert to syntax and video data. The decoder, in step 111, Use the syntax data from the stream to determine the partitions for the frame. The division should match the result of the block division in Step 103. Step 1 The entropy coding / decoding used in 11 will be explained below. Based on the spatial positioning of values in the input image, D blocks from several possible options. Many choices are made during the compression process, such as selecting a partitioning method. Signaling the selection may involve using multiple bins. When this happens, the bin is a binary value treated as a variable (for example, a bin that may change depending on the situation). The value is (T). Entropy coding is clearly good when the encoder is in a particular case. This allows us to discard all undesirable options and leave us with a set of acceptable choices. Each acceptable option is then assigned a codeword. This is based on the number of acceptable choices (for example, one bin for two choices, four bins (For example, two bins for the choices). Then the encoder codes the selected choice. Encode the codeword. This method is used so that the codeword is potentially one of all possible options. In contrast to uniquely showing a choice from a large set, this is a small subset of acceptable choices. The size of the codeword is only desirable to uniquely indicate their selection. Reduce. Then, the decoder, in the same way as the encoder, determines a set of acceptable choices. The selection is decoded by determining the acceptable set of choices. Da can read the codeword and determine the selection made by the encoder.
[0049] In step 113, the decoder performs the decoding of the block. In particular, the decoder The inverse transform is used to generate the residual block. The decoder then processes the residual block and the corresponding Using the corresponding prediction block, the image block is reconstructed according to the division. In step 105, the intra prediction block generated by the encoder and the interpretation block are used. It may include both the measurement block and the image block. The reconstructed image block is then used in step 11. Located within the frame of the video signal reconstructed according to the segmentation data determined in step 1. It can be attached. The syntax for step 113 is also the entropy code discussed above. Signaling may be performed within the bitstream by ing.
[0050] In step 115, the reconstructed video signal is obtained in the same manner as in step 107 of the encoder. Filtering is performed on the frame. For example, a noise suppression filter, Blocking filters, adaptive loop filters, and SAO filters are used in blocking It may be applied to the frame to remove artifacts. The frame is filtered. Then, the video signal is displayed in step 117 for viewing by the end user. It can be output to the game.
[0051] Figure 2 shows an example coding and decoding (codec) for video coding. This is a schematic diagram of the system 200. In particular, the codec system 200 supports the implementation of the operation method 100. It provides a function for coding. The codec system 200 is an encoder and decoder. It is generalized to describe the components used in both. Codec system 200 is The video signal is received as discussed in relation to steps 101 and 103 of the operation method 100. The signal is then divided, resulting in the divided video signal 201. Next, the codec system Tem 200 was considered in relation to steps 105, 107, and 109 of Method 100, When working as a coder, the segmented video signal 201 is coded into a bitstream. It is compressed into a file. When acting as a decoder, the codec system 200 operates according to method 100. As discussed in relation to steps 111, 113, 115, and 117, from the bitstream Generates the output video signal. The codec system 200 includes a general coder control component 211. Transformation, scaling and quantization component 213, intrapicture estimation component 215, Track picture prediction component 217, motion compensation component 219, motion estimation component 221, scale Loop and inverse transform component 229, filter control analysis component 227, in-loop filter configuration component Element 225, decoded picture buffer component 223, and header format and context Context-adaptive binary arithmetic coding (CABAC) It includes component 231. Such components are joined together as shown. In Figure 2, black The solid lines indicate the movement of data being encoded / decoded, while the dashed lines indicate other constituent elements. This shows the movement of control data that controls the basic operation. The components of the codec system 200 are: All may be present in the decoder. The decoder is a component of the codec system 200. It may include busets. For example, the decoder includes an intrapicture prediction component 217, Motion compensation component 219, scaling and inverse transform component 229, in-loop filter component It may also include element 225 and the decoded picture buffer component 223. These components are This will be explained later.
[0052] The segmented video signal 201 is segmented into blocks of pixels by the coding tree. This is a segmented captured video sequence. The coding tree is divided into various segments. Use split mode to subdivide a block of pixels into smaller blocks of pixels. These blocks can then be subdivided into smaller blocks. These may also be called nodes in the coding tree. Larger parent nodes have smaller child nodes. It is divided into subnodes. The number of times a node is subdivided is determined by the depth of the node / coding tree. This is called a divided block, which may be contained within a coding unit (CU). To obtain. For example, CU stands for Lumablock, Red Difference Chroma (Cr)block, and the blue difference chroma (Cb) block has the corresponding syntax for CU. It may be a sub-part of the CTU that is included along with the command. The split mode depends on the split mode used. To divide the node into two, three, or four child nodes with varying shapes This may include binary trees (BT), ternary trees (TT), and quadrant trees (QT) used in the demarcation. The video signal 201 is subjected to general coder control components 211 for compression, conversion, scaling, and Biquantization component 213, intrapicture estimation component 215, filter control analysis component 22 7. This is then transferred to the motion estimation component 221.
[0053] The general coder control component 211 converts the video to a bitstream according to the constraints of the application. It is configured to make decisions related to the coding of images in the scan. For example, general The coder control component 211 optimizes the bitrate / bitstream size versus the quality of reconstruction. Manage optimization. Such decisions involve storage space / bandwidth availability and image resolution. It may be done based on the requirements of. Also, the general coder control component 211 is a buffer To mitigate buffer run and overrun issues, the buffer is adjusted based on the transmission speed. To manage usage, the general coder control component 211 manages usage. It manages classification, prediction, and filtering by other components. For example, General The coder control component 211 increases the complexity of the compression in order to increase the resolution and bandwidth usage. Dynamically increase the resolution, or decrease the complexity of the compression to reduce bandwidth usage. It may be dynamically lowered. Therefore, the general coder control component 211 regenerates the video signal. To strike a balance between build quality and bitrate concerns, the Codec System 200 It controls the other components. The general coder control component 211 controls the operation of the other components. It generates control data to control the decoder. The control data also contains parameters for decoding in the decoder. The header format is encoded into a bitstream to signal the data. The data is then transferred to the CABAC component 231.
[0054] The segmented video signal 201 is used for interpretation of motion estimation components 221 and motion The frame or slice of the segmented video signal 201 is also transmitted to the compensation component 219. This may be divided into multiple video blocks. Motion estimation component 221 and motion compensation component Element 219 provides one or more blocks within one or more reference frames to provide a time prediction. Performs interpredictive coding of received video blocks for the codec. System 200, for example, provides appropriate coding modes for each block of video data. You may run multiple coding passes to select the correct one.
[0055] The motion estimation component 221 and the motion compensation component 219 may be highly integrated, but generally Shown separately for illustrative purposes. Motion estimation performed by motion estimation component 221. This is a process that generates motion vectors to estimate the motion of a video block. The vector is, for example, the displacement of the coded object relative to the prediction block. This may be shown. The prediction block is a block coded in terms of pixel difference. These are blocks that are known to match well. Prediction blocks are also called reference blocks. It's okay if it's discovered. The difference of such pixels is the sum of absolute differences (SAD), the sum of squared differences (SSD), and This may be determined by other measurement criteria for difference. HEVC is CTU, coding tree It uses several coded objects, including blocks (CTBs) and CUs. For example, a CTU can be divided into a CTB, and then the CTB can be included in a CU. It can be divided into CBs. CUs are prediction units containing prediction data for CUs. It can be encoded as a transform unit (PU) and / or a transform unit (TU) containing the transformed residual data. The motion estimation component 221 uses rate distortion analysis as part of the rate distortion optimization process. By using this, motion vectors, PU, and TU are generated. For example, motion estimation configuration Sub-221 has multiple reference blocks and multiple motion vectors related to the current block / frame. You can decide on the reference block, motion vector, etc., that has the best rate distortion characteristics. You may choose this option. The best rate distortion characteristics are related to the quality of video reconstruction (for example, compression). Both the amount of data loss (due to the process) and coding efficiency (e.g., the size of the final encoding) are important factors. To balance the two sides.
[0056] In some cases, the codec system 200 is described in the decoded picture buffer component 223. You may calculate pixel position values that are finer than integer values for the stored reference picture. The video codec system 200 uses a 1 / 4 pixel position of the reference picture, and an 1 / 8 pixel position. The value of the position, or other fractional pixel position, may be interpolated. Therefore, the motion estimation component 221 includes full pixel position and fractional pixel position. You can perform a motion search related to position and output a motion vector with fractional pixel precision. The motion estimation component 221 compares the position of the PU with the position of the predicted block in the reference picture. Movement of video blocks within slices intercoded by PU The function is calculated. The motion estimation component 221 encodes the calculated motion vector. The header format and CABAC component 231 are output as motion data, and motion is compensated. Output to component 219.
[0057] The motion compensation performed by the motion compensation component 219 is determined by the motion estimation component 221. This includes extracting or generating predicted blocks based on a defined motion vector. But that's fine. After all, the motion estimation component 221 and the motion compensation component 219 are used in some cases. And it may be functionally integrated. Receive motion vectors related to the current video block's PU. Then, the motion compensation component 219 may find the predicted block pointed to by the motion vector. Next, the predicted block is obtained from the pixel values of the current video block being coded. By subtracting the pixel value and forming the pixel difference value, a residual video block is formed. Generally, the motion estimation component 221 performs motion estimation related to the luma component and motion compensation. The compensation component 219 is calculated based on the luma component for both the chroma component and the luma component. Uses motion vectors. Prediction blocks and residual blocks are transformed and scaled. It is then transferred to the quantization component 213.
[0058] The segmented video signal 201 consists of intrapicture estimation components 215 and intrapic It is also transmitted to the motion prediction component 217, the motion estimation component 221, and the motion compensation component 219. Similarly, the intrapicture estimation component 215 and the intrapicture prediction component 217 are They may be highly integrated, but are shown separately for conceptual purposes. The estimation component 215 and the intrapicture prediction component 217 are, as described above, between frames. The interpretation is performed by the motion estimation component 221 and the motion compensation component 219. Alternatively, it intrapredicts the current block for the blocks within the current frame. The intrapicture estimation component 215 is used to encode the current block. Determine the intra-prediction mode. In some examples, the intra-picture estimation component 215 is , appropriate for encoding the current block from multiple tested intra prediction modes Select the intra-prediction mode. Then, the selected intra-prediction mode is used for coding. The header format and CABAC component 231 are then forwarded.
[0059] For example, the intrapicture estimation component 215 is a test of various intra-prediction models. Rate distortion values were calculated using rate distortion analysis on the mode, among the tested modes. Then, the intra-prediction mode with the best rate distortion characteristics is selected. Rate distortion analysis is performed. Generally speaking, an encoded block and an encoded block that generates an encoded block The amount of distortion (or error) between the original unencoded block and the encoded block. Determine the bitrate (e.g., number of bits) used to generate the input. The picture estimation component 215 determines which intra prediction mode is best for blocking. To determine whether a distortion value is shown, distortion and regressive values are used for various encoded blocks. The ratio is calculated from the number. Furthermore, the intrapicture estimation component 215 has the lowest rate distortion. Depth blocks of a depth map are created using Depth Modeling Mode (DMM) based on Optimization (RDO). It may be configured to be coded.
[0060] When the intrapicture prediction component 217 is implemented in the encoder, the intrapicture Based on the selected intra-prediction mode determined by the estimation component 215, the prediction is When generating residual blocks from a lock, or when implemented in a decoder, the bitstream You can read the residual block from there. The residual block is the prediction block, represented as a matrix. This includes the difference in values between the original block and the current block. The residual block is then converted and scaled. It is then transferred to the quantization component 213. The intrapicture estimation component 215 and intra The picture prediction component 217 may operate on both the luma and chroma components.
[0061] The transformation, scaling, and quantization component 213 further compresses the residual block. It is composed of the following. The transformation, scaling, and quantization component 213 is discrete to the residual block. Apply a transformation such as the sine transform (DCT), discrete sine transform (DST), or a similar concept transformation, and the remainder Generate video blocks containing differential transform coefficient values. Wavelet transform, integer transform, sub-transformation. A 3D conversion or other type of conversion may also be used. The conversion converts residual information to pixel values. It may be possible to convert from one domain to another, such as the frequency domain. Conversion, scaling, and quantization. Component 213, for example, scales the residual information converted based on frequency. It is further composed of different frequency information at different granularities. This involves applying a scaling factor to the residual information so that it can be childized, which is the reconstructed This may affect the final visual quality of the video. Transformation, scaling, and quantization structures. The component 213 is further configured to quantize the conversion coefficient in order to further reduce the bitrate. The quantization process reduces the bit depth associated with some or all of the coefficients. This is also acceptable. The degree of quantization may be modified by adjusting the quantization parameters. In some examples, the transformation / scaling and quantization component 213 is then quantized. A scan of the matrix containing the transformed coefficients may be performed. The quantized transformed coefficients are bit The header format and CABAC component 231 are transferred to encode into a stream. ru.
[0062] The scaling and inverse transformation component 229 supports motion estimation by transforming and scaling. Apply the inverse operations of the scaling and quantization component 213. Scaling and inverse transform component Element 229 may be a reference block that is a prediction block for another current block, for example. To reconstruct the residual blocks of the pixel region for later use as a lock, reverse the process. Apply kering, inverse transform, and / or inverse quantization. Motion estimation component 221 and / Alternatively, the motion compensation component 219 is used in subsequent block / frame motion estimation. The reference block is calculated by adding the residual block to the corresponding prediction block and returning it. This may be done to mitigate artifacts that occur during scaling, quantization, and transformation. Therefore, a filter is applied to the reconstructed reference block. Otherwise, Artifacts lead to inaccurate predictions when predicting subsequent blocks (approximately (and further artifacts are generated).
[0063] The filter control analysis component 227 and the in-loop filter component 225 are located in the residual block. Apply a filter to the reconstructed image block. For example, scaling And the transformed residual block from the inverse transformation component 229 reconstructs the original image block. To do so from the intrapicture prediction component 217 and / or motion compensation component 219 It may be combined with the corresponding prediction block. Then, the reconstructed image block is A filter may be applied. In some examples, the filter is instead a residual block. It may also be applied to the following. Similar to the other components in Figure 2, the filter control analysis component 227 The filter components 225 within the loop may be highly integrated and implemented together, Shown separately for conceptual purposes. Filter applied to the reconstructed reference block. It is applied to a specific spatial region, and how such a filter is applied can be investigated. It includes multiple parameters for adjustment. The filter control analysis component 227 is such a Analyze the reconstructed reference block to determine where the filter should be applied. Set the corresponding parameters. Such data is filtered for encoding. The header format and CABAC component 231 are forwarded as data. The filter component within the loop... Component 225 applies such a filter based on the filter control data. This includes deblocking filters, noise suppression filters, SAO filters, and adaptive loop filters. Filters may be included. Such filters may be reconstructed depending on the example. Applicable to a pixel block in the spatial / pixel domain or in the frequency domain. It may be used.
[0064] When operating as an encoder, filtered and reconstructed image blocks The residual block and / or prediction block are used in motion estimation, as discussed above. The data is stored in the decoded picture buffer component 223 for later use. When operating, the decoded picture buffer component 223 is reconstructed and filtered. The decoded blocks are stored and transferred to the display as part of the output video signal. The picture buffer component 223 consists of a prediction block, a residual block, and / or a reconstructed block. It may be any memory device capable of storing the image blocks.
[0065] The header format and CABAC component 231 are various configurations of the codec system 200. Code to receive data from an element and send such data to a decoder. Encode into a bitstream. In particular, the header format and CABAC component 231 This is for encoding control data such as general control data and filter control data. It generates various headers. Furthermore, if the prediction data includes intra-prediction and motion data... The residual data in the form of the quantized transformation coefficients are all encoded into a bitstream. The final bitstream is used to reconstruct the original segmented video signal 201. It contains all the information desired by the decoder. Such information is in intra predictive mode. The index table (also called the codeword mapping table), various blocks Definition of the coding context for the 'k', indicator of the most likely intra-prediction mode. This may also include indications of parcel information, etc. Such data may be used in the environment. The information may be encoded by using tropic coding. For example, the information is Context Adaptive Variable Length Coding (CAVLC) g) CABAC, Syntax-Based Context-Adaptive Binary Arithmetic Coding (SBAC: syntax- (based context-adaptive binary arithmetic coding), probability interval entropy (PIPE: Probability interval partitioning entropy) coding, or another entropy It may be encoded using coding techniques. After entropy coding, The processed bitstream is sent to another device (for example, a video decoder). It may be sent, or archived for later transmission or retrieval.
[0066] Figure 3 is a block diagram of an exemplary video encoder 300. Video encoder 300 This performs the encoding function of the codec system 200 and / or the operation method 100. It may be used to carry out steps 101, 103, 105, 107, and / or 109. The coder 300 divides the input video signal and the divided video signal 201 is substantially the same This produces a segmented video signal 301. Then, the segmented video signal 301 is The data is compressed and encoded into a bitstream by the components of the Encoder 300.
[0067] In particular, the segmented video signal 301 is used for intra-picture prediction. It is transferred to component 317. The intrapicture prediction component 317 is an intrapicture estimation component. The component 215 and the intrapicture prediction component 217 may be substantially the same. The decrypted video signal 301 is based on the reference block in the decrypted picture buffer component 323. The motion compensation component 321 is also transferred for interpretation. The estimation component 221 and the motion compensation component 219 may be substantially the same. Prediction blocks and residual blocks from picture prediction component 317 and motion compensation component 321 The block is transferred to the transformation and quantization component 313 for transformation and quantization of the residual block. The transformation and quantization component 313 is converted by the transformation, scaling and quantization component 213. It may be substantially the same as the transformed, quantized residual block and the corresponding preset. The measurement block is an entropy coding structure for coding into a bitstream. It is transferred to component 331 (along with the associated control data). Entropy coding configuration Element 331 may be substantially similar to the header format and CABAC component 231. .
[0068] Furthermore, the transformed and quantized residual blocks and / or corresponding prediction blocks are Transform and quantify to reconstruct the reference block for use by motion compensation component 321. The data is transferred from the child-forming component 313 to the inverse transform and quantization component 329. Inverse transform and quantization Component 329 may be substantially the same as the scaling and inverse transformation component 229. The in-loop filter of the in-loop filter component 325 depends on the residual block and / Alternatively, it is also applied to the reconstructed reference block. The in-loop filter component 325 is The filter control analysis component 227 and the in-loop filter component 225 are substantially similar. It is also acceptable. The in-loop filter component 325 is related to the in-loop filter component 225. It may include multiple filters as considered. Then the filtered blocks The decoded picture block is used as a reference block by the motion compensation component 321. It is stored in the buffer component 323. The decoded picture buffer component 323 is the decoded picture buffer It may be substantially the same as component 223.
[0069] Figure 4 is a block diagram of an exemplary video decoder 400. The video decoder 400 is Performing the decoding function of the codec system 200 and / or step 1 of the operation method 100 It may be used to implement 11, 113, 115, and / or 117. Decoder 400 For example, receiving a bitstream from encoder 300 and displaying it to the end user. To do this, it generates an output video signal that has been reconstructed based on the bitstream.
[0070] The bitstream is received by the entropy decoding component 433. - The decoding component 433 is CAVLC, CABAC, SBAC, PIPE coding, or other input It is configured to implement an entropy decoding method such as entropy coding technology. For example, the entropy decoding component 433 encodes the bitstream as a codeword. Use header information to provide context for interpreting further data. It may be done. The decoded information includes general control data, filter control data, partition information, and motion Video, including data, prediction data, and quantized transformation coefficients from residual blocks. It contains any desired information for decoding the signal. The quantized transformation coefficients are the residual block. Transferred to the inverse transform and quantization component 429 for reconstruction. Inverse transform and quantization Component 429 may be the same as the inverse transform and quantization component 329.
[0071] The reconstructed residual blocks and / or prediction blocks are based on intra-predictive behavior. The image is then transferred to the intrapicture prediction component 417 for reconstruction into an image block. The intrapicture prediction component 417 consists of the intrapicture estimation component 215 and the intrapicture The Kucha prediction component 217 may be substantially the same. In particular, the intrapicture prediction component Element 417 uses prediction mode to identify reference blocks within the frame and performs intraprediction. The residual block is applied to the result to reconstruct the image block. Interpreted image blocks and / or residual blocks and corresponding interpretations The measured data is sent to the decoded picture buffer component 423 via the in-loop filter component 425. These are transferred, and these are respectively the decoded picture buffer component 223 and the in-loop file. The loop component 225 may be substantially the same as the other component 225. The in-loop filter component 425 is restructured Filter the built image blocks, residual blocks, and / or predicted blocks. Such information is stored in the decrypted picture buffer component 423. The reconstructed image block from component 423 is motion-compensated for interpretation. It is transferred to the component 421. The motion compensation component 421 may be substantially the same as the motion estimation component 221 and / or the motion compensation component 219. In particular, the motion compensation component 421 uses the motion vector from the reference block to generate a prediction block and applies the residual block to the result to reconstruct the image block . The resulting reconstructed block is also transferred to the decoded picture buffer component 423 via the in-loop filter component 425 . The decoded picture buffer component 423 continues to store further reconstructed image blocks, and those reconstructed image blocks can be reconstructed into a frame by the partition information. Also, such a frame may be arranged in a sequence. The sequence is output to the display as a reconstructed output video signal . The decoded picture buffer component 423 continues to store further reconstructed image blocks, and those reconstructed image blocks can be reconstructed into a frame by the partition information. Also, such a frame may be arranged in a sequence. The sequence is output to the display as a reconstructed output video signal is output to the display as a reconstructed output video signal. FIG. 5 is a schematic diagram showing an exemplary CVS 500. For example, the CVS 500 may be encoded by an encoder such as the codec system 200 and / or the encoder 300 according to the method 100 . Further, the CVS 500 may be decoded by a decoder such as the codec system 200 and / or the decoder 400. The CVS 500 includes pictures coded in the decoding order 508. The decoding order 508 is the order in which the pictures are positioned within the bitstream
[0072] . FIG. 5 is a schematic diagram showing an exemplary CVS 500. For example, the CVS 500 may be encoded by an encoder such as the codec system 200 and / or the encoder 300 according to the method 100 . Further, the CVS 500 may be decoded by a decoder such as the codec system 200 and / or the decoder 400. The CVS 500 includes pictures coded in the decoding order 508. The decoding order 508 is the order in which the pictures are positioned within the bitstream . Further, the CVS 500 may be decoded by a decoder such as the codec system 200 and / or the decoder 400. The CVS 500 includes pictures coded in the decoding order 508. The decoding order 508 is the order in which the pictures are positioned within the bitstream . Then, the pictures of the CVS 500 are output in the presentation order 510. The presentation order 510 is the order in which the pictures should be displayed by the decoder in order to properly display the resulting video . For example, the pictures of the CVS 500 may generally be positioned in the presentation order 510. However . Then, the pictures of the CVS 500 are output in the presentation order 510. The presentation order 510 is the order in which the pictures should be displayed by the decoder in order to properly display the resulting video . For example, the pictures of the CVS 500 may generally be positioned in the presentation order 510. However . For example, the pictures of the CVS 5 as a whole may be positioned in the presentation order 510. However And certain pictures, for example, use similar pictures to support interpretation. They were moved to different positions to improve coding efficiency by placing them closer together. It is also possible. Moving such a picture in this way results in decoding order 508. In the example shown, the pictures are indexed 508 times in decoding order from 0 to 4. In presentation order 510, the pictures at index 2 and index 3 are the same as the picture at index 0. It has been moved in front of the Kucha.
[0073] CVS 500 includes IRAP picture 502. IRAP picture 502 is a random image related to CVS 500. This is a picture coded by intra-prediction, acting as an access point. In particular, the IRAP picture 502 block is referenced by other blocks of IRAP picture 502. It is coded. IRAP picture 502 is coded without referencing other pictures. Therefore, it can be decrypted without first decrypting any other pictures. The decoder can then begin decoding the CVS 500 in IRAP picture 502. IRAP Picture 502 may refresh the DPB. For example, IRAP Picture 502 The picture presented later is IRAP Picture 502 (for example, Picture) for interpretation. It is not necessary to rely on the previous picture (index 0). Therefore, the picture buffer is IRAP Picture 502 can be refreshed when it is decrypted. This is related to interpretation. Because the coding error cannot be propagated through IRAP picture 502, all It has the effect of stopping such errors. IRAP Picture 502 is a picture of various types It may include the following. For example, an IRAP picture may be coded as an IDR or CRA. Good. IDR starts a new CVS 500 and refreshes the picture buffer intraco It is a picture that has been recorded. CRA will start a new CVS 500 or picture An intracode that acts as a random access point without refreshing the buffer. These are the reading pictures. In this way, 5 reading pictures related to CRA 04 may refer to the previous picture of the CRA, while 50 is a reading picture related to the IDR. 4 does not require referencing the picture before the IDR.
[0074] CVS 500 also includes various non-IRAP pictures. These are Reading Picture 504 and Includes trailing picture 506. Reading picture 504 is decrypted in order 508, IRAP picture Although it is positioned after 502, in presentation order 510 it is positioned before IRAP picture 502. It's a kucha. Trailing picture 506 is an IRAP picture in both the decoded order 508 and the presentation order 510. It is positioned after picture 502. Leading picture 504 and trailing picture 50 Both 6 are coded by interpretation. Trailing picture 506 is, Refer to IRAP Picture 502 or the picture positioned after IRAP Picture 502 to coordinate. Therefore, trailing picture 506 is always decoded by IRAP picture 502. It can be decrypted if it is. Reading picture 504 is a random access skip read This may include reading (RASL) and random access decryptable reading (RADL) pictures. Yes. The RASL picture is coded by reference to the picture before the IRAP picture 502 but is coded at a position after the IRAP picture 502. Since the RASL picture relies on the previous picture, it cannot be decoded when the decoder starts decoding at the IRAP picture 502. Therefore, the RASL picture is skipped and not decoded when the IRAP picture 502 is used as a random access point. However, the RASL picture is decoded and displayed when the encoder uses a previous IRAP picture (before index 0 not shown) as a random access [[ID=1B]]point. The RADL picture is coded by reference to the IRAP picture 502 and / or pictures following the IRAP picture 502, but is positioned before the IRAP picture 502 in presentation order 510. Since the RADL picture does not rely on the picture before the IRAP picture 502, it can be decoded and
[0075] displayed when the IRAP picture 502 is a random access point. Pictures from the CVS 500 may each be stored in an access unit. Further, a picture may be segmented into slices, and the slices may be included in NAL units. A NAL unit is a storage unit that includes a picture parameter set or a slice and a corresponding slice header. A NAL unit is assigned a type to indicate to the decoder the type of data included in the NAL unit. It may be included in IDR (IDR_N_LP) NAL units, CRA NAL units, etc. IDR_W_RADL NA The L unit is an IDR picture in which IRAP picture 502 is associated with RADL reading picture 504. This indicates that it is a chat. The IDR_N_LP NAL unit indicates that IRAP picture 502 is any ready. This indicates that it is an IDR picture not associated with picture 504. (CRA NAL Unit) This is a CRA picture in which IRAP picture 502 may be associated with reading picture 504. This indicates that slices of non-IRAP pictures may also be placed in NAL units. For example, the slice of trailing picture 506 is an intermediate Trailing picture NAL unit indicating that it is a predictively coded picture It may be placed in the IP (TRAIL_NUT). The slice of reading picture 504 corresponds Picture corresponds to type interpredictive coded reading picture 504 The RASL NAL unit type (RASL_NUT) and / or RASL NAL unit type indicate that it is It may be included in the ip(RADL_NUT). Sign the picture slice within the corresponding NAL unit. By sorting, the decoder selects the appropriate decoding method to apply to each picture / slice. Canism can be easily determined.
[0076] Figure 6 shows multiple sub-picture video streams separated from the VR picture video stream 600. This is a schematic diagram showing reams 601, 602, and 603. For example, subpicture video story Each of the streams 601-603 and / or VR picture video stream 600 are coordinated to CVS 500. It may be edited. Therefore, subpicture video streams 601-603 and / or The VR picture video stream 600 is a codec system 200 and / or related to method 100. This may be encoded by an encoder such as encoder 300. Furthermore, subpicture Video streams 601-603 and / or VR picture video stream 600 are codecs It may be decoded by a decoder such as system 200 and / or decoder 400.
[0077] The VR picture video stream 600 includes multiple pictures presented over time. VR coordinates a sphere of video content that can be displayed as if the user were at the center of the sphere. It works by drawing. Each picture contains the entire sphere. Meanwhile, the viewport and Only a portion of the picture known as is is displayed to the user. For example, the user, Based on the user's head movements, the sphere's viewport is selected and displayed on the head-mounted display. You may use a HMD (Head-Mounted Display). This allows you to physically enter a virtual space depicted by video. To give the impression of presence. To achieve this result, each picture in the video sequence This includes the entire sphere of video data for the corresponding moment. However, it includes a small portion of the picture (for example) For example, only a single viewport is displayed to the user. The rest of the picture is... It is discarded without being rendered. Different viewports are dynamically displayed according to the user's head movements. Generally, the entire picture is sent so that it can be selected and displayed.
[0078] In the example shown, the picture in VR picture video stream 600 is available view Each picture can be further subdivided into subpictures based on the port. And the corresponding subpicture is a part of the temporal presentation, in terms of its temporal position (for example, pict (Including the order of ya). Subpicture video streams 601-603 have a consistent subdivision over time. It is generated when applied. Such a consistent subdivision is subpicture video Streams 601-603 are generated, and each stream has a predetermined size, shape, and VR picture. Includes a set of sub-pictures with spatial positions relative to the corresponding picture within video stream 600. Hmm. Furthermore, one set of subpictures within subpicture video streams 601-603 is presented at The time positions in between are different. Therefore, the sub-picture video streams 601-603 are sub A picture can be aligned in the time domain based on its temporal position. Then, at each time Subpictures from the intermediate-position subpicture video streams 601-603 are displayed. Based on predefined spatial positions, the VR picture video stream 600 is reconstructed. They can be merged in the spatial domain. In particular, subpicture video streams 601-603 are Each of these sub-bitstreams can be encoded into separate sub-bitstreams. When merged together, the M has a bitstream that includes the entire set of pictures over time. The resulting bitstream is displayed in the user's currently selected viewport. Based on the data, it can be decoded and sent to a decoder for display.
[0079] One of the problems with VR video is that all of the sub-picture video streams 601-603 are high quality. It is acceptable to send the data to the user in high quality (e.g., high resolution). This is because the decoder Dynamically selects the user's current viewport and the corresponding sub-picture video stream 60 It enables the display of sub-pictures from 1 to 603 in real time. However, the user For example, this allows viewing only a single viewport from subpicture video stream 601. It is possible that sub-picture video streams 602-603 will be discarded. Therefore Therefore, transmitting high-quality sub-picture video streams 602-603 requires a significant amount of bandwidth. It's okay to waste width. To improve coding efficiency, VR video uses multiple videos. It may be encoded into Stream 600, and each video Stream 600 may be of different quality / resolution. It is encoded. In this way, the decoder receives the current subpicture video stream 601 A request can be sent accordingly. In response, an encoder (or intermediate slicer (inter Edit slicer (or other content server) provides higher quality video streams. Higher quality sub-picture video stream from 00 601 and lower quality video You can select lower quality sub-picture video streams 602-603 from stream 600. The encoder then sends such subbitstreams to the decoder. This can be merged into a fully encoded bitstream. The system receives a series of pictures and determines that the current viewport is of higher quality, and other The viewport is of lower quality. Furthermore, the highest quality sub-picture is (head movement (When not available) Generally displayed to the user, lower quality subpictures are generally It was discarded, striking a balance between functionality and coding efficiency.
[0080] The user is looking from sub-picture video stream 601 to sub-picture video stream 602. When transferring, the decoder will have a higher quality new current sub-picture video stream 602. It requests that the data be transmitted in quality. The encoder then uses a merger mechanism accordingly. It can be changed. As mentioned above, the decoder is only for the new CVS 500 in IRAP picture 502. Decryption can begin. Therefore, subpicture video stream 602 is IRAP picture / sa It is displayed at a lower quality until it reaches the sub-picture. Then, the IRAP picture is sub-picture. To begin decoding the higher quality version of the video stream 602, It can be decrypted with quality. This method does not negatively impact the user's viewing experience. Significantly improves shrinkage.
[0081] One concern with the above method is the length of time required to change the resolution. This is based on the length of time it takes to reach the IRAP picture within the video stream. The decoder has different versions of sub-picture video story in non-IRAP pictures. This is because decoding of Mu602 cannot be initiated. To reduce such latency... One way to do this is to include more IRAP pictures. However, this is This leads to an increase in file size. To strike a balance between functionality and coding efficiency, different approaches are used. Viewport / sub-picture video streams 601-603 are IRAP pictures at different frequencies. It may include, for example, a viewport / sub-picture video that is more likely to be seen. Ostream 601-603, other viewport / subpicture video stream 601-60 You may have more than 3 IRAP pictures. For example, in the context of basketball, Viewports / sub-picture video streams 601-603 that show the stand or ceiling are for the user Such viewport / subpicture video is less likely to be seen by Viewpoints related to basketball and / or center court are more relevant than Stream 601-603. Even if the streams 601-603 of the sub-picture video stream contain IRAP pictures more frequently. good.
[0082] This approach leads to further problems, especially for sub-picture video producers who share proof of concept (POC). Subpictures from reem 601-603 are part of a single picture. As mentioned above, Slices from Kucha are included in NAL units based on picture type. Some bi In a decoded system, all NAL units associated with a single picture are , constrained to include the same NAL unit type. Different sub-picture video stories When M601-603 have IRAP pictures at different frequencies, some of the pictures are IRAP subpictures. This includes both chat and non-IRAP subpictures. This means that each individual picture is the same. This violates the constraint that only Ip's NAL units should be used.
[0083] This disclosure states that all NAL units for slices within a picture are the same NAL unit type. This problem is addressed by removing the constraint of using P. The access unit is included in the access unit. By removing this constraint, the access unit The term may include both IRAP NAL unit types and non-IRAP NAL unit types. Furthermore, the picture / access unit is either an IRAP NAL unit type or a non-IRAP NAL unit type. A flag may be encoded to indicate when a mixture with P is included. In some examples, the flag The 'g' is a mixed NAL unit type flag in picture. )(mixed_nalu_types_in_pic_flag). Furthermore, a single mixed picture / access unit The set includes only one type of IRAP NAL unit and one type of non-IRAP NAL unit. However, constraints that require certain conditions may be applied. This is an unintended NAL unit. To prevent the mixing of types from occurring. If such mixing is permitted, the decoder It must be designed to manage such mixtures. This is coding Unnecessarily reduces the hardware complexity required without bringing further benefits to the process. To increase the quality. For example, mixed pictures can be selected from IDR_W_RADL, IDR_N_LP, or CRA_NUT. It may also include one type of IRAP NAL unit. Furthermore, the mixed picture is TRAIL Includes one type of non-IRAP NAL unit selected from _NUT, RADL_NUT, and RASL_NUT. That's fine. An exemplary implementation of this method will be examined in more detail below.
[0084] Figure 7 shows an exemplary bitstory containing a picture with mixed NAL unit types. This is a schematic diagram showing bitstream 700. For example, bitstream 700 is a code related to method 100. Codec system 200 and / or decoder 400 for decoding It can be generated by the bitstream 700 or encoder 300. Furthermore, the bitstream 700 is multiple VR picture merged from multiple sub-picture video streams 601-603 of the video resolution It may include 600 video streams, and each sub-picture video stream may have a different sky Includes CVS 500 in an intermediate position.
[0085] Bitstream 700 is a sequence parameter set (SPS) 710, with multiple picture parameters. Includes a meter set (PPS) 711, multiple slice headers 715, and image data 720. SPS 71 0 is a common value for all pictures in the video sequence contained in bitstream 700. Includes data such as picture size settings, bit depth, code. This may include the parameters of the editing tool, bitrate constraints, etc. PPS 711 is a picture This includes parameters that apply to the entire sequence. Therefore, each picture in the video sequence is You may also refer to PPS 711. Each picture refers to PPS 711, but in some examples, Please note that a single PPS 711 may contain data related to multiple pictures. Multiple similar pictures may be coded with similar parameters. In such cases, even if a single PPS 711 contains data about such similar pictures Good. PPS 711 is a coding tool available for slicing within the corresponding picture. It can show the quantization parameters, offset, etc. Slice header 715 is within the picture. Each slice contains its own unique parameters. Therefore, each slice in the video sequence There may be one slice header 715. The slice header 715 contains slice type information. Picture Order Count (POC), Reference Picture List, Prediction Weights, Tile Entry Point It may include tile entry point, deblocking parameters, etc. Please note that DA715 may also be called a tile group header depending on the context. .
[0086] Image data 720 is a video encoded by interpretation and / or intrapretation. This includes the original data and the corresponding transformed and quantized residual data. For example, video. The sequence contains multiple pictures 721 coded as image data 720. Cha721 is a single frame in a video sequence, and therefore, generally speaking, video sequence When displaying the numbers, they are shown as a single unit. However, subpicture 723 is, It may be displayed to implement certain technologies such as virtual reality. Picture 721 is Refer to PPS 711 respectively. Picture 721 is a subpicture 723, a tile, and / or It may be divided into slices. Subpicture 723 is the coded video sequence This is the spatial region of picture 721 that is consistently applied to the lance. Therefore, subpicture 723 may be displayed by an HMD in the context of VR. Furthermore, the specified POC is Picture 723 is obtained from sub-picture video streams 601-603 of the corresponding resolution. It may also be done. Subpicture 723 may refer to SPS 710. In some systems, Slice 725 is called a tile group that includes tiles. Slice 725 and / or tiles The tile group of the slice refers to slice header 715. Slice 725 is a single NAL unit. An integer number of complete tiles or integers within a tile of picture 721 exclusively contained in the knit It may be defined as a row of consecutive complete CTUs. Thus, slice 725 is CTU It is further divided into a CTU and / or CTB. CTU / CTB are coordinated based on the coding tree. It is further divided into coding blocks. Next, the coding block is the prediction mechanism. It can be encoded / decoded by [this method].
[0087] The parameter set and / or slice 725 are coded into the NAL unit. The AL unit is a syntax structure that includes the indication of the type of data that follows. Also, emulation prevention bytes are inserted in places as needed. It may be defined as a byte containing the data in the form of the input RBSP. More specifically, The NAL unit is the parameter set of picture 721 or slice 725 and the corresponding slice It is a memory unit that includes the IS header 715. In particular, the VCL NAL unit 740 is a slide of picture 721. This is a NAL unit including a chair 725 and a corresponding slice header 715. Furthermore, a non-VCL NAL Unit 730 includes parameter sets such as SPS 710 and PPS 711. The NAL unit of the PPS may be used. For example, both the SPS 710 and PPS 711 are Non-VCL NAL unit 730, SPS NAL unit type (SPS_NUT) 731, and PPS NAL unit Each of these may be included in type (PPS_NUT)732.
[0088] As mentioned above, IRAP pictures such as IRAP picture 502 are included in IRAP NAL unit 745. Obtain. Non-IRAP pictures such as leading picture 504 and trailing picture 506 are , may be included in non-IRAP NAL unit 749. In particular, IRAP NAL unit 745 is IRAP picture or any NAL unit containing slice 725 obtained from subpicture. Non-IRAP N AL unit 749 is any picture that is not an IRAP picture or subpicture (for example, Includes slice 725 obtained from leading picture and trailing picture. This is the NAL unit. The IRAP NAL unit 745 and the non-IRAP NAL unit 749 are both... Since both also include slice data, they are both VCL NAL units 740. In an exemplary embodiment... The IRAP NAL unit 745 does not display IDR pictures or RADL pictures without reading pictures. Slice 725 from the IDR related to the char is IDR_N_LP NAL unit 741 or IDR_w_RADL NAL Each may be included in unit 742. Furthermore, IRAP NAL unit 745 may contain a CRA picture or The slice 725 may be included in CRA_NUT 743. In an exemplary embodiment, non-IRAP NAL Unit 749 is a slide from a RASL picture, RADL picture, or trailing picture. Chair 725 may be included in RASL_NUT 746, RADL_NUT 747, or TRAIL_NUT 748, respectively. In an exemplary embodiment, a complete list of possible NAL units is given, NAL unit type The results are sorted and shown below. [Table 1A] [Table 1B] [Table 1C]
[0089] As mentioned above, VR video streams have sub-pictures with IRAP pictures at different frequencies. It may also include ya723. This is because it is more likely to include spatial areas that the user is less likely to see. Regarding spatial areas where fewer IRAP pictures are used and which users are likely to view frequently: This allows more IRAP pictures to be used. In this way, users Spatial regions that are likely to return frequently can be quickly adjusted to a higher resolution. The law brings picture 721 which includes both IRAP NAL unit 745 and non-IRAP NAL unit 749. In this state, picture 721 is called a mixed picture. This state is a mixed NAL unit within the picture. It can be signaled by the mixed_nalu_types_in_pic_flag (727). ixed_nalu_types_in_pic_flag 727 may also be set to PPS 711. Furthermore, mixed_nalu _types_in_pic_flag 727 refers to each picture 721 that references PPS 711 and has two or more VCL NAL units. It has a 740, and the VCL NAL unit 740 has the same value for NAL unit type (nal_unit_type) When specifying that it is not present, it may be set to be equal to 1. Furthermore, mixed_nalu_type s_in_pic_flag 727 refers to each picture 721 that references PPS 711, and one or more VCL NAL units 740 Each picture 721 of the VCL NAL unit 740 has a nal_unit_type that references PPS 711. When they have the same value, they may be set to be equal to 0.
[0090] Furthermore, when mixed_nalu_types_in_pic_flag 727 is set, sub-picture 721 One or more of the VCL NAL units 740 of the Kucha 723 are all first identification of NAL unit type The value is such that all other VCL NAL units 740 in picture 721 are of the NAL unit type. A constraint may be used that has a different second specific value. For example, the constraint is mixed Picture 721 shows a single type IRAP NAL unit 745 and a single type non-IRAP NAL unit. It may be necessary to include 749. For example, picture 721 may be one or more I DR_N_LP NAL unit 741, one or more IDR_w_RADL NAL units 742, or one It may include multiple CRA_NUT 743s, but any combination of such IRAP NAL units 745 It cannot include combinations. Furthermore, picture 721 contains one or more RASL_NUT 746, and not one. It may include or contain multiple RADL_NUT 747s, or one or more TRAIL_NUT 748s, but This does not include any combination of IRAP NAL units 745.
[0091] In the example implementation, the picture type is used to define the decoding process. Such processes include, for example, identifying pictures by picture order counting (POC). Information derivation, marking of the status of the reference picture in the decoded picture buffer (DPB), D This includes picture output from PB, etc. A picture is all coded pictures. It may be identified by a type based on the NAL unit type, which includes its sub-parts. In the video coding system of the department, the picture type is instantaneous decoding refresh (I May include DR (Digital-Risk) pictures and non-IDR (Individual-Digital-Risk) pictures. Other video coding systems In this case, the picture type is trailing picture, time sublayer access (TSA: te (Primary sub-layer) picture, Step-wise temporal sub-layer access (STSA: step-wise temporal) Sub-layer access pictures, random access decryptable reading (RADL) pictures, Random Access Skip Reading (RASL) for pictures, Broken Link Access (BLA) (broken-link access) Pictures, Instant Random Access Pictures, and Clean Run Dam access pictures may be included. Such picture types are subtitled. Is it a sub-layer referenced picture or a non-sub-layer referenced picture? Further distinctions may be made based on whether it is a sub-layer non-referenced picture (Kucha). BLA pictures include BLAs with reading pictures, BLAs with RADL pictures, and It may be further distinguished as a BLA that does not have a reading picture. IDR pictures are They are further distinguished as IDRs with RADL pictures and IDRs without reading pictures. It may also be used.
[0092] Such picture types are used to implement various video-related features. It may be the case that IDR, BLA, and / or CRA pictures implement IRAP pictures. It may be used for the following purposes. IRAP pictures may offer the following features / benefits. IRAP pictures The presence of kucha may indicate that the decryption process can begin from that picture. The function is that the decoding process continues in the bitstream as long as the IRAP picture is present at that location. This enables the implementation of a random access feature that starts at a specified location. This is not necessarily the beginning of the bitstream. Also, the presence of an IRAP picture is the same as the RASL picture. Excluding the "cha", the coded picture that starts with "IRAP picture" is located before the "IRAP picture". The decoding process is restructured so that the placed picture is coded without any reference whatsoever. To refresh. Therefore, the IRAP picture located within the bitstream is restored. Stop the propagation of error code. Therefore, coding positioned before the IRAP picture. The decryption error of the decrypted picture will be transmitted through the IRAP picture, and will follow the IRAP picture in the decryption order. It cannot be transmitted to the picture.
[0093] IRAP Pictures offers various features, but it comes at a disadvantage in terms of compression efficiency. Therefore, the presence of IRAP pictures can cause a sharp increase in bitrate. Compression effect This disadvantage to the rate has various causes. For example, IRAP pictures are different from non-IRAP pictures. The image used as an interface is represented by significantly more bits than the predicted picture. It is an intra-predicted picture. Furthermore, the presence of an IRAP picture is an intra-predicted picture. This impairs the time prediction used in the context. In particular, IRAP pictures are referenced from DPBs to previous reference pictures. Refresh the decryption process by removing the previous reference picture. This is a reference used for coding the picture that follows the IRAP picture in the decoding order. This reduces the availability of the picture and, therefore, lowers the efficiency of this process.
[0094] IDR pictures have different signaling and derivation processes than other IRAP picture types. You may use Seth. For example, the signaling and derivation processes related to IDR are Instead of deriving the most significant bit (MSB) from the previous key picture, set the MSB portion of the POC to 0. This is also good. Furthermore, the slice header of the IDR picture helps in managing the reference picture. It does not have to include information used for [the purpose]. On the other hand, other [the purpose] such as CRA, trailing, TSA, etc. The Kucha type is a reference that can be used to carry out the marking process of the reference picture. Reference picture set (RPS) or reference picture list, etc. It may include chat information. The marking process for reference pictures is performed on the reference picture in the DPB. Is the status used for reference or not? It is the process of determining whether or not. The existence of IDR means that the decryption process simply determines all references within the DPB. This indicates that the picture will not be used for reference, so the IDR Regarding Kucha, such information does not need to be signaled.
[0095] In addition to picture type, picture identification information obtained through POC is also used in interpretation prediction. For managing the use of illuminated pictures, for outputting pictures from DPB, the scaling of motion vectors. It is used for multiple purposes, such as for rings and weighted predictions. In the department's video coding system, pictures within the DPB are used for short-term reference. Used, used for long-term reference, or not used for reference. It may be marked as such. The picture may be marked as not to be used for reference. Then, the picture can no longer be used for prediction. When not needed for power, pictures can be removed from the DPB. Other video files In the reference system, reference pictures are marked as short-term and long-term. That's fine too. The reference picture can be used when the picture is no longer needed for the reference of the prediction. They may be marked as not to be used for illumination. The conversion may be controlled by the marking process of the decoded reference picture. A typical sliding window process and / or explicit memory management control behavior (MMCO) The process may be used as a marking mechanism for the decrypted reference picture. The sliding window process has a maximum number of reference frames within the SPS (Sliding Window Process). When it is equal to the specified maximum number denoted as s, the short-term reference picture is used for reference. Mark it as something that cannot be decoded. The short-term reference picture is the most recently decoded short-term picture. The char may be stored in a first-in, first-out (FIFO) manner so that it is held in the DPB. Explicit MMCO process It may include multiple MMCO commands. An MMCO command may be one or more short-term commands. Long-term reference pictures may be marked as not being used for reference, and all You may mark the pictures as not to be used for reference, or currently Mark the reference picture or an existing short-term reference picture as the long-term reference, and then, You may assign a long-term picture index to that long-term reference picture.
[0096] In some video coding systems, the marking of reference pictures and The process for outputting and deleting pictures from the DPB is performed after the pictures have been decrypted. It is executed. Other video coding systems use RPS for managing reference pictures. Use the most radical difference between the RPS mechanism and the MMCO / sliding window process. The key difference is that, for each specific slice, the RPS is the current picture or any later. The objective is to provide the complete set of reference pictures used by the subsequent pictures. This means all PPB should hold all P for current or future use by the picture. A complete set of kucha is signaled in RPS. This is relative to DPB. Unlike the MMCO / sliding window system, where only changes are signaled, the RPS mechanics... According to Zum, in order to maintain the correct status of the referenced pictures in the DPB, the decoding order is important. Information from the picture is not needed. The order of picture decoding and the operation of DPB are determined by RPS. Taking advantage of its benefits, some video coding systems have been modified to improve error tolerance. It will be updated. In some video coding systems, picture marking, and The buffer's operation, which includes both outputting and deleting decoded pictures from the DPB, is currently It may be applied after the decryption. In other video coding systems First, the RPS is decoded from the slice header of the current picture, and then the picture The marking and buffering behavior may be applied before decoding the current picture. .
[0097] In VVC, the method for managing reference pictures can be summarized as follows: List 0 Two reference picture lists, denoted as "Call List 1," are directly signaled and derived. These are based on the RPS or sliding window + MMCO process discussed above. No. The marking of a reference picture is not the active entry in the reference picture list. It uses both the active entry and the reference picture lists 0 and 1 directly, Only active entries are used as the reference index in CTU interprediction. It may be done. Information for deriving the two reference picture lists is SPS, PPS, and SLA It is signaled by syntax elements and syntax structures within the header. A predefined RPL structure is used by referencing it within the slice header in the SPS. Signaling occurs internally. The two reference picture lists are bidirectional interprediction (B) slices. All ties, including unidirectional interprediction (P) slices and intraprediction (I) slices. It is generated with respect to the slice of the picture. The two reference picture lists are the first reference picture list. It can be built without using a pre-processing process or a reference picturelist modification process. The Long-Term Reference Picture (LTRP) is identified by the Point of Consciousness (POC) LSB. The Delta POC MSB cycle (d The elta POC MSB cycle is signaled regarding LTRP as determined for each picture. That's fine.
[0098] To code a video image, the image is first divided into sections, and these sections become a bitstory. It is coded into a m. Various picture classification methods are available. For example, The image is divided into regular slices, dependent slices, and tiles. They can be distinguished by, and / or wavefront parallel processing (WPP). For simplicity, HEVC is When dividing slices into CTB groups for video coding, the usual slices Only chairs, dependent slices, tiles, WPPs, and combinations thereof may be used. This constrains the encoder. Such a division is size matching of the Maximum Transmission Unit (MTU). It can be applied to support parallel processing and reduced end-to-end latency. The MTU represents the maximum amount of data that can be transmitted in a single packet. If the payload exceeds the MTU, the payload undergoes a process called fragmentation. It is split into two packets by the 's'.
[0099] A regular slice, also simply called a slice, is caused by loop filtering behavior. Despite some interdependencies, other normal slices within the same picture and These are segmented parts of an image that can be reconstructed independently. Each normal slice is It is then encapsulated in its own Network Abstraction Layer (NAL) unit for transmission. Furthermore, predictions within a picture that cross slice boundaries (intra-sample prediction, motion information prediction, The interdependence of coding mode prediction and entropy coding is independent of each other. It may be disabled to support construction. Such independent reconstructions are parallel. It supports parallelization. For example, parallelization based on normal slicing requires minimal inter-processor interaction. Alternatively, inter-core communication can be used. However, since each normal slice is independent, Each slice is associated with a separate slice header. Typical slice usage This is due to the bit cost of the slice header for each slice and the slice boundaries. The lack of cross-sectional predictions can lead to significant coding overhead. Regular slices are used to support matching regarding MTU size requirements. This may also be done. In particular, a normal slice may be encapsulated in a separate NAL unit and independently When coded, each normal slice is divided into multiple packets. To avoid splitting, it should be smaller than the MTU of the MTU scheme. Therefore Therefore, the purpose of parallelization and MTU size matching is the ray of slices in the picture You may impose demands that contradict what is stated by the out.
[0100] Dependent slices are similar to regular slices but have a shortened slice header, and This enables the demarcation of image tree block boundaries without compromising prediction within the code. In other words, a dependent slice is a slice that has been fragmented into multiple NAL units. This makes it possible for a portion of a normal slice to be completely encoded within the entire normal slice. End-to-end delay reduced by allowing transmission before completion. It brings about.
[0101] The picture may be divided into tile groups / slices and tiles. The tiles are A sequence of CTUs covering a rectangular area of a picture. Tile group / slice This includes several tiles of the picture. Raster scan tile group mode and long The square tile group mode may be used to generate tiles. In tile group mode, tile groups are the last tile scans of picture tiles. Includes a sequence of tiles. In rectangular tile group mode, tile group This includes several tiles of a picture that collectively form a rectangular area of the picture. The tiles within a rectangular tile group are in the order of the tile group's raster scan. For example, tiles have horizontal and vertical boundaries that generate columns and rows of tiles. It may also be a segmented portion of the image generated by the boundary. The tile is a rastasc. The coding order may be from right to left and from top to bottom. The scan order of the CTB is It is limited to within the tile. Therefore, a CTB in the first tile must move to the next CTB before moving to the next tile. Coding is done in star scan order. Like regular slices, tiles are within the picture. This impairs the interdependence of predictions and the interdependence of entropy decoding. However, the tiles are individual Each NAL unit does not necessarily have to be included, and therefore the tiles are MTU size matching. It does not have to be used for that purpose. Each tile is processed by one processor / core. This is possible and can be used for in-picture prediction between processing units that decode neighboring tiles. The inter-processor / inter-core communication used is shared (when neighboring tiles are in the same slice). Carrying the slice header and the reconstructed sample and metadata It may be limited to performing sharing related to tile filtering. Two or more tiles When included in a slice, each of the other elements is the offset of the first entry point of the slice. The byte offset of the entry point for the file is signaled within the slice header. This may be done. With respect to each slice and tile, the following conditions apply, namely, 1) within the slice 1) All coded tree blocks belong to the same tile, and 2) Tie All coded tree blocks in a slice belong to the same slice. At least one of these conditions must be met.
[0102] In WPP, the image is divided into single rows of CTB. Entropy decoding and prediction mechanism Nism may use data from other rows of the CTB. By parallel decoding of rows of the CTB... This enables parallel processing. For example, the current line can be decoded in parallel with the previous line. However, the decoding of the current row is delayed by 2 CTB from the decoding process of the previous row. The extension is the CTB above the current CTB in the current row and the upper right before the current CTB is coded. This ensures that data related to CTB is available. This method graphically When represented, it appears as a wavefront. This shifted start is at most the same as the row of CTB that the image contains. Enables parallelization up to a certain number of processors / cores. Neighborhood tree blocks within a picture. Since in-picture prediction is allowed between the lines, the process to enable in-picture prediction is Inter-service / inter-core communication can be quite extensive. WPP segmentation takes into account the size of the NAL unit. Therefore, WPP does not support MTU size matching. However, as requested To perform MTU size matching, a normal slice is used for specific coding It can be used in conjunction with WPP, although it involves overhead. Finally, wavefront segment (wavefr The segment may contain exactly one row of CTB. Furthermore, when using WPP and When a slice begins in a row of a center back (CTB), the slice should end in the same row of the same CTB.
[0103] Tiles may also include motion-constrained tile sets. Motion-constrained tilesets (MCTS) are those in which the associated motion vectors are full samples within the MCTS (full- A fractional sample refers to a location where only the full sample location within the MCTS is needed for interpolation. This is a tile set designed to be constrained to point to a pull (fractional-sample). Furthermore, motion vectors for predicting temporal motion vectors derived from blocks outside MCTS The use of candidate tiles is not permitted. In this way, each MCTS is not included in the MCTS. It may be decrypted independently without any presence. Temporal MCTS Supplemental Enhancement Information (SEI) The cement information message indicates the presence of MCTS in the bitstream, and the MCTS It may be used for signaling. The MCTS SEI message conforms to the MCTS set. To generate a bitstream (as part of the semantics of the SEI message) This provides supplementary information that can be used in the extraction of MCTS sub-bitstreams (as defined). The report defines several sets of MCTS and the MCTS subbitstream extraction process. The Alternative Video Parameter Set (VPS) and Sequence Parameter Set (SPS) used within it, and the bytes of the raw byte sequence payload (RBSP) of the Picture Parameter Set (PPS) Includes several sets of extracted information, including the MCTS subbitstream extraction process. When extracting subbitstreams, the parameter set (VPS, SPS, and PPS) is: The slice header may be rewritten or replaced, and (first_slice_segme Slice address related syntax (including nt_in_pic_flag and slice_segment_address) Different in the subbitstream from which one or all of the x elements have been extracted. The value may be used and therefore may be updated.
[0104] VR applications, also known as 360-degree video applications, allow you to view a portion of a complete sphere. You may also display only a subset of the entire picture. Dynamic Adaptive Streaming Over Highway (DASH) The 360 delivery mechanism, which relies on the viewport via the Hypertext Transfer Protocol, Lowering the bitrate and supporting 360-degree video delivery via a streaming mechanism. It may be used for the purpose of cubemap projection (CMP: cu By using bemap projection, the sphere / projected picture is divided into multiple MCTS. Two or more bitstreams may be encoded at different spatial resolutions or qualities. When delivering data to the decoder, use a bitstream with higher resolution / quality. The CTS is sent for the viewport that will be displayed (for example, the front viewport). MCTS from lower resolution / quality bitstreams will be sent for other viewports. These MCTSs are packed in a specific way and then sent for decryption. It is sent to the device. The viewport seen by the user creates a positive viewing experience. Therefore, it is expected to be represented by high-resolution / quality MCTS. When you rotate to view a viewport (for example, the left or right viewport), the display The system will fetch high-resolution / high-quality MCTS for the new viewport. During the short period of time, the view comes from a lower resolution / quality viewport. The user's head is different. When you turn your head to view the viewport, the user's head orientation changes and the view There is a delay between when the higher resolution / quality representation of the port is seen and when it is not. This delay is How quickly does the system fetch higher resolution / quality MCTS for that viewport? It depends on whether it can be done, and furthermore, how fast it can be fetched, IR It depends on the AP period. The IRAP period is the interval between two IRAP occurrences. This delay is Since the MCTS of the new viewport may only be decryptable from IRAP pictures, the IRAP period Related to.
[0105] For example, if the IRAP period is coded every second, then the following applies: The best-case scenario for the delay is when the system begins fetching the new segment / IRAP cycle. Network when the user's head turns to see the new viewport just before the user moves. This is the same as round-trip delay. In this scenario, the system is due to the new viewport. You can immediately request a higher resolution / quality MCTS, and therefore the smallest buff The firing delay can be set to almost zero, and the sensor delay is small and negligible. Assuming this is possible, the only delay is the network round-trip delay, and the network round-trip delay This is the fetch request delay plus the requested MCTS transmission time. The round-trip delay can be, for example, about 200 milliseconds. The worst-case scenario for delay is The system has already made a request for the next segment, and the user's head is in a new viewport. This is the IRAP period when changing orientation to view, plus the network round-trip delay. The system will be more frequent so that the IRAP cycle becomes shorter in order to improve the worst-case scenario above. It can be encoded using IRAP pictures, which reduces the overall delay. This is because it reduces the amount of data. However, this method reduces compression efficiency, so the bandwidth requirements are significantly higher. Yes.
[0106] In the exemplary implementation, subpictures of the same coded picture are different This allows the value of nal_unit_type to be included. This mechanism is explained as follows: The picture may be divided into subpictures. The subpictures are equal to 0 tile_ A set of rectangles in a tile group / slice that starts with a tile group that has a group_address. Each sub-picture may refer to the corresponding PPS, and therefore a separate tile. It may have the following divisions. The presence of subpictures may be indicated in the PPS. Each subpicture The 'cha' is treated like a picture during the decryption process, crossing the boundaries of subpictures. In-loop filtering may always be disabled. The width and height of the subpicture are: The unit may be specified as Luma CTU size. The position of subpictures within a picture is Gunnaring is not required, but it may be derived using the following rule: Subpict The CTU within the picture is large enough to include subpictures within the picture boundary. Take the next unoccupied position in the raster scan order. To decode each subpicture. The reference picture is decoded from the reference picture in the picture buffer to the current subpicture and It is generated by extracting the area to be located. The extracted area is decoded. It is a sub-picture, and therefore a sub-picture of the same size and in the same position within the picture. Interpretation is performed between chats. In such cases, within the picture being coded Allowing different nal_unit_type values is a subpicture derived from random access pictures. Subpictures derived from chat and non-random access pictures can be processed without significant difficulty (for example) For example, it would be possible to merge them into the same coded picture (without VCL-level modifications). This also applies to coding based on MCTS.
[0107] Allowing different nal_unit_type values within a coded picture is a different approach from other systems. It may be beneficial in Nario. For example, a user can view 360-degree video content. You may view certain areas more frequently than other areas. MCTS / Subpicture-based view - Coding efficiency and average equivalent quality in port-dependent 360-degree video distribution To create a better trade-off with viewport switching latency, other For areas that are more frequently seen than those areas, more frequent IRAP pictures are coded. It is possible. The viewport switching latency of equivalent quality is from the first viewport to the second. When switching to the second viewport, the presentation quality of the second viewport is the same as that of the first viewport. This is the latency experienced by the user before achieving a comparable presentation quality.
[0108] Another implementation involves the derivation of POCs and management of reference pictures, including mixed NALs within pictures. Use the following solution for unit type support: Mixed IRAP subpict To specify whether or not a picture with a subpicture and a non-IRAP subpicture may exist , flags that are directly or indirectly referenced by tile groups (sps_mixed_tile_gr The `oups_in_pic_flag` is present in the parameter set. NAL units including IDR tile groups. Regarding knitting, when deriving the POC for the picture, whether or not the POC MSB is reset is important. To specify this, the flag (poc_msb_reset_flag) must be present in the corresponding tile group header. A variable called PicRefreshFlag is defined and associated with the picture. The lag should be such that the POC derivation and DPB state are refreshed when the picture is decoded. Specifies whether or not the current If the group is included in the first access unit of the bitstream, PicRefreshF The lag is set to be equal to 1. Instead, the current tile group is an IDR tile. If it is a group, PicRefreshFlag is sps_mixed_tile_groups_in_pic_flag ? poc_ms b_reset_flag: Set to equal to 1. Otherwise, the current tile group is CR If it is an A-tile group, the following applies: The current access unit is code If it is the first access unit of the sequence being refreshed, PicRefreshFlag is equal to 1. It is configured as follows: The access unit is the end of sequence NAL unit. The variable HandleCraAsFirstPicInCvsFlag that immediately follows or is related to the set is equal to 1. When set to this, the current access unit is the first of the sequence to be coded It is an access unit. Otherwise, PicRefreshFlag is set to be equal to 0. (For example, the current tile group is the first access unit of the bitstream) Not affiliated with the IRAP Tile Group isn't it).
[0109] When PicRefreshFlag is equal to 1, the value of POC MSB (PicOrderCntMsb) is related to the picture. It is reset to equal to 0 during the derivation of the POC. Reference Picture Set (RPS) or reference picture set Information used for reference picture management, such as the Picture List (RPL), corresponds to the NAL unit. Signaling occurs within the tile group / slice header regardless of the tile type. The crucifix is constructed at the beginning of decoding each tile group, regardless of the NAL unit type. The reference picture lists are RefPicList
[0000] and RefPicList
[0001] for the RPL method, RefPicList0[ ] and RefPicList1[ ] for the RPS method, or related to the picture. It may also include a similar list containing reference pictures for predictive operation. PicRefreshFl When ag is equal to 1, during the marking process of the reference picture, all reference pictures in the DPB The kucha is marked as not to be used for reference.
[0110] Such implementations are associated with specific problems. For example, nal_unit_t in a picture. When mixing of type values is not allowed, and when determining whether a picture is an IRAP picture When the output and the derivation of the variable NoRaslOutputFlag are described at the picture level, the decoder is: These derivations can be performed after receiving the first VCL NAL unit of any picture. And, due to the support for mixed NAL unit types in the picture, the decoder is above We must wait for the arrival of other VCL NAL units before performing the derivation. Worst case scenario In total, the decoder must wait for the arrival of the last VCL NAL unit in the picture. Furthermore, in such systems, the POC MSB is reset when the POC for the picture is derived. To specify whether or not to use a flag, a flag is used in the tile group header of the IDR NAL unit. It may be mixed. This mechanism has the following problems: Mixed CRA NAL unit For net-type and non-IRAP NAL unit types, this mechanism provides support. It is not done. Furthermore, this information is stored within the tile group / slice header of the VCL NAL unit. Gunnering means that the IRAP (IDR or CRA) NAL unit is in the picture and the non-IRAP NAL unit is in the picture. When there is a change in the status of whether or not to mix with the bitstream, the bitstream is extracted or This requires that the value be changed during the merge. Such rewriting of the slice header is This occurs whenever a user requests video, and therefore requires significant hardware resources. It requires a specific IRAP NAL unit type and a specific non-IRAP NAL unit. Aside from the mixing of different NAL unit types, there are several other mixed types within the picture. Such a combination is permitted. Such flexibility does not provide support for practical use cases. On the other hand, these complicate the codec design, which unnecessarily increases the complexity of the decoder. Therefore, it increases the cost of the associated implementation.
[0111] In general, this disclosure relates to running subpictures or MCTS in video coding. This document describes technologies for supporting dam access. For more details, see sub-pictures. Used to support random access based on or MCTS, within pictures This document describes an improved design for supporting mixed NAL unit types. This is based on the VVC standard, but also applies to other video / media codec specifications.
[0112] To solve the above problem, the following exemplary implementation is disclosed. Such an implementation is individual They can be applied individually or in combination. In one example, each picture is mixed with the pictures. This is associated with an indication of whether it contains the specified nal_unit_type value. The indication is signaled within PPS. This indication is used for all references. By marking the picture as not to be used for reference, the POC MSB Supports determining whether a reset should be performed and / or whether the DPB should be reset. When an indication is signaled within PPS, changes in values within PPS are merged or This may be done during a separate extraction. However, this is the extraction of such a bitstream. Or when the PPS is rewritten and replaced by other mechanisms during the merger. It is permissible.
[0113] Alternatively, this indication is signaled within the tile group header. It may be required that this be the same for all tile groups of the picture. In this case, the value is during the extraction of the sub-bitstream of the MCTS / subpicture sequence. This may need to be changed. Alternatively, this indication can be changed to NAL Unit It is signaled within the header, but the same applies to all tile groups in the picture. It may be required to do something. However, in this case, the value is the sequence of MCTS / subpicture. It may need to be modified during the extraction of the sub-bitstream. Alternatively, The indication is used for all VCL NAL of the picture when used for the picture. Define additional VCL NAL unit types such that the unit has the same NAL unit type value. Signaling may be performed by interpreting. However, in this case, the N of the VCL NAL unit The value of the AL unit type is the extraction of the subbitstream of the MCTS / subpicture sequence. This may need to be changed during the broadcast. Alternatively, this indication is pict When used for the picture, all VCL NAL units in the picture are the same NAL unit type. By defining additional IRAP VCL NAL unit types that have a value of P, the signal It may be ringed. However, in this case, the value of the NAL unit type of the VCL NAL unit is The MCTS / subpicture sequence needs to be modified during the extraction of the subbitstream. This is possible. Alternatively, at least one VCL N of any of the IRAP NAL unit types Each picture with an AL unit contains a value of the mixed NAL unit type. It may be associated with an indication of whether or not.
[0114] Furthermore, only mixed IRAP NAL unit types and non-IRAP NAL unit types are permitted. This allows for mixing of nal_unit_type values within a picture in a limited way. Such restrictions may apply. For any particular picture, all VCL NAL units The set has the same NAL unit type, or some VCL NAL units have a specific IRAP N Either the AL unit type has the remainder having a specific non-IRAP VCL NAL unit type. In other words, any particular picture's VCL NAL unit is two or more IRAP units. It is not possible to have a NAL unit type, and it is not possible to have two or more non-IRAP NAL unit types. This is not possible. The picture does not contain the value of nal_unit_type which is mixed with the picture, VCL NAL A unit may be considered an IRAP picture only if it has the IRAP NAL unit type. For all IRAP NAL units that do not belong to an IRAP picture (including IDRs), the POC MSB is It does not need to be reset. All IRAP NAL units that do not belong to IRAP pictures (including IDRs) Regarding knitting, the DPB is not reset, and therefore all reference pictures are referenced. Marking it as something not to be used for is not performed. TemporalId is, If at least one of the VCL NAL units of the Kucha is an IRAP NAL unit, then the picture is related It may be set to be equal to 0.
[0115] The following are specific implementations of one or more of the above embodiments. IRAP picture is mixed_nal The value of u_types_in_pic_flag is equal to 0, and each VCL NAL unit is IDR_W_RADL and RSV_IRAP_ Codes containing VCL13 and having nal_unit_types within the range from IDR_W_RADL to RSV_IRAP_VCL13 It may be defined as a printed picture. Exemplary PPS syntax and s The mantics are as follows: [Table 2] mixed_nalu_types_in_pic_flag indicates that each picture referencing a PPS has multiple VCL NAL units To specify that these NAL units do not have the same value for nal_unit_type Set to equal to 0. mixed_nalu_types_in_pic_flag refers to each pic that references PPS. To specify that the VCL NAL unit of the ja has the same value for nal_unit_type, it is equal to 0. It will be set up in that way.
[0116] The syntax for an exemplary tile group / slice header is as follows: [Table 3A] [Table 3B]
[0117] The semantics of an exemplary NAL unit header are as follows: Any specific pin Regarding the VCL NAL unit of Kucha, one of the following two conditions is met: All VCL NAL units have the same value for nal_unit_type. Some VCL NAL units , values for specific IRAP NAL unit types (i.e., including IDR_W_RADL and RSV_IRAP_VCL13) While all have a nal_unit_type value within the range from IDR_W_RADL to RSV_IRAP_VCL13, all Other VCL NAL units are specific non-IRAP VCL NAL unit types (i.e., TRAIL_ NUT and RSV_VCL_7 are included within the range from TRAIL_NUT to RSV_VCL_7 or RSV_VCL14 and (including RSV_VCL15, and having a nal_unit_type value within the range from RSV_VCL14 to RSV_VCL15) The value obtained by subtracting 1 from nuh_temporal_id_plus1 is the time identifier (tempor) for the NAL unit. Specify the identifier (al). The value of nuh_temporal_id_plus1 is not equal to 0.
[0118] The variable TemporalId is derived as follows: TemporalId = nuh_temporal_id_plus1 - 1 (7-1)
[0119] For VCL NAL units where nal_unit_type is a picture, IDR_W_RADL and RSV_IRAP_VCL13 When it includes and is within the range from IDR_W_RADL to RSV_IRAP_VCL13, other VCs of the picture Regardless of the value of nal_unit_type of the L NAL unit, TemporalId is all VCL of the picture. Equals 0 for NAL units. The value of TemporalId is equal to all VCL NAs of the access unit. The same applies to the L unit. The T unit of the coded picture or access unit. The value of emporalId is the VCL NAL unit of the coded picture or access unit. This is the value of TemporalId.
[0120] An exemplary decryption process for a coded picture is as follows: The process operates as follows with respect to the current picture CurrPic: Recovery of the NAL unit The numbers are shown in detail herein. The following decoding process is performed on the tile group header. Use the syntax elements of the layer and higher layers. Picture order The variables and functions related to the und are derived as shown in detail herein. This is called only for the first tile group / slice of the picture. At the beginning of the decoding process for each group / slice, the reference picture list is constructed. The decoding process involves reference picture list 0 (RefPicList
[0000] ) and reference picture list Called to derive 1(RefPicList
[0001] ). The current picture is an IDR picture. In that case, the decryption process for constructing the reference picture list then proceeds for the bitstream It may be called for the purpose of checking compatibility, but in the current picture or decryption order, the current picture It may not be necessary for decrypting the picture that follows the 'kucha'.
[0121] The decoding process for constructing the reference picture list is as follows: This process This is called at the beginning of the decoding process for each tile group. The reference picture is used for reference. It is addressed by the reference index. The reference index is the reference picture list. This is an index to the I-tile group. When decrypting the I-tile group, the reference picture list is the tile. It is not used in the decoding of P-tile group data. When decoding P-tile groups, Only reference picture list 0 (RefPicList
[0000] ) is used in decoding the tile group data. It is used. When decoding a B-tile group, reference picturelist 0 and reference picturelist Both T1 (RefPicList
[0001] ) and T1 are used in decoding the tile group data. At the beginning of the decoding process for the tile group, the reference picture list RefPicList
[0000] The RefPicList
[0001] is derived. The reference picture list is the marking of the reference picture. Used in or in decoding tile group data. All IDR pictures For tile groups or non-IDR picture I-tile groups, RefPicList
[0000] and RefPicList
[0001] may be derived for the purpose of checking the bitstream conformance. However, their derivation is based on the current picture or the picture that follows the current picture in the decoding order. Not necessary for decryption. Regarding P-tile groups, RefPicList
[0001] is the bit It may be derived for the purpose of checking the suitability of the trim, but the derivation may be based on the current picture or It is not necessary to decrypt the picture that follows the current picture in the decryption order.
[0122] Figure 8 is a schematic diagram of an exemplary video coding device 800. Device 800 implements the examples / embodiments disclosed herein as described herein. It is suitable for the following. The video coding device 800 is connected to the downstream port 820, up Stream port 850, and / or upstream and / or downstream over the network. A transceiver unit including a transmitter and / or receiver for transmitting data in the stream. Includes Tx / Rx)810. The video coding device 800 is a logical unit for processing data. A processor 830 including a t-pack and / or central processing unit (CPU), and a data storage system It further includes memory 832. The video coding device 800 is electrically, optically, or wirelessly Upstream port 850 for data communication via wireless communication network and / or electrical, optical-electrical (OE) components coupled to downstream port 820, electrical- Optical (EO) components and / or wireless communication components may also be included. Video code The input device 800 is an input and / or input device for transmitting data to and from the user. It may also include an output (I / O) device 860. The I / O device 860 displays video data. Output devices such as displays for outputting audio data and speakers for outputting audio data. It may be included. I / O device 860 is an input device such as a keyboard, mouse, or trackball. A chair and / or a corresponding input for interacting with such output devices. The center face may also be included.
[0123] Processor 830 is implemented through hardware and software. 30 refers to one or more CPU chips, cores (for example, as a multi-core processor), and fields Programmable gate arrays (FPGAs), application-specific integrated circuits (ASICs), and digital signals It may be implemented as a processor (DSP). Processor 830 has downstream port 82 0, communicates with Tx / Rx 810, upstream port 850, and memory 832. Processor 83 0 includes coding module 814. Coding module 814 is CVS 500, VR pin You may use the videostream 600 and / or bitstream 700. Implement the embodiments disclosed herein, such as those described in Acts 100, 900, and 1000. Coding module 814 may be used in any other way described herein. Canism may also be implemented. Furthermore, coding module 814 is a codec system A 200 motor, an encoder 300, and / or a decoder 400 may be implemented. For example, a decoder The 814 picture module allows both IRAP NAL units and non-IRAP NAL units to view the picture. This indicates that it includes a single type of IRAP NAL unit and a single type of non-IRAP NAL unit. You can set a flag in the PPS to restrict such pictures to only include the 't'. Therefore, when coding the video data, the coding module 814 The video coding device 800 provides additional functionality and / or coding efficiency. Therefore, the coding module 814 is the function of the video coding device 800. To enhance capabilities and address issues inherent in video coding technology. Furthermore, coding mode Joule 814 results in the conversion of the video coding device 800 to a different state. Alternative Specifically, the coding module 814 is stored in memory 832 and implemented by processor 830. As instructions to be executed (for example, a computer program product stored on a non-temporary medium) It can be implemented as follows.
[0124] Memory 832 is for disk, tape drive, solid state drive, read-only. Memory (ROM), Random Access Memory (RAM), Flash Memory, Tri-Valued Associative Memory (TCAM): (Terminal content-addressable memory), static random access memory (SRAM) It includes one or more memory locations. Memory 832 is selected when the program is executed. To store such programs and instructions that are read during program execution Overflow data storage device for storing data It may be used as an orage device.
[0125] Figure 9 shows the merging of multiple sub-picture video streams 601-603 with multiple video resolutions. Bitstreams such as VR picture video stream 600 and bitstream 700. Video sequences such as CVS 500 containing pictures with mixed NAL unit types This is a flowchart of an exemplary method 900 for encoding S. Method 900 is performed by executing Method 100. The codec system 200, encoder 300, and / or video coding device It may be used with encoders such as the S800.
[0126] Method 900 involves an encoder that processes a video sequence containing multiple pictures, such as VR pictures. It receives and, for example, encodes that video sequence into a bitstream based on user input. It may start when you decide to make it numbered. In step 901, the current picture is different The encoder determines whether it contains multiple subpictures of a certain type. The IPU is at least one slice of the picture that includes part of the IRAP subpicture and non-IRA It may include at least one slice of a picture that contains part of a P NAL subpicture. In step 903, the encoder uses a bitstream to process slices of the subpictures of the picture. Encode into multiple VCL NAL units within the same room. There may be one or more such VCL NAL units. It may include an IRAP NAL unit and one or more non-IRAP NAL units. For example, The encoding step involves converting sub-bitstreams of different resolutions into a single bitstream to be transmitted to the decoder. This may include merging into the bitstream.
[0127] In step 905, the encoder encodes the PPS into a bitstream and the flags into bitstream. Encode into PPS within the stream. For example, encoding PPS can be done even if Then, in accordance with the merging of subbitstreams, the already encoded PPS includes the flag value. This may include making changes to the value of the NAL unit type. It may be set to the first value when it is the same for all VCL NAL units. The flag relates to a VCL NAL unit that contains one or more subpictures of a picture. The value of the first NAL unit type is a VCL NAL unit containing one or more subpictures of the picture. If it differs from the value of the second NAL unit type related to knitting, it may be set to the second value. For example, the value of the first NAL unit type is that the picture contains an IRAP subpicture. It may also be shown that the value of the second NAL unit type includes non-IRAP subpictures. It may also indicate that. Furthermore, the value of the first NAL unit type is IDR_W_RADL, IDR_N_LP, or may be equal to one of the CRA_NUTs. Furthermore, the value of the second NAL unit type is TR It may be equal to one of AIL_NUT, RADL_NUT, or RASL_NUT. As a specific example, The lag may also be mixed_nalu_types_in_pic_flag. In a specific example, mixed_nalu _types_in_pic_flag refers to each picture that references a PPS containing the flag, and has two or more VCL NAL units. It may be set to equal to 1 to specify that it has a set. Furthermore, flat The group consists of all VCL NAL units associated with the corresponding picture of type NAL unit (nal_uni Specify that they do not have the same value for t_type). In another specific example, mixed_nal u_types_in_pic_flag is a flag that refers to one or more VCL NAL units for each picture that references a PPS containing the flag. It may be set to equal to 0 to specify that it has a flag. This means that all VCL NAL units in the corresponding picture have the same value for nal_unit_type. Specify.
[0128] In step 907, the encoder transmits a bitstream to the decoder. You can remember this.
[0129] Figure 10 shows the merging of multiple sub-picture video streams 601-603 with multiple video resolutions. Bitstreams such as VR picture video stream 600 and bitstream 700. Video sequencers such as CVS 500 that include pictures with NAL unit types mixed from M This is a flowchart of an exemplary method 1000 for decoding a signal. Method 1000 is an implementation of Method 100. Sometimes codec system 200, decoder 400, and / or video coding device It may be used with decoders such as the S800.
[0130] Method 1000 is, for example, a coding that represents a video sequence as a result of Method 900. Step may be started when the decoder begins receiving the bitstream of the data. In step 1001, the decoder receives the bitstream. The bitstream is a bitstream. Includes multiple subpictures and flags related to the chat. In a specific example, the bitst The reem may include PPS containing flags. Furthermore, the subpicture may contain multiple VCL NAL units. It is included in the knit. For example, slices related to subpictures are included in the VCL NAL unit. It is included.
[0131] In step 1003, when the flag is set to the first value, the decoder performs NAL unit It was determined that the set type value is the same for all VCL NAL units associated with the picture. Determined. Furthermore, when the flag is set to the second value, the decoder determines the sub-value of the picture. The value of the first NAL unit type for a VCL NAL unit that contains one or more pictures A second NAL unit relating to a VCL NAL unit that includes one or more subpictures of a picture. It is determined that the value is different from the type value. For example, the value of the first NAL unit type is picture It may also indicate that includes an IRAP subpicture, and the value of the second NAL unit type is pict It may also indicate that the first NAL unit type includes non-IRAP subpictures. The value may be equal to one of IDR_W_RADL, IDR_N_LP, or CRA_NUT. Furthermore, the The value of NAL unit type 2 is equal to one of TRAIL_NUT, RADL_NUT, or RASL_NUT. It doesn't have to be. For example, the flag could be mixed_nalu_types_in_pic_flag. The `mixed_nalu_types_in_pic_flag` flag indicates that each picture referencing a PPS has two or more VCL NAL units. It has a set and the VCL NAL unit does not have the same value for NAL unit type (nal_unit_type). When specifying this, it may be set to be equal to 1. Also, mixed_nalu_types_in_p ic_flag is a flag that each picture referencing PPS has one or more VCL NAL units and references PPS When the VCL NAL units of each picture have the same value for nal_unit_type, it is equal to 0. It can also be used as a setting.
[0132] In step 1005, the decoder determines the subpicture based on the NAL unit type value. One or more of these may be decoded. Also, in step 1007, the decoder decodes Transfer one or more of the subpictures to display as part of the video sequence. You may do so.
[0133] Figure 11 shows the merging of multiple sub-picture video streams 601-603 with multiple video resolutions. Bitstreams such as VR picture video stream 600 and bitstream 700. Video sequences such as CVS 500 containing pictures with mixed NAL unit types This is a schematic diagram of an exemplary system 1100 for coding. System 1100 is a coding system. - Deck system 200, encoder 300, decoder 400, and / or video coding It may be implemented by an encoder and decoder such as device 800. Furthermore, sys Tem 1100 may be used when implementing methods 100, 900, and / or 1000.
[0134] System 1100 includes a video encoder 1102. The video encoder 1102 is used to process pictures. Determination module 1101 for determining whether it contains multiple sub-pictures of different types. The video encoder 1102 includes multiple subpictures of a picture within the bitstream. It further includes an encoding module 1103 for encoding to VCL NAL units. The numbering module 1103 contains all VCL NAL units whose NAL unit type values are associated with the picture. It is set to the first value when it is the same for knitting, and one of the subpictures of the picture. The value of the first NAL unit type for VCL NAL units containing one or more is a sub-piece of the picture. The value of the second NAL unit type for a VCL NAL unit containing one or more of the Kucha is different. This is for encoding the flag, which is set to the second value when this happens, into a bitstream. The video encoder 1102 stores the bitstream to be transmitted to the decoder. It further includes a memory module 1105. The video encoder 1102 converts the bitstream into video. The system further includes a transmit module 1107 for sending data to the decoder 1110. 1102 may be further configured to perform any of the steps of method 900.
[0135] System 1100 also includes video decoder 1110. Video decoder 1110 is associated with pictures. The receiving module for receiving a bitstream containing multiple subpictures and flags. Subpictures, including rule 1111, are contained within multiple VCL NAL units. VideoDecode When the flag is set to the first value, the value of the NAL unit type is set to picture. A determination module 1113 determines that all related VCL NAL units are the same. Furthermore, it also includes. In addition, when the flag is set to the second value, the determination module 1113 The first NAL unit relating to a VCL NAL unit containing one or more subpictures of a picture Regarding VCL NAL units where the type value includes one or more subpictures of the picture. This is to determine if it differs from the value of the second NAL unit type. Video decoder 1110 This is a method for decrypting one or more subpictures based on the value of the NAL unit type. The module further includes module 1115. The video decoder 1110 processes the decoded video sequence. A transfer module for transferring one or more subpictures to be displayed as part of a larger image. The following is further included: The video decoder 1110 performs one of the steps of method 1000. It may be further configured as follows.
[0136] The first component is a line, trace, or other between the first component and the second component. When there are no intermediary components other than the medium, it is directly connected to the second component. The element is a connection, trace, or other medium between the first and second components. When there is an intermediary component, it is indirectly connected to the second component. The variations of "bi" include both direct and indirect combinations. (Term: "approximately") The use of means a range including ±10% of the following number unless otherwise stated. ru.
[0137] The steps of the exemplary methods described herein may not necessarily be performed in the order described. It is not required to be carried out, and the order of steps in such a method is merely illustrative. It should also be understood that this is how it should be understood. Similarly, further steps are taken in this way. The method may include a particular step in a manner consistent with various embodiments of this disclosure. In this context, it may be omitted or combined.
[0138] Although several embodiments are given in this disclosure, the disclosed systems and methods , embodied in many other specific forms without exceeding the spirit or scope of this disclosure It will be understood that this is also good. These examples are illustrative and not limiting. The intent should be limited to the details given herein. For example, various elements or components are combined or integrated into another system. It is permissible for certain features to be omitted or not implemented at all.
[0139] In addition, they are described as being separate or distinct in various embodiments, The illustrated technologies, systems, subsystems, and methods do not exceed the scope of this disclosure. Not combined with or integrated with other systems, components, technologies, or methods. Other examples of changes, substitutions, and modifications may be identified by those skilled in the art. This may be done without departing from the spirit and scope disclosed herein. It's possible. [Explanation of Symbols]
[0140] 100 How it works 200 coding and decoding (codec) systems 201 segmented video signal 211 General Coda Control Components 213 Transformation, Scaling, and Quantization Components 215 Intrapicture Estimation Components 217 Intrapicture Prediction Components 219 Motion compensation components 221 Motion Estimation Components 223 Decoded Picture Buffer Components 225 In-loop filter components 227 Filter Control Analysis Components 229 Scaling and Inverse Transformation Components 231 Header Format and Context-Adaptive Binary Arithmetic Coding (CABAC) Configuration element 300 video encoders 301 segmented video signals 313 Transformation and Quantization Components 317 Intrapicture Prediction Components 321 Motion compensation components 323 Decoded Picture Buffer Components 325 In-loop filter components 329 Inverse Transform and Quantization Components 331 Entropy Coding Components 400 video decoders 417 Intrapicture Prediction Components 421 Motion compensation components 423 Decoded Picture Buffer Components 425 In-loop filter components 429 Inverse Transform and Quantization Components 433 Entropy Decoding Components 500 CVS 502 IRAP Picture 504 Reading Pictures 506 Trailing Picture 508 Decryption order 510 Presentation order 600 VR picture video streams 601 Subpicture Video Stream 602 Subpicture video stream 603 Subpicture video stream 700 bitstream 710 Sequence Parameter Set (SPS) 711 Picture Parameter Set (PPS) 715 Slice Header 720 image data 721 Pictures 723 Sub-picture 725 slices 727 Mixed NAL unit type flag in picture (mixed_nalu_types_in_pic_flag) 730 Non-VCL NAL Unit 731 SPS NAL Unit Type (SPS_NUT) 732 PPS NAL unit type (PPS_NUT) 740 VCL NAL Unit 741 IDR_N_LP NAL Unit 742 IDR_w_RADL NAL Unit 743 CRA_NUT 745 IRAP NAL Unit 746 RASL_NUT 747 RADL_NUT 748 TRAIL_NUT 749 Non-IRAP NAL Unit 800 video coding devices 810 Transceiver Unit (Tx / Rx) 814 Coding Modules 820 Downstream Ports 830 Processor 832 memory 850 Upstream Port 860 Input and / or Output (I / O) Devices 900 ways 1000 ways 1100 System 1101 Judgment Module 1102 Video Encoder 1103 Encoding Module 1105 Memory Module 1107 Transmitter Module 1110 Video Decoder 1111 Receiver Module 1113 Judgment Module 1115 Decryption Module 1117 Transfer Module
Claims
1. A method implemented in a decoder, A step of receiving a bitstream containing encoded data and flags for a plurality of pictures, wherein the encoded data for the plurality of pictures is contained in a plurality of video coding layer (VCL) network abstraction layer (NAL) units, the flag equal to a first value specifies that all VCL NAL units of the picture have the same NAL unit type value, the flag equal to a second value specifies that one or more VCL NAL units of the picture have the first NAL unit type value and all other VCL NAL units of the picture have the second NAL unit type value, the bitstream further contains a NAL unit header, the NAL unit header contains nuh_temporal_id_plus1, and nuh_temporal_id_plus1 minus 1 specifies the time identifier of the NAL unit, the value of nuh_temporal_id_plus1 is not equal to 0, The steps include: analyzing the flag from the bitstream; A method comprising the step of decoding one or more of the pictures based on the flag.
2. The method according to claim 1, wherein the bitstream includes a picture parameter set (PPS) containing the flag.
3. The method according to claim 1 or 2, wherein the value of the first NAL unit type is equal to an instantaneous decryption refresh (IDR) with a random access decryptable reading picture (IDR_W_RADL), an IDR without a reading picture (IDR_N_LP), or a clean random access (CRA) NAL unit type (CRA_NUT).
4. The method according to any one of claims 1 to 3, wherein the value of the second NAL unit type is equal to the trailing picture NAL unit type (TRAIL_NUT).
5. The method according to any one of claims 1 to 4, wherein the flag is mixed_nalu_types_in_pic_flag.
6. A method implemented in an encoder, The steps include encoding multiple pictures into multiple video coding layer (VCL) network abstraction layer (NAL) units within a bitstream, A step of encoding a flag into the bitstream, wherein a flag equal to a first value specifies that all VCL NAL units of the picture have the same NAL unit type value, and a flag equal to a second value specifies that one or more VCL NAL units of the picture have the first NAL unit type value, and all other VCL NAL units of the picture have the second NAL unit type value, A method comprising the steps of encoding a NAL unit header into the bitstream, wherein the NAL unit header includes nuh_temporal_id_plus1, and the value obtained by subtracting 1 from nuh_temporal_id_plus1 specifies the time identifier of the NAL unit, and the value of nuh_temporal_id_plus1 is not equal to 0.
7. The method according to claim 6, further comprising the step of encoding a picture parameter set (PPS) into the bitstream, wherein the flags are encoded in the PPS.
8. The method according to claim 6 or 7, wherein the value of the first NAL unit type is equal to an instantaneous decryption refresh (IDR) with a random access decryptable reading picture (IDR_W_RADL), an IDR without a reading picture (IDR_N_LP), or a clean random access (CRA) NAL unit type (CRA_NUT).
9. The method according to any one of claims 6 to 8, wherein the value of the second NAL unit type is equal to the trailing picture NAL unit type (TRAIL_NUT).
10. The method according to any one of claims 6 to 9, wherein the flag is mixed_nalu_types_in_pic_flag.
11. A video coding device comprising a processing circuit configured to perform the method described in any one of claims 1 to 6 or any one of claims 7 to 10.
12. A non-temporary computer-readable medium containing a computer program product for use by a video coding device, wherein the computer program product contains computer-executable instructions stored in the non-temporary computer-readable medium that, when executed by a processor, cause the video coding device to perform the method according to any one of claims 1 to 6 or any one of claims 7 to 10.
13. Receiving means for receiving a bitstream comprising encoded data and flags for a plurality of pictures, wherein the encoded data for the plurality of pictures comprises a plurality of video coding layer (VCL) network abstraction layer (NAL) units, the flag equal to a first value specifies that all VCL NAL units of the picture have the same NAL unit type value, the flag equal to a second value specifies that one or more VCL NAL units of the picture have the first NAL unit type value and all other VCL NAL units of the picture have the second NAL unit type value, the bitstream further comprises a NAL unit header, the NAL unit header comprises nuh_temporal_id_plus1, the value obtained by subtracting 1 from nuh_temporal_id_plus1 specifies a time identifier for the NAL unit, and the value of nuh_temporal_id_plus1 is not equal to 0, the receiving means Analysis means for analyzing the flag from the bitstream, A decoder comprising decoding means for decoding one or more of the pictures based on the aforementioned flag.
14. Encoding means, Encode multiple pictures into multiple Video Coding Layer (VCL) Network Abstraction Layer (NAL) units within a bitstream. A flag, wherein a flag equal to a first value specifies that all VCL NAL units of the picture have the same NAL unit type value, and a flag equal to a second value specifies that one or more VCL NAL units of the picture have the first NAL unit type value and all other VCL NAL units of the picture have the second NAL unit type value, the flag is encoded into the bitstream, An encoder comprising encoding means for encoding a NAL unit header into a bitstream, wherein the NAL unit header includes nuh_temporal_id_plus1, and the value obtained by subtracting 1 from nuh_temporal_id_plus1 specifies a time identifier for the NAL unit, and the value of nuh_temporal_id_plus1 is not equal to 0.
15. A device for storing a bitstream, comprising at least one storage medium and at least one communication interface, The at least one communication interface is configured to receive or transmit the bitstream, and the at least one storage medium is configured to store the bitstream. The bitstream includes encoded data and flags for a plurality of pictures, wherein the encoded data for the plurality of pictures is contained in a plurality of video coding layer (VCL) network abstraction layer (NAL) units, the flag equal to a first value specifies that all VCL NAL units of the picture have the same NAL unit type value, the flag equal to a second value specifies that one or more VCL NAL units of the picture have the first NAL unit type value and all other VCL NAL units of the picture have the second NAL unit type value, the bitstream further includes a NAL unit header, the NAL unit header includes nuh_temporal_id_plus1, and nuh_temporal_id_plus1 minus 1 specifies the time identifier of the NAL unit, and the value of nuh_temporal_id_plus1 is not equal to 0, for the device.
16. A method for storing a bitstream, The steps of receiving or transmitting the bitstream through a communication interface, The step includes storing the bitstream in one or more storage media, The bitstream comprises encoded data and flags for a plurality of pictures, wherein the encoded data for the plurality of pictures comprises a plurality of video coding layer (VCL) network abstraction layer (NAL) units, the flag equal to a first value specifies that all VCL NAL units of the picture have the same NAL unit type value, the flag equal to a second value specifies that one or more VCL NAL units of the picture have the first NAL unit type value and all other VCL NAL units of the picture have the second NAL unit type value, the bitstream further comprises a NAL unit header, the NAL unit header comprises nuh_temporal_id_plus1, and the value obtained by subtracting 1 from nuh_temporal_id_plus1 specifies the time identifier of the NAL unit, the value of nuh_temporal_id_plus1 is not equal to 0, in this manner.
17. A device for transmitting a bitstream, At least one storage medium configured to store at least one bitstream, wherein the at least one bitstream includes encoded data and flags for a plurality of pictures, the encoded data for the plurality of pictures is contained in a plurality of video coding layer (VCL) network abstraction layer (NAL) units, the flag equal to a first value specifies that all VCL NAL units of the picture have the same NAL unit type value, the flag equal to a second value specifies that one or more VCL NAL units of the picture have the first NAL unit type value and all other VCL NAL units of the picture have the second NAL unit type value, the bitstream further includes a NAL unit header, the NAL unit header includes nuh_temporal_id_plus1, and nuh_temporal_id_plus1 minus 1 specifies a time identifier for the NAL unit, the value of nuh_temporal_id_plus1 is not equal to 0, and At least one processor configured to acquire one or more bitstreams from the at least one storage medium, A device including a transmitter configured to transmit one or more bitstreams.
18. A method for transmitting a bitstream, A step of storing at least one bitstream in at least one storage medium, wherein the at least one bitstream includes encoded data and flags for a plurality of pictures, the encoded data for the plurality of pictures is contained in a plurality of video coding layer (VCL) network abstraction layer (NAL) units, the flag equal to a first value specifies that all VCL NAL units of the picture have the same NAL unit type value, the flag equal to a second value specifies that one or more VCL NAL units of the picture have the first NAL unit type value and all other VCL NAL units of the picture have the second NAL unit type value, the bitstream further includes a NAL unit header, the NAL unit header includes nuh_temporal_id_plus1, and nuh_temporal_id_plus1 minus 1 specifies the time identifier of the NAL unit, the value of nuh_temporal_id_plus1 is not equal to 0, The steps include obtaining one or more bitstreams from the at least one storage medium, A method comprising the step of transmitting one or more bitstreams.
Citation Information
Patent Citations
Video codec allowing sub-picture or region wise random access and concept for video composition using the same
WO2020157287A1