Coding method and encoder supporting mixed NAL unit types in one picture
By receiving and decoding a code stream containing flags and multiple sub-images in the decoder, and processing images according to flags and unit type values, the problems of insufficient compression ratio and mixed image processing in the prior art are solved, and efficient video decoding and dynamic resolution changes are achieved.
Patent Information
- Application Number
- CN202211188702.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Priority Date
- 2019-04-10
- Filing Date
- 2020-03-11
- Publication Date
- 2025-05-06
- Estimated Expiration
- 2040-03-11
AI Technical Summary
The existing video decoding technology is difficult to effectively improve the compression ratio during compression and decompression, especially in the case of limited network resources and high video quality requirements, and cannot effectively handle the situation where mixed images include IRAP and non-IRAP areas.
By implementing the method in the decoder, a code stream containing a flag and a plurality of sub-images is received, a video decoding layer network abstract layer unit type where the sub-image is located is determined according to the flag, and a sub-image is decoded according to these type values. At the same time, the image is constrained using a hybrid image flag to include IRAP and non-IRAP NAL unit types, supporting dynamic resolution changes.
It improves the efficiency of video decoding, reduces the use of network resources, memory resources and processing resources, and supports streaming VR videos, which can transmit low-resolution sub-image code streams without affecting the user experience.
Smart Images

Figure CN115550659B_ABST
Abstract
Description
[0001] This application is a divisional application. The application number of the original application is 202080019678.8, and the original application date is March 11, 2020. The entire contents of the original application are incorporated into this application by reference. Technical Field
[0002] The present invention relates generally to video coding, and more particularly to coding of sub-pictures of pictures in video coding. Background Art
[0003] Even in the case of short videos, a large amount of video data needs to be described, which can cause difficulties when the data is to be sent or otherwise transmitted in a communication network with limited bandwidth capacity. Therefore, video data is usually compressed before being sent in modern telecommunications networks. Due to limited memory resources, the size of the video can also become a problem when storing the video in a storage device. Video compression devices usually use software and / or hardware on the source side to decode the video data before sending or storing, thereby reducing the amount of data required to represent the digital video image. The compressed data is then received on the destination side by a video decompression device for decoding the video data. With limited network resources and a growing demand for higher video quality, there is a need for improved compression and decompression techniques that can increase the compression ratio with little impact on image quality. Summary of the invention
[0004] In one embodiment, the present invention includes a method implemented in a decoder, the method comprising: a receiver of the decoder receives a code stream including a flag and multiple sub-images associated with an image, wherein the sub-images are contained in multiple video coding layer (VCL) network abstraction layer (NAL) units; the processor determines, based on the value of the flag, that the VCL NAL units where one or more sub-images of the sub-images of the image are located all have a first specific NAL unit type value, and other VCL NAL units of the image all have different second specific NAL unit type values; the processor decodes one or more sub-images of the sub-image based on the first specific NAL unit type value or the second specific NAL unit type value.
[0005] An image can be divided into multiple sub-images. These sub-images can be encoded into different sub-streams, which can then be merged into a stream for transmission to a decoder. For example, sub-images can be used in virtual reality (VR) applications. In a specific example, a user can only view a portion of a VR image at any time. Therefore, different sub-images can be transmitted at different resolutions, so that more bandwidth can be allocated to sub-images that may be displayed, and sub-images that are unlikely to be displayed can be compressed to improve decoding efficiency. In addition, a video stream can be encoded by using intra-random access point (IRAP) images. IRAP images are encoded based on intra-frame prediction and can be decoded without reference to other images. Non-IRAP images can be encoded based on inter-frame prediction and can be decoded by reference to other images. Non-IRAP images are more compressed than IRAP images. However, the video sequence must start decoding with an IRAP image because the IRAP image contains enough data to be decoded without reference to other images. IRAP images can be used for sub-images to support dynamic resolution changes. Therefore, the video system can transmit more IRAP images for sub-images that are more likely to be viewed (for example, based on the user's current viewing angle) and fewer IRAP images for sub-images that are less likely to be viewed, so as to further improve decoding efficiency. However, the sub-images are part of the same image. Therefore, the scheme can obtain an image that contains both IRAP sub-images and non-IRAP sub-images. Some video systems cannot handle mixed images that have both IRAP and non-IRAP areas. The present invention includes a flag indicating whether the image is a mixed image and therefore contains both IRAP and non-IRAP components. In addition, the flag constrains the image so that the mixed image contains exactly two NAL unit types: one IRAP type and one non-IRAP type. Based on the flag, the decoder can process different sub-images in different ways when decoding so as to correctly decode and display the image / sub-image. The flag can be stored in the PPS and can be called mixed_nalu_types_in_pic_flag. Therefore, the disclosed mechanism enables other functions to be implemented. In addition, the disclosed mechanism supports dynamic resolution changes when using sub-image code streams. Therefore, the disclosed mechanism enables the transmission of low-resolution sub-image code streams when streaming VR videos without significantly affecting the user experience. Therefore, the disclosed mechanism improves decoding efficiency, so that the encoder and decoder use less network resources, memory resources and / or processing resources.
[0006] Optionally, according to any one of the above aspects, in another implementation of the aspect, the first specific NAL unit type value indicates that the image contains a single type of intra-random access point (IRAP) sub-image, and the second specific NAL unit type value indicates that the image contains a single type of non-IRAP sub-image.
[0007] Optionally, according to any one of the above aspects, in another implementation manner of the aspect, the code stream includes a picture parameter set (PPS) including the flag.
[0008] Optionally, according to any of the above aspects, in another implementation of the aspect, the first specific NAL unit type value is equal to an instantaneous decoding refresh (IDR) with a decodable random access preceding picture (IDR_W_RADL), an IDR without a preceding picture (IDR_N_LP), or a clean random access (CRA) NAL unit type (CRA_NUT).
[0009] Optionally, according to any of the above aspects, in another implementation of the aspect, the second specific NAL unit type value is equal to the trailing image NAL unit type (TRAIL_NUT), the decodable random access leading image NAL unit type (RADL_NUT), or the skipped random access leading (random access skipped leading, RASL) image NAL unit type (RASL_NUT).
[0010] Optionally, according to any of the above aspects, in another implementation manner of the aspect, the flag is mixed_nalu_types_in_pic_flag.
[0011] Optionally, according to any of the above aspects, in another implementation of the aspect, when each image referring to the PPS has more than one VCL NAL unit and the NAL unit type (nal_unit_type) values of the VCL NAL units are different, mixed_nalu_types_in_pic_flag is equal to 1; when each image referring to the PPS has one or more VCL NAL units and the nal_unit_type values of the VCL NAL units of each image referring to the PPS are the same, mixed_nalu_types_in_pic_flag is equal to 0.
[0012] In one embodiment, the present invention includes a method implemented in an encoder, the method comprising: the processor determines that a picture contains multiple sub-pictures of different types; the processor encodes the sub-pictures of the picture into multiple video coding layer (VCL) network abstraction layer (NAL) units in a bitstream; the processor encodes a flag into the bitstream, the flag being set to indicate that the VCL NAL units containing one or more of the sub-pictures of the picture all have a first specific NAL unit type value, and other VCL NAL units of the picture all have a different second specific NAL unit type value; a memory coupled to the processor stores the bitstream for sending to a decoder.
[0013] An image can be divided into multiple sub-images. These sub-images can be encoded into different sub-streams, which can then be merged into a stream for transmission to a decoder. For example, sub-images can be used in virtual reality (VR) applications. In a specific example, a user can only view a portion of a VR image at any time. Therefore, different sub-images can be transmitted at different resolutions, so that more bandwidth can be allocated to sub-images that may be displayed, and sub-images that are unlikely to be displayed can be compressed to improve decoding efficiency. In addition, a video stream can be encoded by using intra-random access point (IRAP) images. IRAP images are encoded based on intra-frame prediction and can be decoded without reference to other images. Non-IRAP images can be encoded based on inter-frame prediction and can be decoded by reference to other images. Non-IRAP images are more compressed than IRAP images. However, the video sequence must start decoding with an IRAP image because the IRAP image contains enough data to be decoded without reference to other images. IRAP images can be used for sub-images to support dynamic resolution changes. Therefore, the video system can transmit more IRAP images for sub-images that are more likely to be viewed (for example, based on the user's current viewing angle) and fewer IRAP images for sub-images that are less likely to be viewed, so as to further improve decoding efficiency. However, the sub-images are part of the same image. Therefore, the scheme can obtain an image that contains both IRAP sub-images and non-IRAP sub-images. Some video systems cannot handle mixed images that have both IRAP and non-IRAP areas. The present invention includes a flag indicating whether the image is a mixed image and therefore contains both IRAP and non-IRAP components. In addition, the flag constrains the image so that the mixed image contains exactly two NAL unit types: one IRAP type and one non-IRAP type. Based on the flag, the decoder can process different sub-images in different ways when decoding so as to correctly decode and display the image / sub-image. The flag can be stored in the PPS and can be called mixed_nalu_types_in_pic_flag. Therefore, the disclosed mechanism enables other functions to be implemented. In addition, the disclosed mechanism supports dynamic resolution changes when using sub-image code streams. Therefore, the disclosed mechanism enables the transmission of low-resolution sub-image code streams when streaming VR videos without significantly affecting the user experience. Therefore, the disclosed mechanism improves decoding efficiency, so that the encoder and decoder use less network resources, memory resources and / or processing resources.
[0014] Optionally, according to any of the above aspects, in another implementation of the aspect, the first specific NAL unit type value indicates that the image contains a single type of IRAP sub-image, and the second specific NAL unit type value indicates that the image contains a single type of non-IRAP sub-image.
[0015] Optionally, according to any of the above aspects, in another implementation manner of the aspect, it also includes encoding a PPS into the bitstream, wherein the flag is encoded into the PPS.
[0016] Optionally, according to any of the above aspects, in another implementation of the aspect, the first specific NAL unit type value is equal to IDR_W_RADL, IDR_N_LP or CRA_NUT.
[0017] Optionally, according to any of the above aspects, in another implementation of the aspect, the second specific NAL unit type value is equal to TRAIL_NUT, RADL_NU or RASL_NU.
[0018] Optionally, according to any of the above aspects, in another implementation manner of the aspect, the flag is mixed_nalu_types_in_pic_flag.
[0019] Optionally, according to any of the above aspects, in another implementation of the aspect, when each image referring to the PPS has more than one VCL NAL unit and the nal_unit_type values of the VCL NAL units are different, mixed_nalu_types_in_pic_flag is equal to 1; when each image referring to the PPS has one or more VCL NAL units and the nal_unit_type values of the VCL NAL units of each image referring to the PPS are the same, mixed_nalu_types_in_pic_flag is equal to 0.
[0020] In one embodiment, the present invention includes a video decoding device, which includes: a processor, a receiver coupled to the processor, a memory coupled to the processor, and a transmitter coupled to the processor, wherein the processor, the receiver, the memory, and the transmitter are used to execute the method described in any one of the above aspects.
[0021] In one embodiment, the present invention includes a non-transitory computer-readable medium, which includes a computer program product for use by a video decoding device, wherein the computer program product includes computer executable instructions stored in the non-transitory computer-readable medium, and when the computer executable instructions are executed by a processor, the video decoding device performs the method according to any one of the above aspects.
[0022] In one embodiment, the present invention includes a decoder, comprising: a receiving module for receiving a code stream including a flag and multiple sub-images associated with an image, wherein the multiple sub-images are contained in multiple VCL NAL units; a determining module for determining, based on the value of the flag, that the VCL NAL units in which one or more of the sub-images of the image are located all have a first specific NAL unit type value, and that other VCL NAL units of the image all have different second specific NAL unit type values; a decoding module for decoding one or more of the sub-images according to the first specific NAL unit type value or the second specific NAL unit type value; and a forwarding module for forwarding one or more of the sub-images for display as part of a decoded video sequence.
[0023] Optionally, according to any of the above aspects, in another implementation manner of the aspect, the decoder is further used to execute the method according to any of the above aspects.
[0024] In one embodiment, the present invention includes an encoder, comprising: a determination module, used to determine that an image contains multiple sub-images of different types; an encoding module, used to: encode the sub-images of the image into multiple VCL NAL units in a bitstream; encode a flag into the bitstream, wherein the flag is set to indicate that the VCL NAL units where one or more sub-images of the sub-images of the image are located all have a first specific NAL unit type value, and other VCL NAL units of the image all have different second specific NAL unit type values; and a storage module, used to store the bitstream for sending to a decoder.
[0025] Optionally, according to any of the above aspects, in another implementation manner of the aspect, the encoder is also used to execute the method according to any of the above aspects.
[0026] For the sake of clarity, any of the embodiments described above may be combined with any one or more of the other embodiments described above to create new embodiments within the scope of the present invention.
[0027] These and other features will become more clearly understood from the following detailed description taken in conjunction with the accompanying drawings and claims. BRIEF DESCRIPTION OF THE DRAWINGS
[0028] For a more complete understanding of the present invention, reference is made to the following brief description taken in conjunction with the accompanying drawings and detailed description, wherein like reference numerals represent like parts.
[0029] Figure 1 A flow chart of an exemplary method of decoding a video signal.
[0030] Figure 2 is a schematic diagram of an exemplary encoding and decoding (codec) system for video coding.
[0031] Figure 3 is a schematic diagram of an exemplary video encoder.
[0032] Figure 4 is a schematic diagram of an exemplary video decoder.
[0033] Figure 5 is a schematic diagram of an exemplary decoded video sequence.
[0034] Figure 6 A schematic diagram of multiple sub-image video streams divided from a virtual reality (VR) image video stream.
[0035] Figure 7 is a schematic diagram of an exemplary codestream containing a picture with mixed network abstraction layer (NAL) unit types.
[0036] Figure 8 is a schematic diagram of an exemplary video decoding device.
[0037] Fig. 9 A flow chart of an exemplary method for encoding a video sequence containing pictures with mixed NAL unit types into a bitstream.
[0038] Fig.10 A flow chart of an exemplary method for decoding a video sequence containing pictures with mixed NAL unit types from a codestream.
[0039] Fig.11 A schematic diagram of an exemplary system for encoding a video sequence containing pictures with mixed NAL unit types into a bitstream. DETAILED DESCRIPTION
[0040] First, it should be understood that although illustrative implementations of one or more embodiments are provided below, the disclosed systems and / or methods can be implemented using any number of techniques, whether currently known or existing. The present invention should not be limited to the illustrative implementations, drawings, and techniques described below, including the exemplary designs and implementations illustrated and described herein, but may be modified within the scope of the appended claims and the full scope of their equivalents.
[0041] This document uses the following abbreviations: coded video sequence (CVS), decoded picture buffer (DPB), instantaneous decoding refresh (IDR), intra random access point (IRAP), least significant bit (LSB), most significant bit (MSB), network abstraction layer (NAL), picture order count (POC), raw byte sequence payload (RBSP), sequence parameter set (SPS), and working draft (WD).
[0042] Many video compression techniques can be used to reduce the size of video files while minimizing data loss. For example, video compression techniques can include performing spatial (e.g., intra-frame) prediction and / or temporal (e.g., inter-frame) prediction to reduce or remove data redundancy in a video sequence. For block-based video decoding, a video slice (e.g., a video image or a portion of a video image) can be divided into video blocks, which can also be referred to as tree blocks, coding tree blocks (CTB), coding tree units (CTU), coding units (CU) and / or coding nodes. Video blocks in an intra-frame coding (I) slice of an image are encoded using spatial prediction for reference samples in neighboring blocks in the same image. Video blocks in an inter-frame unidirectional prediction (P) slice or bidirectional prediction (B) slice of an image can use spatial prediction for reference samples in neighboring blocks in the same image, or use temporal prediction for reference samples in other reference images. A picture (picture / image) can be referred to as a frame, and a reference image can be referred to as a reference frame. Spatial or temporal prediction produces a prediction block representing an image block. The residual data represents the pixel difference between the original image block and the predicted block. Therefore, the inter-coded block is encoded based on the motion vector pointing to the block of reference samples constituting the predicted block and the residual data representing the difference between the coded block and the predicted block. The intra-coded block is encoded according to the intra-coding mode and the residual data. For further compression, the residual data can be transformed from the pixel domain to the transform domain to produce residual transform coefficients, which can be quantized. The quantized transform coefficients are initially arranged in a two-dimensional array. The quantized transform coefficients can be scanned to generate a one-dimensional vector of transform coefficients. Entropy coding can be applied to achieve further compression. This video compression technique is described in detail below.
[0043] To ensure that the encoded video can be correctly decoded, the video is encoded and decoded according to the corresponding video coding standards. Video coding standards include international telecommunication union (ITU) standardization sector, ITU-T H.261, international organization for standardization / international electrotechnical commission (ISO / IEC) motion picture experts group (MPEG)-1 part 2, ITU-T H.262 or ISO / IEC MPEG-2 part 2, ITU-T H.263, ISO / IEC MPEG-4 part 2, advanced video coding (AVC) (also known as ITU-T H.264 or ISO / IEC MPEG-4 part 10), and high efficiency video coding (HEVC) (also known as ITU-T H.265 or MPEG-H part 2). AVC includes scalable video coding (SVC), multiview video coding (MVC), multiview video coding plus depth (MVC+D), and three-dimensional (3D) AVC (3D-AVC) and other extensions. HEVC includes scalable HEVC (SHVC), multiview HEVC (MV-HEVC), and 3D HEVC (3D-HEVC) and other extensions. The joint video experts team (JVET) of ITU-T and ISO / IEC has begun to develop a video coding standard called versatile video coding (VVC). VVC is included in the working draft (WD), which includes JVET-M1001-v6, which provides algorithm descriptions, encoder-side descriptions of VVC WD, and reference software.
[0044] A video decoding system can encode video by using IRAP pictures and non-IRAP pictures. IRAP pictures are pictures coded according to intra-frame prediction and are used as random access points for a video sequence. In intra-frame prediction, blocks of a picture are coded by referring to other blocks in the same picture. This is in stark contrast to non-IRAP pictures that use inter-frame prediction. In inter-frame prediction, blocks of the current picture are coded by referring to other blocks in a reference picture different from the current picture. Since IRAP pictures are coded without referring to other pictures, IRAP pictures can be decoded without first decoding any other pictures. Therefore, a decoder can start decoding a video sequence at any IRAP picture. In contrast, non-IRAP pictures are coded with reference to other pictures, so a decoder cannot usually start decoding a video sequence at a non-IRAP picture. IRAP pictures also refresh the DPB. This is because IRAP pictures are the starting point of a CVS, and pictures in a CVS do not refer to pictures in previous CVSs. Therefore, IRAP pictures can also stop decoding errors related to inter-frame prediction because such errors cannot propagate through IRAP pictures. However, from a data size perspective, IRAP pictures are significantly larger than non-IRAP pictures. Therefore, a video sequence usually includes many non-IRAP pictures, among which a smaller number of IRAP pictures are interspersed to balance decoding efficiency and functionality. For example, a 60-frame CVS may include 1 IRAP picture and 59 non-IRAP pictures.
[0045] In some cases, the video decoding system can be used to decode virtual reality (VR) videos, which can also be referred to as 360-degree videos. VR videos can include video content of a sphere that is displayed as if the user is at the center of the sphere. Only a portion of the sphere (called the perspective) is displayed to the user. For example, the user can use a head mounted display (HMD) that selects and displays the perspective of the sphere based on the user's head movements. This creates the effect of physically existing in the virtual space depicted by the video. To achieve this result, each image of the video sequence includes the entire sphere of video data at the corresponding moment. However, only a small portion of the image (e.g., a single perspective) is displayed to the user. The rest of the image is discarded and not presented. The entire image is usually transmitted so that different perspectives can be dynamically selected and displayed in response to the user's head movements. The video file size of this approach is very large.
[0046] To improve decoding efficiency, some systems divide the image into sub-images. A sub-image is a defined spatial region of an image. Each sub-image contains the corresponding view of the image. Video can be encoded at two or more resolutions. Each resolution is encoded into a different sub-stream. When a user streams a VR video, the decoding system can merge the sub-streams into one stream for transmission based on the current view used by the user. Specifically, the current view is obtained from the high-resolution sub-stream, and the unwatched view is obtained from the low-resolution stream. In this way, the highest quality video is displayed to the user and the low-quality video is discarded. If the user selects a new view, a lower resolution video is displayed to the user. The decoder can request a higher resolution video from the new view. The encoder can then change the fusion process accordingly. Once the IRAP image arrives, the decoder can start decoding the higher resolution video sequence at the new view. This approach significantly improves video compression without negatively affecting the user's viewing experience.
[0047] One problem with the above approach is that the length of time required to change the resolution is based on the length of time before an IRAP picture arrives. This is because the decoder cannot start decoding different video sequences at a non-IRAP picture as described above. One way to reduce this delay is to include more IRAP pictures. However, this will increase the file size. In order to balance functionality with decoding efficiency, different views / sub-images can include IRAP images at different frequencies. For example, views that are more likely to be viewed can have more IRAP images than other views. For example, in a basketball environment, views associated with the basket and / or midfield can include a greater frequency of IRAP images than views used to view the stands or ceiling because users are less likely to view views used to view the stands or ceiling.
[0048] This approach leads to other problems. Specifically, the sub-images containing the view are part of a single image. When different sub-images have IRAP images with different frequencies, some images include both IRAP sub-images and non-IRAP sub-images. This is a problem because images are stored in the bitstream using NAL units. A NAL unit is a storage unit that contains a parameter set or slice of an image and the corresponding slice header. An access unit is a unit that contains an entire image. Therefore, an access unit contains all NAL units associated with an image. The NAL unit also contains a type that indicates the type of image that includes the slice. In some video systems, all NAL units associated with a single image (for example, contained in the same access unit) need to be of the same type. Therefore, when an image includes both IRAP sub-images and non-IRAP sub-images, the NAL unit storage mechanism may stop operating correctly.
[0049] This document discloses a mechanism for adjusting the NAL storage scheme to support images that include both IRAP sub-images and non-IRAP sub-images. This in turn allows VR videos to include different IRAP sub-image frequencies from different perspectives. In a first example, this document discloses a flag indicating whether an image is a mixed image. For example, the flag may indicate that an image contains both IRAP sub-images and non-IRAP sub-images. Based on the flag, the decoder can handle different types of sub-images in different ways when decoding in order to correctly decode and display the image / sub-image. The flag may be stored in a picture parameter set (PPS) and may be referred to as mixed_nalu_types_in_pic_flag.
[0050] In a second example, a flag indicating whether a picture is a mixed picture is disclosed herein. For example, the flag may indicate that the picture contains both IRAP sub-pictures and non-IRAP sub-pictures. In addition, the flag constrains the picture so that the mixed picture contains exactly two NAL unit types: one IRAP type and one non-IRAP type. For example, a picture may contain an IRAP NAL unit, the IRAP NAL unit including one and only one of an instantaneous decoding refresh (IDR) with a decodable random access leading picture (IDR_W_RADL), an IDR without a leading picture (IDR_N_LP), or a clean random access (CRA) NAL unit type (CRA_NUT). In addition, a picture may contain a non-IRAP NAL unit, the non-IRAP NAL unit including one and only one of a trailing picture NAL unit type (TRAIL_NUT), a decodable random access leading picture NAL unit type (RADL_NUT), or a skipped random access leading (RASL) picture NAL unit type (RASL_NUT). According to this flag, the decoder can process different sub-pictures in different ways when decoding in order to correctly decode and display the picture / sub-picture. This flag can be stored in the PPS and can be called mixed_nalu_types_in_pic_flag.
[0051] Figure 1Flowchart of an exemplary method 100 of operation for decoding a video signal. Specifically, the video signal is encoded at the encoder side. The encoding process compresses the video signal using various mechanisms to reduce the video file size. The smaller file size is advantageous for sending the compressed video file to the user while reducing the associated bandwidth overhead. The decoder then decodes the compressed video file to reconstruct the original video signal for display to the end user. The decoding process generally corresponds to the encoding process so that the decoder can consistently reconstruct the video signal.
[0052] In step 101, a video signal is input to an encoder. For example, the video signal can be an uncompressed video file stored in a memory. In another example, the video file can be captured by a video capture device (e.g., a video camera) and encoded to support live broadcast of the video. The video file can include an audio component and a video component. The video component includes a series of image frames that, when viewed in sequence, produce a visual effect of motion. These frames include pixels represented by light (referred to herein as luminance components (or luminance samples)) and color (referred to as chrominance components (or chrominance samples)). In some examples, these frames can also include depth values to support three-dimensional viewing.
[0053] In step 103, the video is segmented into blocks. Segmentation includes subdividing the pixels in each frame into square and / or rectangular blocks for compression. For example, in high efficiency video coding (HEVC) (also known as H.265 and MPEG-H Part 2), the frames can be first divided into coding tree units (CTUs), which are blocks of a predefined size (e.g., 64×64 pixels). The CTU includes luma samples and chroma samples. The CTU can be divided into blocks using a coding tree, and then the blocks are recursively subdivided until a configuration that supports further encoding is obtained. For example, the luma component of a frame can be subdivided until each block includes a relatively uniform luma value. In addition, the chroma component of a frame can be subdivided until each block includes a relatively uniform color value. Therefore, the segmentation mechanism varies depending on the content of the video frame.
[0054] In step 105, various compression mechanisms are used to compress the image blocks segmented in step 103. For example, inter-frame prediction and / or intra-frame prediction can be used. Inter-frame prediction is intended to take advantage of the fact that objects in a common scene tend to appear in consecutive frames. Therefore, there is no need to repeatedly describe the blocks depicting the object in the reference frame in adjacent frames. Specifically, an object (such as a table) can remain in a constant position in multiple frames. Therefore, the table is only described once, and adjacent frames can re-reference the reference frame. Pattern matching mechanisms can be used to match objects in multiple frames. In addition, due to object movement or camera movement, etc., a moving object can be represented by multiple frames. As a specific example, a video can show a car moving on the screen through multiple frames. Motion vectors can be used to describe this movement. A motion vector is a two-dimensional vector that provides an offset from the coordinates of an object in a frame to the coordinates of the object in a reference frame. Therefore, inter-frame prediction can encode the image blocks in the current frame as a set of motion vectors, representing the offset relative to the corresponding blocks in the reference frame.
[0055] Intra-frame prediction encodes blocks in a common frame. Intra-frame prediction takes advantage of the fact that luminance components and chrominance components tend to cluster in a frame. For example, a piece of green in part of a tree is often adjacent to several similar pieces of green. Intra-frame prediction uses a variety of directional prediction modes (e.g., 33 modes in HEVC), plane mode, and direct current (DC) mode. The directional mode indicates that the samples of the current block are similar / identical to those of the neighboring blocks in the corresponding direction. The plane mode indicates that a series of blocks on a row / column (e.g., a plane) can be interpolated based on the neighboring blocks at the edge of the row. In fact, the plane mode represents the smooth transition of light / color on the row / column by using a relatively constant slope of the changing value. The DC mode is used for boundary smoothing, indicating that the block is similar / identical to the average value associated with the samples of all neighboring blocks, which are associated with the angular direction of the directional prediction mode. Therefore, the intra-frame prediction block can represent the image block as various relational prediction mode values instead of the actual value. In addition, the inter-frame prediction block can represent the image block as a motion vector value instead of the actual value. In either case, the prediction block may not accurately represent the image block in some cases. All difference values are stored in the residual block. The residual block can be transformed to further compress the file.
[0056] In step 107, various filtering techniques can be applied. In HEVC, filters are applied according to an in-loop filtering scheme. The block-based prediction described above may create a blocky image on the decoder side. In addition, a block-based prediction scheme can encode a block and then reconstruct the coded block for later use as a reference block. The in-loop filtering scheme iteratively applies a noise suppression filter, a deblocking filter, an adaptive loop filter, and a sample adaptive offset (SAO) filter to a block / frame. These filters reduce such block artifacts so that the encoded file can be accurately reconstructed. In addition, these filters reduce artifacts in the reconstructed reference block, making it less likely that the artifact will produce other artifacts in subsequent blocks encoded according to the reconstructed reference block.
[0057] In step 109, once the video signal has been segmented, compressed and filtered, the resulting data is encoded into a bitstream. The bitstream includes the above data and any indicative data required to support appropriate reconstruction of the video signal at the decoder side. For example, such data may include segmentation data, prediction data, residual blocks, and various flags that provide decoding instructions to the decoder. The bitstream can be stored in a memory so that it can be sent to the decoder upon request. The bitstream can also be broadcast and / or multicast to multiple decoders. Creating a bitstream is an iterative process. Therefore, steps 101, 103, 105, 107 and 109 can be performed continuously and / or simultaneously in multiple frames and blocks. Figure 1 The order shown is presented for clarity and ease of description and is not intended to limit the video coding process to a particular order.
[0058] In step 111, the decoder receives the code stream and starts the decoding process. Specifically, the decoder uses an entropy decoding scheme to convert the code stream into corresponding syntax data and video data. In step 111, the decoder uses the syntax data in the code stream to determine the segmented parts of the frame. The segmentation should match the block segmentation result in step 103. The entropy coding / entropy decoding used in step 111 is described here. The encoder makes many choices during the compression process, such as selecting a block segmentation scheme from multiple possible choices based on the spatial positioning of values in one or more input images. Indicating the exact choice may use a large number of bits. The bits used in this article are binary values that are considered variables (for example, bit values that can vary depending on the context). Entropy coding allows the encoder to discard any options that are obviously not suitable for a particular situation, leaving a set of usable options. Then, a codeword is assigned to each usable option. The length of the codeword depends on the number of usable options (for example, one bit corresponds to two options, two bits correspond to three or four options, etc.). Then, the encoder encodes the codeword of the selected option. This scheme reduces the size of the codeword because the codeword is as large as would be needed to uniquely represent an option from a small subset of available options, rather than a potentially large set of all possible options. The decoder then decodes the selection by determining the set of available options in a similar manner to the encoder. By determining the set of available options, the decoder can read the codeword and determine the selection made by the encoder.
[0059] In step 113, the decoder performs block decoding. Specifically, the decoder performs an inverse transform to generate a residual block. Then, the decoder uses the residual block and the corresponding prediction block to reconstruct the image block according to the segmentation. The prediction block may include an intra-frame prediction block and an inter-frame prediction block generated by the encoder side in step 105. Then, the reconstructed image block is positioned in the frame of the reconstructed video signal according to the segmentation data determined in step 111. The syntax of step 113 can also be indicated in the bitstream by entropy coding described above.
[0060] In step 115, the frames of the reconstructed video signal are filtered at the encoder side in a manner similar to step 107. For example, a noise suppression filter, a deblocking filter, an adaptive loop filter, and a SAO filter may be applied to the frames to remove blocking artifacts. Once the frames are filtered, in step 117, the video signal may be output to a display for viewing by an end user.
[0061] Figure 21 is a schematic diagram of an exemplary encoding and decoding (codec) system 200 for video decoding. Specifically, the codec system 200 can implement the operating method 100. The codec system 200 broadly describes the components used in an encoder and a decoder. The codec system 200 receives a video signal and segments the video signal, as described in steps 101 and 103 of the operating method 100, thereby generating segmented video signals 201. Then, when acting as an encoder, the codec system 200 compresses the segmented video signals 201 into encoded bitstreams, as described in steps 105, 107, and 109 of the method 100. When acting as a decoder, the codec system 200 generates an output video signal from the bitstream, as described in steps 111, 113, 115, and 117 of the operating method 100. The codec system 200 includes a general decoder control component 211, a transform scaling and quantization component 213, an intra-frame estimation component 215, an intra-frame prediction component 217, a motion compensation component 219, a motion estimation component 221, a scaling and inverse transform component 229, a filter control analysis component 227, an in-loop filter component 225, a decoded image buffer component 223, and a header formatting and context adaptive binary arithmetic coding (context adaptive binary arithmetic coding, CABAC) component 231. These components are coupled as shown. Figure 2 In the figure, the black lines represent the movement of data to be encoded / decoded, and the dotted lines represent the movement of control data that controls the operation of other components. The components of the codec system 200 can all be used in an encoder. A decoder can include a subset of the components of the codec system 200. For example, a decoder can include an intra-frame prediction component 217, a motion compensation component 219, a scaling and inverse transform component 229, an in-loop filter component 225, and a decoded image buffer component 223. These components are described here.
[0062] The segmented video signal 201 is a captured video sequence that has been segmented into pixel blocks by a coding tree. The coding tree uses various partitioning modes to subdivide the pixel blocks into smaller pixel blocks. These blocks can then be further subdivided into smaller blocks. These blocks can be referred to as nodes on the coding tree. Larger parent nodes are divided into smaller child nodes. The number of times a node is subdivided is called the depth of the node / coding tree. In some cases, a coding unit (CU) may include a partitioned block. For example, a CU may be a sub-part of a CTU, which includes a luminance block, a red difference chrominance (Cr) block, and a blue difference chrominance (Cb) block and corresponding syntax instructions for the CU. The partitioning mode may include a binary tree (BT), a triple tree (TT), and a quad tree (QT), which are used to partition a node into two, three, or four child nodes of different shapes, respectively, according to the partitioning mode used. The segmented video signal 201 is forwarded to the general decoder control component 211, the transform scaling and quantization component 213, the intra-frame estimation component 215, the filter control analysis component 227 and the motion estimation component 221 for compression.
[0063] The universal decoder control component 211 is used to make decisions related to encoding the images of the video sequence into the bitstream according to the application constraints. For example, the universal decoder control component 211 manages the optimization of the bitrate / bitstream size relative to the reconstruction quality. Such decisions can be made based on the storage space / bandwidth availability and the image resolution request. The universal decoder control component 211 also manages the buffer utilization according to the transmission speed to alleviate the buffer under-load and overload problems. In order to solve these problems, the universal decoder control component 211 manages the segmentation, prediction and filtering performed by other components. For example, the universal decoder control component 211 can dynamically increase the compression complexity to increase the resolution and increase the bandwidth utilization, or reduce the compression complexity to reduce the resolution and bandwidth utilization. Therefore, the universal decoder control component 211 controls other components of the codec system 200 to balance the video signal reconstruction quality and bitrate issues. The universal decoder control component 211 creates control data that controls the operation of other components. The control data is also forwarded to the header formatting and CABAC component 231 to be encoded into the codestream, indicating parameters for decoding in the decoder.
[0064] The segmented video signal 201 is also sent to the motion estimation component 221 and the motion compensation component 219 for inter-frame prediction. The frames or slices of the segmented video signal 201 can be divided into multiple video blocks. The motion estimation component 221 and the motion compensation component 219 perform inter-frame prediction decoding on the received video blocks relative to one or more blocks in one or more reference frames to provide temporal prediction. The codec system 200 can perform multiple decoding processes to select an appropriate decoding mode for each video data block, etc.
[0065] The motion estimation component 221 and the motion compensation component 219 can be highly integrated, but are shown separately for conceptual purposes. The motion estimation performed by the motion estimation component 221 is the process of generating motion vectors, which are used to estimate the motion of video blocks. For example, a motion vector can represent the displacement of a decoding object relative to a prediction block. A prediction block is a block that is found to be highly matched to a block to be decoded in terms of pixel difference. A prediction block may also be referred to as a reference block. Such pixel differences can be determined by the sum of absolute difference (SAD), the sum of square difference (SSD), or other difference metrics. HEVC uses several decoding objects, including CTU, coding tree block (CTB), and CU. For example, a CTU can be divided into multiple CTBs, which can then be divided into multiple CBs, and multiple CBs are used to be included in a CU. A CU can be encoded as a prediction unit (PU) including prediction data and / or a transform unit (TU) including transform residual data of a CU. The motion estimation component 221 uses rate-distortion analysis as part of a rate-distortion optimization process to generate motion vectors, PUs, and TUs. For example, the motion estimation component 221 can determine multiple reference blocks, multiple motion vectors, etc. for the current block / frame, and can select reference blocks, motion vectors, etc. with the best rate-distortion characteristics. The best rate-distortion characteristics balance the quality of video reconstruction (e.g., the amount of data lost due to compression) with decoding efficiency (e.g., the size of the final code).
[0066] In some examples, the codec system 200 can calculate the value of the sub-integer pixel position of the reference image stored in the decoded image buffer component 223. For example, the video codec system 200 can interpolate the values of the quarter pixel position, the eighth pixel position, or other fractional pixel position of the reference image. Therefore, the motion estimation component 221 can perform a motion search relative to the full pixel position and the fractional pixel position, and output a motion vector with fractional pixel accuracy. The motion estimation component 221 calculates the motion vector of the PU of the video block in the inter-frame decoding slice by comparing the position of the PU with the position of the prediction block of the reference image. The motion estimation component 221 outputs the calculated motion vector as motion data to the header formatting and CABAC component 231 for encoding, and outputs the motion to the motion compensation component 219.
[0067] The motion compensation performed by the motion compensation component 219 may involve obtaining or generating a prediction block based on the motion vector determined by the motion estimation component 221. In addition, in some examples, the motion estimation component 221 and the motion compensation component 219 may be functionally integrated. After receiving the motion vector of the PU of the current video block, the motion compensation component 219 may locate the prediction block pointed to by the motion vector. Then, the pixel difference values are formed by subtracting the pixel values of the prediction block from the pixel values of the decoded current video block, thereby forming a residual video block. Typically, the motion estimation component 221 performs motion estimation with respect to the luma component, and the motion compensation component 219 uses the motion vector calculated based on the luma component for the chroma component and the luma component. The prediction block and the residual block are forwarded to the transform scaling and quantization component 213.
[0068] The segmented video signal 201 is also sent to the intra-frame estimation component 215 and the intra-frame prediction component 217. Like the motion estimation component 221 and the motion compensation component 219, the intra-frame estimation component 215 and the intra-frame prediction component 217 can be highly integrated, but are shown separately for conceptual purposes. The intra-frame estimation component 215 and the intra-frame prediction component 217 perform intra-frame prediction on the current block relative to the block in the current frame, replacing the inter-frame prediction performed between frames by the motion estimation component 221 and the motion compensation component 219 as described above. Specifically, the intra-frame estimation component 215 determines an intra-frame prediction mode for encoding the current block. In some examples, the intra-frame estimation component 215 selects an appropriate intra-frame prediction mode from a plurality of tested intra-frame prediction modes to encode the current block. The selected intra-frame prediction mode is then forwarded to the header formatting and CABAC component 231 for encoding.
[0069] For example, the intra-frame estimation component 215 uses a rate-distortion analysis of various tested intra-frame prediction modes to calculate rate-distortion values and selects the intra-frame prediction mode with the best rate-distortion characteristics among the tested modes. The rate-distortion analysis generally determines the amount of distortion (or error) between the coded block and the original uncoded block that is encoded to generate the coded block, as well as the code rate (e.g., the number of bits) used to generate the coded block. The intra-frame estimation component 215 calculates a ratio based on the distortion and rate of the various coded blocks to determine which intra-frame prediction mode exhibits the best rate-distortion value for the block. In addition, the intra-frame estimation component 215 can be used to decode the depth block of the depth map using a depth modeling mode (DMM) according to rate-distortion optimization (RDO).
[0070] When implemented on an encoder, the intra prediction component 217 can generate a residual block from the prediction block according to the selected intra prediction mode determined by the intra estimation component 215, or when implemented on a decoder, read the residual block from the code stream. The residual block includes the difference in values between the prediction block and the original block, represented as a matrix. The residual block is then forwarded to the transform scaling and quantization component 213. The intra estimation component 215 and the intra prediction component 217 can operate on the luminance component and the chrominance component.
[0071] The transform scaling and quantization component 213 is used to further compress the residual block. The transform scaling and quantization component 213 applies a transform, such as a discrete cosine transform (DCT), a discrete sine transform (DST), or a conceptually similar transform to the residual block, thereby generating a video block including residual transform coefficient values. Wavelet transforms, integer transforms, subband transforms, or other types of transforms may also be used. The transform may convert the residual information from a pixel value domain to a transform domain, such as a frequency domain. The transform scaling and quantization component 213 is also used to scale the transformed residual information according to frequency, etc. This scaling involves applying a scaling factor to the residual information so as to quantize different frequency information at different granularities, which may affect the final visual quality of the reconstructed video. The transform scaling and quantization component 213 is also used to quantize the transform coefficients to further reduce the bit rate. The quantization process may reduce the bit depth associated with some or all coefficients. The degree of quantization may be modified by adjusting the quantization parameter. In some examples, the transform scaling and quantization component 213 may then scan the matrix including the quantized transform coefficients. The quantized transform coefficients are forwarded to the header formatting and CABAC component 231 for encoding into the codestream.
[0072] The scaling and inverse transform component 229 applies the inverse operation of the transform scaling and quantization component 213 to support motion estimation. The scaling and inverse transform component 229 applies inverse scaling, inverse transform and / or inverse quantization to reconstruct the residual block in the pixel domain, for example, for later use as a reference block, which can become a prediction block for another current block. The motion estimation component 221 and / or the motion compensation component 219 can calculate the reference block by adding the residual block back to the corresponding prediction block for motion estimation of subsequent blocks / frames. Filters are used to reconstruct the reference block to reduce artifacts generated during scaling, quantization and transformation. Such artifacts may cause inaccurate predictions (and generate other artifacts) when predicting subsequent blocks.
[0073] The filter control analysis component 227 and the in-loop filter component 225 apply filters to the residual block and / or the reconstructed image block. For example, the transformed residual block of the scaling and inverse transform component 229 can be merged with the corresponding prediction block of the intra-frame prediction component 217 and / or the motion compensation component 219 to reconstruct the original image block. The filter can then be applied to the reconstructed image block. In some examples, the filter can be applied to the residual block instead. As Figure 2 The filter control analysis component 227 and the in-loop filter component 225 are highly integrated with other components in the image processing unit and can be implemented together, but are shown separately for conceptual purposes. The filter used to reconstruct the reference block is used for a specific spatial area and includes multiple parameters to adjust the way such filters are applied. The filter control analysis component 227 analyzes the reconstructed reference block to determine the location where such filters should be applied and sets the corresponding parameters. Such data is forwarded to the header formatting and CABAC component 231 for encoding as filter control data. The in-loop filter component 225 applies such filters according to the filter control data. The filters may include deblocking filters, noise suppression filters, SAO filters, and adaptive loop filters. Such filters can be used in the spatial domain / pixel domain (e.g., for reconstructed pixel blocks) or the frequency domain according to the example.
[0074] When acting as an encoder, the filtered reconstructed image blocks, residual blocks and / or prediction blocks are stored in the decoded image buffer component 223 for later use in motion estimation, as described above. When acting as a decoder, the decoded image buffer component 223 stores the reconstructed and filtered blocks and forwards them to the display as part of the output video signal. The decoded image buffer component 223 can be any storage device capable of storing prediction blocks, residual blocks and / or reconstructed image blocks.
[0075] The header formatting and CABAC component 231 receives data from various components of the codec system 200 and encodes such data into a coded bitstream for transmission to the decoder. Specifically, the header formatting and CABAC component 231 generates various headers to encode control data (such as general control data and filter control data). In addition, prediction data (including intra-frame prediction) and motion data, as well as residual data in the form of quantized transform coefficient data are encoded in the bitstream. The final bitstream includes all the information required for the decoder to reconstruct the original segmented video signal 201. Such information may also include an intra-frame prediction mode index table (also called a codeword mapping table), a definition of the coding context of various blocks, a representation of the most likely intra-frame prediction mode, a representation of segmentation information, etc. Such data can be encoded using entropy coding. For example, the information may be encoded using context adaptive variable length coding (CAVLC), CABAC, syntax-based context-adaptive binary arithmetic coding (SBAC), probability interval partitioning entropy (PIPE) coding, or other entropy coding techniques. After entropy coding, the encoded bitstream may be sent to another device (e.g., a video decoder) or archived for subsequent transmission or retrieval.
[0076] Figure 3 is a block diagram of an exemplary video encoder 300. The video encoder 300 may be used to implement the encoding function of the codec system 200 and / or to implement steps 101, 103, 105, 107, and / or 109 of the operating method 100. The encoder 300 segments the input video signal to generate segmented video signals 301, wherein the segmented video signals 301 are substantially similar to the segmented video signals 201. Then, the segmented video signals 301 are compressed by the components of the encoder 300 and encoded into a bitstream.
[0077] Specifically, the segmented video signal 301 is forwarded to the intra prediction component 317 for intra prediction. The intra prediction component 317 may be substantially similar to the intra estimation component 215 and the intra prediction component 217. The segmented video signal 301 is also forwarded to the motion compensation component 321 for inter prediction based on the reference block in the decoded image buffer component 323. The motion compensation component 321 may be substantially similar to the motion estimation component 221 and the motion compensation component 219. The prediction blocks and residual blocks of the intra prediction component 317 and the motion compensation component 321 are forwarded to the transform and quantization component 313 for transforming and quantizing the residual blocks. The transform and quantization component 313 may be substantially similar to the transform scaling and quantization component 213. The transformed and quantized residual blocks and the corresponding prediction blocks (together with the relevant control data) are forwarded to the entropy coding component 331 for encoding into the bitstream. The entropy coding component 331 may be substantially similar to the header formatting and CABAC component 231.
[0078] The transformed and quantized residual block and / or the corresponding prediction block are also forwarded from the transform and quantization component 313 to the inverse transform and quantization component 329 to be reconstructed into a reference block for use by the motion compensation component 321. The inverse transform and quantization component 329 can be substantially similar to the scaling and inverse transform component 229. According to the example, the in-loop filter in the in-loop filter component 325 is also used for the residual block and / or the reconstructed reference block. The in-loop filter component 325 can be substantially similar to the filter control analysis component 227 and the in-loop filter component 225. As described with respect to the in-loop filter component 225, the in-loop filter component 325 can include a plurality of filters. The filtered block is then stored in the decoded image buffer component 323 for use as a reference block by the motion compensation component 321. The decoded image buffer component 323 can be substantially similar to the decoded image buffer component 223.
[0079] Figure 4 is a block diagram of an exemplary video decoder 400. The video decoder 400 may be used to implement the decoding function of the codec system 200 and / or implement steps 111, 113, 115, and / or 117 of the operation method 100. For example, the decoder 400 receives a bitstream from the encoder 300, and generates a reconstructed output video signal according to the bitstream for display to an end user.
[0080] The code stream is received by entropy decoding component 433. Entropy decoding component 433 is used to implement entropy decoding schemes, such as CAVLC, CABAC, SBAC, PIPE decoding or other entropy decoding techniques. For example, entropy decoding component 433 can use header information to provide context to explain other data encoded as code words in the code stream. Decoding information includes any information required to decode the video signal, such as general control data, filter control data, segmentation information, motion data, prediction data and quantized transform coefficients in the residual block. The quantized transform coefficients are forwarded to the inverse transform and quantization component 429 to reconstruct the residual block. The inverse transform and quantization component 429 can be similar to the inverse transform and quantization component 329.
[0081] The reconstructed residual block and / or prediction block is forwarded to the intra prediction component 417 to be reconstructed into an image block according to the intra prediction operation. The intra prediction component 417 can be similar to the intra estimation component 215 and the intra prediction component 217. Specifically, the intra prediction component 417 uses the prediction mode to locate the reference block in the frame, and uses the residual block for the result to reconstruct the intra prediction image block. The reconstructed intra prediction image block and / or residual block and the corresponding inter prediction data are forwarded to the decoded image buffer component 423 through the in-loop filter component 425. The decoded image buffer component 423 and the in-loop filter component 425 can be substantially similar to the decoded image buffer component 223 and the in-loop filter component 225, respectively. The in-loop filter component 425 filters the reconstructed image block, residual block and / or prediction block, and such information is stored in the decoded image buffer component 423. The reconstructed image block of the decoded image buffer component 423 is forwarded to the motion compensation component 421 for inter prediction. The motion compensation component 421 can be substantially similar to the motion estimation component 221 and / or the motion compensation component 219. Specifically, the motion compensation component 421 uses the motion vector of the reference block to generate a prediction block, and uses the residual block for the result to reconstruct the image block. The resulting reconstructed block can also be forwarded to the decoded image buffer component 423 through the in-loop filter component 425. The decoded image buffer component 423 continues to store other reconstructed image blocks, which can be reconstructed into frames through segmentation information. Such frames can also be located in a sequence. The sequence is output to the display as a reconstructed output video signal.
[0082] Figure 55 is a schematic diagram of an exemplary CVS 500. For example, according to method 100, CVS 500 can be encoded by an encoder, such as by codec system 200 and / or encoder 300. In addition, CVS 500 can be decoded by a decoder, such as by codec system 200 and / or decoder 400. CVS 500 includes images decoded in decoding order 508. Decoding order 508 is the order in which images are positioned in the bitstream. Then, the images of CVS 500 are output in presentation order 510. Presentation order 510 is the order in which the decoder should display the images so that the resulting video is displayed correctly. For example, the images of CVS 500 can generally be positioned in presentation order 510. However, some images can be moved to different positions to improve decoding efficiency, such as by placing similar images closer together to support inter-frame prediction. Moving such images in this way results in decoding order 508. In the example shown, the images are indexed from 0 to 4 in decoding order 508. In presentation order 510 , images at index 2 and index 3 have been moved in front of the image at index 0 .
[0083] CVS 500 includes IRAP image 502. IRAP image 502 is an image encoded according to intra-frame prediction and is used as a random access point of CVS 500. Specifically, blocks of IRAP image 502 are encoded by referring to other blocks of IRAP image 502. Since IRAP image 502 is encoded without referring to other images, IRAP image 502 can be decoded without first decoding any other images. Therefore, the decoder can start decoding CVS 500 at IRAP image 502. In addition, IRAP image 502 can cause DPB to be refreshed. For example, images presented after IRAP image 502 can be inter-frame predicted without relying on images before IRAP image 502 (e.g., image index 0). Therefore, once IRAP image 502 is decoded, the image buffer can be refreshed. This has the effect of stopping any coding errors related to inter-frame prediction because such errors cannot be propagated through IRAP image 502. IRAP image 502 can include various types of images. For example, IRAP image can be encoded as IDR or CRA. IDR is an intra-frame coded picture that starts a new CVS 500 and flushes the picture buffer. CRA is an intra-frame coded picture that acts as a random access point and does not start a new CVS 500 or flush the picture buffer. In this way, the previous picture 504 associated with CRA can refer to the picture before CRA, while the previous picture 504 associated with IDR may not refer to the picture before IDR.
[0084] CVS500 also includes various non-IRAP images. These images include leading images 504 and trailing images 506. Leading images 504 are images that are located after IRAP images 502 in decoding order 508, but before IRAP images 502 in presentation order 510. In decoding order 508 and presentation order 510, trailing images 506 are located after IRAP images 502. Both leading images 504 and trailing images 506 are encoded according to inter-frame prediction. The trailing images 506 are encoded with reference to IRAP images 502 or images that are located after IRAP images 502. Therefore, once IRAP images 502 are decoded, trailing images 506 can always be decoded. Leading images 504 can include random access skipped leading (RASL) and decodable random access leading (RADL) images. RASL images are encoded by referring to images before IRAP images 502, but are encoded at a position after IRAP images 502. Since RASL pictures depend on previous pictures, RASL pictures cannot be decoded when the decoder starts decoding at IRAP picture 502. Therefore, when IRAP picture 502 is used as a random access point, RASL pictures are skipped and not decoded. However, when the decoder uses the previous IRAP picture (before index 0 and not shown) as a random access point, the RASL picture is decoded and displayed. RADL pictures are encoded with reference to IRAP picture 502 and / or pictures after IRAP picture 502, but are positioned before IRAP picture 502 in presentation order 510. Since RADL pictures do not depend on pictures before IRAP picture 502, RADL pictures can be decoded and displayed when IRAP picture 502 is a random access point.
[0085] Each of the images in CVS500 can be stored in an access unit. In addition, the image can be divided into slices, and the slices can be included in the NAL unit. The NAL unit is a storage unit that contains a parameter set or slice of the image and a corresponding slice header. The NAL unit is assigned a type to indicate to the decoder the type of data contained in the NAL unit. For example, the slice of the IRAP image 502 can be contained in an IDR (IDR_W_RADL) NAL unit with RADL, an IDR (IDR_N_LP) NAL unit without a preceding image, a CRA NAL unit, etc. The IDR_W_RADL NAL unit indicates that the IRAP image 502 is an IDR image associated with a RADL preceding image 504. The IDR_N_LP NAL unit indicates that the IRAP image 502 is an IDR image that is not associated with any preceding image 504. The CRA NAL unit indicates that the IRAP image 502 is a CRA image that can be associated with a preceding image 504. The slice of a non-IRAP image can also be located in a NAL unit. For example, the slices of the trailing picture 506 may be placed in a trailing picture NAL unit type (TRAIL_NUT), which indicates that the trailing picture 506 is an inter-frame prediction coded picture. The slices of the preceding picture 504 may be included in a RASL NAL unit type (RASL_NUT) and / or a RADL NAL unit type (RADL_NUT), both of which may indicate that the corresponding picture is an inter-frame prediction coded preceding picture 504 of the corresponding type. By indicating the slices of the picture in the corresponding NAL unit, the decoder can easily determine the appropriate decoding mechanism to apply to each picture / slice.
[0086] Figure 6 Schematic diagram of a plurality of sub-image video streams 601, 602, and 603 divided from a VR image video stream 600. For example, each of the sub-image video streams 601-603 and / or the VR image video stream 600 may be encoded in the CVS 500. Therefore, according to the method 100, the sub-image video streams 601-603 and / or the VR image video stream 600 may be encoded by an encoder, such as the codec system 200 and / or the encoder 300. In addition, the sub-image video streams 601-603 and / or the VR image video stream 600 may be decoded by a decoder, such as the codec system 200 and / or the decoder 400.
[0087] The VR image video stream 600 includes multiple images presented over time. Specifically, VR operates by encoding the video content of a sphere, which can be displayed as if the user is at the center of the sphere. Each image includes the entire sphere. At the same time, only a portion of the image (called the perspective) is displayed to the user. For example, a user can use a head mounted display (HMD) that selects and displays the perspective of the sphere based on the user's head movement. This produces the effect of physically existing in the virtual space depicted by the video. To achieve this result, each image of the video sequence includes the entire video data sphere at the corresponding moment. However, only a small portion of the image (e.g., a single perspective) is displayed to the user. The rest of the image is discarded and not presented. The entire image is usually transmitted so that different perspectives can be dynamically selected and displayed in response to the user's head movement.
[0088] In the example shown, the images of the VR image video stream 600 can be subdivided into sub-images according to the available viewing angle. Therefore, each image and the corresponding sub-image include a temporal position (e.g., image order) as part of the temporal presentation. The sub-image video streams 601-603 are created when the subdivision is applied consistently over a period of time. This consistent subdivision creates sub-image video streams 601-603, wherein each stream contains a set of sub-images having a predetermined size, shape, and spatial position relative to the corresponding image in the VR image video stream 600. In addition, the temporal position of the sub-image set in the sub-image video streams 601-603 varies with the presentation time. Therefore, the sub-images in the sub-image video streams 601-603 can be aligned in the time domain according to the temporal position. Then, the sub-images in the sub-image video streams 601-603 at each temporal position can be fused in the spatial domain according to the predefined spatial position to reconstruct the VR image video stream 600 for display. Specifically, each of the sub-image video streams 601-603 can be encoded into a different sub-code stream. When these sub-streams are fused together, a single stream is produced that includes the entire set of images over time. The resulting stream can be transmitted to a decoder for decoding and display based on the user's currently selected viewing angle.
[0089] One of the problems with VR video is that all sub-image video streams 601-603 can be transmitted to the user at high quality (e.g., high resolution). This allows the decoder to dynamically select the user's current perspective and display the sub-images in the corresponding sub-image video streams 601-603 in real time. However, the user can only watch a single perspective, such as from the sub-image video stream 601, while the sub-image video streams 602-603 are discarded. Therefore, transmitting the sub-image video streams 602-603 at high quality wastes a lot of bandwidth. To improve decoding efficiency, VR video can be encoded into multiple video streams 600, each of which is encoded at a different quality / resolution. In this way, the decoder can transmit a request for the current sub-image video stream 601. In response, the encoder (or mid-striper or other content server) can select a higher quality sub-image video stream 601 from the higher quality video stream 600 and a lower quality sub-image video stream 602-603 from the lower quality video stream 600. The encoder can then merge these sub-streams into a complete encoded stream for transmission to the decoder. In this way, the decoder receives a series of images, wherein the current view quality is high and the other view quality is low. In addition, the highest quality sub-image is usually displayed to the user (when not moving the head), and the low quality sub-image is usually discarded, which balances the function and decoding efficiency.
[0090] In the case where the user switches from viewing the sub-image video stream 601 to the sub-image video stream 602, the decoder requests that the new current sub-image video stream 602 be transmitted at a higher quality. The encoder can then change the fusion mechanism accordingly. As described above, the decoder can only start decoding the new CVS 500 at the IRAP picture 502. Therefore, the sub-image video stream 602 is displayed at a lower quality before the IRAP picture / sub-image arrives. The IRAP picture can then be decoded at a higher quality to start decoding the higher quality version of the sub-image video stream 602. This approach significantly improves video compression without negatively affecting the user's viewing experience.
[0091] One problem with the above method is that the length of time required to change the resolution is based on the length of time before the IRAP picture arrives in the video stream. This is because the decoder cannot start decoding different versions of the sub-image video stream 602 at a non-IRAP picture. One way to reduce this delay is to include more IRAP pictures. However, this will increase the file size. In order to balance functionality and decoding efficiency, different perspective / sub-image video streams 601-603 can include IRAP pictures at different frequencies. For example, a perspective / sub-image video stream 601-603 that is more likely to be viewed can have more IRAP pictures than other perspective / sub-image video streams 601-603. For example, in a basketball environment, the perspective / sub-image video stream 601-603 associated with the basket and / or the midfield includes a greater frequency of IRAP pictures than the perspective / sub-image video stream 601-603 used to view the stands or ceiling, because the user is less likely to view the perspective / sub-image video stream 601-603 used to view the stands or ceiling.
[0092] This approach can lead to other problems. Specifically, the sub-images in the sub-image video streams 601-603 that share a POC are part of a single image. As described above, the slices of the image are included in the NAL unit based on the image type. In some video decoding systems, all NAL units associated with a single image are restricted to include the same NAL unit type. When different sub-image video streams 601-603 have IRAP images of different frequencies, some images include both IRAP sub-images and non-IRAP sub-images. This eliminates the constraint that each image can only use the same type of NAL unit.
[0093] The present invention solves this problem by eliminating the constraint that all NAL units of the strips in the image use the same NAL unit type. For example, the image is contained in an access unit. By eliminating this constraint, the access unit can include both the IRAPNAL unit type and the non-IRAP NAL unit type. In addition, a flag can be encoded to indicate when the image / access unit includes a mixture of IRAP NAL unit types and non-IRAP NAL unit types. In some examples, the flag is a mixed NAL unit type (mixed_nalu_types_in_pic_flag) in the image flag. In addition, a constraint can be applied to require that a single mixed image / access unit can only contain one type of IRAP NAL unit and one type of non-IRAP NAL unit. This can prevent unexpected NAL unit type mixing. If such mixing can be performed, the decoder must be designed to manage such mixing. This will unnecessarily increase the required hardware complexity without providing other benefits to the encoding process. For example, a mixed image can include an IRAP NAL unit of a type selected from IDR_W_RADL, IDR_N_LP or CRA_NUT. Furthermore, the hybrid picture may include a type of non-IRAP NAL unit selected from TRAIL_NUT, RADL_NUT and RASL_NUT.Exemplary implementations of this scheme are discussed in detail below.
[0094] Figure 7 Schematic diagram of an exemplary bitstream 700 including an image with mixed NAL unit types. For example, according to method 100, the bitstream 700 may be generated by the codec system 200 and / or the encoder 300, and decoded by the codec system 200 and / or the decoder 400. In addition, the bitstream 700 may include a VR image video stream 600 merged from multiple sub-image video streams 601-603 at multiple video resolutions, wherein each sub-image video stream includes a CVS 500 located at a different spatial position.
[0095] The code stream 700 includes a sequence parameter set (SPS) 710, one or more picture parameter sets (PPS) 711, multiple slice headers 715, and image data 720. The SPS 710 includes sequence data common to all images in the video sequence of the code stream 700. These data may include image size, bit depth, coding tool parameters, bit rate limits, etc. The PPS 711 contains parameters applied to the entire image. Therefore, each image in the video sequence can refer to the PPS 711. It should be noted that although each image refers to the PPS 711, in some examples, a single PPS 711 can contain data for multiple images. For example, multiple similar images can be encoded according to similar parameters. In this case, a single PPS 711 can contain data for these similar images. The PPS 711 can represent coding tools, quantization parameters, offsets, etc. that can be used for the slices in the corresponding image. The slice header 715 contains specific parameters for each slice in the image. Therefore, each slice in the video sequence can have a slice header 715. The slice header 715 may include slice type information, picture order count (POC), reference picture list, prediction weight, partition entry point, deblocking filter parameters, etc. It should be noted that in some contexts, the slice header 715 may also be referred to as a partition group header.
[0096] The image data 720 includes video data encoded according to inter-frame prediction and / or intra-frame prediction and corresponding transformed and quantized residual data. For example, a video sequence includes a plurality of images 721 encoded as image data 720. Image 721 is a single frame of a video sequence and is therefore typically displayed as a single unit when the video sequence is displayed. However, sub-images 723 may be displayed to implement certain technologies such as virtual reality. Each of the images 721 references PPS 711. Image 721 may be divided into sub-images 723, blocks, and / or strips. Sub-images 723 are spatial regions of images 721 that are consistently applied to a decoded video sequence. Accordingly, in a VR environment, sub-images 723 may be displayed by an HMD. In addition, sub-images 723 with a specified POC may be obtained from sub-image video streams 601-603 of corresponding resolutions. Sub-images 723 may reference SPS 710. In some systems, strips 725 are referred to as block groups containing blocks. The slice 725 and / or the partition group of the partition refer to the slice header 715. The slice 725 can be defined as an integer number of complete partitions within the block of the image 721 or an integer number of continuous complete CTU rows within the partition of the image 721, which are exclusively contained in a single NAL unit. Therefore, the slice 725 is further divided into CTUs and / or CTBs. The CTU / CTB is further divided into coding blocks according to the coding tree. Then, the coding blocks can be encoded / decoded according to the prediction mechanism.
[0097] Parameter sets and / or slices 725 are encoded in NAL units. NAL units can be defined as a syntax structure that contains a representation of the data type to be followed and bytes containing the data in the form of RBSP (interspersed with emulation prevention bytes if necessary). More specifically, a NAL unit is a storage unit that contains a parameter set or slice 725 of an image 721 and a corresponding slice header 715. Specifically, a VCL NAL unit 740 is a NAL unit that contains a slice 725 of an image 721 and a corresponding slice header 715. In addition, a non-VCL NAL unit 730 contains a parameter set, such as an SPS 710 and a PPS 711. Several types of NAL units can be used. For example, SPS 710 and PPS 711 can be included in an SPS NAL unit type (SPS_NUT) 731 and a PPS NAL unit type (PPS_NUT) 732, respectively, both of which are non-VCL NAL units 730.
[0098] As described above, an IRAP image such as the IRAP image 502 may be included in an IRAP NAL unit 745. Non-IRAP images such as the preceding image 504 and the succeeding image 506 may be included in a non-IRAP NAL unit 749. Specifically, the IRAP NAL unit 745 is any NAL unit containing a slice 725 obtained from an IRAP image or a sub-image. The non-IRAP NAL unit 749 is any NAL unit containing a slice 725 obtained from any image that is not an IRAP image or a sub-image (e.g., a preceding image or a succeeding image). Both the IRAP NAL unit 745 and the non-IRAP NAL unit 749 contain slice data and are therefore both VCL NAL units 740. In an exemplary embodiment, the IRAP NAL unit 745 may include a slice 725 of an IDR image without a preceding image or an IDR image associated with a RADL image in an IDR_N_LP NAL unit 741 or an IDR_w_RADL NAL unit 742, respectively. Additionally, the IRAP NAL unit 745 may include a slice 725 of a CRA picture in the CRA_NUT 743. In an exemplary embodiment, the non-IRAP NAL unit 749 may include a slice 725 of a RASL picture, a RADL picture, or a trailing picture belonging to a RASL_NUT 746, a RADL_NUT 747, or a TRAIL_NUT 748, respectively. In an exemplary embodiment, a complete list of possible NAL units, sorted by NAL unit type, is shown below.
[0099]
[0100]
[0101]
[0102] As described above, the VR video stream may include sub-images 723 with IRAP images of different frequencies. This allows spatial areas that are less likely to be viewed by the user to use fewer IRAP images, while spatial areas that the user may frequently view use more IRAP images. In this way, spatial areas that the user may switch back to regularly can be quickly adjusted to a higher resolution. When this method causes the image 721 to include an IRAP NAL unit 745 and a non-IRAP NAL unit 749, the image 721 is referred to as a mixed image. This condition may be indicated by a mixed NAL unit type (mixed_nalu_types_in_pic_flag) 727 in the image flag. The mixed_nalu_types_in_pic_flag 727 may be set in the PPS 711. In addition, when each image 721 representing the reference PPS 711 has more than one VCL NAL unit 740 and the NAL unit type (nal_unit_type) values of the VCL NAL units 740 are not the same, the mixed_nalu_types_in_pic_flag 727 may be set to 1. In addition, when each picture 721 referring to the PPS 711 has one or more VCL NAL units 740 and the nal_unit_type values of the VCL NAL units 740 of each picture 721 referring to the PPS 711 are the same, the mixed_nalu_types_in_pic_flag 727 may be set to 0.
[0103] In addition, a constraint may be employed such that when mixed_nalu_types_in_pic_flag 727 is set, the VCL NAL units 740 of one or more sub-pictures 723 of the picture 721 all have a first specific NAL unit type value, and the other VCL NAL units 740 in the picture 721 all have a different second specific NAL unit type value. For example, the constraint may require that the mixed picture 721 contain a single type of IRAP NAL unit 745 and a single type of non-IRAP NAL unit 749. For example, the picture 721 may include one or more IDR_N_LP NAL units 741, one or more IDR_w_RADL NAL units 742, or one or more CRA_NUTs 743, but not any combination of such IRAP NAL units 745. In addition, the picture 721 may include one or more RASL_NUTs 746, one or more RADL_NUTs 747, or one or more TRAIL_NUTs 748, but not any combination of such IRAP NAL units 745.
[0104] In an exemplary implementation, the picture type is used to define the decoding process. Such processes include, for example, deriving picture identification through picture order count (POC), identifying reference picture status in decoded picture buffer (DPB), outputting pictures from DPB, etc. Pictures can be based on the NAL unit type containing all coded pictures or their sub-parts, and identified by type. In some video decoding systems, picture types may include instantaneous decoding refresh (IDR) pictures and non-IDR pictures. In other video decoding systems, picture types may include post-pictures, temporal sub-layer access (TSA) pictures, step-wise temporal sub-layer access (STSA) pictures, decodable random access leading (RADL) pictures, skipped random access leading (RASL) pictures, broken-link access (BLA) pictures, instantaneous random access pictures, and pure random access pictures. This type of picture can also be distinguished according to whether the picture is a sub-layer reference picture or a sub-layer non-reference picture. BLA pictures can also be distinguished into BLA pictures with preceding pictures, BLA pictures with RADL pictures, and BLA pictures without preceding pictures. IDR pictures can also be distinguished into IDR pictures with RADL pictures and IDR pictures without preceding pictures.
[0105] Such image types can be used to implement various video-related functions. For example, IRAP images can be implemented using IDR, BLA and / or CRA images. IRAP images can provide the following functions / benefits. The presence of an IRAP image can indicate that the decoding process can be started from this image. This function makes it possible to implement a random access function, in which the decoding process starts at a specified position in the codestream as long as the IRAP image exists at that position. Such a position is not required at the beginning of the codestream. The presence of an IRAP image also refreshes the decoding process, so that the encoded images starting from the IRAP image (excluding RASL images) are encoded without reference to the images preceding the IRAP image. Therefore, the IRAP image located in the codestream stops the propagation of decoding errors. Therefore, decoding errors of the encoded images preceding the IRAP image cannot propagate through the IRAP image and enter the images following the IRAP image in the decoding order.
[0106] IRAP images provide a variety of functions, but they have an impact on compression efficiency. For example, the presence of IRAP images may cause a surge in bit rate. There are many reasons that affect compression efficiency. For example, IRAP images are intra-frame prediction images represented by much more bits than inter-frame prediction images used as non-IRAP images. In addition, the presence of IRAP images causes the temporal prediction used in inter-frame prediction to fail. Specifically, IRAP images refresh the decoding process by deleting previous reference images from the DPB. Deleting previous reference images reduces the availability of reference images when used to encode images that follow the IRAP image in decoding order, thereby reducing the efficiency of the process.
[0107] IDR pictures may use different indication and derivation processes than other IRAP picture types. For example, the indication and derivation process associated with IDR may set the most significant bit (MSB) portion of the POC to 0 instead of deriving the MSB from the previous key picture. In addition, the slice header of the IDR picture may not contain information to assist in reference picture management. At the same time, other picture types, such as CRA, post-position, TSA, etc., may contain reference picture information such as a reference picture set (RPS) or a reference picture list, which may be used to implement the reference picture identification process. The reference picture identification process is the process of determining whether the status of a reference picture in the DPB is used for reference or not used for reference. For IDR pictures, such information may not be indicated because the presence of an IDR indicates that the decoding process should simply identify all reference pictures in the DPB as not used for reference.
[0108] In addition to the image type, the image identification of the POC is also used for a variety of purposes, such as for managing reference images in inter-frame prediction, for outputting images from the DPB, for motion vector scaling, for weighted prediction, etc. For example, in some video decoding systems, images in the DPB can be identified as being used for short-term reference, for long-term reference, or for reference. Once an image is identified as not being used for reference, the image can no longer be used for prediction. When such an image is no longer needed to be output, the image can be deleted from the DPB. In other video decoding systems, reference images can be identified as short-term and long-term. When a reference image is no longer used for prediction reference, the reference image can be identified as not being used for reference. The transition between these states can be controlled by the reference image identification process for decoding. An implicit sliding window process and / or an explicit memory management control operation (MMCO) process can be used as a reference image identification mechanism for decoding. When the number of reference frames is equal to the specified maximum number (expressed as max_num_ref_frames in SPS), the sliding window process identifies the short-term reference image as not being used for reference. Short-term reference pictures are stored in a first-in, first-out manner so that the most recently decoded short-term picture is saved in the DPB. The explicit MMCO process may include multiple MMCO commands. The MMCO command may identify one or more short-term or long-term reference pictures as not used for reference, identify all pictures as not used for reference, or identify the current reference picture or an existing short-term reference picture as a long-term reference picture, and then assign a long-term picture index to the long-term reference picture.
[0109] In some video decoding systems, the reference picture identification operation and the process of outputting and deleting pictures from the DPB are performed after the picture is decoded. Other video decoding systems use RPS for reference picture management. The most fundamental difference between the RPS mechanism and the MMCO / sliding window process is that, for each specific slice, the RPS provides a complete set of reference pictures for use by the current picture or any subsequent picture. Therefore, the complete set of all pictures that should be stored in the DPB for use by the current picture or subsequent pictures is indicated in the RPS. This is different from the MMCO / sliding window scheme, which only indicates relative changes to the DPB. With the RPS mechanism, information from the previous pictures in the decoding order is not required to maintain the correct state of the reference pictures in the DPB. In some video decoding systems, in order to take advantage of RPS and improve error resilience, the order of picture decoding and DPB operations is changed. In some video decoding systems, picture identification and buffering operations, including outputting and deleting decoded pictures from the DPB, can be applied after the current picture is decoded. In other video coding systems, the RPS is first decoded from the slice header of the current picture, and then picture identification and buffering operations may be applied before decoding the current picture.
[0110] In VVC, the reference picture management method can be summarized as follows. Two reference picture lists, denoted as List 0 and List 1, are directly indicated and derived. They are not based on the RPS or sliding window plus MMCO process, as described above. Reference picture identification is directly based on reference picture list 0 and list 1, using active and inactive entries in the reference picture list, while in inter prediction of CTU, only active entries can be used as reference indexes. The information used to derive the two reference picture lists is indicated by the syntax elements and syntax structures in the SPS, PPS and slice header. The predefined RPL structure is indicated in the SPS for use by reference in the slice header. Two reference picture lists are generated for all types of slices including bidirectional inter prediction (B) slices, unidirectional inter prediction (P) slices and intra prediction (I) slices. The two reference picture lists can be constructed without using the reference picture list initialization process or the reference picture list modification process. The long-term reference picture (LTRP) is identified by the POC LSB. An incremental POC MSB period can be indicated for the LTRP, as determined on a picture-by-picture basis.
[0111] In order to encode a video image, the image is first segmented and the resulting parts are encoded into a bitstream. There are various image segmentation schemes. For example, an image can be segmented into regular strips, non-independent strips, blocks, and / or segmented according to wavefront parallel processing (WPP). For simplicity, HEVC restricts the encoder so that only regular strips, non-independent strips, blocks, WPP, and combinations thereof can be used when segmenting strips into CTB groups for video decoding. This segmentation can be used to support maximum transfer unit (MTU) size matching, parallel processing, and reduce end-to-end delay. MTU represents the maximum amount of data that can be sent in a single message. If the message payload exceeds the MTU, the payload is divided into two messages by a process called segmentation.
[0112] A regular strip, also referred to as a strip, is a portion obtained after the image is divided. It can be reconstructed independently of other regular strips in the same image, but there is still mutual dependence due to the existence of loop filtering operations. Each regular strip is encapsulated in its own network abstraction layer (NAL) unit for transmission. In addition, intra-frame prediction (intra-frame sample prediction, motion information prediction, decoding mode prediction) and entropy coding dependencies across strip boundaries can be disabled to support independent reconstruction. This independent reconstruction supports parallel operations. For example, parallelization based on regular strips reduces inter-processor or inter-core communication. However, since each regular strip is independent, each strip is associated with a separate strip header. Due to the bit cost of the strip header of each strip and the lack of prediction across strip boundaries, the use of regular strips will cause large decoding overhead. In addition, regular strips can be used to support MTU size matching requirements. Specifically, since regular strips are encapsulated in separate NAL units and can be decoded independently, each regular strip should be smaller than the MTU in the MTU scheme to avoid dividing the strip into multiple messages. Therefore, the stripe layouts in the image will contradict each other in order to achieve parallelization and MTU size matching.
[0113] Dependent slices are similar to regular slices, but have shortened slice headers and can be split on picture tree block boundaries without affecting intra prediction. Therefore, dependent slices split regular slices into multiple NAL units, which reduces end-to-end latency by sending part of a regular slice before the encoding of the entire regular slice is completed.
[0114] The image can be divided into block groups / strips and blocks. A block is a sequence of CTUs covering a rectangular area of the image. A block group / strip contains multiple blocks of the image. Blocks can be created using raster scan block group mode and rectangular block group mode. In raster scan block group mode, a block group includes a sequence of blocks in a block raster scan of the image. In rectangular block group mode, a block group contains multiple blocks of the image, which together constitute a rectangular area of the image. The blocks in a rectangular block group are arranged in the block raster scan order of the block group. For example, a block can be a portion obtained by dividing an image along horizontal and vertical boundaries, and the horizontal and vertical boundaries produce block columns and block rows. Blocks can be decoded in raster scan order (from right to left, from top to bottom). The scan order of CTBs is performed within the block. Therefore, the CTBs in the first block are decoded in raster scan order before proceeding to the CTBs in the next block. Similar to conventional strips, blocks eliminate intra-frame prediction dependencies and entropy decoding dependencies. However, blocks may not be included in each NAL unit, and therefore, blocks may not be used for MTU size matching. Each block can be processed by one processor / core, and the inter-processor / inter-core communication for intra-frame prediction between processing units for decoding adjacent blocks can be limited to sending a shared slice header (when adjacent blocks are in the same slice), and sharing reconstructed samples and metadata related to loop filtering. When a slice includes more than one block, the entry point byte offset of each block can also be indicated in the slice header except for the first block in the slice. For each slice and block, at least one of the following conditions should be met: (1) all coding tree blocks in the slice belong to the same block; (2) all coding tree blocks in the block belong to the same slice.
[0115] In WPP, the image is partitioned into a single row of CTBs. Entropy decoding and prediction mechanisms can use data from CTBs in other rows. Parallel processing is supported by decoding CTB rows in parallel. For example, the current row can be decoded in parallel with the previous row. However, the decoding of the current row is delayed by two CTBs from the decoding process of the previous rows. This delay ensures that data related to the above CTB and the right side of the current CTB in the current row are available before the current CTB is decoded. When represented graphically, this approach is represented as a wavefront. This interleaved start can be parallelized using as many processors / cores as there are CTB rows included in the image. Since intra-frame prediction can be performed between adjacent treeblock rows within the image, it is important to use inter-processor / inter-core communication to implement intra-frame prediction. WPP partitioning does not take into account the NAL unit size. Therefore, WPP does not support MTU size matching. However, regular slices can be used in conjunction with WPP to achieve MTU size matching as needed, which requires some decoding overhead. Finally, a wavefront segment can contain exactly one CTB row. Furthermore, when using WPP, when a stripe starts in a CTB row, the stripe should end in the same CTB row.
[0116] Tiles may also include motion constrained tilesets. A motion constrained tileset (MCTS) is a set of tiles that constrains the associated motion vectors to point to integer sample positions within the MCTS and fractional sample positions that only require integer sample positions within the MCTS for interpolation. In addition, motion vector candidates cannot be used for temporal motion vector prediction derived from blocks outside the MCTS. In this way, each MCTS can be decoded independently without including tiles in the MCTS. Temporal MCTS supplemental enhancement information (SEI) messages can be used to indicate the presence of MCTS in the codestream and indicate the MCTS. The MCTS SEI message provides supplementary information that can be used for MCTS sub-codestream extraction (expressed as part of the SEI message semantics) to generate a consistent codestream for the MCTS. The information includes multiple extraction information sets, each of which defines multiple MCTS sets, and includes raw bytes sequence payload (RBSP) bytes of replacement video parameter sets (VPS), sequence parameter sets (SPS), and picture parameter sets (PPS) used in the MCTS substream extraction process. Because one or all of the slice address-related syntax elements (including first_slice_segment_in_pic_flag and slice_segment_address) can use different values in the extracted substream, the parameter sets (VPS, SPS, PPS) can be rewritten or replaced when extracting the substream according to the MCTS substream extraction process, and the slice header is updated.
[0117] VR applications, also known as 360-degree video applications, can display only a portion of a complete sphere, and therefore only a subset of the entire image. Perspective-dependent 360 transmission based on the dynamic adaptive streaming over hypertext transfer protocol (DASH) mechanism can be used to reduce the bit rate and support the transmission of 360-degree video through a streaming mechanism. This mechanism divides the sphere / projected image into multiple MCTSs using a cubemap projection (CMP) or the like. Two or more code streams can be encoded with different spatial resolutions or qualities. When the data is transmitted to the decoder, a MCTS of a higher resolution / quality code stream is transmitted for the perspective to be displayed (e.g., the front perspective). MCTS of a lower resolution / quality code stream is transmitted for other perspectives. These MCTSs are packaged in a certain way and sent to the receiver for decoding. The perspective that the user is expected to see is represented by a high-resolution / quality MCTS to provide a positive viewing experience. When the user turns his head to see another perspective (e.g., left or right perspective), the content is displayed at a lower resolution / quality perspective for a short time while the system obtains the high-resolution / quality MCTS of the new perspective. When the user turns his head to see another view, there is a delay between the time the user turns his head and the time the higher resolution / quality representation of the view is seen. This delay depends on how fast the system can fetch a higher resolution / quality MCTS for that view, which in turn depends on the IRAP cycle. The IRAP cycle is the interval between two IRAP occurrences. This delay is related to the IRAP cycle because the MCTS for the new view can only be decoded starting from the IRAP image.
[0118] For example, if the IRAP cycle is encoded once every second, the following applies. If the user turns his head to see a new view just before the system starts acquiring a new segment / IRAP cycle, the best case scenario for latency is the same as the round-trip delay. In this case, the system is able to immediately request a higher resolution / quality MCTS for the new view, so the only delay is the network round-trip delay, which is the delay of the fetch request plus the transmission time of the requested MCTS, assuming that the minimum buffer delay can be set to approximately zero and the sensor latency is small and negligible. For example, the network round-trip delay is about 200 milliseconds. If the user turns his head to see a new view just after the system requests the next segment, the worst case scenario is the IRAP cycle plus the network round-trip delay. The codestream can be encoded with more frequent IRAP images, making the IRAP cycle shorter, to improve the worst case above, as this reduces the overall latency. However, this approach increases bandwidth requirements due to reduced compression efficiency.
[0119] In an exemplary implementation, sub-images of the same coded image may contain different nal_unit_type values. The mechanism is described as follows. An image may be divided into sub-images. A sub-image is a set of rectangular tile groups / strips, beginning with a tile group whose tile_group_address is equal to 0. Each sub-image may reference a corresponding PPS and may therefore have a separate tile partitioning scheme. The presence of a sub-image may be indicated in the PPS. During decoding, each sub-image is treated as an image. In-loop filtering across sub-image boundaries may always be disabled. The sub-image width and height may be expressed in units of luma CTU size. The position of the sub-image in the image may not be indicated but may be derived using the following rules. The sub-image is located at the next such unoccupied position in the image in CTU raster scan order that is large enough to include the sub-image within the image boundary. The reference image used to decode each sub-image is generated by extracting a region collocated with the current sub-image from the reference image in the decoded image buffer. The extracted region is the decoded sub-image, so inter-frame prediction occurs between sub-images of the same size and position within the image. In this case, having different nal_unit_type values in the coded picture enables sub-pictures originating from random access pictures and sub-pictures originating from non-random access pictures to be merged into the same coded picture without substantial difficulty (e.g., without VCL level modifications). This benefit also applies to MCTS-based coding.
[0120] In other cases, it is beneficial to have different nal_unit_type values in the encoded pictures. For example, a user may view some areas of a 360-degree video content more frequently than other areas. In order to establish a better trade-off between decoding efficiency and average comparable quality view switching latency in MCTS / sub-image based view-dependent 360-degree video transmission, more frequent IRAP pictures may be encoded for areas that are viewed more frequently than other areas. The comparable quality view switching latency is the latency experienced by a user when switching from a first view to a second view until the presentation quality of the second view reaches a presentation quality comparable to that of the first view.
[0121] Another implementation adopts the following scheme to support mixed NAL unit types in an image, including POC derivation and reference picture management. There is a flag (sps_mixed_tile_groups_in_pic_flag) in the parameter set, which is directly or indirectly referenced by the tile group to indicate whether there is an image with mixed IRAP and non-IRAP sub-images. For NAL units including IDR tile groups, there is a flag (poc_msb_reset_flag) in the corresponding tile group header to indicate whether the POC MSB is reset in the POC derivation of the image. A variable PicRefreshFlag is defined, which is associated with the image. This flag indicates whether the POC derivation and DPB status should be refreshed when decoding the image. The value of PicRefreshFlag is derived as follows: If the current tile group is included in the first access unit in the codestream, PicRefreshFlag is set to 1. Otherwise, if the current tile group is an IDR tile group, PicRefreshFlag is set to sps_mixed_tile_groups_in_pic_flag? poc_msb_reset_flag: 1. Otherwise, if the current block group is a CRA block group, the following applies. If the current access unit is the first access unit of the decoding sequence, PicRefreshFlag is set to 1. The current access unit is the first access unit of the decoding sequence when the access unit immediately follows the end-of-sequence NAL unit or the associated variable HandleCraAsFirstPicInCvsFlag is set to 1. Otherwise, PicRefreshFlag is set to 0 (for example, the current block group does not belong to the first access unit in the codestream and is not an IRAP block group).
[0122] When PicRefreshFlag is equal to 1, the value of the POC MSB (i.e., PicOrderCntMsb) is reset to 0 during POC derivation for a picture. Regardless of the corresponding NAL unit type, information used for reference picture management, such as a reference picture set (RPS) or a reference picture list (RPL), is indicated in the partition group / slice header. Regardless of the NAL unit type, a reference picture list is constructed at the start of decoding of each partition group. The reference picture list may include RefPicList[0] and RefPicList[1] for the RPL method, RefPicList0[] and RefPicList1[] for the RPS method, or a similar list containing reference pictures used for inter-frame prediction operations of a picture. When PicRefreshFlag is equal to 1, during the reference picture identification process, all reference pictures in the DPB are identified as not used for reference.
[0123] There are some problems with this type of implementation. For example, when nal_unit_type values cannot be mixed in an image, when the derivation of whether an image is an IRAP image and the derivation of the variable NoRaslOutputFlag are described at the image level, the decoder can perform these derivations after receiving the first VCL NAL unit of any image. However, due to the support of mixed NAL unit types in an image, the decoder must wait for the arrival of other VCL NAL units of the image before performing the above derivation. In the worst case, the decoder must wait for the arrival of the last VCL NAL unit of the image. In addition, such a system can indicate a flag in the block group header of the IDRNAL unit to indicate whether the POC MSB is reset in the POC derivation of the image. The mechanism has the following problems. The mechanism does not support the case of mixing CRA NAL unit types and non-IRAP NAL unit types. In addition, when the state of whether the IRAP (IDR or CRA) NAL unit in the image is mixed with the non-IRAP NAL unit changes, indicating this information in the block group / slice header of the VCL NAL unit will require changing the value during code stream extraction or fusion. This rewriting of the slice header needs to be done every time a user requests a video, thus requiring a lot of hardware resources. Furthermore, in addition to a mix of specific IRAP NAL unit types and specific non-IRAP NAL unit types, there can be some other mix of different NAL unit types in the picture. This flexibility cannot be used for practical applications and complicates the design of the codec, which unnecessarily increases the complexity of the decoder and thus the associated implementation cost.
[0124] This invention generally describes techniques for supporting sub-picture or MCTS-based random access in video decoding. More specifically, this invention describes an improved design for supporting mixed NAL unit types in images, which is used to support sub-picture or MCTS-based random access. The description of these techniques is based on the VVC standard, but is also applicable to other video / media codec specifications.
[0125] In order to solve the above-mentioned problems, the following exemplary implementation is disclosed. Such implementation can be applied alone or in combination. In one example, each image is associated with a representation of whether the image contains a mixed nal_unit_type value. The representation is indicated in the PPS. The representation can be used to determine whether to reset the POC MSB and / or reset the DPB by identifying all reference images as not used for reference. When the representation is indicated in the PPS, the change of the value in the PPS can be made during fusion or separate extraction. However, this is acceptable because during such code stream extraction or fusion, the PPS is rewritten and replaced by other mechanisms.
[0126] Alternatively, the representation may be indicated in the partition group header, but the representation must be the same for all partition groups of a picture. However, in this case, the value may need to be changed during the substream extraction of the MCTS / sub-image sequence. Alternatively, the representation may be indicated in the NAL unit header, but the representation must be the same for all partition groups of a picture. However, in this case, the value may need to be changed during the substream extraction of the MCTS / sub-image sequence. Alternatively, the representation may be indicated by defining such other VCL NAL unit types that, when used for a picture, the NAL unit type value of all VCL NAL units of the picture shall be the same. However, in this case, the NAL unit type value of the VCL NAL units may need to be changed during the substream extraction of the MCTS / sub-image sequence. Alternatively, the representation may be indicated by defining such other IRAP VCL NAL unit types that, when used for a picture, the NAL unit type value of all VCL NAL units of the picture shall be the same. However, in this case, the NAL unit type value of the VCL NAL units may need to be changed during the substream extraction of the MCTS / sub-image sequence. Alternatively, each picture with at least one VCL NAL unit with any IRAP NAL unit type may be associated with an indication of whether the picture contains mixed NAL unit type values.
[0127] In addition, constraints may be applied such that mixing of nal_unit_type values in a picture is allowed in a limited way by mixing only IRAP and non-IRAP NAL unit types. For any particular picture, the NAL unit type of all VCL NAL units should be the same, or some VCL NAL units have a specific IRAP NAL unit type and the rest have a specific non-TRAP VCL NAL unit type. In other words, the VCL NAL units of any particular picture cannot have more than one IRAP NAL unit type, nor more than one non-IRAP NAL unit type. A picture may be considered an IRAP picture only if the picture does not contain mixed nal_unit_type values and the VCL NAL units have IRAP NAL unit types. The POC MSB may not be reset for any IRAP NAL unit (including IDR) that does not belong to an IRAP picture. The DPB is not reset for any TRAP NAL unit (including IDR) that does not belong to an IRAP picture, so all reference pictures are not identified as not used for reference. If at least one VCL NAL unit of a picture is an IRAP NAL unit, the TemporalId of the picture may be set to 0.
[0128] The following is a specific implementation of one or more of the above aspects. An IRAP picture may be defined as a coded picture whose mixed_nalu_types_in_pic_flag value is equal to 0 and the nal_unit_type of each VCL NAL unit is in the range of IDR_W_RADL to RSV_IRAP_VCL13 (including the end value). An example of PPS syntax and semantics is as follows.
[0129]
[0130]
[0131] mixed_nalu_types_in_pic_flag is set to 0 to indicate that each picture that references the PPS has multiple VCL NAL units and the nal_unit_type values of these NAL units are different. mixed_nalu_types_in_pic_flag is set to 0 to indicate that the nal_unit_type values of the VCL NAL units of each picture that references the PPS are the same.
[0132] An example chunk group / strip header syntax is as follows:
[0133]
[0134]
[0135] Exemplary NAL unit header semantics are as follows. For the VCL NAL unit of any particular image, one of the following two conditions should be met. The nal_unit_type values of all VCL NAL units are the same. Some VCL NAL units have specific IRAP NAL unit type values (i.e., nal_unit_type values in the range of IDR_W_RADL to RSV_IRAP_VCL13 (including the end value)), while all other VCL NAL units have specific non-IRAP VCL NAL unit types (i.e., nal_unit_type values in the range of TRAIL_NUT to RSV_VCL_7 (including the end value), or in the range of RSV_VCL14 to RSV_VCL15 (including the end value)). nuh_temporal_id_plus1-1 represents the temporal identifier of the NAL unit. The value of nuh_temporal_id_plus1 should not be equal to 0.
[0136] The derivation process of the variable TemporalId is as follows:
[0137] TemporalId=nuh_temporal_id_9lus1-1 (7-1)
[0138] When nal_unit_type is in the range of IDR_W_RADL to RSV_IRAP_VCL13, inclusive, for the VCL NAL units of a picture, the TemporalId of all VCL NAL units of the picture shall be equal to 0, regardless of the value of nal_unit_type of other VCL NAL units of the picture. The value of TemporalId shall be the same for all VCL NAL units of an access unit. The value of TemporalId of a coded picture or access unit is the value of TemporalId of the VCL NAL unit of the coded picture or access unit.
[0139] An exemplary decoding process for a coded image is as follows. The decoding process for the current image CurrPic is as follows. This document specifies the decoding of NAL units. The following decoding process uses syntax elements at and above the partition group header layer. Variables and functions related to the image sequence number are derived as specified in this document. This is only called for the first partition group / slice of the image. At the beginning of the decoding process for each partition group / slice, the decoding process for constructing a reference image list is called to derive reference image list 0 (RefPicList[0]) and reference image list 1 (RefPicList[1]). If the current image is an IDR image, the decoding process for constructing a reference image list can be called for the purpose of bitstream consistency checking, but may not be necessary for decoding the current image or images after the current image in decoding order.
[0140] The decoding process for building the reference picture list is as follows. This process is called at the beginning of the decoding process for each partition group. Reference pictures are addressed by reference indices. The reference index is an index into the reference picture list. When decoding an I partition group, the reference picture list is not used to decode the partition group data. When decoding a P partition group, only reference picture list 0 (RefPicList[0]) is used to decode the partition group data. When decoding a B partition group, both reference picture list 0 and reference picture list 1 (RefPicList[1]) are used to decode the partition group data. At the beginning of the decoding process for each partition group, reference picture lists RefPicList[0] and RefPicList[1] are derived. These two reference picture lists are used to identify reference pictures or decode partition group data. For any partition group of an IDR picture or an I partition group of a non-IDR picture, RefPicList[0] and RefPicList[1] may be derived for codestream consistency checking, but RefPicList[0] and RefPicList[1] may not be derived for decoding the current picture or pictures that follow the current picture in decoding order. For P partition groups, RefPicList[1] may be derived for codestream consistency checking, but RefPicList[1] may not be derived for decoding the current picture or pictures that follow the current picture in decoding order.
[0141] Figure 8Schematic diagram of an exemplary video decoding device 800. As described herein, the video decoding device 800 is suitable for implementing the disclosed examples / embodiments. The video decoding device 800 includes a downstream port 820, an upstream port 850 and / or a transceiver unit (Tx / Rx) 810, and the transceiver unit (Tx / Rx) 810 includes a transmitter and / or a receiver for transmitting data upstream and / or downstream through a network. The video decoding device 800 also includes: a processor 830 including a logic unit and / or a central processing unit (CPU) for processing data; and a memory 832 for storing data. The video decoding device 800 may also include electrical components, optical-to-electrical (OE) components, electrical-to-optical (EO) components, and / or wireless communication components coupled to the upstream port 850 and / or the downstream port 820 for transmitting data through an electrical, optical or wireless communication network. The video decoding device 800 may also include an input and / or output (I / O) device 860 for communicating data with a user. The I / O device 860 may include an output device, such as a display for displaying video data, a speaker for outputting audio data, etc. The I / O device 860 may also include an input device, such as a keyboard, a mouse, a trackball, etc., and / or a corresponding interface for interacting with these output devices.
[0142] The processor 830 is implemented by hardware and software. The processor 830 can be implemented as one or more CPU chips, cores (e.g., multi-core processors), field-programmable gate arrays (FPGAs), application specific integrated circuits (ASICs), and digital signal processors (DSPs). The processor 830 communicates with the downstream port 820, Tx / Rx 810, the upstream port 850, and the memory 832. The processor 830 includes a decoding module 814. The decoding module 814 implements the disclosed embodiments described herein, such as methods 100, 900, and 1000, which can use CVS 500, VR image video stream 600, and / or code stream 700. The decoding module 814 can also implement any other method / mechanism described herein. In addition, the decoding module 814 can implement the codec system 200, the encoder 300, and / or the decoder 400. For example, the decoding module 814 can set a flag in the PPS to indicate when a picture contains both IRAP NAL units and non-IRAP NAL units, and limit such pictures to contain only a single type of IRAP NAL unit and a single type of non-IRAP NAL unit. Therefore, the decoding module 814 enables the video decoding device 800 to have other functions and / or higher decoding efficiency when decoding video data. Therefore, the decoding module 814 improves the functionality of the video decoding device 800 and solves problems unique to the field of video decoding. In addition, the decoding module 814 can transform the video decoding device 800 to different states. Alternatively, the decoding module 814 can be implemented as instructions stored in the memory 832 and executed by the processor 830 (e.g., a computer program product stored in a non-transitory medium).
[0143] The memory 832 includes one or more memory types, such as a disk, a tape drive, a solid-state drive, a read only memory (ROM), a random access memory (RAM), a flash memory, a ternary content-addressable memory (TCAM), a static random-access memory (SRAM), etc. The memory 832 can be used as an overflow data storage device to store programs when such programs are selected for execution, and to store instructions and data read during program execution.
[0144] Fig. 91 is a flowchart of an exemplary method 900 for encoding a video sequence (e.g., CVS 500) containing images with mixed NAL unit types into a bitstream (e.g., a bitstream 700 including a VR image video stream 600 fused from a plurality of sub-image video streams 601-603 at a plurality of video resolutions). The encoder (e.g., the codec system 200, the encoder 300, and / or the video decoding device 800) may use the method 900 when executing the method 100.
[0145] Method 900 may start when an encoder receives a video sequence including multiple images (eg, VR images) and determines to encode the video sequence into a bitstream based on user input, etc. In step 901 , the encoder determines that an image of the video sequence includes multiple sub-images of different types.
[0146] In step 903, the encoder encodes the sub-images of the image into a plurality of VCL NAL units in the bitstream.
[0147] In step 905, the encoder encodes the PPS into the bitstream. The encoder also encodes a flag into the PPS and thus into the bitstream. The flag is set to indicate that the VCL NAL units in which one or more sub-images of the sub-images of the image are located all have a first specific NAL unit type value, and the other VCL NAL units of the image all have different second specific NAL unit type values. For example, the first specific NAL unit type value may indicate that the image contains a single type of IRAP sub-image. In a specific example, the first specific NAL unit type value may be equal to one of IDR_W_RADL, IDR_N_LP, or CRA_NUT. Therefore, the image may have any number of IRAP sub-images, but all IRAP sub-images should be of the same type (e.g., one and only one of IDR_W_RADL, IDR_N_LP, or CRA_NUT). In addition, the second specific NAL unit type value may indicate that the image contains a single type of non-IRAP sub-image. In a specific example, the second specific NAL unit type value may be equal to TRAIL_NUT, RADL_NUT, or RASL_NUT. Thus, a picture may have any number of non-IRAP sub-pictures, but all non-IRAP sub-pictures should be of the same type (e.g., one and only one of TRAIL_NUT, RADL_NUT, or RASL_NUT).
[0148] In step 907, the encoder stores the bitstream including the flag for sending to the decoder. In some examples, the flag is mixed_nalu_types_in_pic_flag. In a specific example, when each picture representing the reference PPS has more than one VCL NAL unit and the NAL unit type (nal_unit_type) values of the VCL NAL units are not the same, mixed_nalu_types_in_pic_flag can be set to 1. In addition, when each picture of the reference PPS has one or more VCL NAL units and the nal_unit_type values of the VCL NAL units of each picture of the reference PPS are the same, mixed_nalu_types_in_pic_flag can be set to 0.
[0149] Fig.10 1 is a flowchart of an exemplary method 1000 for decoding a video sequence (e.g., CVS 500) containing images with mixed NAL unit types from a bitstream (e.g., a bitstream 700 including a VR image video stream 600 fused from a plurality of sub-image video streams 601-603 at a plurality of video resolutions). A decoder (e.g., the codec system 200, the decoder 400, and / or the video decoding device 800) may use the method 1000 when executing the method 100.
[0150] For example, after method 900 ends, method 1000 may start when a decoder begins to receive a code stream representing encoded data of a video sequence. In step 1001, the decoder receives a code stream. The code stream includes a flag and a plurality of sub-images associated with an image. The sub-image is divided into strips, which are contained in VCL NAL units. Therefore, each of the plurality of sub-images is also contained in a plurality of VCL NAL units. The code stream may also include a PPS. In some examples, the PPS includes a flag. In a specific example, the flag may be mixed_nalu_types_in_pic_flag. In addition, when each image representing a reference PPS has more than one VCL NAL unit and the nal_unit_type values of the VCL NAL units are not the same, mixed_nalu_types_in_pic_flag may be set to 1. In addition, when each image of the reference PPS has one or more VCL NAL units and the nal_unit_type values of the VCL NAL units of each image of the reference PPS are the same, mixed_nalu_types_in_pic_flag may be set to 0.
[0151] In step 1003, the decoder determines, based on the value of the flag, that the VCL NAL units in which one or more sub-images of the sub-images of the image are located all have a first specific NAL unit type value, and the other (the remaining) VCL NAL units of the image (the remaining sub-images) all have different second specific NAL unit type values. For example, the first specific NAL unit type value may indicate that the image contains a single type of IRAP sub-image. In a specific example, the first specific NAL unit type value may be equal to one of IDR_W_RADL, IDR_N_LP, or CRA_NUT. Therefore, the image may have any number of IRAP sub-images, but all IRAP sub-images should be of the same type (e.g., one and only one of IDR_W_RADL, IDR_N_LP, or CRA_NUT). In addition, the second specific NAL unit type value may indicate that the image contains a single type of non-IRAP sub-image. In a specific example, the second specific NAL unit type value may be equal to TRAIL_NUT, RADL_NUT, or RASL_NUT. Thus, a picture may have any number of non-IRAP sub-pictures, but all non-IRAP sub-pictures should be of the same type (e.g., one and only one of TRAIL_NUT, RADL_NUT, or RASL_NUT).
[0152] In step 1005, the decoder decodes one or more sub-pictures in the sub-picture according to the first specific NAL unit type value and the second specific NAL unit type value.
[0153] In step 1007, one or more of the sub-images are forwarded for display as part of the decoded video sequence.
[0154] Fig.11 1 is a schematic diagram of an exemplary system 1100 for encoding a video sequence (e.g., CVS 500) containing images with mixed NAL unit types into a bitstream (e.g., a bitstream 700 including a VR image video stream 600 fused from a plurality of sub-image video streams 601-603 at a plurality of video resolutions). The system 1100 may be implemented by an encoder and a decoder (e.g., the encoding and decoding system 200, the encoder 300, the decoder 400, and / or the video decoding device 800). In addition, the system 1100 may be used to implement the methods 100, 900, and / or 1000.
[0155] The system 1100 includes a video encoder 1102. The video encoder 1102 includes a determination module 1101, which is used to determine that an image contains multiple sub-images of different types. The video encoder 1102 also includes an encoding module 1103, which is used to encode the sub-images of the image into multiple VCL NAL units in a bitstream. The encoding module 1103 is also used to encode a flag into the bitstream, and the flag is set to indicate that the VCL NAL units where one or more sub-images of the sub-images of the image are located have a first specific NAL unit type value, and other VCL NAL units of the image have different second specific NAL unit type values. The video encoder 1102 also includes a storage module 1105, which is used to store a bitstream for sending to a decoder. The video encoder 1102 also includes a sending module 1107, which is used to send the bitstream to a video decoder 1110. The video encoder 1102 can also be used to perform any step in the method 900.
[0156] The system 1100 further includes a video decoder 1110. The video decoder 1110 includes a receiving module 1111 for receiving a bitstream including a flag and a plurality of sub-images associated with an image, wherein the plurality of sub-images are contained in a plurality of VCL NAL units. The video decoder 1110 further includes a determining module 1113 for determining, based on the value of the flag, that the VCL NAL units in which one or more sub-images of the sub-images of the image are located all have a first specific NAL unit type value, and that the other VCL NAL units of the image all have different second specific NAL unit type values. The video decoder 1110 further includes a decoding module 1115 for decoding one or more sub-images of the sub-images based on the first specific NAL unit type value and the second specific NAL unit type value. The video decoder 1110 further includes a forwarding module 1117 for forwarding one or more sub-images of the sub-images so as to be displayed as part of a decoded video sequence. The video decoder 1110 can also be used to perform any step in the method 1000.
[0157] A first component is directly coupled to a second component when there is no intermediate component between the first component and the second component except a line, trace or other medium. A first component is indirectly coupled to a second component when there is an intermediate component between the first component and the second component except a line, trace or other medium. The term "coupled" and its synonyms include direct coupling and indirect coupling. Unless otherwise specified, the term "approximately" means a range of ±10% of the number that follows it.
[0158] It should also be understood that the steps of the exemplary methods set forth herein do not necessarily need to be performed in the order described, and the order of the steps of these methods should be understood to be merely exemplary. Similarly, in methods consistent with various embodiments of the present invention, these methods may include other steps, and some steps may be omitted or combined.
[0159] Although several embodiments have been provided in the present invention, it is understood that the disclosed systems and methods may be embodied in many other specific forms without departing from the spirit and scope of the present invention. The present examples are considered to be illustrative rather than restrictive, and are not intended to be limited to the details given herein. For example, various elements or components may be combined or integrated in another system, or certain features may be omitted or not implemented.
[0160] In addition, the techniques, systems, subsystems and methods described and shown as discrete or separate in various embodiments may be combined or integrated with other systems, components, techniques or methods without departing from the scope of the present invention. Other examples of changes, substitutions and modifications may be determined by those skilled in the art, and changes, substitutions and modifications may be made without departing from the spirit and scope of the present invention.
Claims
1. A coding method, characterized in that: The method comprises: Divide the image into multiple sub-images; encode the multiple sub-images into multiple video decoding layer VCL network abstraction layer NAL units in the bitstream; A flag is encoded into the picture parameter set PPS of the code stream. When the flag is set to a first value, it indicates that all VCL NAL units of the image have the same NAL unit type. When the flag is set to a second value, it indicates that one or more VCL NAL units of the multiple VCL NAL units of the image have a first specific NAL unit type value, and other VCL NAL units have a second specific NAL unit type value. The first specific NAL unit type value is a specific value in the intra random access point IRAP NAL unit type, and the second specific NAL unit type value is a specific value in the non-IRAP NAL unit type.
2. The method according to claim 1, characterized in that The first specific NAL unit type value is equal to one of an immediate decoding refresh IDR with a decodable random access preceding picture (IDR_W_RADL), an IDR without preceding pictures (IDR_N_LP), or a pure random access CRA NAL unit type (CRA_NUT).
3. The method according to claim 1 or 2, characterized in that: The second specific NAL unit type value is equal to one of a trailing picture NAL unit type (TRAIL_NUT), a decodable random access preceding picture NAL unit type (RADL_NUT), or a skipped random access preceding RASL picture NAL unit type (RASL_NUT).
4. The method according to claim 1 or 2, characterized in that: The flag is mixed_nalu_types_in_pic_flag.
5. The method according to claim 1 or 2, characterized in that: The method further includes: sending the code stream to a terminal device; or storing the code stream in at least one storage medium.
6. The method according to claim 1 or 2, characterized in that: The code stream includes a plurality of sub-image video streams with different resolutions.
7. The method according to claim 1 or 2, characterized in that: The code stream includes sub-image video streams of different viewing angles, and the method further includes: The code stream is sent to a decoding end so that the decoding end decodes and displays the target sub-image video stream according to the viewing angle selected by the user.
8. A video encoding device, characterized in that: include: A processor and a memory, wherein the memory stores program instructions, and when the program instructions are executed on the processor, the video encoding device executes the method according to any one of claims 1 to 7.
9. A non-transitory computer-readable medium, characterized in that A computer program product for use in a video decoding device, wherein the computer program product comprises computer executable instructions stored in the non-transitory computer readable medium, and when the computer executable instructions are executed by a processor, the video decoding device performs the method according to any one of claims 1 to 7.
10. An encoder, characterized in that: include: A division module, used for dividing an image into a plurality of sub-images; Encoding module for: Encoding the multiple sub-images into multiple video coding layer VCL network abstraction layer NAL units in a bitstream; A flag is encoded into the picture parameter set PPS of the code stream. When the flag is set to a first value, it indicates that all VCL NAL units of the image have the same NAL unit type. When the flag is set to a second value, it indicates that one or more VCL NAL units of the multiple VCL NAL units of the image have a first specific NAL unit type value, and other VCL NAL units have a second specific NAL unit type value. The first specific NAL unit type value is a specific value in the intra random access point IRAP NAL unit type, and the second specific NAL unit type value is a specific value in the non-IRAP NAL unit type.
11. The encoder according to claim 10, characterized in that The first specific NAL unit type value is equal to one of an immediate decoding refresh IDR with a decodable random access preceding picture (IDR_W_RADL), an IDR without preceding pictures (IDR_N_LP), or a pure random access CRA NAL unit type (CRA_NUT).
12. The encoder according to claim 10 or 11, characterized in that The second specific NAL unit type value is equal to one of a trailing picture NAL unit type (TRAIL_NUT), a decodable random access preceding picture NAL unit type (RADL_NUT), or a skipped random access preceding RASL picture NAL unit type (RASL_NUT).
13. The encoder according to claim 10 or 11, characterized in that: The flag is mixed_nalu_types_in_pic_flag.
14. The encoder according to claim 10 or 11, characterized in that Also includes: A sending module, used to send the code stream to a terminal device; and / or The storage module is used to store the code stream.
15. The encoder according to claim 10 or 11, characterized in that The code stream includes sub-image video streams of different viewing angles or the code stream includes multiple sub-image video streams of different resolutions, and the encoder further includes: The sending module sends the code stream to the decoding end so that the decoding end decodes and displays the target sub-image video stream according to the viewing angle selected by the user.
16. A device for storing a code stream, characterized in that: The device includes a receiver and a memory; The receiver is used to receive a code stream, and the memory is used to store the code stream; The code stream includes a flag and coded data of multiple sub-images associated with the image, the flag is contained in a picture parameter set PPS of the code stream, and the coded data of the sub-images is contained in multiple video decoding layer VCL network abstraction layer NAL units; when the flag is set to a first value, it indicates that all VCL NAL units of the image have the same NAL unit type, and when the flag is set to a second value, it indicates that one or more VCL NAL units of the multiple VCL NAL units of the image have a first specific NAL unit type value, and other VCL NAL units have a second specific NAL unit type value, the first specific NAL unit type value is a specific value in the intra random access point IRAP NAL unit type, and the second specific NAL unit type value is a specific value in the non-IRAP NAL unit type.
17. A device for encoding video data to obtain a code stream, characterized in that: including a processor and memory, The processor is used to execute program instructions to perform the encoding method according to any one of claims 1 to 7 to generate a bit stream; The memory is used to store the code stream.