Stripe entry points in video coding

By deducing the number of entry points in the strip in the video decoder to determine the offset of the data subset, the problem of low video compression and decompression efficiency in the prior art is solved, and a higher compression ratio and lower resource utilization are achieved.

CN120111247APending Publication Date: 2025-06-06HUAWEI TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510112316.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2019-04-10
Filing Date
2020-04-09
Publication Date
2025-06-06

AI Technical Summary

Technical Problem

Existing video decoding technologies are difficult to effectively improve the compression ratio during compression and decompression, especially when network resources are limited, resulting in limited video quality and transmission efficiency.

Method used

Efficient decoding of video strips is achieved by deriving the number of entry points in the strip in the decoder without relying on the indicated values ​​in the code stream.

Benefits of technology

This method reduces redundant data in the video stream, improves decoding efficiency, reduces the utilization of memory and network resources, while maintaining image quality.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120111247A_ABST
    Figure CN120111247A_ABST
Patent Text Reader

Abstract

The invention discloses a video coding mechanism. The mechanism comprises: receiving a code stream comprising a slice; deriving the number of entry points (NumEntry Points) in the strip, and deriving the number of entry points (NumEntry Points) in the strip; determining an offset of a subset of the encoded stripe data, where a subset index value of the subset ranges from 0 to NumEntrryPoints, and determining an offset of the subset of the encoded stripe data, where a subset index value of the subset ranges from 0 to NumEnryPoints; decoding the slice according to an offset of the encoded slice data subset; the slice is forwarded for display as part of the decoded video sequence.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] This application is a divisional application. The application number of the original application is 202080027784.0 and the original application date is April 9, 2020. The entire contents of the original application are incorporated into this application by reference.

[0002] Cross-reference to related applications

[0003] This patent application claims the benefit of U.S. Provisional Patent Application No. 62 / 832,128, filed by Ye-Kui Wang on April 10, 2019, entitled “Video Coding Improvements,” which is incorporated herein by reference. Technical Field

[0004] The present disclosure relates generally to video coding, and more particularly to determining entry points for coded data in a slice of a picture in video coding. Background Art

[0005] Even in the case of shorter videos, a large amount of video data is required to describe it, which can cause difficulties when the data is to be streamed or otherwise transmitted across a communication network with limited bandwidth capacity. Therefore, video data is often compressed before being sent across modern telecommunications networks. Since memory resources may be limited, the size of the video may also become an issue when storing the video on a storage device. Video compression devices typically use software and / or hardware on the source side to decode the video data before transmission or storage, thereby reducing the amount of data required to represent the digital video image. A video decompression device that decodes the video data then receives the compressed data at the destination side. With limited network resources and a growing demand for higher video quality, there is a need for improved compression and decompression techniques that can increase the compression ratio with little impact on image quality. Summary of the invention

[0006] In one embodiment, the present invention includes a method implemented in a decoder. The method includes: a receiver of the decoder receives a code stream including a slice; a processor of the decoder derives the number of entry points (NumEntryPoints) in the slice; the processor determines an offset of a subset of coded slice data, wherein the subset index value of the subset ranges from 0 to NumEntryPoints; the processor decodes the slice according to the offset of the coded slice data subset; the processor forwards the slice to be displayed as part of a decoded video sequence. In some video decoding systems, the encoder indicates (signal) num_entry_point_offsets values ​​for each slice. Then, the decoder sets the NumEntryPoints value based on the code stream. The value is used to determine the offset of each group of coded data in the slice. Then, the offset can be used to decode these groups. This example uses a mechanism for determining the NumEntryPoints value without indicating it in the code stream. NumEntryPoints can then be used to obtain offsets for a subset of slices (e.g., coding tree unit (CTU) rows), and then the subset can be reconstructed to reconstruct the slices. For example, these mechanisms can be used to derive the NumEntryPoints of a slice when the slice is decoded according to wavefront parallel processing (WPP). NumEntryPoints can be derived based on the size of the WPP / CTU row, the size of the WPP / CTU column, the number of CTUs in the slice, and / or the CTU and / or slice address. This method deletes the indication value of each slice header, thereby deleting the indication value of each slice in the video stream. An image may include multiple slices. In addition, a video sequence may include thousands of images. Therefore, deleting num_entry_point_offsets from the bitstream significantly reduces the video stream size, improves decoding efficiency, and reduces the utilization of memory resources and network resources on the encoder and decoder sides.

[0007] Optionally, according to any of the above aspects, in another implementation manner of the aspect, the NumEntryPoints are derived when the strip is decoded according to wavefront parallel processing (WPP).

[0008] Optionally, according to any of the above aspects, in another implementation manner of the aspect, the code stream includes a slice header, and the slice header does not include a value corresponding to the number of entry point offsets in the slice.

[0009] Optionally, according to any of the above aspects, in another implementation of the aspect, the NumEntryPoints is derived based on the row size in the stripe.

[0010] Optionally, according to any of the above aspects, in another implementation of the aspect, the NumEntryPoints is derived based on the column size in the stripe.

[0011] Optionally, according to any of the above aspects, in another implementation manner of the aspect, the NumEntryPoints is derived based on the number of coding tree units (CTU) in the slice.

[0012] Optionally, according to any of the above aspects, in another implementation of the aspect, the NumEntryPoints is derived based on the addresses in the stripe.

[0013] Optionally, according to any of the above aspects, in another implementation of the aspect, the NumEntryPoints is derived based on the size of the stripe.

[0014] In one embodiment, the present invention includes a method implemented in an encoder, wherein the method includes: a processor of the encoder obtains a reference stripe of a coded reference image; the processor derives NumEntryPoints in the reference stripe; the processor determines an offset of a subset of coded data in the reference stripe based on the NumEntryPoints; the processor decodes the reference stripe of the coded reference image based on the offset of the subset of coded data in the reference stripe; the processor encodes a current stripe into a bitstream based on the reference stripe; a memory coupled to the processor stores the bitstream for transmission to a decoder. In some video decoding systems, the encoder indicates (signal) num_entry_point_offsets values ​​for each stripe. Then, the decoder sets the NumEntryPoints value based on the bitstream. The value is used to determine the offset of each group of coded data in the stripe. Then, the offsets can be used to decode these groups. This example uses a mechanism for determining the NumEntryPoints value without indicating it in the bitstream. NumEntryPoints can then be used to obtain offsets for a subset of slices (e.g., coding tree unit (CTU) rows), and then the subset can be reconstructed to reconstruct the slices. For example, these mechanisms can be used to derive the NumEntryPoints of a slice when the slice is decoded according to wavefrontparallel processing (WPP). NumEntryPoints can be derived based on the size of the WPP / CTU row, the size of the WPP / CTU column, the number of CTUs in the slice, and / or the CTU and / or slice address. This method deletes the indication value of each slice header, thereby deleting the indication value of each slice in the video stream. An image may include multiple slices. In addition, a video sequence may include thousands of images. Therefore, removing num_entry_point_offsets from the bitstream significantly reduces the size of the video stream, improves decoding efficiency, and reduces the utilization of memory resources and network resources on the encoder and decoder sides.

[0015] Optionally, according to any of the above aspects, in another implementation manner of the aspect, the NumEntryPoints is derived when the reference slice is decoded according to WPP.

[0016] Optionally, according to any of the above aspects, in another implementation manner of the aspect, the code stream includes a slice header, and the slice header does not include a value corresponding to the number of entry point offsets in the reference slice.

[0017] Optionally, according to any of the above aspects, in another implementation of the aspect, the NumEntryPoints is derived based on the row size in the reference strip.

[0018] Optionally, according to any of the above aspects, in another implementation of the aspect, the NumEntryPoints is derived based on the column size in the reference strip.

[0019] Optionally, according to any of the above aspects, in another implementation of the aspect, the NumEntryPoints is derived based on the number of CTUs in the reference strip.

[0020] Optionally, according to any of the above aspects, in another implementation manner of the aspect, the NumEntryPoints is derived based on the address in the reference strip.

[0021] Optionally, according to any of the above aspects, in another implementation manner of the aspect, the NumEntryPoints is derived based on the size of the reference strip.

[0022] In one embodiment, the present invention includes a video decoding device, comprising: a processor, a receiver coupled to the processor, and a memory coupled to the processor; a transmitter coupled to the processor, wherein the processor, receiver, memory and transmitter are used to execute the method described in any one of the above aspects.

[0023] In one embodiment, the present invention includes a non-transitory computer-readable medium, including a computer program product for use by a video decoding device, wherein the computer program product includes computer-executable instructions stored in the non-transitory computer-readable medium, and when a processor executes the computer-executable instructions, the video decoding device performs the method described in any of the above aspects.

[0024] In one embodiment, the present invention includes a decoder, wherein the decoder includes: a receiving module for receiving a code stream including a stripe; a deriving module for deriving NumEntryPoints in the stripe; a determining module for determining an offset of a subset of encoded stripe data, wherein a subset index value of the subset ranges from 0 to NumEntryPoints; a decoding module for decoding the stripe according to the offset of the subset of encoded stripe data; and a forwarding module for forwarding the stripe to be displayed as part of a decoded video sequence.

[0025] Optionally, according to any of the above aspects, in another implementation manner of the aspect, the decoder is also used to execute the method described in any of the above aspects.

[0026] In one embodiment, the present invention includes an encoder, wherein the encoder includes: an acquisition module, used to acquire a reference strip of an encoded reference image; a derivation module, used to derive NumEntryPoints in the reference strip; a determination module, used to determine an offset of an encoded data subset in the reference strip according to the NumEntryPoints; a decoding module, used to: decode the reference strip of the encoded reference image according to the offset of the encoded data subset in the reference strip; encode a current strip into a code stream according to the reference strip; and a storage module, used to store the code stream for sending to a decoder.

[0027] Optionally, according to any of the above aspects, in another implementation manner of the aspect, the encoder is also used to execute the method described in any of the above aspects.

[0028] For clarity of description, any of the above-described embodiments may be combined with any or more of the other above-described embodiments to create new embodiments within the scope of the present invention.

[0029] These and other features will be more clearly understood from the following detailed description taken in conjunction with the accompanying drawings and claims. BRIEF DESCRIPTION OF THE DRAWINGS

[0030] For a more thorough understanding of the present invention, reference is now made to the following brief description taken in conjunction with the accompanying drawings and detailed description, wherein like reference numerals represent like parts.

[0031] Figure 1 A flow chart of an exemplary method of decoding a video signal.

[0032] Figure 2 is a schematic diagram of an exemplary codec (codec) system for video coding.

[0033] Figure 3 is a schematic diagram of an exemplary video encoder.

[0034] Figure 4 is a schematic diagram of an exemplary video decoder.

[0035] Figure 5 A schematic diagram of an exemplary mechanism of wavefront parallel processing (WPP).

[0036] Figure 6 A schematic diagram of an exemplary code stream.

[0037] Figure 7 is a schematic diagram of an exemplary video decoding device.

[0038] Figure 8 The present invention is a flowchart of an exemplary method for encoding a video sequence into a bitstream without indicating the number of entry points (NumEntryPoints) of a slice.

[0039] Fig. 9 Flowchart of an exemplary method for deriving NumEntryPoints to decode a video sequence when NumEntryPoints is not included in a codestream.

[0040] Fig.10 is a diagram of an exemplary system for decoding a video sequence into a bitstream without indicating NumEntryPoints. DETAILED DESCRIPTION

[0041] First, it should be understood that although illustrative implementations of one or more embodiments are provided below, the systems and / or methods disclosed herein may be implemented using any number of techniques, whether currently known or existing. The present invention should in no way be limited to the illustrative implementations, drawings, and techniques described below, including the exemplary designs and implementations illustrated and described herein, but may be modified within the scope of the appended claims and their full scope of equivalents.

[0042] The following definitions of terms are as follows, unless used in a contrary context herein. Specifically, the following definitions are intended to more clearly describe the present invention. However, terms may be described differently in different contexts. Therefore, the following definitions should be considered as supplementary information and should not be considered to limit any other definitions provided herein for the descriptions of these terms.

[0043] A bitstream is a series of bits that include video data that is compressed for transmission between an encoder and a decoder. An encoder is a device that compresses video data into a bitstream using an encoding process. A decoder is a device that reconstructs video data from a bitstream for display using a decoding process. A picture is a complete image that is intended to be displayed to a user in whole or in part at a corresponding moment in a video sequence. A reference picture is an image that includes reference samples that can be used when decoding other images by reference according to inter-frame prediction. A decoded picture is a representation of an image that is decoded according to inter-frame prediction or intra-frame prediction, included in a single access unit in a bitstream, and includes a complete set of coding tree units (CTUs) of the image. A slice is a partition of an image that includes an integer number of complete blocks in the image or an integer number of consecutive complete CTU rows within a block, where a slice and all subdivisions are included in only a single network abstraction layer (NAL) unit. A reference slice is a slice of a reference image that includes reference samples or can be used when decoding other slices by reference according to inter-frame prediction. A slice header is a part of a coded slice that includes data elements related to all blocks represented in the slice or CTU rows within a block. An entry point is a bit position in the bitstream that includes the first bit of video data for the corresponding subset of the coded slice. An offset is the bit distance between a known bit position and an entry point. A subset is a subdivision of a set. For example, when a slice is a set, a block, a CTU / coding tree block (CTB) row, or a CTU / CTB is a subset of a set. A coding tree unit (CTU) is a group of samples of a predefined size that can be partitioned by a coding tree. A CTU row is a group of CTUs that extend horizontally between the left slice boundary and the right slice boundary. A CTU column is a group of CTUs extending vertically between the upper and lower stripe boundaries. A CTB is a portion of a CTU that includes only luma samples, only red difference chroma samples, or only blue difference chroma samples. A CTB row / CTB column is a CTU row / column that includes only luma samples, only red difference chroma samples, or only blue difference chroma samples. It should be noted that CTU and CTB can be used interchangeably in many cases.Wavefront parallel processing (WPP) is a mechanism to delay decoding of CTU / CTB rows of a slice so that all rows are decoded simultaneously by different threads. A slice address is an identifiable location of a slice or a sub-portion thereof.

[0044] The following abbreviations are used in this document: Coded Video Sequence (CVS), Decoded Picture Buffer (DPB), Instantaneous Decoding Refresh (IDR), Intra-Random Access Point (IRAP), Joint Video Experts Team (JVET), Least Significant Bit (LSB), Most Significant Bit (MSB), Network Abstraction Layer (NAL), Picture Order Count (POC), Raw Byte Sequence Payload (RBSP), Real-Time Transport Protocol (RTP), Sequence Parameter Set (SPS), Versatile Video Coding (VVC), Working Draft (WD), and Wavefront Parallel Processing (WPP).

[0045] Many video compression techniques can be used to reduce the size of video files while minimizing data loss. For example, video compression techniques can include performing spatial (e.g., intra-frame) prediction and / or temporal (e.g., inter-frame) prediction to reduce or remove data redundancy in a video sequence. For block-based video decoding, a video slice (e.g., a video image or a portion of a video image) can be divided into video blocks, which can also be referred to as tree blocks, coding tree blocks (CTB), coding tree units (CTU), coding units (CU) and / or coding nodes. Video blocks in an intra-frame decoding (I) slice of an image are decoded using spatial prediction of reference samples in adjacent blocks within the same image. Video blocks in an inter-frame decoding unidirectional prediction (P) or bidirectional prediction (B) slice of an image can be decoded using spatial prediction of reference samples in adjacent blocks within the same image, or using temporal prediction of reference samples in other reference images. A picture (picture / image) can be referred to as a frame, and a reference picture can be referred to as a reference frame. Spatial or temporal prediction produces a prediction block representing an image block. The residual data represents the pixel difference between the original image block and the predicted block. Accordingly, the inter-frame decoding block is encoded according to the motion vector pointing to the block of reference samples constituting the predicted block and the residual data representing the difference between the coded block and the predicted block, and the intra-frame decoding block is encoded according to the intra-frame decoding mode and the residual data. For further compression, the residual data can be transformed from the pixel domain to the transform domain to produce residual transform coefficients, which can then be quantized. The quantized transform coefficients can initially be arranged in a two-dimensional array. The quantized transform coefficients can be scanned to produce a one-dimensional vector of transform coefficients. Entropy decoding can be applied to achieve a greater degree of compression. These video compression techniques will be discussed in more detail below.

[0046] In order to ensure that the encoded video can be accurately decoded, the video is encoded and decoded according to the corresponding video coding standards. These video coding standards include International Telecommunication Union Telecommunication Standardization Sector (ITU-T) H.261, International Organization for Standardization / International Electrotechnical Commission (ISO / IEC) Motion Picture Experts Group (MPEG)-1 Part 2, ITU-T H.262 or ISO / IEC MPEG-2 Part 2, ITU-TH.263, ISO / IEC MPEG-4 Part 2, Advanced Video Coding (AVC) (also known as ITU-T H.264 or ISO / IEC MPEG-4 Part 10), and High Efficiency Video Coding (HEVC) (also known as ITU-T H.265 or MPEG-H Part 2). AVC includes Scalable Video Coding (SVC), Multiview Video Coding (MVC), Multiview Video Coding plus Depth (MVC+D), and three-dimensional (3D) AVC (3D-AVC) and other extended versions. HEVC includes Scalable HEVC (SHVC), Multiview HEVC (MV-HEVC), 3D HEVC (3D-HEVC) and other extended versions. The joint video experts team (JVET) of ITU-T and ISO / IEC has begun to develop a video coding standard called Versatile Video Coding (VVC). VVC is included in the Working Draft (WD), which includes JVET-N1001-v1.

[0047] The video decoding system can encode the video using many different mechanisms. For example, the video decoding system can divide the image into slices. In some examples, the slices are divided into blocks. Then, according to the example, the slices and / or the blocks of the slices can be divided into CTU rows. Then, the CTU is subdivided into coding blocks according to the coding tree, and the coding blocks are encoded according to the prediction mechanism. Then, the encoded coding blocks are included in the code stream. In WPP, CTU rows and their partitions are decoded simultaneously. For example, according to the example, WPP can use a first thread to decode one or two CTUs in the first row, and then use a second thread to start decoding the CTUs in the second row. Once one / two CTUs in the second row are decoded, the third thread can start decoding the CTUs in the third row, and so on. Some prediction schemes decode by referring to blocks located above or to the left of the current block. By setting a CTU delay between starting the next row of decoding, WPP can ensure that blocks located above or to the left of the current block are decoded before decoding the current block. This solves the problem that this data may not be available when the encoder starts decoding the current block.

[0048] Therefore, a large number of bits are included in the code stream. In addition, these bits are respectively associated with the corresponding slices or sub-parts thereof. An offset array can be included in the code stream to help the decoder decode the code stream and find relevant data. Each offset represents the position of the corresponding data in the code stream. The offset of the first bit of the video data of the corresponding subset of the coded slice is called the entry point. In the case of WPP, the subset is a CTU row. Therefore, in WPP, there may be as many entry points in the slice as there are CTU rows in the slice. In some video decoding systems, num_entry_point_offsets is included in the slice header. num_entry_point_offsets is a parameter indicating the number of entry point offsets included in the slice, thus indicating to the decoder the number of offsets that should be obtained from the offset array. The decoder can use the num_entry_point_offsets in the slice header to obtain the relevant offset and start decoding the corresponding CTU. Although this method is effective, num_entry_point_offsets includes multiple bits and is indicated in each slice header. Since each picture may consist of several slices and a video sequence may consist of thousands of pictures, num_entry_point_offsets is indicated many times and may have a significant aggregate impact on the amount of data in the video sequence.

[0049] This article discloses a mechanism for deriving the value of the number of entry points (NumEntryPoints) on the decoder side. NumEntryPoints can then be used to obtain the offset of a slice subset (e.g., a CTU row), and then the subset is reconstructed to reconstruct the slice. In this way, num_entry_point_offsets can be deleted from each slice header of the codestream. These mechanisms can be used on the decoder side and / or can be used at the hypothetical reference decoder (HRD) on the encoder. Since an image can include multiple slices and a video can include thousands of images, removing this value from the codestream significantly reduces the size of the video stream, improves decoding efficiency, and reduces the utilization of memory resources and network resources on the encoder and decoder sides. For example, when the slice is decoded according to WPP, these mechanisms can be used to derive the NumEntryPoints of the slice. NumEntryPoints can be derived based on the size of the WPP / CTU row, the size of the WPP / CTU column, the number of CTUs in the slice, and / or the CTU and / or slice address.

[0050] Figure 1 Flowchart of an exemplary method 100 of operation for decoding a video signal. Specifically, the video signal is encoded at the encoder side. The encoding process compresses the video signal using various mechanisms to reduce the size of the video file. The smaller file size facilitates the transmission of the compressed video file to the user while reducing the associated bandwidth overhead. The decoder then decodes the compressed video file to reconstruct the original video signal for display to the end user. The decoding process is generally the same as the encoding process, which facilitates the decoder to reconstruct the video signal in the same manner.

[0051] In step 101, a video signal is input into an encoder. For example, the video signal can be an uncompressed video file stored in a memory. For another example, the video file can be captured by a video capture device (e.g., a camera) and encoded to support real-time streaming of the video. The video file can include an audio component and a video component at the same time. The video component includes a series of image frames that, when viewed in sequence, produce a visual effect of motion. These frames include pixels represented by light (referred to herein as brightness components (or brightness samples)) and colors (referred to as chrominance components (or color samples)). In some examples, the frame may also include a depth value to support three-dimensional viewing.

[0052] In step 103, the video is segmented into blocks. Segmentation includes subdividing the pixels in each frame into square and / or rectangular blocks for compression. For example, in High Efficiency Video Coding (HEVC) (also known as H.265 and MPEG-H Part 2), the frames can be first divided into coding tree units (CTUs), which are blocks of a predefined size (e.g., 64 pixels × 64 pixels). The CTU includes luma samples and chroma samples. The CTU can be divided into blocks using a coding tree, and then the blocks are recursively subdivided until a configuration structure that supports further encoding is obtained. For example, the luma component of the frame can be subdivided until each block includes a relatively uniform luminance value. In addition, the chroma component of the frame can be subdivided until each block includes a relatively uniform color value. Therefore, the content of the video frame is different, and the segmentation mechanism is different.

[0053] In step 105, various compression mechanisms are used to compress the image blocks segmented in step 103. For example, inter-frame prediction and / or intra-frame prediction may be used. Inter-frame prediction is intended to take advantage of the fact that objects in a common scene tend to appear in consecutive frames. Therefore, there is no need to repeatedly describe blocks depicting objects in a reference frame in adjacent frames. An object, such as a table, may remain in a constant position in multiple frames. Therefore, the table is only described once, and adjacent frames can refer back to the reference frame. Pattern matching mechanisms may be used to match objects across multiple frames. In addition, moving objects may be represented across multiple frames due to, for example, object movement or camera movement. In a specific example, a video may show a car moving on the screen across multiple frames. Motion vectors may be used to describe such movement. A motion vector is a two-dimensional vector that provides an offset from the coordinates of an object in one frame to the coordinates of the object in a reference frame. Therefore, inter-frame prediction may encode image blocks in a current frame as a set of motion vectors, representing the offsets between image blocks in the current frame and corresponding blocks in a reference frame.

[0054] Intra prediction encodes blocks in a common frame. Intra prediction exploits the fact that luminance and chrominance components tend to be clustered in one frame. For example, a patch of green in part of a tree tends to be adjacent to several similar patches of green. Intra prediction uses a variety of directional prediction modes (e.g., 33 modes in HEVC), plane mode, and direct current (DC) mode. Directional mode indicates that the samples of the current block are similar / identical to the samples of the neighboring blocks in the corresponding direction. Planar mode indicates that a series of blocks on a row / column (e.g., a plane) can be interpolated based on the neighboring blocks at the edge of the row. In fact, the plane mode represents the smooth transition of luminance / color between rows / columns by using a relatively constant slope in the changing value. The DC mode is used for boundary smoothing, indicating that the average values ​​of the samples of all neighboring blocks related to the angular direction of the block and the directional prediction mode are similar / identical. Therefore, the intra prediction block can represent the image block as various relationship prediction mode values ​​instead of the actual value. In addition, the inter prediction block can represent the image block as a motion vector value instead of the actual value. In both cases, the prediction block may not fully represent the image block in some cases. Any differences are stored in the residual block. The residual block can be transformed to further compress the file.

[0055] In step 107, various filtering techniques can be applied. In HEVC, filters are applied according to an in-loop filtering scheme. The block-based prediction discussed above can create a blocky image in the decoder. In addition, a block-based prediction scheme can encode the block and then reconstruct the encoded block for later use as a reference block. The in-loop filtering scheme iteratively applies a noise suppression filter, a deblocking filter, an adaptive loop filter, and a sample adaptive offset (SAO) filter to the block / frame. These filters reduce these block artifacts so that the encoded file can be accurately reconstructed. In addition, these filters reduce the reconstruction of the reference block artifacts, making it less likely that the artifacts will produce other artifacts in subsequent blocks encoded based on the reconstructed reference block.

[0056] In step 109, once the video signal is segmented, compressed and filtered, the resulting data is encoded into a bitstream. The bitstream includes the above data and any indicative data that is expected to support appropriate reconstruction of the video signal in the decoder. For example, these data may include segmentation data, prediction data, residual blocks, and various flags that provide decoding instructions to the decoder. The bitstream can be stored in a memory and sent to a decoder upon request. The bitstream can also be broadcast and / or multicast to multiple decoders. The creation of the bitstream is an iterative process. Therefore, steps 101, 103, 105, 107 and 109 can occur continuously and / or simultaneously over multiple frames and blocks. Figure 1The order shown is presented for clarity and ease of discussion and is not intended to limit the video coding process to a particular order.

[0057] In step 111, the decoder receives the code stream and starts the decoding process. Specifically, the decoder uses an entropy decoding scheme to convert the code stream into corresponding syntax and video data. In step 111, the decoder uses the syntax data in the code stream to determine the segmentation of the frame. The segmentation should match the result of the block segmentation in step 103. Now describe the entropy coding / decoding used in step 111. The encoder makes many choices during the compression process, such as selecting a block segmentation scheme from multiple possible options based on the spatial positioning of the values ​​in the input image. A large number of binary bits can be used to indicate the exact option. The binary bits used in this article are binary values ​​that are treated as variables (for example, bit values ​​that can vary according to context). Entropy coding helps the encoder discard any options that are obviously not suitable for a particular situation, leaving a set of usable options. Then, a code word is assigned to each usable option. The length of the code word is based on the number of allowable options (for example, one binary bit is used for two options, and two binary bits are used for three to four options). Then, the encoder encodes the code word of the selected option. This scheme reduces the size of the codeword because the codeword size is as large as desired to uniquely indicate one option in a small subset of available options, rather than uniquely indicating an option in a possibly large set of all possible options. The decoder then decodes the option by determining the set of available options in a similar manner to the encoder. By determining the set of available options, the decoder can read the codeword and determine the selection made by the encoder.

[0058] In step 113, the decoder performs block decoding. Specifically, the decoder performs an inverse transform to generate a residual block. The decoder then uses the residual block and the corresponding prediction block to reconstruct the image block according to the segmentation. The prediction block may include an intra-frame prediction block and an inter-frame prediction block generated by the encoder in step 105. The reconstructed image block is then placed in a frame of the reconstructed video signal according to the segmentation data determined in step 111. The syntax of step 113 can also be indicated in the bitstream by entropy coding discussed above.

[0059] In step 115, filtering is performed on the frame of the reconstructed video signal in a manner similar to that of the encoder in step 107. For example, a noise suppression filter, a deblocking filter, an adaptive loop filter, and a SAO filter may be applied to the frame to remove blocking artifacts. Once the frame is filtered, the video signal may be output to a display in step 117 for viewing by an end user.

[0060] Figure 21 is a schematic diagram of an exemplary encoding and decoding (codec) system 200 for video decoding. Specifically, the codec system 200 is capable of implementing the operating method 100. Broadly speaking, the codec system 200 is used to describe components used in an encoder and a decoder. As discussed with respect to steps 101 and 103 in the operating method 100, the codec system 200 receives a video signal and segments the video signal to produce segmented video signals 201. Then, when acting as an encoder, the codec system 200 compresses the segmented video signals 201 into encoded bitstreams, as discussed with respect to steps 105, 107, and 109 in the method 100. When acting as a decoder, the codec system 200 generates an output video signal from the bitstream, as described in conjunction with steps 111, 113, 115, and 117 in the operating method 100. The codec system 200 includes a general decoder control component 211, a transform scaling quantization component 213, an intra-frame estimation component 215, an intra-frame prediction component 217, a motion compensation component 219, a motion estimation component 221, a scaling and inverse transform component 229, a filter control analysis component 227, an in-loop filter component 225, a decoded image buffer component 223, and a header format and context adaptive binary arithmetic coding (CABAC) component 231. These components are coupled as shown. Figure 2 In the figure, the black lines represent the movement of the data to be encoded / decoded, and the dotted lines represent the movement of the control data that controls the operation of other components. The components in the codec system 200 can all be present in the encoder. The decoder can include a subset of the components in the codec system 200. For example, the decoder can include an intra-frame prediction component 217, a motion compensation component 219, a scaling and inverse transform component 229, an in-loop filter component 225, and a decoded image buffer component 223. These components are now described.

[0061] The segmented video signal 201 is a captured video sequence that has been segmented into pixel blocks by a coding tree. The coding tree uses various partitioning modes to subdivide the pixel blocks into smaller pixel blocks. These blocks can then be further subdivided into smaller blocks. The blocks can be referred to as nodes on the coding tree. A larger parent node is divided into smaller child nodes. The number of times a node is subdivided is called the depth of the node / coding tree. In some cases, the divided blocks may be included in a coding unit (coding unit, CU). For example, a CU may be a sub-part of a CTU, including a luminance block, a red difference chrominance (Cr) block, and a blue difference chrominance (Cb) block and corresponding syntax instructions for the CU. The partitioning mode may include a binary tree (BT), a triple tree (TT), and a quad tree (QT) for dividing a node into two, three, or four child nodes of different shapes, respectively, depending on the partitioning mode used. The segmented video signal 201 is forwarded to the general decoder control component 211, the transform scaling and quantization component 213, the intra-frame estimation component 215, the filter control analysis component 227 and the motion estimation component 221 for compression.

[0062] The universal decoder control component 211 is used to make decisions related to encoding the images of the video sequence into the bitstream according to the application constraints. For example, the universal decoder control component 211 manages the optimization of the bitrate / bitstream size relative to the reconstruction quality. These decisions can be made based on the storage space / bandwidth availability and the image resolution request. The universal decoder control component 211 also manages the utilization of the buffer according to the transmission speed to alleviate the buffer under-load and over-load problems. In order to manage these problems, the universal decoder control component 211 manages the segmentation, prediction and filtering performed by other components. For example, the universal decoder control component 211 can dynamically increase the compression complexity to increase the resolution and bandwidth utilization, or reduce the compression complexity to reduce the resolution and bandwidth utilization. Therefore, the universal decoder control component 211 controls other components of the codec system 200 to balance the video signal reconstruction quality and bitrate issues. The universal decoder control component 211 creates control data to control the operation of other components. The control data is also forwarded to the header format and CABAC component 231 to be encoded in the bitstream, thereby indicating the parameters decoded in the decoder.

[0063] The segmented video signal 201 is also sent to the motion estimation component 221 and the motion compensation component 219 for inter-frame prediction. The frames or slices of the segmented video signal 201 can be divided into multiple video blocks. The motion estimation component 221 and the motion compensation component 219 perform inter-frame prediction decoding on the received video blocks according to one or more blocks in one or more reference frames to provide temporal prediction. The codec system 200 can perform multiple decoding processes to select an appropriate decoding mode for each video data block, etc.

[0064] The motion estimation component 221 and the motion compensation component 219 can be highly integrated, but are described separately for conceptual purposes. The motion estimation performed by the motion estimation component 221 is the process of generating motion vectors, which are used to estimate the motion of video blocks. For example, a motion vector may indicate the displacement of a coding object relative to a prediction block. A prediction block is a block that is found to closely match the block to be coded in terms of pixel difference. A prediction block may also be referred to as a reference block. This pixel difference may be determined by the sum of absolute difference (SAD), the sum of square difference (SSD), or other difference metrics. HEVC uses several coding objects, including CTUs, coding tree blocks (CTBs), and CUs. For example, a CTU may be divided into multiple CTBs, which may then be divided into multiple CBs and included in a CU. A CU may be encoded as a prediction unit (PU) including prediction data and / or a transform unit (TU) including transform residual data of a CU. The motion estimation component 221 generates motion vectors, PUs, and TUs using rate-distortion analysis as part of a rate-distortion optimization process. For example, the motion estimation component 221 may determine multiple reference blocks, multiple motion vectors, etc. for the current block / frame, and may select a reference block, motion vector, etc. having an optimal rate-distortion characteristic. The optimal rate-distortion characteristic balances the quality of video reconstruction (e.g., the amount of data lost due to compression) and decoding efficiency (e.g., the size of the final encoding).

[0065] In some examples, the codec system 200 may calculate values ​​for sub-integer pixel positions of a reference image stored in the decoded image buffer component 223. For example, the video codec system 200 may interpolate values ​​for quarter-pixel positions, eighth-pixel positions, or other fractional pixel positions of the reference image. Thus, the motion estimation component 221 may perform motion searches with respect to integer pixel positions and fractional pixel positions, and output motion vectors with fractional pixel precision. The motion estimation component 221 calculates the motion vector of a PU of a video block in an inter-coded slice by comparing the position of the PU with the position of a prediction block of the reference image. The motion estimation component 221 outputs the calculated motion vector as motion data to the header format and CABAC component 231 for encoding, and outputs the motion to the motion compensation component 219.

[0066] The motion compensation performed by the motion compensation component 219 may involve obtaining or generating a prediction block based on the motion vector determined by the motion estimation component 221. Likewise, in some examples, the motion estimation component 221 and the motion compensation component 219 may be functionally integrated. After receiving the motion vector of the PU of the current video block, the motion compensation component 219 may locate the prediction block to which the motion vector points. Then, by subtracting the pixel values ​​of the prediction block from the pixel values ​​of the current video block being encoded, a pixel difference value is generated, thereby forming a residual video block. Typically, the motion estimation component 221 performs motion estimation on the luma component, and the motion compensation component 219 uses the motion vector calculated based on the luma component for the chroma component and the luma component. The prediction block and the residual block are forwarded to the transform scaling and quantization component 213.

[0067] The segmented video signal 201 is also sent to the intra-frame estimation component 215 and the intra-frame prediction component 217. Like the motion estimation component 221 and the motion compensation component 219, the intra-frame estimation component 215 and the intra-frame prediction component 217 can be highly integrated, but are described separately for conceptual purposes. The intra-frame estimation component 215 and the intra-frame prediction component 217 perform intra-frame prediction on the current block based on the block in the current frame, replacing the inter-frame prediction performed between frames by the motion estimation component 221 and the motion compensation component 219 as described above. Specifically, the intra-frame estimation component 215 determines the intra-frame prediction mode for encoding the current block. In some examples, the intra-frame estimation component 215 selects an appropriate intra-frame prediction mode from a plurality of tested intra-frame prediction modes to encode the current block. The selected intra-frame prediction mode is then forwarded to the header format and CABAC component 231 for encoding.

[0068] For example, the intra-frame estimation component 215 uses rate-distortion analysis of various tested intra-frame prediction modes to calculate rate-distortion values ​​and selects the intra-frame prediction mode with the best rate-distortion characteristics among the tested modes. Rate-distortion analysis generally determines the amount of distortion (or error) between the coded block and the original uncoded block encoded to generate the coded block and the code rate (e.g., number of bits) used to generate the coded block. The intra-frame estimation component 215 determines which intra-frame prediction mode obtains the best rate-distortion value for the block based on the distortion and rate calculation ratios of the various coded blocks. In addition, the intra-frame estimation component 215 can be used to encode the depth block of the depth map using a depth modeling mode (DMM) according to rate-distortion optimization (RDO).

[0069] When implemented on an encoder, the intra prediction component 217 may generate a residual block from the prediction block according to the selected intra prediction mode determined by the intra estimation component 215, or read the residual block from the bitstream when implemented on a decoder. The residual block includes the difference in values ​​between the prediction block and the original block, represented as a matrix. The residual block is then forwarded to the transform scaling and quantization component 213. The intra estimation component 215 and the intra prediction component 217 may perform operations on the luma component and the chroma component.

[0070] The transform scaling and quantization component 213 is used to further compress the residual block. The transform scaling and quantization component 213 applies a transform such as a discrete cosine transform (DCT), a discrete sine transform (DST), or a conceptually similar transform to the residual block to produce a video block including residual transform coefficient values. Wavelet transforms, integer transforms, subband transforms, or other types of transforms may also be used. The transform may transform the residual information from a pixel value domain to a transform domain, such as a frequency domain. The transform scaling and quantization component 213 is also used to scale the transformed residual information according to frequency, etc. This scaling involves applying a scaling factor to the residual information so as to quantize different frequency information at different granularities, which may affect the final visual quality of the reconstructed video. The transform scaling and quantization component 213 is also used to quantize the transform coefficients to further reduce the bit rate. The quantization process may reduce the bit depth associated with some or all of the coefficients. The degree of quantization may be modified by adjusting the quantization parameter. In some examples, the transform scaling and quantization component 213 may then scan the matrix including the quantized transform coefficients. The quantized transform coefficients are forwarded to the header format and CABAC component 231 for encoding into the codestream.

[0071] The scaling and inverse transform component 229 performs the inverse operations of the transform scaling and quantization component 213 to support motion estimation. The scaling and inverse transform component 229 performs inverse scaling, inverse transform and / or inverse quantization to reconstruct a residual block in the pixel domain, for example, for subsequent use as a reference block, which may become a prediction block for another current block. The motion estimation component 221 and / or the motion compensation component 219 may calculate a reference block by adding the residual block to the corresponding prediction block for motion estimation of subsequent blocks / frames. A filter is applied to the reconstructed reference block to reduce artifacts generated during the scaling, quantization and transform processes. These artifacts may produce inaccurate predictions (and generate other artifacts) when predicting subsequent blocks.

[0072] The filter control analysis component 227 and the in-loop filter component 225 apply filters to the residual block and / or the reconstructed image block. For example, the transformed residual block in the scaling and inverse transform component 229 can be combined with the corresponding prediction block in the intra-frame prediction component 217 and / or the motion compensation component 219 to reconstruct the original image block. Then, the filter can be applied to the reconstructed image block. In some examples, the filter can be applied to the residual block. Figure 2 Like other components in the, the filter control analysis component 227 and the in-loop filter component 225 are highly integrated and can be implemented together, but are described separately for conceptual purposes. The filters applied to the reconstructed reference block are applied to specific spatial regions and include multiple parameters to adjust how these filters are applied. The filter control analysis component 227 analyzes the reconstructed reference block to determine where these filters should be applied and sets the corresponding parameters. These data are forwarded to the header format and CABAC component 231 for encoding as filter control data. The in-loop filter component 225 applies these filters according to the filter control data. The filters may include deblocking filters, noise suppression filters, SAO filters, and adaptive loop filters. These filters can be applied in the spatial / pixel domain (e.g., on the reconstructed pixel block) or in the frequency domain according to the example.

[0073] When operating as an encoder, the filtered reconstructed image blocks, residual blocks and / or prediction blocks are stored in the decoded image buffer component 223 for later motion estimation as described above. When operating as a decoder, the decoded image buffer component 223 stores the reconstructed blocks and the filtered blocks and forwards the reconstructed blocks and the filtered blocks to the display as part of the output video signal. The decoded image buffer component 223 can be any memory device capable of storing prediction blocks, residual blocks and / or reconstructed image blocks.

[0074] The header format and CABAC component 231 receives data from various components of the codec system 200 and encodes the data into an encoded bitstream for sending to the decoder. Specifically, the header format and CABAC component 231 generates various headers to encode control data such as overall control data and filter control data. In addition, prediction data including intra-frame prediction and motion data and residual data in the form of quantized transform coefficient data are encoded into the bitstream. The final bitstream includes all the information that the decoder wants to reconstruct the original segmented video signal 201. This information can also include an intra-frame prediction mode index table (also called a codeword mapping table), a definition of the coding context of various blocks, an indication of the most likely intra-frame prediction mode, an indication of segmentation information, etc. This data can be encoded by entropy decoding techniques. For example, the information may be encoded using context adaptive variable length coding (CAVLC), CABAC, syntax-based context-adaptive binary arithmetic coding (SBAC), probability interval partitioning entropy (PIPE) coding or other entropy coding techniques. After entropy coding, the encoded bitstream may be used to send to another device (e.g., a video decoder) or archived for later transmission or retrieval.

[0075] Figure 3 is a block diagram of an exemplary video encoder 300. The video encoder 300 may be used to implement the encoding function of the codec system 200 and / or to implement steps 101, 103, 105, 107, and / or 109 of the operating method 100. The encoder 300 segments the input video signal to generate segmented video signals 301 that are substantially similar to the segmented video signal 201. Then, the segmented video signals 301 are compressed and encoded into a bitstream by components of the encoder 300.

[0076] Specifically, the segmented video signal 301 is forwarded to the intra prediction component 317 for intra prediction. The intra prediction component 317 may be substantially similar to the intra estimation component 215 and the intra prediction component 217. The segmented video signal 301 is also forwarded to the motion compensation component 321 for inter prediction based on the reference block in the decoded image buffer 323. The motion compensation component 321 may be substantially similar to the motion estimation component 221 and the motion compensation component 219. The prediction blocks and residual blocks in the intra prediction component 317 and the motion compensation component 321 are forwarded to the transform and quantization component 313 for transforming and quantizing the residual blocks. The transform and quantization component 313 may be substantially similar to the transform scaling and quantization component 213. The transformed and quantized residual blocks and the corresponding prediction blocks (and related control data) are forwarded to the entropy coding component 331 for encoding into the bitstream. The entropy coding component 331 may be substantially similar to the header format and CABAC component 231.

[0077] The transformed and quantized residual block and / or the corresponding prediction block are also forwarded from the transform and quantization component 313 to the inverse transform and quantization component 329 to be reconstructed as a reference block for use by the motion compensation component 321. The inverse transform and quantization component 329 may be substantially similar to the scaling and inverse transform component 229. According to an example, the in-loop filter in the in-loop filter component 325 is also applied to the residual block and / or the reconstructed reference block. The in-loop filter component 325 may be substantially similar to the filter control analysis component 227 and the in-loop filter component 225. As discussed with respect to the in-loop filter component 225, the in-loop filter component 325 may include a plurality of filters. The filtered block is then stored in the decoded image buffer component 323 for use as a reference block by the motion compensation component 321. The decoded image buffer component 323 may be substantially similar to the decoded image buffer component 223.

[0078] Figure 4 is a block diagram of an exemplary video decoder 400. The video decoder 400 may be used to implement the decoding function of the codec system 200 and / or implement steps 111, 113, 115, and / or 117 of the operation method 100. For example, the decoder 400 receives a bitstream from the encoder 300, and generates a reconstructed output video signal based on the bitstream for display to an end user.

[0079] The code stream is received by an entropy decoding component 433. Entropy decoding component 433 is used to implement an entropy decoding scheme, such as CAVLC, CABAC, SBAC, PIPE decoding or other entropy decoding techniques. For example, entropy decoding component 433 can use header information to provide context to interpret other data encoded as code words in the code stream. The decoded information includes any information required to decode the video signal, such as overall control data, filter control data, segmentation information, motion data, prediction data, and quantized transform coefficients of residual blocks. The quantized transform coefficients are forwarded to the inverse transform and quantization component 429 to be reconstructed as residual blocks. The inverse transform and quantization component 429 can be substantially similar to the inverse transform and quantization component 329.

[0080] The reconstructed residual block and / or prediction block are forwarded to the intra prediction component 417 to be reconstructed into an image block according to the intra prediction operation. The intra prediction component 417 may be similar to the intra estimation component 215 and the intra prediction component 217. Specifically, the intra prediction component 417 uses the prediction mode to locate the reference block in the frame and applies the residual block to the result to reconstruct the intra prediction image block. The reconstructed intra prediction image block and / or residual block and the corresponding inter prediction data are forwarded to the decoded image buffer component 423 through the in-loop filter component 425, and the decoded image buffer component 423 and the in-loop filter component 425 may be substantially similar to the decoded picture buffer component 223 and the in-loop filter component 225, respectively. The in-loop filter component 425 filters the reconstructed image block, residual block and / or prediction block, and this information is stored in the decoded image buffer component 423. The reconstructed image block in the decoded image buffer component 423 is forwarded to the motion compensation component 421 for inter prediction. The motion compensation component 421 may be substantially similar to the motion estimation component 221 and / or the motion compensation component 219. Specifically, the motion compensation component 421 uses the motion vector in the reference block to generate a prediction block and applies the residual block to the result to reconstruct the image block. The resulting reconstructed block may also be forwarded to the decoded image buffer component 423 via the in-loop filter component 425. The decoded image buffer component 423 continues to store other reconstructed image blocks, which may be reconstructed into frames using the segmentation information. These frames may also be arranged in sequence. The sequence is output to the display screen as a reconstructed output video signal.

[0081] Figure 5 5 is a schematic diagram of an exemplary mechanism of WPP 500. For example, WPP 500 may encode and / or decode slice 501 as part of method 100. Therefore, WPP 500 may be used by codec system 200, encoder 300 and / or decoder 400.

[0082] As shown, WPP 500 is applied to slice 501. Slice 501 is a partition of an image. Specifically, an image can be partitioned into one or more slices 501. Slice 501 may include an integer number of consecutive complete CTU rows 521, 522, 523, 524, and 525. In addition, slice 501 and all sub-divisions (e.g., CTU 511 and 512) are included in only a single NAL unit. WPP 500 can be applied when slice 501 includes an integer number of consecutive complete CTU rows 521, 522, 523, 524, and 525. Optionally, in some examples, slice 501 may include one or more partitions, and these partitions may include an integer number of consecutive complete CTU rows 521, 522, 523, 524, and 525, respectively. Therefore, a slice 501 is defined in the VVC standard as an integer number of complete partitions in a picture or an integer number of consecutive complete CTU rows 521 to 525 within a partition, and these partitions or CTU rows are included in only a single NAL unit.

[0083] The slice 501 includes CTUs 511 and 512. CTUs 511 / 512 are a group of samples of a predefined size that can be divided into coding blocks by a coding tree. For example, CTUs 511 and 512 can be arranged in CTU rows 521 to 525 and CTU columns 516. CTU rows 521 to 525 are a group of CTUs 511 / 512 extending horizontally between the left boundary of the slice 501 and the right boundary of the slice 501. CTU columns 516 are a group of CTUs 511 / 512 extending vertically between the upper boundary of the slice 501 and the lower boundary of the slice 501.

[0084] WPP 500 may employ multiple computing threads operating in parallel to decode CTU 511 / 512. In the example shown, CTU 511 has been decoded and CTU 512 has not yet been decoded. For example, a first thread may begin decoding CTU row 521 for the first time. Once CTU 511 has been decoded in the first CTU row 521, a second thread may begin decoding CTU row 522. Once CTU 511 has been decoded in the second CTU row 522, a third thread may begin decoding CTU row 523. Once CTU 511 has been decoded in the third CTU row 523, a fourth thread may begin decoding CTU row 524. Once CTU 511 has been decoded in the fourth CTU row 524, a fifth thread may begin decoding the fifth CTU row 525. This produces Figure 5The pattern shown. Other threads can be used as needed. This mechanism creates a pattern with a wavefront appearance, so it is called WPP 500. Some video decoding mechanisms decode the current CTU 512 based on the decoded CTU 511 located above or to the left of the current CTU 512. WPP 500 leaves a CTU 511 decoding delay between starting each thread to ensure that these CTUs 511 have been decoded when they reach any current CTU 512 to be decoded. It should be noted that systems using the VVC standard can use a decoding delay of one CTU 511, while other systems such as HEVC can use a decoding delay of two CTUs 511. Any other CTU 511 decoding delay value can also be used within the scope of the present invention.

[0085] CTU 511 is decoded into the codestream according to CTU rows 521 to 525. For example, CTU row 521 may be included in the codestream, CTU row 521 is followed by CTU row 522, CTU row 522 is followed by CTU row 523, CTU row 523 is followed by CTU row 524, CTU row 524 is followed by CTU row 525, and so on. Each CTU row 521 to 525 may be an independently addressable subset of the slice 501 in the codestream. For example, each CTU row 521 to 525 may be addressed at entry point 517. Entry point 517 is a bit position in the codestream that includes the first bit of the video data of the corresponding subset of the slice 501 after the slice 501 is encoded. When WPP 500 is adopted, entry point 517 is a bit position that includes the first bit of the corresponding CTU row 521 to 525. For example, entry point 517 of CTU row 523 includes the first bit of video data of CTU row 523 and is located after the last bit of video data of CTU row 522. Therefore, number of entry points (NumEntryPoints) 518 is the number of entry points 517 of CTU rows 521-525.

[0086] In some video coding systems, the bit offsets of the entry points 517 may be indicated in an array in a slice header associated with the slice 501. In addition, the encoder may determine NumEntryPoints 518 and indicate this in the slice header, for example in a num_entry_point_offsets parameter. The decoder may then use the number of entry points 517 to determine how many bit offsets should be obtained from the array to obtain all relevant entry points 517 to start decoding. The num_entry_point_offsets parameter includes a number of bits and is indicated in each slice header. Since each picture may include several slices 501 and a video sequence may include thousands of pictures, num_entry_point_offsets may be indicated multiple times and may have a significant aggregate impact on the amount of data in the video sequence.

[0087] The present invention includes various mechanisms for deriving NumEntryPoints 518 at the decoder side (and / or at the HRD on the encoder side) without indicating such data in the codestream. Thus, num_entry_point_offsets can be omitted from the codestream, significantly improving overall codestream compression. For example, when slice 501 is decoded according to WPP 500, these mechanisms can be used to derive NumEntryPoints 518 for slice 501. NumEntryPoints 518 can be derived based on the size of CTU rows 521 to 525, the size of CTU columns 516, the number of CTUs 511 / 512 in slice 501, and / or the addresses of CTUs 511 / 512 and / or slice 501.

[0088] For example, NumEntryPoints 518 may be derived as follows:

[0089]

[0090] Wherein, !entropy_coding_sync_enabled_flag indicates that WPP 500 is not in use, NumBricksInCurrSlice is the number of blocks in slice 501, numBrickSpecificCtuRowsInSlice is the number of CTU rows 521 to 525 in slice 501, and BrickHeight[SliceBrickIdx[i]] is the height of CTU rows 521 to 525 (and therefore also the height of CTU column 516). It should be noted that entropy_coding_sync_enabled_flag can be used to indicate whether WPP 500 is used to decode slice 501. For example, when WPP 500 is used, entropy_coding_sync_enabled_flag can be set to 1, and when WPP 500 is not used, entropy_coding_sync_enabled_flag can be set to 0. In some decoding systems, ! represents a not value. Therefore, ! entropy_coding_sync_enabled_flag indicates that WPP 500 is not used in the if clause of the above pseudo code. Therefore, WPP is used in the else clause of the above pseudo code.

[0091] The example checks whether slice 501 uses WPP 500. If WPP 500 is not used, NumEntryPoints 518 is one less than the number of tiles in slice 501. Otherwise, WPP 500 is used. In this case, the example determines the number of CTU rows 521-525 by adding the height of each of CTU rows 521-525 to the height of slice 501 (and therefore the height of CTU column 516). NumEntryPoints 518 is one less than the number of CTU rows 521-525.

[0092] Figure 6 is a schematic diagram of an exemplary code stream 600. For example, the code stream 600 may be generated by the codec system 200 and / or the encoder 300, and decoded by the codec system 200 and / or the decoder 400 according to the method 100. In addition, the code stream 600 may include a slice 501 decoded according to the WPP 500.

[0093] The code stream 600 includes an SPS 610, a plurality of picture parameter sets (PPS) 611, a plurality of slice headers 615, and image data 620. The SPS 610 includes sequence data common to all images in the coded video sequence included in the code stream 600. These data may include image size, bit depth, decoding tool parameters, bit rate limits, etc. The PPS 611 includes parameters applied to the entire image. Therefore, each image in the video sequence may refer to the PPS 611. It should be noted that although each image refers to the PPS 611, in some examples, a single PPS 611 may include data for multiple images. For example, multiple similar images may be encoded according to similar parameters. In this case, a single PPS 611 may include data for such similar images. The PPS 611 may indicate decoding tools, quantization parameters, offsets, etc. that may be used for the slices in the corresponding image. The slice header 615 includes parameters specific to each slice in the image. Therefore, each slice in the video sequence may have a slice header 615. The slice header 615 may include slice type information, POC, reference picture list, prediction weight, partition entry point, deblocking filter parameters, etc. It should be noted that in some contexts, the slice header 615 may also be referred to as a partition group header.

[0094] Image data 620 includes video data encoded according to inter-frame prediction and / or intra-frame prediction and corresponding transformed and quantized residual data. For example, a video sequence includes multiple images 621. Image 621 is a complete image that is expected to be fully or partially displayed to the user at a corresponding moment in the video sequence. Image 621 may be included in a single access unit (access unit, AU). Image 621 includes one or more slices 623. Slice 623 may be defined as an integer number of complete blocks in image 621 or an integer number of continuous complete CTU rows (for example, within a block), which are included only in a single NAL unit. For example, slice 623 may be substantially similar to slice 501. Slice 623 is also divided into CTU 625 and / or coding tree block (coding tree block, CTB). CTU 625 is a set of samples of a predefined size that can be partitioned by a coding tree. For example, CTU 625 in code stream 600 has been encoded and can therefore be substantially similar to CTU511. CTB is a subset of CTU 625, including the luma component or chroma component of CTU. According to the coding tree, CTU 625 / CTB is further divided into coding blocks. The coding blocks can then be encoded / decoded according to the prediction mechanism.

[0095] As described above, NumEntryPoints 518 can be determined on the decoder side. Therefore, the slice header 615 does not include a num_entry_point_offsets parameter or other parameters indicating the value of NumEntryPoints 518. When num_entry_point_offsets is deleted from each slice header 615, the codestream 600 is compressed. This reduces the utilization of memory resources and network resources on the encoder and decoder sides.

[0096] The above information is described in more detail below. In the video codec specification, images need to be identified for a variety of purposes. These uses include use as reference images in inter-frame prediction, for outputting images from the DPB, for scaling of motion vectors, for weighted prediction, etc. For example, images can be identified by picture order counts (POC). In some video decoding systems, images in the DPB can be marked as "for short-term reference", "for long-term reference" or "not for reference". Once an image is marked as not for reference, the image can no longer be used for prediction. When it is no longer necessary to output such an image, the image can be deleted from the DPB.

[0097] Some video decoding systems use short-term and long-term reference images. When an image is no longer needed for prediction reference, the reference image can be marked as not used for reference. The transition between these three states (short-term reference, long-term reference, not used for reference) of the image is controlled by the decoded reference image marking process. For example, an implicit sliding window process and / or an explicit memory management control operation (MMCO) process can be used to mark reference images. When the number of reference frames is equal to the maximum number of reference frames that can be stored in the SPS (max_num_ref_frames), the sliding window process marks the short-term reference image as not used for reference. Short-term reference images are stored in a first-in-first-out manner so that the most recently decoded short-term image is saved in the DPB. The explicit MMCO process can include multiple MMCO commands. The MMCO command can mark one or more short-term or long-term reference images as "not used for reference", mark all images as "not used for reference", or mark the current reference image or an existing short-term reference image as "long-term reference", and assign a long-term image index to the long-term reference image. In some video coding systems, the reference picture marking operation and the process of outputting and deleting pictures from the DPB are performed after the pictures are decoded.

[0098] Other video decoding systems use reference picture sets (RPS) for reference picture management. One difference between the RPS process and the MMCO / sliding window process is that each slice is provided with a complete set of reference pictures for use by the current picture or any subsequent picture. Therefore, a complete set of all pictures that should be retained in the DPB for use by the current or future pictures is indicated. This is different from the scheme that only indicates relative changes to the DPB. The RPS mechanism does not need to obtain information from earlier pictures in the decoding order to maintain the correct state of the reference pictures in the DPB. In order to give full play to the advantages of RPS and improve the error resistance, the image decoding and DPB operations have been modified accordingly. In some video decoding systems, picture marking and buffering operations, including outputting and deleting decoded pictures from the DPB, are usually applied after decoding the current picture. In other video decoding systems, RPS first decodes from the slice header of the current picture. Then, picture marking and buffering operations are usually applied before decoding the current picture.

[0099] Other video decoding systems manage reference pictures according to two reference picture lists denoted as reference picture list 0 and reference picture list 1. In this method, the reference picture list of a picture can be directly constructed without using the reference picture list initialization process and the reference picture list modification process. In addition, reference picture marking is performed directly according to the two reference picture lists. Examples of syntax and semantics related to reference picture management are as follows:

[0100] The syntax of the sequence parameter set RBSP is as follows:

[0101] seq_parameter_set_rbsp(){ Descriptors ... log2_max_pic_order_cnt_lsb_minus4 ue(v) sps_max_dec_pic_buffering_minus1 ue(v) long_term_ref_pics_flag u(1) sps_idr_rpl_present_flag u(1) rpl1_same_as_rpl0_flag u(1) for(i=0;i<!rpl1_same_as_rpl0_flag?2:1;i++){ num_ref_pic_lists_in_sps[i] ue(v) for(j=0;j<num_ref_pic_lists_in_sps[i];j++) ref_pic_list_struct(i,j) } ...

[0102] The following is an example of the syntax of the picture parameter set RBSP:

[0103] pic_parameter_set_rbsp(){ Descriptors ... for(i=0;i<2;i++) num_ref_idx_default_active_minus1[i] ue(v) rpl1_idx_present_flag u(1) ...

[0104] An example of a common stripe header syntax is as follows:

[0105]

[0106]

[0107] An example of the reference picture list structure syntax is as follows:

[0108]

[0109]

[0110] The semantics of the sequence parameter set RBSP are as follows: log2_max_pic_order_cnt_lsb_minus4 represents the value of the variable MaxPicOrderCntLsb used for the picture order numbering during decoding:

[0111] MaxPicOrderCntLsb=2 (log2 _ max _ pic _ order _ cnt _ lsb _ minus4+4) (7-7)

[0112] The value range of log2_max_pic_order_cnt_lsb_minus4 can be 0 to 12 (inclusive). sps_max_dec_pic_buffering_minus1+1 indicates the maximum size of the decoded picture buffer required for the CVS, in units of picture storage buffers. sps_max_dec_pic_buffering_minus1 can range from 0 to MaxDpbSize–1 (inclusive), where MaxDpbSize is the same as the value specified elsewhere. long_term_ref_pics_flag can be set to 0 to indicate that long-term reference pictures (LTRPs) are not used for inter-frame prediction of any decoded pictures in the CVS. long_term_ref_pics_flag can be set to 1 to indicate that the LTRP can be used for inter-frame prediction of one or more decoded pictures in the CVS. sps_idr_rpl_present_flag can be set to 1 to indicate that the reference picture list syntax element is present in the slice header of the IDR picture. sps_idr_rpl_present_flag may be set to 0, indicating that the reference picture list syntax element is not present in the slice header of the IDR picture.

[0113] rpl1_same_as_rpl0_flag may be set to 1 to indicate that the syntax structures num_ref_pic_lists_in_sps[1] and ref_pic_list_struct(1,rplsIdx) do not exist and the following apply: The value of num_ref_pic_lists_in_sps[1] is inferred to be equal to the value of num_ref_pic_lists_in_sps[0]. The value of each syntax element in ref_pic_list_struct(1,rplsIdx) is inferred to be equal to the value of the corresponding syntax element in ref_pic_list_struct(0,rplsIdx), where rplsIdx ranges from 0 to num_ref_pic_lists_in_sps[0]–1. num_ref_pic_lists_in_sps[i] represents the number of ref_pic_list_struct(listIdx,rplsIdx) syntax structures included in the SPS with listIdx equal to i. The value of num_ref_pic_lists_in_sps[i] can range from 0 to 64 (including the end value). For each listIdx value (equal to 0 or 1), the decoder should allocate memory for num_ref_pic_lists_in_sps[i]+1 ref_pic_list_struct(listIdx,rplsIdx) syntax structures, because a ref_pic_list_struct(listIdx,rplsIdx) syntax structure can be directly indicated in the slice header of the current picture.

[0114] The semantics of picture parameter set RBSP are as follows: num_ref_idx_default_active_minus1[i]+1, when i=0, indicates the inferred value of variable NumRefIdxActive[0] for P or B slice with num_ref_idx_active_override_flag=0, when i=1, indicates the inferred value of NumRefIdxActive[1] for B slice with num_ref_idx_active_override_flag=0. The value range of num_ref_idx_default_active_minus1[i] shall be 0 to 14 (inclusive). rpl1_idx_present_flag may be set to 0, indicating that ref_pic_list_sps_flag[1] and ref_pic_list_idx[1] are not present in the slice header. rpl1_idx_present_flag may be set to 1, indicating that ref_pic_list_sps_flag[1] and ref_pic_list_idx[1] may be present in the slice header.

[0115] Examples of generic slice header semantics are as follows: slice_pic_order_cnt_lsb represents the picture order number of the current picture modulo MaxPicOrderCntLsb. The length of the slice_pic_order_cnt_lsb syntax element is (log2_max_pic_order_cnt_lsb_minus4+4) bits. The value range of slice_pic_order_cnt_lsb is 0 to MaxPicOrderCntLsb–1 (inclusive). ref_pic_list_sps_flag[i] can be set to 1 to indicate that the reference picture list i of the current slice is derived from one of the syntax structures ref_pic_list_struct(listIdx, rplsIdx) with listIdx equal to i in the active SPS. ref_pic_list_sps_flag[i] may be set to 0, indicating that the reference picture list i of the current slice is derived from the syntax structure ref_pic_list_struct(listIdx, rplsIdx) with listIdx equal to i, where the syntax structure ref_pic_list_struct(listIdx, rplsIdx) is directly included in the slice header of the current picture. When num_ref_pic_lists_in_sps[i] is equal to 0, the value of ref_pic_list_sps_flag[i] is inferred to be equal to 0. When rpl1_idx_present_flag is equal to 0, the value of ref_pic_list_sps_flag[1] is inferred to be equal to the value of ref_pic_list_sps_flag[0].

[0116] ref_pic_list_idx[i] represents the index in the list of ref_pic_list_struct(listIdx,rplsIdx,ltrpFlag) syntax structures with listIdx equal to i included in the active SPS, of the ref_pic_list_struct(listIdx,rplsIdx,) syntax structures for deriving reference picture list i for the current picture with listIdx equal to i. The syntax element ref_pic_list_idx[i] is represented by Ceil(Log2(num_ref_pic_lists_in_sps[i])) bits. If ref_pic_list_idx[i] does not exist, the value of ref_pic_list_idx[i] is inferred to be equal to 0. The value range of ref_pic_list_idx[i] is 0 to num_ref_pic_lists_in_sps[i]–1 (inclusive). When ref_pic_list_sps_flag[i] is equal to 1 and num_ref_pic_lists_in_sps[i] is equal to 1, then the value of ref_pic_list_idx[i] is inferred to be equal to 0. When ref_pic_list_sps_flag[i] is equal to 1 and rpl1_idx_present_flag is equal to 0, then the value of ref_pic_list_idx[1] is inferred to be equal to the value of ref_pic_list_idx[0].

[0117] The variable RplsIdx[i] is derived as follows:

[0118] RplsIdx[i]=ref_pic_list_sps_flag[i]? ref_pic_list_idx[i]:

[0119] num_ref_pic_lists_in_sps[i](7-40)

[0120] slice_poc_lsb_lt[i][j] represents the value modulo MaxPicOrderCntLsb of the picture order number of the j-th LTRP entry in the i-th reference picture list. The length of the syntax element slice_poc_lsb_lt[i][j] is log2_max_pic_order_cnt_lsb_minus4+4 bits.

[0121] The variable PocLsbLt[i][j] is derived as follows:

[0122] PocLsbLt[i][j]=ltrp_in_slice_header_flag[i][RplsIdx[i]]? (7-41)

[0123] slice_poc_lsb_lt[i][j]:rpls_poc_lsb_lt[listIdx][RplsIdx[i]][j]

[0124] delta_poc_msb_present_flag[i][j] may be set to 1 to indicate the presence of delta_poc_msb_cycle_lt[i][j]. delta_poc_msb_present_flag[i][j] may be set to 0 to indicate the absence of delta_poc_msb_cycle_lt[i][j]. Assume that prevTid0Pic is the previous picture in decoding order whose TemporalId is equal to 0 and is not a skipped random access skipped leading (RASL) picture or a decodable random access decodable leading (RADL) picture. Assume that setOfPrevPocVals is the set consisting of: the PicOrderCntVal of prevTid0Pic; the PicOrderCntVal of each picture referenced by an entry in RefPicList[0] of prevTid0Pic and an entry in RefPicList[1]; the PicOrderCntVal of each picture that follows prevTid0Pic in decoding order and precedes the current picture in decoding order. When there are multiple values ​​in setOfPrevPocVals whose values ​​are equal to PocLsbLt[i][j] modulo MaxPicOrderCntLsb, the value of delta_poc_msb_present_flag[i][j] shall be equal to 1.

[0125] delta_poc_msb_cycle_lt[i][j] indicates the value of the variable FullPocLt[i][j] as follows:

[0126]

[0127] delta_poc_msb_cycle_lt[i][j] shall be in the range of 0 to 2(32–log2_max_pic_order_cnt_lsb_minus4–4), inclusive. If delta_poc_msb_cycle_lt[i][j] is not present, delta_poc_msb_cycle_lt[i][j] is inferred to be equal to 0. num_ref_idx_active_override_flag may be set to 1 to indicate that the syntax element num_ref_idx_active_minus1[0] is present in P and B slices, and the syntax element num_ref_idx_active_minus1[1] is present in B slices. num_ref_idx_active_override_flag may be set to 0 to indicate that the syntax elements num_ref_idx_active_minus1[0] and num_ref_idx_active_minus1[1] are not present. If num_ref_idx_active_override_flag is not present, the value of num_ref_idx_active_override_flag is inferred to be 1. num_ref_idx_active_minus1[i] is used to derive the variable NumRefIdxActive[i]. num_ref_idx_active_minus1[i] may range from 0 to 14 (inclusive). For i equal to 0 or 1, when the current slice is a B slice, num_ref_idx_active_override_flag is 1, and num_ref_idx_active_minus1[i] is not present, num_ref_idx_active_minus1[i] is inferred to be 0. When the current slice is a P slice, num_ref_idx_active_override_flag is 1, and num_ref_idx_active_minus1[0] is not present, num_ref_idx_active_minus1[0] is inferred to be 0.

[0128] The variable NumRefIdxActive[i] is derived as follows:

[0129]

[0130]

[0131] The value of NumRefIdxActive[i]–1 indicates the maximum reference index of reference picture list i that can be used to decode the slice. When the value of NumRefIdxActive[i] is equal to 0, the reference index of reference picture list i can be used to decode the slice. The variable CurrPicIsOnlyRef indicates that the current decoded picture is the only reference picture of the current slice, which is derived as follows:

[0132] CurrPicIsOnlyRef=sps_cpr_enabled_flag&&(slice_type==P)&&(7-44)

[0133] (num_ref_idx_active_minus1[0]==0)

[0134] The semantics of the reference picture list structure are exemplified as follows: The ref_pic_list_struct(listIdx, rplsIdx) syntax structure can exist in the SPS or in the slice header. Depending on whether the syntax structure is included in the slice header or the SPS, the following applies. If the syntax structure ref_pic_list_struct(listIdx, rplsIdx) exists in the slice header, then the syntax structure represents the reference picture list listIdx of the current image (e.g., an image including a slice). Otherwise (e.g., present in the SPS), the syntax structure ref_pic_list_struct(listIdx, rplsIdx,) represents a candidate list of the reference picture list listIdx. The term "current picture" in the semantics refers to (1) each picture having one or more slices, where the one or more slices include ref_pic_list_idx[listIdx] equal to the index in the list of syntax structure ref_pic_list_struct(listIdx,rplsIdx,) included in the active SPS, and (2) each picture in the CVS, where the CVS has the same SPS as the active SPS.

[0135] num_ref_entries[listIdx][rplsIdx] indicates the number of entries in the syntax structure ref_pic_list_struct(listIdx,rplsIdx). The value range of num_ref_entries[listIdx][rplsIdx] is 0 to sps_max_dec_pic_buffering_minus1+14 (inclusive). ltrp_in_slice_header_flag[listIdx][rplsIdx] may be set to 0, indicating that the POCLSB of the LTRP entry in the syntax structure ref_pic_list_struct(listIdx,rplsIdx) exists in the syntax structure ref_pic_list_struct(listIdx,rplsIdx). ltrp_in_slice_header_flag[listIdx][rplsIdx] may be set to 1, indicating that the POC LSB of the LTRP entry in the syntax structure ref_pic_list_struct(listIdx,rplsIdx) does not exist in the syntax structure ref_pic_list_struct(listIdx,rplsIdx). st_ref_pic_flag[listIdx][rplsIdx][i] may be set to 1, indicating that the i-th entry in the syntax structure ref_pic_list_struct(listIdx,rplsIdx) is a short term reference picture (STRP) entry. st_ref_pic_flag[listIdx][rplsIdx][i] may be set to 0, indicating that the i-th entry in the syntax structure ref_pic_list_struct(listIdx,rplsIdx) is a LTRP entry. If st_ref_pic_flag[listIdx][rplsIdx][i] is not present, the value of st_ref_pic_flag[listIdx][rplsIdx][i] is inferred to be equal to 1.

[0136] The variable NumLtrpEntries[listIdx][rplsIdx] can be derived as follows:

[0137] for(i=0,NumLtrpEntries[listIdx][rplsIdx]=0;i <num_ref_entries[listIdx][rplsIdx];i++)

[0138] if(!st_ref_pic_flag[listIdx][rplsIdx][i])(7-86)

[0139] NumLtrpEntries[listIdx][rplsIdx]++

[0140] When the i-th entry is the first STRP entry in the syntax structure ref_pic_list_struct(listIdx,rplsIdx), abs_delta_poc_st[listIdx][rplsIdx][i] represents the absolute difference between the picture sequence number values ​​of the current picture and the picture referenced by the i-th entry, or when the i-th entry is a STRP entry but not the first STRP table entry in the syntax structure ref_pic_list_struct(listIdx,rplsIdx), abs_delta_poc_st[listIdx][rplsIdx][i] represents the absolute difference between the picture sequence number values ​​of the picture referenced by the i-th entry in the syntax structure ref_pic_list_struct(listIdx,rplsIdx) and the picture referenced by the previous STRP entry. The value range of abs_delta_poc_st[listIdx][rplsIdx][i] can be 0 to 2 15 –1 (inclusive). strp_entry_sign_flag[listIdx][rplsIdx][i] may be set to 1 to indicate that the value of the i-th entry in the syntax structure ref_pic_list_struct(listIdx,rplsIdx) is greater than or equal to 0. strp_entry_sign_flag[listIdx][rplsIdx] may be set to 0 to indicate that the value of the i-th entry in the syntax structure ref_pic_list_struct(listIdx,rplsIdx) is less than 0. If strp_entry_sign_flag[i][j] is not present, strp_entry_sign_flag[i][j] is inferred to be equal to 1.

[0141] The list DeltaPocSt[listIdx][rplsIdx] is derived as follows:

[0142]

[0143] rpls_poc_lsb_lt[listIdx][rplsIdx][i] represents the value of the picture order number of the picture referenced by the i-th entry in the syntax structure ref_pic_list_struct (listIdx, rplsIdx) modulo MaxPicOrderCntLsb. The length of the syntax element rpls_poc_lsb_lt[listIdx][rplsIdx][i] is log2_max_pic_order_cnt_lsb_minus4+4 bits.

[0144] In order to decode a video image, the image is first segmented and each segment is decoded into a bitstream. There are many image segmentation schemes. For example, an image can be segmented into regular strips, non-independent strips, tiles, and / or segmented according to wavefront parallel processing (WPP). For simplicity, HEVC restricts the encoder so that only regular strips, non-independent strips, tiles, WPP, and combinations thereof can be used when segmenting strips into CTB groups for video coding. This segmentation method can be used to support maximum transfer unit (MTU) size matching, parallel processing, and reduce end-to-end delay. MTU represents the maximum amount of data that can be transmitted by a single data packet. If the data packet payload exceeds the MTU, the data packet payload is divided into two data packets through a fragmentation process.

[0145] A regular strip, also referred to as a strip, is a portion obtained after the image is divided. It can be reconstructed independently of other regular strips in the same image, but there is still mutual dependence due to the existence of loop filtering operations. Each regular strip is encapsulated in its own Network Abstraction Layer (NAL) unit for transmission. In addition, to support reconstruction, intra-frame prediction (intra-frame sample prediction, motion information prediction, decoding mode prediction) and entropy decoding dependencies across strip boundaries can be disabled. This independent reconstruction supports parallelization. For example, parallelization based on regular strips uses minimal inter-processor or inter-core communication. However, since each regular strip is independent, each strip is associated with a separate strip header. Since each strip has a strip header bit cost overhead and lacks prediction across strip boundaries, the use of regular strips will generate a large amount of decoding overhead. In addition, regular strips can be used to support matching MTU size requirements. Specifically, since regular strips are encapsulated in separate NAL units and can be decoded independently, each regular strip needs to be smaller than the MTU in the MTU scheme to avoid splitting the strip into multiple data packets. Therefore, in order to achieve parallelization MTU size matching, the stripe layouts in the image will contradict each other.

[0146] A non-independent slice is similar to a regular slice, but the slice header is shortened, and the image tree block boundary can be segmented without destroying inter-frame prediction. Accordingly, a non-independent slice allows a regular slice to be dispersed into multiple NAL units, so that a part of the regular slice can be sent out before the encoding of the entire regular slice is completed, thereby reducing end-to-end delay.

[0147] A tile is a partition in an image formed by horizontal and vertical boundaries that form columns and rows of tiles. Tiles can be decoded in raster scan order (from right to left, from top to bottom). The scan order of CTBs is the order in which the scan is performed within a tile. Accordingly, the CTBs in the first tile are decoded in raster scan order before proceeding to the CTBs in the next tile. Similar to conventional slices, tiles eliminate the reliance on inter-frame prediction and entropy decoding. However, tiles may not be included in a single NAL unit, and therefore, tiles may not be used for MTU size matching. Each tile can be processed by one processor / core, and the inter-processor / inter-core communication used for inter-frame prediction between processing units for decoding adjacent tiles can be limited to transmitting a common slice header (when adjacent tiles are in the same slice) and sharing reconstructed samples and metadata related to loop filtering. When a slice includes multiple tiles, the entry point byte offset of each tile other than the first entry point offset in the slice can be indicated in the slice header. For each slice and partition, at least one of the following conditions must be met: (1) all coding tree blocks in a slice belong to the same partition; (2) all coding tree blocks in a partition belong to the same slice.

[0148] In WPP, the image is partitioned into a single row of CTBs. Entropy decoding and prediction mechanisms can use data from CTBs in other rows. Parallel processing is achieved by parallel decoding of CTB rows. For example, the current row can be decoded in parallel with the previous row. However, according to the example, the decoding of the current row is delayed by one or two CTBs compared to the decoding process of the previous rows. This delay ensures that data related to the CTBs above and to the right of the current CTB in the current row are available before the current CTB is decoded. When represented graphically, this approach is displayed as a wavefront. This staggered start decoding can be parallelized using as many processors / cores as the CTB rows included in the image. Since inter-frame prediction is allowed between adjacent tree block rows within the image, a large amount of inter-processor / inter-core communication may be required to achieve intra-frame prediction. WPP segmentation does not take into account the NAL unit size. Therefore, WPP does not support MTU size matching. However, conventional strips can be used in conjunction with WPP with a certain decoding overhead to achieve MTU size matching as needed.

[0149] In one example, a video decoding system may segment an image using stripes, blocks, bricks, and WPP. An image may be divided into one or more block rows and one or more block columns. A block may be a sequence of CTUs covering a rectangular area of ​​an image. A block is divided into one or more bricks, each block including several CTU rows within the block. A block that is not segmented into multiple blocks may also be referred to as a brick. A block is a true subset of a block and is therefore not referred to as a block. A strip may include multiple blocks of an image, or multiple blocks of a block. Two strip modes may be supported, namely a raster scan strip mode and a rectangular strip mode. In the raster scan strip mode, a strip contains a sequence of blocks in a raster scan of a block of an image. In the rectangular strip mode, a strip may include multiple bricks of an image, which together constitute a rectangular area of ​​the image. The individual bricks in a rectangular strip are arranged in the brick raster scan order of the strip. single_tile_in_pic_flag may be set to 1, indicating that there is only one tile in a picture and the tile is no longer divided into bricks. entropy_coding_sync_enabled_flag may be set to 1, indicating that WPP is in use.

[0150] The following is an example of the syntax of the picture parameter set RBSP:

[0151]

[0152]

[0153] An example of a common stripe header syntax is as follows:

[0154]

[0155] The semantics of the picture parameter set RBSP are as follows: single_tile_in_pic_flag can be set to 1, indicating that there is only one tile reference PPS in each picture. single_tile_in_pic_flag can be set to 0, indicating that there are multiple tiles reference PPS in each picture. In the case where there is no further block division within the tile, the entire tile can be called a brick. When a picture contains only one tile and no further block division is required, the picture can be called a block. For codestream consistency, the value of single_tile_in_pic_flag should be the same for all PPSs activated in the CVS.

[0156] uniform_tile_spacing_flag may be set to 0, indicating that tile column boundaries and tile row boundaries are uniformly distributed across the picture, and are indicated using the syntax elements tile_cols_width_minus1 and tile_rows_height_minus1. uniform_tile_spacing_flag may be set to 0, indicating that tile column boundaries and tile row boundaries may or may not be uniformly distributed across the picture, and are indicated using the syntax elements num_tile_columns_minus1 and num_tile_rows_minus1 and a list of syntax elements tile_column_width_minus1[i] and tile_row_height_minus1[i]. If uniform_tile_spacing_flag is not present, the value of uniform_tile_spacing_flag is inferred to be equal to 0. When uniform_tile_spacing_flag is equal to 1, tile_cols_width_minus1+1 represents the width of the tile columns excluding the rightmost tile columns of the picture, in units of CTBs. The value of tile_cols_width_minus1 can range from 0 to PicWidthInCtbsY–1 (inclusive). If tile_cols_width_minus1 does not exist, the value of tile_cols_width_minus1 is inferred to be equal to PicWidthInCtbsY–1.

[0157] When uniform_tile_spacing_flag is equal to 1, tile_rows_height_minus1+1 indicates the width of the tile rows excluding the tile rows below the picture, in units of CTBs. The value of tile_rows_height_minus1 shall be in the range of 0 to PicHeightInCtbsY–1, inclusive. If tile_rows_width_minus1 is not present, the value of tile_rows_width_minus1 is inferred to be equal to PicHeightInCtbsY–1. When uniform_tile_spacing_flag is equal to 0, num_tile_columns_minus1+1 indicates the number of tile columns used to partition the picture. The value of num_tile_columns_minus1 shall be in the range of 0 to PicWidthInCtbsY–1, inclusive. If single_tile_in_pic_flag is equal to 1, the value of num_tile_columns_minus1 is inferred to be equal to 0. Otherwise, if uniform_tile_spacing_flag is equal to 1, the value of num_tile_columns_minus1 is inferred. When uniform_tile_spacing_flag is equal to 0, num_tile_rows_minus1+1 indicates the number of tile rows used to partition the image. The value of num_tile_rows_minus1 should range from 0 to PicHeightInCtbsY–1, inclusive. If single_tile_in_pic_flag is equal to 1, the value of num_tile_rows_minus1 is inferred to be 0. Otherwise, if uniform_tile_spacing_flag is equal to 1, the value of num_tile_rows_minus1 is inferred. The variable NumTilesInPic may be set to (num_tile_columns_minus1+1)*(num_tile_rows_minus1+1).

[0158] When single_tile_in_pic_flag is equal to 0, NumTilesInPic shall be greater than 1. tile_column_width_minus1[i]+1 represents the width of the i-th tile column in CTBs. tile_row_height_minus1[i]+1 represents the height of the i-th tile row in CTBs. brick_splitting_present_flag may be set to 1 to indicate that one or more tiles of the image referenced by the PPS may be split into two or more bricks. brick_splitting_present_flag may be set to 0 to indicate that the tiles of the image referenced by the PPS are not split into two or more bricks. brick_split_flag[i] may be set to 1 to indicate that the i-th tile is split into two or more bricks. brick_split_flag[i] may be set to 0 to indicate that the i-th tile is not split into two or more bricks. If brick_split_flag[i] does not exist, the value of brick_split_flag[i] is inferred to be equal to 0.

[0159] uniform_brick_spacing_flag[i] can be set to 1, indicating that the block boundaries are uniformly distributed on the i-th partition, and is indicated using the syntax element brick_height_minus1[i]. uniform_brick_spacing_flag[i] can be set to 0, indicating that the block boundaries may or may not be uniformly distributed on the i-th partition, and is indicated using the syntax element num_brick_rows_minus1[i] and the syntax element list brick_row_height_minus1[i][j]. If uniform_brick_spacing_flag[i] is not present, the value of uniform_brick_spacing_flag[i] is inferred to be equal to 1. When uniform_brick_spacing_flag[i] is equal to 1, brick_height_minus1[i] represents the width of the block row excluding the lower block in the i-th partition, in units of CTBs. If brick_height_minus1 is present, the value of brick_height_minus1 shall be in the range of 0 to RowHeight[i]–2 (inclusive). If brick_height_minus1[i] is not present, the value of brick_height_minus1[i] is inferred to be equal to RowHeight[i]–1. When uniform_brick_spacing_flag[i] is equal to 0, num_brick_rows_minus1[i]+1 represents the number of blocks used to split the i-th block. If num_brick_rows_minus1[i] is present, the value of num_brick_rows_minus1[i] shall be in the range of 1 to RowHeight[i]–1 (inclusive). If brick_split_flag[i] is equal to 0, the value of num_brick_rows_minus1[i] is inferred to be equal to 0. Otherwise, if uniform_brick_spacing_flag[i] is equal to 1, the value of num_brick_rows_minus1[i] is inferred. When uniform_tile_spacing_flag is equal to 0, brick_row_height_minus1[i][j]+1 represents the height of the jth brick of the i-th tile in CTBs.

[0160] By calling the CTB raster and block scan conversion process, the following variables can be derived, and if uniform_tile_spacing_flag is equal to 1, the values ​​of num_tile_columns_minus1 and num_tile_rows_minus1 are inferred, and for each i in the range of 0 to NumTilesInPic–1 (including the end values), if uniform_brick_spacing_flag[i] is equal to 1, the value of num_brick_rows_minus1[i] is inferred. Variables including the following can be derived: List RowHeight[j] represents the height of the j-th tile row in CTB units, where j ranges from 0 to num_tile_rows_minus1 (including the end value); List CtbAddrRsToBs[ctbAddrRs] represents the conversion from the CTB address in the CTB raster scan of the image to the CTB address in the block scan, where ctbAddrRs ranges from 0 to PicSizeInCtbsY–1 (including the end value); List CtbAddrBsToRs[ctbAddrBs] represents the conversion from the CTB address in the block scan to the CTB address in the CTB raster scan of the image, where ctbAddrBs ranges from 0 to PicSize InCtbsY–1 (inclusive); the list BrickId[ctbAddrBs] represents the conversion from the CTB address in the block scan to the block ID, where the value range of ctbAddrBs is 0 to PicSizeInCtbsY–1 (inclusive); the list NumCtusInBrick[brickIdx] represents the conversion from the block index to the number of CTUs in the block, where the value range of brickIdx is 0 to NumBricksInPic–1 (inclusive); the list FirstCtbAddrBs[brickIdx] represents the conversion from the block ID to the CTB address in the block scan of the first CTB in the block, where the value range of brickIdx is 0 to NumBricksInPic–1.

[0161] single_brick_per_slice_flag may be set to 1, indicating that each slice referencing the PPS includes one brick. single_brick_per_slice_flag may be set to 0, indicating that a slice referencing the PPS may include multiple bricks. If single_brick_per_slice_flag does not exist, the value of single_brick_per_slice_flag is inferred to be equal to 1. rect_slice_flag may be set to 0, indicating that the blocks within each slice are arranged in raster scan order and the slice information is not indicated in the PPS. rect_slice_flag may be set to 1, indicating that the blocks within each slice cover a rectangular area of ​​the image and the slice information is indicated in the PPS. If single_brick_per_slice_flag is equal to 1, the value of single_brick_per_slice_flag is inferred to be equal to 1. num_slices_in_pic_minus1+1 indicates the number of slices in each picture referencing the PPS. The value of num_slices_in_pic_minus1 can range from 0 to NumBricksInPic–1 (inclusive). If num_slices_in_pic_minus1 does not exist and single_brick_per_slice_flag is equal to 1, the value of num_slices_in_pic_minus1 is inferred to be equal to NumBricksInPic–1.

[0162] top_left_brick_idx[i] represents the block index of the block located at the top left corner of the i-th slice. For any i not equal to j, the value of top_left_brick_idx[i] shall not be equal to the value of top_left_brick_idx[j]. If top_left_brick_idx[i] does not exist, the value of top_left_brick_idx[i] is inferred to be equal to i. The length of the syntax element top_left_brick_idx[i] is Ceil(Log2(NumBricksInPic) bits. bottom_right_brick_idx_delta[i] represents the difference between the block index of the block located at the bottom right corner of the i-th slice and top_left_brick_idx[i]. If single_brick_per_slice_flag is equal to 1, the value of bottom_right_brick_idx_delta[i] is inferred to be equal to 0. The length of the syntax element bottom_right_brick_idx_delta[i] is Ceil(Log2(NumBricksInPic–top_left_brick_idx[i])) bits.

[0163] Code stream consistency may require that a slice includes multiple complete blocks or only a subset of a block. The variables NumBricksInSlice[i] and BricksToSliceMap[j] represent the number of blocks in the i-th slice and the mapping relationship between blocks and slices, which are derived as follows:

[0164]

[0165] loop_filter_across_bricks_enabled_flag can be set to 1, indicating that the in-loop filtering operation can be performed across block boundaries in the image of the reference PPS. loop_filter_across_bricks_enabled_flag can be set to 0, indicating that the in-loop filtering operation is not performed across block boundaries in the image of the reference PPS. In-loop filtering operations include deblocking filtering operations, sample adaptive offset filtering operations, and adaptive loop filtering operations. If loop_filter_across_bricks_enabled_flag does not exist, it is inferred that the value of loop_filter_across_bricks_enabled_flag is equal to 1. loop_filter_across_slices_enabled_flag can be set to 1, indicating that the in-loop filtering operation can be performed across slice boundaries in the image of the reference PPS. loop_filter_across_slice_enabled_flag can be set to 0, indicating that the in-loop filtering operation is not performed across slice boundaries in the image of the reference PPS. In-loop filtering operations include deblocking filtering operations, sample adaptive offset filtering operations, and adaptive loop filtering operations. If loop_filter_across_slices_enabled_flag is not present, the value of loop_filter_across_slices_enabled_flag is inferred to be equal to 0.

[0166] signalled_slice_id_flag may be set to 1 to indicate that a slice ID for each slice is indicated. signalled_slice_id_flag may be set to 0 to indicate that a slice ID is not indicated. If rect_slice_flag is equal to 0, the value of signalled_slice_id_flag is inferred to be equal to 0. signalled_slice_id_length_minus1+1 indicates the number of bits used to represent the syntax element slice_id[i] (if present) and the syntax element slice_address in the slice header. The value of signalled_slice_id_length_minus1 may range from 0 to 15 (inclusive). If signalled_slice_id_length_minus1 does not exist, the value of signalled_slice_id_length_minus1 is inferred to be equal to Ceil(Log2(num_slices_in_pic_minus1+1))–1.

[0167] slice_id[i] represents the slice ID of the i-th slice. The length of the syntax element slice_id[i] is (signalled_slice_id_length_minus1+1) bits. If slice_id[i] does not exist, the value of slice_id[i] is inferred to be equal to i for each i in the range of 0 to num_slices_in_pic_minus1. entropy_coding_sync_enabled_flag can be set to 1, indicating that the specific synchronization process of context variables is called before decoding the CTU of the first CTB of a row of CTBs in each block in each picture that includes the reference PPS, and the specific storage process of context variables is called after decoding the CTU of the first CTB of a row of CTBs in each block in each picture that includes the reference PPS. entropy_coding_sync_enabled_flag may be set to 0, indicating that a specific synchronization process for context variables need not be called before decoding a CTU including the first CTB of a row of CTBs in each block in each picture that references a PPS, and a specific storage process for context variables need not be called after decoding a CTU including the first CTB of a row of CTBs in each block in each picture that references a PPS. Codestream conformance may require that the value of entropy_coding_sync_enabled_flag should be the same for all PPSs activated within a CVS.

[0168] An example of general slice header semantics is as follows: slice_address represents the slice address of the slice. If slice_address does not exist, the value of slice_address is inferred to be equal to 0. If rect_slice_flag is equal to 0, the following applies: The slice address is the block ID. The length of slice_address is Ceil(Log2(NumBricksInPic)) bits. The value range of slice_address can be 0 to NumBricksInPic–1 (including the end value). Otherwise (rect_slice_flag is equal to 1), the following applies: The slice address is the slice ID of the slice. The length of slice_address is signalled_slice_id_length_minus1+1 bits. If signalled_slice_id_flag is equal to 0, the value range of slice_address is 0 to num_slices_in_pic_minus1 (including the end value). Otherwise, the value range of slice_address is 0 to 2 (signalled _slice _ id _ length _ minus1+1) –1 (inclusive).

[0169] Codestream consistency may require that the following constraints apply. The value of slice_address may not be equal to the value of slice_address of any other decoded slice NAL unit of the same decoded picture. The slices of an image may be arranged in increasing order of their slice_address values. The shape of the slices in an image may ensure that each block, when decoded, has its entire left and entire top borders including the picture boundary or including the boundary of one or more previously decoded blocks. If num_bricks_in_slice_minus1 exists, it indicates the number of blocks in the slice minus 1. The value of num_bricks_in_slice_minus1 may range from 0 to NumBricksInPic–1 (inclusive). If rect_slice_flag is equal to 0 and single_brick_per_slice_flag is equal to 1, then the value of num_bricks_in_slice_minus1 is inferred to be equal to 0.

[0170] The variable NumBricksInCurrSlice represents the number of blocks in the current stripe, and SliceBrickIdx[i] represents the block index of the i-th block in the current stripe, which can be derived as follows:

[0171]

[0172] num_entry_point_offsets is used to represent the variable NumEntryPoints, which represents the number of entry points in the current strip, and is derived as follows:

[0173] NumEntryPoints=entropy_coding_sync_enabled_flag? num_entry_point_offsets:

[0174] NumBricksInCurrSlice–1(7-60)

[0175] offset_len_minus1+1 represents the length of the syntax element entry_point_offset_minus1[i] in bits. offset_len_minus1 can take a value in the range of 0 to 31 (inclusive). entry_point_offset_minus1[i]+1 represents the i-th entry point offset in bytes and is represented by (offset_len_minus1+1) bits. The slice data following the slice header consists of (NumEntryPoints+1) subsets, with subset index values ​​ranging from 0 to NumEntryPoints (inclusive). The first byte of the slice data is byte 0. If present, the anti-aliasing bytes present in the slice data portion of a decoded slice NAL unit are counted as part of the slice data for subset identification purposes. Subset 0 includes bytes 0 to entry_point_offset_minus1[0] (including the end value) of the decoded slice segment data, and subset k includes bytes firstByte[k] to lastByte[k] (including the end value) of the decoded slice data, where k ranges from 1 to NumEntryPoints-1 (including the end value), and firstByte[k] and lastByte[k] are defined as follows:

[0176] firstByte[k]=∑ k n=1 (entry_point_offset_minus1[n-1]+1)(7-61)

[0177] lastByte[k]=firstByte[k]+entry_point_offset_minus1[k](7-61)

[0178] The last subset (whose subset index is equal to NumEntryPoints) includes the remaining bytes of the coded slice data.

[0179] If entropy_coding_sync_enabled_flag is equal to 0, each subset may include all coded bits of all CTUs in the same block in the slice, and the number of subsets (e.g., the value of NumEntryPoints+1) shall be equal to the number of blocks in the slice. If entropy_coding_sync_enabled_flag is equal to 1, each subset k (k ranges from 0 to NumEntryPoints, including the end value) may include all coded bits of all CTUs in a CTU row in the block, and the number of subsets (e.g., the value of NumEntryPoints+1) may be equal to the total number of block-specific luma CTU rows in the slice.

[0180] The above examples may include one or more issues. For example, the syntax element ltrp_in_slice_header_flag in ref_pic_list_struct() may be conditional on long_term_ref_pics_flag being equal to 1. Therefore, when long_term_ref_pics_flag is equal to 1, the flag is present in each candidate reference image list structure, and there may be multiple of them. When long_term_ref_pics_flag is equal to 1, a larger SPS may be generated. As another example, the combination of the condition and the corresponding inference of the syntax element num_ref_idx_active_override_flag may be problematic. For example, for an I slice, the value of num_ref_idx_active_override_flag may be equal to 1, and therefore, the syntax element num_ref_idx_active_minus1[0], num_ref_entries[0][RplsIdx[0]] may be indicated for the I slice to be greater than 1, which is not required. As another example, when WPP is in use (eg, when entropy_coding_sync_enabled_flag is equal to 1), the syntax element num_entry_point_offsets may be indicated for use in deriving the variable NumEntryPoints. However, NumEntryPoints may be derived without these indications.

[0181] In general, this disclosure describes various techniques for improving video decoding, including improving the indication of reference picture lists and improving the indication of entry points when using WPP. The description of these techniques is based on the developing VVC standard, but may also be applicable to other video / media codec specifications.

[0182] One or more of the above problems can be solved as follows. In one exemplary aspect, a method for decoding a video bitstream is disclosed, wherein a flag is included in an SPS. The flag indicates that a POC LSB of a long term reference picture (LTRP) entry in a reference picture list structure exists in a slice header or in a reference picture list structure. For example, the flag may be ltrp_in_slice_header_flag. In another exemplary aspect, a method for decoding a video bitstream is disclosed, wherein the bitstream includes a flag indicating whether the number of active entries in a reference picture list is explicitly indicated, and when the flag is not present, the value of the flag is inferred to be equal to 0 when the slice is an I slice, and the value of the flag is inferred to be equal to 1 when the slice is a B or P slice. For example, the flag may be num_ref_idx_active_override_flag. In another exemplary aspect, a method for decoding a video bitstream is disclosed, wherein the bitstream includes a plurality of partitions in an image, a wavefront parallel processing function is applied, and the number of entry points in a slice is inferred without explicitly indicating the number of entry points. For example, the number of entry points may be inferred to be the sum of the number of CTU rows in all blocks included in the slice. As another example, the number of entry points may be represented by a variable NumEntryPoints.

[0183] The syntax of the sequence parameter set RBSP is as follows:

[0184]

[0185]

[0186] An example of a common stripe header syntax is as follows:

[0187]

[0188]

[0189] Reference picture list structure syntax

[0190]

[0191] An example of sequence parameter set RBSP semantics is as follows: long_term_ref_pics_flag can be set to 0, indicating that LTRP is not used for inter-frame prediction of any decoded picture in the CVS. long_term_ref_pics_flag can be set to 1, indicating that LTRP can be used for inter-frame prediction of one or more decoded pictures in the CVS. ltrp_in_slice_header_flag can be set to 0, indicating that the POC LSB of the LTRP entry in each syntax structure ref_pic_list_struct(listIdx,rplsIdx) exists in the syntax structure ref_pic_list_struct(listIdx,rplsIdx). ltrp_in_slice_header_flag can be set to 1, indicating that the POC LSB of the LTRP entry in each syntax structure ref_pic_list_struct(listIdx,rplsIdx) does not exist in the syntax structure ref_pic_list_struct(listIdx,rplsIdx).

[0192] The general slice header semantics are as follows: slice_poc_lsb_lt[i][j] represents the value modulo MaxPicOrderCntLsb of the picture order number of the jth LTRP entry in the i-th reference picture list. The length of the syntax element slice_poc_lsb_lt[i][j] can be (log2_max_pic_order_cnt_lsb_minus4+4) bits. The variable PocLsbLt[i][j] can be derived as follows:

[0193] PocLsbLt[i][j]=ltrp_in_slice_header_flag? (7-41)

[0194] slice_poc_lsb_lt[i][j]:rpls_poc_lsb_lt[listIdx][RplsIdx[i]][j]

[0195] num_ref_idx_active_override_flag may be equal to 1, indicating that the syntax element num_ref_idx_active_minus1[0] is present in P slices and B slices, and that the syntax element num_ref_idx_active_minus1[1] is present in B slices. num_ref_idx_active_override_flag may be set to 0, indicating that the syntax elements num_ref_idx_active_minus1[0] and num_ref_idx_active_minus1[1] are not present. If num_ref_idx_active_override_flag is not present, the value of num_ref_idx_active_override_flag is inferred as follows: If slice_type is equal to B or P, the value of num_ref_idx_active_override_flag may be inferred to be equal to 1. Otherwise (slice_type is equal to I), the value of num_ref_idx_active_override_flag may be inferred to be equal to 0.

[0196] The semantics of the reference picture list structure are exemplified as follows: num_ref_entries[listIdx][rplsIdx] may represent the number of entries in the syntax structure ref_pic_list_struct(listIdx, rplsIdx). The value range of num_ref_entries[listIdx][rplsIdx] may be between 0 and sps_max_dec_pic_buffering_minus1+14 (including the end value).

[0197] In another example, the general stripe header syntax is as follows:

[0198] slice_header(){ Descriptors ... ue(v) if(rect_slice_flag||NumBricksInPic>1) slice_address u(v) if(!rect_slice_flag&&!single_brick_per_slice_flag) num_bricks_in_slice_minus1 ... if(NumEntryPoints>0){ offset_len_minus1 ue(v) for(i=0;i<NumEntryPoints;i++) entry_point_offset_minus1[i] u(v) } ... }

[0199] The general stripe header semantics is exemplified as follows: The variable NumEntryPoints represents the number of entry points in the current stripe, which is derived as follows:

[0200]

[0201] offset_len_minus1+1 represents the length of the syntax element entry_point_offset_minus1[i] in bits. The value range of offset_len_minus1 can be 0 to 31 (including the end value).

[0202] Figure 7Schematic diagram of an exemplary video decoding device 700. The video decoding device 700 is suitable for implementing the disclosed examples / embodiments described herein. The video decoding device 700 includes a downstream port 720, an upstream port 750 and / or a transceiver unit (Tx / Rx) 710, the transceiver unit including a transmitter and / or a receiver for transmitting data upstream and / or downstream through a network. The video decoding device 700 also includes a processor 730 and a memory 732 for storing data, the processor 730 including a logic unit and / or a central processing unit (CPU) for processing data. The video decoding device 700 may also include an electrical component, an optical-to-electrical (OE) component, an electrical-to-optical (EO) component, and / or a wireless communication component coupled to the upstream port 750 and / or the downstream port 720 for data communication through an electrical communication network, an optical communication network, or a wireless communication network. The video encoding device 700 may also include an input and / or output (I / O) device 760 for sending data to and from a user. The I / O device 760 may include an output device, such as a display for displaying video data, a speaker for outputting audio data. The I / O device 760 may also include an input device such as a keyboard, a mouse, a trackball, and / or a corresponding interface for interacting with such an output device.

[0203] The processor 730 is implemented by hardware and software. The processor 730 can be implemented as one or more CPU chips, cores (for example, as a multi-core processor), field-programmable gate arrays (FPGAs), application specific integrated circuits (ASICs), and digital signal processors (DSPs). The processor 730 communicates with the downstream port 720, Tx / Rx710, upstream port 750, and memory 732. The processor 730 includes a decoding module 714. The decoding module 714 implements the embodiments disclosed herein, such as methods 100, 800, and 900, which may include WPP 500 and / or code stream 600. The decoding module 714 may also implement any other methods / mechanisms described herein. In addition, the decoding module 714 may implement the encoding and decoding system 200, the encoder 300, and / or the decoder 400. For example, the decoding module 714 can determine the value of NumEntryPoints 518 without receiving the num_entry_point_offsets parameter in the slice header of the codestream. Therefore, the decoding module 714 enables the video decoding device 700 to provide other functions and / or decoding efficiency when decoding video data. Therefore, the decoding module 714 improves the function of the video decoding device 700 and solves problems with video decoding technology. In addition, the decoding module 714 transforms the video decoding device 700 to different states. Alternatively, the decoding module 714 can be implemented as instructions stored in the memory 732 and executed by the processor 730 (e.g., a computer program product stored on a non-transitory medium).

[0204] The memory 732 includes one or more memory types, such as a disk, a tape drive, a solid-state drive, a read only memory (ROM), a random access memory (RAM), a flash memory, a ternary content-addressable memory (TCAM), a static random-access memory (SRAM), etc. The memory 732 can be used as an overflow data storage device to store programs when they are selected for execution, as well as to store instructions and data read when the programs are executed.

[0205] Figure 8Flow chart of an exemplary method 800 for encoding a video sequence into a code stream (e.g., code stream 600 encoded using WPP 500) without indicating NumEntryPoints of slices in the video sequence. Method 800 may be performed by an encoder (e.g., codec system 200, encoder 300, and / or video decoding device 700) when performing method 100.

[0206] Method 800 may start with: an encoder receiving a video sequence including a plurality of images, and determining to encode the video sequence into a bitstream according to user input, etc. In step 801, the encoder may encode an image including a stripe into a bitstream. The encoder may also determine that an image and / or a stripe should be used as an encoded reference image and / or a reference stripe of another image and / or a stripe, respectively. For clarity, this image / strip is referred to as the current image / current stripe hereinafter. When it is determined that the reference image will be used as a reference image, the reference image may be forwarded to the HRD for decoding. It should be noted that the encoder may encode a stripe header into the bitstream. In one example, the stripe header (e.g., for the reference stripe and / or the current stripe) does not include a value corresponding to the number of entry point offsets in the stripe. For example, the stripe header does not include the parameter num_entry_point_offsets. In step 803, the HRD of the encoder obtains the reference stripe of the encoded reference image from the bitstream or the memory.

[0207] In step 805, the HRD of the encoder derives NumEntryPoints in the reference strip in the reference strip according to the size of the reference strip. For example, NumEntryPoints may be derived when the reference strip is decoded according to WPP. In one example, NumEntryPoints may also be derived according to the size of the row in the reference strip. In one example, NumEntryPoints may also be derived according to the size of the column in the reference strip. In one example, NumEntryPoints may also be derived according to the number of CTUs in the reference strip. In one example, NumEntryPoints may also be derived according to the address (e.g., CTU address) in the reference strip. Then, the HRD of the encoder may determine the offset of the coded data subset in the reference strip according to NumEntryPoints. For example, the HRD may obtain an offset array. Then, the HRD may obtain the number of offsets from the array according to NumEntryPoints. In one example, the subset is a CTU row. Accordingly, the obtained offset may represent the entry point of each CTU row in the strip.

[0208] For example, NumEntryPoints can be derived as follows:

[0209]

[0210] where !entropy_coding_sync_enabled_flag indicates that WPP 500 is not in use, NumBricksInCurrSlice is the number of slices in slice 501, numBrickSpecificCtuRowsInSlice is the number of CTU rows 521 to 525 in slice 501, and BrickHeight[SliceBrickIdx[i]] is the height of CTU rows 521 to 525 (and therefore also the height of CTU column 516).

[0211] In step 807, the HRD on the processor side decodes the reference strip of the encoded reference image according to the offset of the encoded data subset in the reference strip. For example, the HRD may obtain data of each CTU row starting from a bit in a memory or a bitstream, as indicated by the corresponding offset. Then, the HRD may decode each CTU of the reference strip according to the obtained data.

[0212] In step 809, the encoder may encode the current slice into a bitstream according to the reference slice by using inter-frame prediction, etc. In step 811, the encoder may also store the bitstream for transmission to the decoder.

[0213] Fig. 9 Flow chart of an exemplary method 900 for deriving NumEntryPoints to decode a video sequence when NumEntryPoints is not included in a codestream (e.g., codestream 600 encoded using WPP 500). Method 900 may be performed by a decoder (e.g., codec system 200, decoder 400, and / or video decoding device 700) when performing method 100.

[0214] Method 900 may begin when a decoder begins receiving a codestream representing decoded data of a video sequence, for example as a result of method 800. In step 901, the decoder receives a codestream including a slice. For example, the slice may be a slice 501 decoded according to WPP 500. The codestream also includes a slice header associated with the slice. In one example, the slice header does not include a value corresponding to the number of entry point offsets in the slice. For example, the slice header does not include a parameter num_entry_point_offsets.

[0215] In step 903, the decoder derives NumEntryPoints in the slice based on the size of the slice. For example, NumEntryPoints may be derived when the slice is decoded according to WPP. In one example, NumEntryPoints may also be derived based on the size of the rows in the slice. In one example, NumEntryPoints may also be derived based on the size of the columns in the slice. In one example, NumEntryPoints may also be derived based on the number of CTUs in the slice. In one example, NumEntryPoints may also be derived based on an address in the slice (e.g., a CTU address).

[0216] For example, NumEntryPoints can be derived as follows:

[0217]

[0218] where !entropy_coding_sync_enabled_flag indicates that WPP 500 is not in use, NumBricksInCurrSlice is the number of slices in slice 501, numBrickSpecificCtuRowsInSlice is the number of CTU rows 521 to 525 in slice 501, and BrickHeight[SliceBrickIdx[i]] is the height of CTU rows 521 to 525 (and therefore also the height of CTU column 516).

[0219] In step 905, the decoder may determine the offset of the subset of the encoded stripe data included in the stripe according to the NumEntryPoints derived in step 903. For example, the decoder may obtain an offset array. Then, the decoder may obtain the number of offsets from the array according to the NumEntryPoints derived in step 903. For example, the offsets of the encoded stripe data may be associated with subset index values, respectively, and these values ​​may range from 0 to NumEntryPoints. Then, the decoder may iteratively obtain each offset of the corresponding subset index value until NumEntryPoints is reached. In one example, the subset is a CTU row. Accordingly, the obtained offset may represent an entry point for each CTU row in the stripe.

[0220] In step 907, the decoder decodes the slice according to the offset of the coded slice data subset. For example, the decoder can obtain the data of each CTU row starting from the bit in the code stream, as shown by the corresponding offset. The decoder can then decode each CTU of the slice according to the obtained data. Then, in step 909, the decoder can forward the slice to be displayed as part of the decoded video sequence.

[0221] Fig.10 1 is a schematic diagram of an exemplary system 1000 for decoding a video sequence into a code stream (e.g., a code stream 600 decoded according to WPP500) without indicating NumEntryPoints. The system 1000 may be implemented by an encoder and a decoder (e.g., the encoding and decoding system 200, the encoder 300, the decoder 400, and / or the video decoding device 700). In addition, the system 1000 may be used to implement the method 100, the method 800, and / or the method 900.

[0222] The system 1000 includes a video encoder 1002. The video encoder 1002 includes an acquisition module 1001 for acquiring a reference strip of a coded reference image. The video encoder 1002 also includes a derivation module 1003 for deriving NumEntryPoints in the reference strip according to the size of the reference strip. The video encoder 1002 also includes a determination module 1004 for determining the offset of the coded data subset in the reference strip according to NumEntryPoints. The video encoder 1002 also includes a decoding module 1005 for decoding the reference strip of the coded reference image according to the offset of the coded data subset in the reference strip. The decoding module 1005 is also used to encode the current strip into a bitstream according to the reference strip. The video encoder 1002 also includes a storage module 1006 for storing the bitstream for sending to a decoder. The video encoder 1002 also includes a sending module 1007 for sending the bitstream to a video decoder 1010. The video encoder 1002 may also be used to perform any step of the method 800 .

[0223] The system 1000 also includes a video decoder 1010. The video decoder 1010 includes: a receiving module 1011 for receiving a code stream including a slice. The video decoder 1010 also includes a determining module 1013 for deriving NumEntryPoints in the slice based on the size of the slice. The video decoder 1010 also includes a determining module 1015 for determining an offset of a subset of encoded data, wherein the subset index value ranges from 0 to NumEntryPoints. The video decoder 1010 also includes a decoding module 1017 for decoding the slice based on the offset of the subset of encoded data in the slice. The video decoder 1010 also includes a forwarding module 1019 for forwarding the slice to be displayed as part of a decoded video sequence. The video decoder 1010 can also be used to perform any step of the method 900.

[0224] A first component is directly coupled to a second component when there is no other intermediate component between the first component and the second component except a line, trace or other medium. A first component is indirectly coupled to a second component when there is another intermediate component other than a line, trace or other medium between the first component and the second component. The term "coupled" and its variations include direct coupling and indirect coupling. Unless otherwise specified, the term "approximately" is used to refer to ±10% of the number described below.

[0225] It should also be understood that the steps of the exemplary methods set forth herein do not necessarily need to be performed in the order described, and the order of the steps of these methods should be understood to be merely exemplary. Similarly, in methods consistent with various embodiments of the present invention, these methods may include other steps, and some steps may be omitted or combined.

[0226] Although the present invention provides a plurality of specific embodiments, it should be understood that the disclosed systems and methods may also be embodied in a variety of other specific forms without departing from the spirit or scope of the present invention. The examples of the present invention should be illustrative rather than restrictive, and the present invention is not limited to the details given in this text. For example, various elements or components may be combined or integrated in another system, or certain features may be omitted or not implemented.

[0227] In addition, without departing from the scope of the present invention, the techniques, systems, subsystems and methods described and illustrated as discrete or separate in various embodiments may be combined or integrated with other systems, components, techniques or methods. Other changes, substitutions and replacement examples are obvious to those skilled in the art, and all do not depart from the spirit and scope disclosed herein.

Claims

1. A method implemented in a decoder, It is characterized in that The method comprises: receiving an encoded video stream, the video stream comprising data representing a picture and a picture parameter set (PPS), wherein the picture is divided into at least one slice, the PPS comprising num_ref_idx_default_active_minus1[i] represented as a first identifier, wherein the first identifier represents an inferred value of a variable NumRefIdxActive[0] for a P or B slice with num_ref_idx_active_override_flag=0 when i=0, and represents an inferred value of NumRefIdxActive[1] for a B slice with num_ref_idx_active_override_flag=0 when i=1, wherein num_ref_idx_active_override_flag=0 indicates that syntax elements num_ref_idx_active_minus1[0] and num_ref_idx_active_minus1[1] do not exist, wherein num_ref_idx_active_minus1[i] is used to derive the variable NumRefIdxActive[i]; Derived the number of entry points (NumEntryPoints) in the stripe according to the addresses in the stripe; Determine an offset of a subset of coded stripe data, wherein a subset index value of the subset ranges from 0 to NumEntryPoints; The slice is decoded based on the PPS and an offset of the coded slice data subset.

2. The method according to claim 1, It is characterized in that The NumEntryPoints are derived when the slice is decoded according to wavefront parallel processing (WPP).

3. The method according to claim 1 or 2, It is characterized in that The codestream includes a slice header, and the slice header does not include a value corresponding to the number of entry point offsets in the slice.

4. The method according to any one of claims 1 to 3, It is characterized in that The NumEntryPoints are derived from the size of the rows in the stripe.

5. The method according to any one of claims 1 to 4, It is characterized in that The NumEntryPoints are derived from the size of the columns in the stripe.

6. The method according to any one of claims 1 to 5, It is characterized in that The NumEntryPoints is derived according to the number of coding tree units (CTUs) in the slice.

7. The method according to any one of claims 1 to 6, It is characterized in that The NumEntryPoints are derived from the addresses in the stripe.

8. The method according to any one of claims 1 to 7, It is characterized in that The NumEntryPoints are derived based on the size of the stripe.

9. The method according to any one of claims 1 to 8, It is characterized in that The num_ref_idx_active_override_flag being 1 indicates that the syntax element num_ref_idx_active_minus1[0] exists in the P slice and the B slice, and the syntax element num_ref_idx_active_minus1[1] exists in the B slice.

10. A method implemented in an encoder, It is characterized in that The method comprises: Obtaining a reference slice of a coded reference image; Derived the number of entry points (NumEntryPoints) in the reference stripe according to the address in the reference stripe; Determine an offset of a subset of encoded data in the reference slice according to the NumEntryPoints; decoding the reference slice of the encoded reference picture according to the offset of the subset of encoded data in the reference slice; determining num_ref_idx_default_active_minus1[i] represented as a first identifier, wherein the first identifier represents, when i=0, an inferred value of a variable NumRefIdxActive[0] for a P or B slice with num_ref_idx_active_override_flag=0, and when i=1, an inferred value of NumRefIdxActive[1] for a B slice with num_ref_idx_active_override_flag=0, wherein num_ref_idx_active_override_flag=0 indicates that syntax elements num_ref_idx_active_minus1[0] and num_ref_idx_active_minus1[1] do not exist, wherein num_ref_idx_active_minus1[i] is used to derive the variable NumRefIdxActive[i]; encoding the first representation into a picture parameter set (PPS); According to the reference slice, the current slice is encoded into a bitstream; and the PPS is encoded into the bitstream.

11. The method according to claim 10, It is characterized in that The NumEntryPoints are derived when the reference slice is decoded according to wavefront parallel processing (WPP).

12. The method according to claim 10 or 11, It is characterized in that The codestream includes a slice header, and the slice header does not include a value corresponding to the number of entry point offsets in the reference slice.

13. The method according to any one of claims 10 to 12, It is characterized in that The NumEntryPoints are derived from the row size in the reference stripe.

14. The method according to any one of claims 10 to 13, It is characterized in that The NumEntryPoints are derived from the size of the columns in the stripe.

15. The method according to any one of claims 10 to 14, It is characterized in that The NumEntryPoints is derived according to the number of coding tree units (CTUs) in the reference slice.

16. The method according to any one of claims 10 to 15, It is characterized in that The NumEntryPoints are derived from the addresses in the reference stripe.

17. The method according to any one of claims 10 to 16, It is characterized in that The NumEntryPoints is derived based on the size of the reference strip.

18. The method according to any one of claims 10 to 17, It is characterized in that The num_ref_idx_active_override_flag being 1 indicates that the syntax element num_ref_idx_active_minus1[0] exists in the P slice and the B slice, and the syntax element num_ref_idx_active_minus1[1] exists in the B slice.

19. A video decoding device, It is characterized in that include: A processor, a receiver coupled to the processor, a memory coupled to the processor, and a transmitter coupled to the processor, wherein the processor, the receiver, and the transmitter are configured to execute the method according to any one of claims 1 to 18.

20. A non-transitory computer readable medium, It is characterized in that A video code stream is stored, and the video code stream is obtained by executing the method according to any one of claims 10 to 17.

21. A decoder, It is characterized in that include: A receiving module, configured to receive an encoded video stream, the video stream comprising data representing an image and a picture parameter set (PPS), wherein the image is divided into at least one slice, the PPS comprising num_ref_idx_default_active_minus1[i] represented as a first identifier, wherein the first identifier represents an inferred value of a variable NumRefIdxActive[0] for a P or B slice with num_ref_idx_active_override_flag=0 when i=0, and an inferred value of NumRefIdxActive[1] for a B slice with num_ref_idx_active_override_flag=0 when i=1, wherein num_ref_idx_active_override_flag=0 indicates that syntax elements num_ref_idx_active_minus1[0] and num_ref_idx_active_minus1[1] do not exist, wherein num_ref_idx_active_minus1[i] is used to derive the variable NumRefIdxActive[i]; A derivation module, configured to derive the number of entry points (NumEntryPoints) in the stripe according to the addresses in the stripe; A determination module, configured to determine an offset of a subset of the encoded stripe data, wherein a subset index value of the subset ranges from 0 to NumEntryPoints; A decoding module is used to decode the slice according to the PPS and the offset of the encoded slice data subset.

22. The decoder according to claim 21, It is characterized in that The decoder is further configured to perform the method according to any one of claims 1 to 9.

23. An encoder, It is characterized in that include: An acquisition module, used for acquiring a reference strip of a coded reference image; A derivation module, configured to derive the number of entry points (NumEntryPoints) in the reference stripe according to the address in the reference stripe; a determination module, configured to determine an offset of a subset of coded data in the reference slice according to the NumEntryPoints, and to determine num_ref_idx_default_active_minus1[i] represented as a first identifier, wherein the first identifier represents an inferred value of a variable NumRefIdxActive[0] for a P or B slice with num_ref_idx_active_override_flag=0 when i=0, and an inferred value of NumRefIdxActive[1] for a B slice with num_ref_idx_active_override_flag=0 when i=1, wherein num_ref_idx_active_override_flag=0 indicates that syntax elements num_ref_idx_active_minus1[0] and num_ref_idx_active_minus1[1] do not exist, wherein num_ref_idx_active_minus1[i] is used to derive the variable NumRefIdxActive[i]; Decoding module, used for: encoding the first representation into a picture parameter set (PPS); decoding the reference slice of the encoded reference picture according to the offset of the subset of encoded data in the reference slice; According to the reference slice, the current slice is encoded into a bitstream; and the PPS is encoded into the bitstream.

24. The encoder according to claim 23, It is characterized in that The encoder is further configured to perform the method according to any one of claims 10 to 18.