Encoder, decoder, and corresponding method for simplifying signaling slice header syntax elements
The method enhances video coding efficiency by deriving the number of tiles in a current slice from the slice header, addressing challenges in compression and decompression, and improving image quality.
Patent Information
- Application Number
- JP2024111741
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Priority Date
- 2020-02-28
- Filing Date
- 2024-07-11
- Publication Date
- 2025-06-17
- Estimated Expiration
- 2041-03-01
AI Technical Summary
Existing video coding technologies face challenges in efficiently compressing and decompressing video data, particularly in managing signaling slice header syntax elements, which affects compression efficiency and image quality.
The proposed method involves a decoding apparatus and method that derive the number of tiles in a current slice from the slice header under specific conditions, allowing for the reconstruction of the current slice. This is achieved by obtaining a parameter from the slice header when the slice address is not the address of the last tile in the image, enabling enhanced compression efficiency.
This approach improves compression efficiency by optimizing the representation of signaling slice headers, allowing for better utilization of limited network resources while maintaining image quality.
Smart Images

Figure 0007693914000041 
Figure 0007693914000042 
Figure 0007693914000043
Abstract
Description
Technical Field
[0001] Embodiments of the present application (disclosure) generally relate to the field of image processing, and more specifically, to simplifying signaling slice header syntax elements.
Background Art
[0002] Video coding (video encoding and decoding) is used in a wide range of digital video applications such as, for example, broadcast digital TV, video transmission via the Internet and mobile networks, or real-time conversational applications such as video chat, video conferencing, DVDs and Blu-ray discs, video content acquisition and editing systems, and camcorders for security applications.
[0003] Even for relatively short videos, the amount of video data required can be substantial, and as a result, difficulties can arise when data is to be streamed or otherwise communicated over a communication network with limited bandwidth capacity. Therefore, video data is generally compressed before being communicated over modern telecommunications networks. The size of the video can also be a problem when the video is stored on a storage device, as memory resources may be limited. Video compression devices often use software and / or hardware at the source to encode the video data before transmission or storage, thereby reducing the amount of data required to represent the digital video image. The compressed data is then received at the destination by a video decompression device that decodes the video data. Improved compression and decompression techniques that improve the compression ratio without sacrificing much or any image quality are desirable because network resources are limited and the demand for higher video quality is increasing.
Summary of the Invention
[0004] Embodiments of the present disclosure provide an apparatus and method for encoding and decoding as recited in the independent claims.
[0005] The present invention provides the following.
[0006] A method for decoding an image from a video bitstream implemented by a decoding device, the bitstream including a slice header of a current slice and data representing the current slice, the method comprising: obtaining a parameter (e.g., um_tiles_in_slice_minus1) used to derive the number of tiles in the current slice from the slice header when a condition is satisfied, the condition including that the slice address (e.g., slice_address) of the current slice is not the address of the last tile in the image in which the current slice is located; and reconstructing the current slice using the number of tiles in the current slice and the data representing the current slice.
[0007] In the method, as described above, the fact that the slice address of the current slice is the address of the last tile in the image may include determining that the value obtained by subtracting the slice address of the current slice from the number of tiles in the image is equal to 1.
[0008] In the method, as described above, the fact that the slice address of the current slice is not the address of the last tile in the image may include determining that the value obtained by subtracting the slice address of the current slice from the number of tiles in the image (e.g., NumTilesInPic) is greater than 1.
[0009] Thus, according to the present invention, the presence of the image header structure in the slice header can be used to control the presence of the slice address and the number of tiles within the slice indication. If there is a single slice within an image, the slice address must be equal to the first tile within the image, and the number of tiles within the slice must be equal to the number of tiles within the image. Thus, this can enhance the compression efficiency.
[0010] In the method as described above, it can be inferred that the value of the parameter of the current slice is equal to the default value when the condition is not satisfied.
[0011] In the method as described above, the default value may be equal to 0.
[0012] In the method as described above, the slice address may be in tile units.
[0013] In the method as described above, the condition may further include a step of determining that the current slice is in the raster scan mode.
[0014] In the method as described above, the step of reconstructing the current slice using the number of tiles within the current slice may include a step of determining the scan order of the coding tree units within the current slice using the number of tiles within the current slice, and a step of reconstructing the coding tree units within the current slice using the scan order.
[0015] The present invention further provides a method for encoding a video bitstream implemented by an encoding device, where the bitstream includes a slice header of a current slice and data representing the current slice. The method includes encoding a parameter used to derive the number of tiles in the current slice from the slice header when a condition is satisfied, where the condition includes that the slice address of the current slice is not the address of the last tile in the image where the current slice is located, and reconstructing the current slice using the number of tiles in the current slice and the data representing the current slice.
[0016] The present invention further provides a device for decoding an image from a video bitstream, where the bitstream includes a slice header of a current slice and data representing the current slice. The device includes an acquisition unit configured to acquire a parameter used to derive the number of tiles in the current slice from the slice header when a condition is satisfied, where the condition includes that the slice address of the current slice is not the address of the last tile in the image where the current slice is located, and a reconstruction unit configured to reconstruct the current slice using the number of tiles in the current slice and the data representing the current slice. The present invention further provides a device for encoding an image from a coded video bitstream, where the bitstream includes a slice header of a current slice and data representing the current slice. The device includes an encoding unit configured to encode a parameter used to derive the number of tiles in the current slice from the slice header when a condition is satisfied, where the condition includes that the slice address of the current slice is not the address of the last tile in the image where the current slice is located, and a reconstruction unit configured to reconstruct the current slice using the number of tiles in the current slice and the data representing the current slice.
[0017] The present invention further provides an encoder including a processing circuit for executing the method for encoding a video bitstream as described above.
[0018] The present invention further provides a decoder including a processing circuit for executing the method for decoding a video bitstream as described above.
[0019] The present invention further provides a computer program product including program code for executing the method for encoding a video bitstream as described above or the method for decoding a video bitstream as described above when executed on a computer or a processor respectively.
[0020] The present invention further provides a decoder including one or more processors and a non-transitory computer-readable storage medium coupled to the processors and storing programming for execution by the processors, the programming configuring the decoder to execute a method for decoding a video bitstream as described above when executed by the processors.
[0021] The present invention further provides an encoder including one or more processors and a non-transitory computer-readable storage medium coupled to the processors and storing programming for execution by the processors, the programming configuring the encoder to execute a method for encoding a video bitstream as described above when executed by the processors.
[0022] The present invention further provides a non-transitory computer-readable medium holding program code for causing a computer device to execute the method for encoding a video bitstream as described above or the method for decoding a video bitstream as described above when executed by the computer device.
[0023] The present invention further provides a non-transitory storage medium including a video bitstream, where the bitstream includes a slice header of a current slice and data representing the current slice. Here, the slice header includes a slice address of the current slice. Here, when a condition is satisfied, the slice header further includes a parameter used to derive the number of tiles within the current slice from the slice header, and the condition includes that the slice address of the current slice is not the address of the last tile within the image where the current slice is located.
[0024] Details of one or more embodiments are set forth in the accompanying drawings and the description below. Other features, objects, and advantages will be apparent from the description, drawings, and claims.
Brief Description of the Drawings
[0025] In the following, embodiments of the present invention will be described in more detail with reference to the accompanying figures and drawings.
Figure 1A
Figure 1B
Figure 2
Figure 3
Figure 4
Figure 5
Figure 6
Figure 7
Figure 8
Figure 9
Figure 10
Figure 11
Figure 12
Figure 13
Figure 14
Figure 15
Figure 16
Figure 17
DETAILED DESCRIPTION OF THE INVENTION
[0026] In the following description, reference is made to the accompanying drawings which form a part hereof and which illustrate specific aspects of embodiments of the present disclosure or specific aspects in which embodiments of the present invention may be used. It is understood that embodiments of the present invention may be used in other aspects and may include structural or logical changes not shown in the figures. Accordingly, the following detailed description is not to be taken in a limiting sense, and the scope of the present invention is defined by the appended claims.
[0027] For example, it is understood that the disclosures related to the described method may also apply to corresponding devices or systems configured to perform the method, and vice versa. For example, when one or more of the steps of a particular method are described, the corresponding device may include one or more units, such as functional units, for performing the one or more described method steps (e.g., one unit for performing the one or more steps, or multiple units each performing one or more of the multiple steps), even if such one or more units are not explicitly described or shown in the drawings. On the other hand, for example, when a particular device is described based on one or more units, such as functional units, the corresponding method may include one or more steps (e.g., one step for performing the functions of the one or more units, or multiple steps each performing one or more of the functions of the multiple units) for performing the functions of the one or more units, even if such one or more steps are not explicitly described or shown in the drawings. Further, it is understood that the features of the various exemplary embodiments and / or aspects described herein may be combined with each other unless otherwise specifically noted.
[0028] Video coding typically refers to the processing of a series of images that form a video or video sequence. In the field of video coding, the terms "frame" or "image" may be used synonymously with the term "image" instead. Video coding (or generally coding) includes two parts: video encoding and video decoding. Video encoding is performed on the source side and typically involves processing the original video image (e.g., by compression) to reduce the amount of data required to represent the video image (for more efficient storage and / or transmission). Video decoding is performed on the destination side and typically involves the reverse process compared to the encoder to reconstruct the video image. Embodiments referring to the "coding" of a video image (or generally an image) are to be understood as relating to the "encoding" or "decoding" of the video image or each video sequence. The combination of the encoding part and the decoding part is also referred to as a codec (coding and decoding).
[0029] In the case of lossless video coding, the original video image can be reconstructed, i.e., the reconstructed video image has the same quality as the original video image (assuming no transmission loss or other data loss during storage or transmission). In the case of irreversible video coding, further compression, e.g., by quantization, is performed to reduce the amount of data representing the video image, but this cannot be fully reconstructed at the decoder, i.e., the quality of the reconstructed video image is degraded or deteriorated compared to the quality of the original video image.
[0030] Some video coding standards belong to the group of "irreversible hybrid video codecs" (i.e., they combine spatial and temporal prediction in the sample domain with 2D transform coding for applying quantization in the transform domain). Each image of a video sequence is typically partitioned into a set of non-overlapping blocks, and coding is typically performed at the block level. In other words, in the encoder, the video generates prediction blocks using, for example, spatial (intra-picture) prediction and / or temporal (inter-picture) prediction, subtracts the prediction blocks from the current block (the block being currently processed / to be processed), obtains a residual block, transforms the residual block, and quantizes the residual block in the transform domain to reduce (compress) the amount of data to be transmitted, typically processed (i.e., encoded) at the block (video block) level. On the other hand, in the decoder, the reverse process compared to the encoder is applied to the encoded or compressed blocks to reconstruct the current block for representation. Further, the encoder repeats the decoder processing loop, so that both generate the same prediction (e.g., intra and inter prediction) and / or reconstruction for the processing of subsequent blocks, i.e., for coding.
[0031] Embodiments of a video coding system 10, a video encoder 20, and a video decoder 30 are described below with reference to FIGS. 1A to 3.
[0032] FIG. 1A is a schematic block diagram showing an exemplary coding system 10, e.g., a video coding system 10 (or simply coding system 10) that may use the techniques of the present disclosure. The video encoder 20 (or simply encoder 20) and the video decoder 30 (or simply decoder 30) of the video coding system 10 represent examples of devices that may be configured to perform techniques according to various examples described in the present disclosure.
[0033] As shown in FIG. 1A, the coding system 10 includes a source device 12 configured to provide encoded image data 21, for example, to a destination device 14 for decoding the encoded image data 21.
[0034] The source device 12 includes an encoder 20 and may additionally, i.e., optionally, include an image source 16, a preprocessor (or preprocessing unit) 18, for example, an image preprocessor 18, and a communication interface or communication unit 22.
[0035] The image source 16 may include any type of image capture device, such as a camera that captures real-world images, and / or any type of image generation device, such as a computer graphics processor that generates computer-animated images, or any other device that acquires and / or provides real-world images, computer-generated images (e.g., screen content, virtual reality (VR) images), and / or any combination thereof (e.g., augmented reality (AR) images). The image source may be any type of memory or storage that stores any of the aforementioned images.
[0036] To distinguish from the processing performed by the preprocessor 18 and the preprocessing unit 18, the image or image data 17 may also be referred to as raw image or raw image data 17.
[0037] The preprocessor 18 is configured to receive (raw) image data 17 and perform preprocessing on the image data 17 to obtain preprocessed image 19 or preprocessed image data 19. The preprocessing performed by the preprocessor 18 may include, for example, trimming, color format conversion (e.g., from RGB to YCbCr), color correction, or noise removal. It can be understood that the preprocessing unit 18 may be an optional component.
[0038] The video encoder 20 is configured to receive the pre - processed image data 19 and provide the encoded image data 21 (further details are described below, for example, based on FIG. 2).
[0039] The communication interface 22 of the source device 12 is configured to receive the encoded image data 21 via the communication channel 13 and transmit the encoded image data 21 (or any further processed version thereof) for storage or direct reconstruction to another device, such as the destination device 14 or any other device.
[0040] The destination device 14 includes a decoder 30 (e.g., a video decoder 30), and in addition, i.e., optionally, may include a communication interface or communication unit 28, a post - processor 32 (or post - processing unit 32), and a display device 34.
[0041] The communication interface 28 of the destination device 14 is configured to receive the encoded image data 21 (or any further processed version thereof) from, for example, directly from the source device 12 or any other source, such as a storage device, such as an encoded image data storage device, and provide the encoded image data 21 to the decoder 30.
[0042] The communication interface 22 and the communication interface 28 are configured to transmit or receive the encoded image data 21 or the encoded data 13 between the source device 12 and the destination device 14 via a direct communication link, such as a direct wired or wireless connection, or via any type of network, such as a wired or wireless network or any combination thereof, or via any type of private and public network or any combination thereof.
[0043] The communication interface 22 may be configured to process the encoded image data 21, for example, by packaging the encoded image data 21 into an appropriate format, such as packets, and / or using any kind of transmission encoding or processing for transmission via a communication link or communication network.
[0044] The communication interface 28 that forms the counterpart of the communication interface 22 may be configured to receive the transmitted data and process the transmitted data using any kind of corresponding transmission decoding or processing and / or unpackaging to obtain the encoded image data 21.
[0045] Both the communication interface 22 and the communication interface 28 may be configured as a unidirectional communication interface or a bidirectional communication interface as indicated by the arrow of the communication channel 13 that faces from the source device 12 to the destination device 14 in FIG. 1A, and may be configured to, for example, transmit and receive messages, for example, set up a connection, and confirm and exchange any other information related to the communication link and / or data transmission, for example, the transmission of encoded image data.
[0046] The decoder 30 is configured to receive the encoded image data 21 and provide the decoded image data 31 or the decoded image 31 (further details will be described below, for example, based on FIG. 3 or FIG. 5).
[0047] The post-processor 32 of the destination device 14 is configured to post-process the decoded image data 31 (also referred to as reconstructed image data), for example, the decoded image 31, to obtain post-processed image data 33, for example, the post-processed image 33. The post-processing performed by the post-processing unit 32 may include, for example, color format conversion (e.g., from YCbCr to RGB), color correction, trimming, or resampling, or any other processing for the purpose of preparing the decoded image data 31 for display by, for example, the display device 34.
[0048] The display device 34 of the destination device 14 is configured to receive the post-processed image data 33 for displaying an image to, for example, a user or viewer. The display device 34 may be any type of display for representing the reconstructed image, such as an integrated or external display or monitor, or may include this. The display may include, for example, a liquid crystal display (LCD), an organic light emitting diode (OLED) display, a plasma display, a projector, a micro LED display, a liquid crystal on silicon (LCoS), a digital light processor (DLP), or any other type of display.
[0049] FIG. 1A shows the source device 12 and the destination device 14 as separate devices, but embodiments of the devices may include both of them or both of their functions, i.e., the source device 12 or corresponding functions and the destination device 14 or corresponding functions. In such embodiments, the source device 12 or corresponding functions and the destination device 14 or corresponding functions may be implemented using the same hardware and / or software, or by separate hardware and / or software or any combination thereof.
[0050] As will be apparent to those skilled in the art based on this description, the functions of different units or the presence and (exact) partitioning of functions within the source device 12 and / or the destination device 14, as shown in FIG. 1A, can vary depending on the actual devices and applications.
[0051] The encoder 20 (e.g., a video encoder 20) or the decoder 30 (e.g., a video decoder 30) or both the encoder 20 and the decoder 30 may be implemented via a processing circuit as shown in FIG. 1B, such as one or more microprocessors, digital signal processors (DSPs), application specific integrated circuits (ASICs), field programmable gate arrays (FPGAs), discrete logic, hardware, dedicated video coding, or any combination thereof. The encoder 20 may be implemented via the processing circuit 46 to embody various modules described in relation to the encoder 20 of FIG. 2 and / or any other encoder system or subsystem described herein. The decoder 30 may be implemented via the processing circuit 46 to embody various modules described in relation to the decoder 30 of FIG. 3 and / or any other decoder system or subsystem described herein. The processing circuit may be configured to perform various operations as will be described later. As shown in FIG. 5, when the technology is implemented partially in software, the device may store instructions for the software in a suitable non-transitory computer-readable storage medium and execute the instructions using one or more processors within the hardware to perform the technology of the present disclosure. Either the video encoder 20 or the video decoder 30 may be integrated, for example, as part of a combined encoder / decoder (codec) within a single device, as shown in FIG. 1B.
[0052] Source device 12 and destination device 14 may include any of a wide range of devices, such as any type of handheld or stationary device, for example, a notebook or laptop computer, a mobile phone, a smartphone, a tablet or tablet computer, a camera, a desktop computer, a set-top box, a television, a display device, a digital media player, a video gaming console, a video streaming device (such as a content service server or a content delivery server, etc.), a broadcast receiver device, a broadcast transmitter device, etc., and may use no operating system or any type of operating system. In some cases, source device 12 and destination device 14 may support wireless communication. Thus, source device 12 and destination device 14 may be wireless communication devices.
[0053] In some cases, the video coding system 10 shown in FIG. 1A is merely an example, and the techniques of the present disclosure can be applied to video coding settings (such as video encoding or video decoding) that do not necessarily include any data communication between an encoding device and a decoding device. In other examples, data is obtained from local memory and streamed via a network. The video encoding device may encode and store the data in memory and / or the video decoding device may decode and obtain the data from memory. In some examples, encoding and decoding are performed by devices that do not communicate with each other but simply encode data into memory and / or obtain and decode data from memory.
[0054] For the sake of convenience in explanation, embodiments of the present invention are described herein by referring to, for example, the reference software of High Efficiency Video Coding (HEVC) or Versatile Video Coding (VVC), and the next-generation video coding standard developed by the Joint Collaborative Team on Video Coding (JCT-VC) of ITU-T Video Coding Experts Group (VCEG) and ISO / IEC Moving Picture Experts Group (MPEG). Those skilled in the art will understand that the embodiments of the present invention are not limited to HEVC or VVC.
[0055] [Encoder and Encoding Method] FIG. 2 shows a schematic block diagram of an exemplary video encoder 20 configured to implement the technology of the present disclosure. In the example of FIG. 2, the video encoder 20 includes an input 201 (or input interface 201), a residual calculation unit 204, a conversion processing unit 206, a quantization unit 208, an inverse quantization unit 210, an inverse conversion processing unit 212, a reconstruction unit 214, a loop filter unit 220, a decoded picture buffer (DPB) 230, a mode selection unit 260, an entropy encoding unit 270, and an output 272 (or output interface 272). The mode selection unit 260 may include an inter prediction unit 244, an intra prediction unit 254, and a partitioning unit 262. The inter prediction unit 244 may include a motion estimation unit and a motion compensation unit (not shown). The video encoder 20 shown in FIG. 2 may also be referred to as a hybrid video encoder or a video encoder by a hybrid video codec.
[0056] The residual calculation unit 204, the conversion processing unit 206, the quantization unit 208, and the mode selection unit 260 may be referred to as forming the forward signal path of the encoder 20. On the other hand, the inverse quantization unit 210, the inverse conversion processing unit 212, the reconstruction unit 214, the buffer 216, the loop filter 220, the decoded picture buffer (DPB) 230, the inter prediction unit 244, and the intra prediction unit 254 may be referred to as forming the reverse signal path of the video encoder 20. The reverse signal path of the video encoder 20 corresponds to the signal path of the decoder (see the video decoder 30 in FIG. 3). The inverse quantization unit 210, the inverse conversion processing unit 212, the reconstruction unit 214, the loop filter 220, the decoded picture buffer (DPB) 230, the inter prediction unit 244, and the intra prediction unit 254 are also referred to as forming the "built-in decoder" of the video encoder 20.
[0057] [Image and Image Segmentation (Image and Block)] The encoder 20 may be configured to receive, for example, via the input 201, an image 17 (or image data 17), for example, an image from a series of images forming a video or video sequence. The received image or image data may be the preprocessed image 19 (or preprocessed image data 19). For the sake of brevity, the image 17 is referred to in the following description. The image 17 may also be referred to as the current image or the image to be coded (especially in video coding, to distinguish the current image from other images, for example, images previously encoded and / or decoded in the same video sequence, i.e., the video sequence that also includes the current image). (Digital) images are or can be considered as two-dimensional arrays or matrices of samples having intensity values. Samples within the array can also be referred to as pixels (abbreviation for picture elements) or pels. The number of samples in the horizontal and vertical directions (or axes) of the array or image defines the size and / or resolution of the image. To represent color, typically three color components are utilized, i.e., the image may be represented by or may include three sample arrays. In the RGB format or color space, the image includes corresponding red, green, and blue sample arrays. However, in video coding, each pixel is typically represented in a luminance and chrominance format or color space, such as YCbCr, which includes a luminance component indicated by Y (where L may also be used instead) and two chrominance components indicated by Cb and Cr. The luminance (or simply luma) component Y represents brightness or intensity of gray levels (such as in a grayscale image), and the two chrominance (or simply chroma) components Cb and Cr represent chrominance or color information components. Thus, an image in YCbCr format includes a luminance sample array of luminance sample values (Y) and two chrominance sample arrays of chrominance values (Cb and Cr). An image in RGB format can be converted or transformed to the YCbCr format and vice versa, and this process is also known as color conversion or transformation. If the image is monochrome, the image may include only a luminance sample array. Thus, the image can be, for example, an array of luma samples in monochrome format or an array of luma samples and two corresponding arrays of chroma samples in color formats of 4:2:0, 4:2:2, and 4:4:4.
[0058] An embodiment of the video encoder 20 may include an image segmentation unit (not shown in FIG. 2) configured to segment the image 17 into a plurality of (typically non-overlapping) image blocks 203. These blocks may also be referred to as root blocks, macroblocks (H.264 / AVC), or coding tree blocks (CTB) or coding tree units (CTU) (H.265 / HEVC and VVC). The image segmentation unit may use the same block size for all images of the video sequence and the corresponding grid defining the block size, or may be configured to vary the block size between images or subsets or groups of images to segment each image into corresponding blocks.
[0059] In a further embodiment, the video encoder may be configured to directly receive the blocks 203 of the image 17, for example, one, some, or all of the blocks forming the image 17. The image block 203 may also be referred to as the current image block or the image block to be coded.
[0060] Similar to the image 17 here, the image block 203 is also a two-dimensional array or matrix of samples that is smaller in size than the image 17 but has intensity values (sample values), or can be regarded as such. In other words, the block 203 may include, for example, one sample array (e.g., a luma array in the case of a monochrome image 17, or a luma or chroma array in the case of a color image), or three sample arrays (e.g., a luma and two chroma arrays in the case of a color image 17), or any other number and / or type of array depending on the color format applied. The number of samples in the horizontal and vertical directions (or axes) of the block 203 defines the size of the block 203. Thus, the block may be, for example, an M×N (M columns × N rows) array of samples, or an M×N array of transform coefficients.
[0061] An embodiment of the video encoder 20 as shown in FIG. 2 may be configured to encode the image 17 block by block. For example, encoding and prediction may be performed for each block 203.
[0062] The embodiment of the video encoder 20 shown in FIG. 2 may be further configured to segment and / or encode an image by using slices (also referred to as video slices), where the image may be segmented or encoded using one or more slices (usually without overlap), and each slice may include one or more blocks (e.g., CTUs).
[0063] The embodiment of the video encoder 20 as shown in FIG. 2 may be further configured to segment and / or encode an image by using tile groups (also referred to as video tile groups) and / or tiles (also referred to as video tiles). The image may be segmented or encoded using one or more tile groups (usually without overlap). Each tile group may include, for example, one or more blocks (e.g., CTUs) or one or more tiles. Each tile may be, for example, rectangular in shape and may include one or more blocks (e.g., CTUs), such as complete blocks or partial blocks.
[0064] [Residual calculation] The residual calculation unit 204 may be configured to calculate a residual block 205 (also referred to as residual 205) based on the image block 203 and the prediction block 265 (more details regarding the prediction block 265 will be provided later), for example, by subtracting the sample values of the prediction block 265 from the sample values of the image block 203 for each sample (pixel by pixel) to obtain a residual block 205 in the sample region.
[0065] [Transformation] The conversion processing unit 206 may be configured to apply a conversion, such as a discrete cosine transform (DCT) or a discrete sine transform (DST), to the sample values of the residual block 205 to obtain conversion coefficients 207 in the conversion domain. The conversion coefficients 207 may also be referred to as conversion residual coefficients and represent the residual block 205 in the conversion domain.
[0066] The conversion processing unit 206 may be configured to apply an integer approximation of DCT / DST such as the conversion specified in H.265 / HEVC. Compared with the orthogonal DCT transform, such an integer approximation is typically scaled by a specific coefficient. To preserve the norm of the residual block processed by the forward and inverse transforms, an additional scaling coefficient is applied as part of the conversion process. The scaling coefficient is typically selected based on specific constraints such as a scaling coefficient that is a power of two with respect to a shift operation, the bit depth of the conversion coefficients, and the trade-off between accuracy and implementation cost. For example, a specific scaling coefficient may be specified for the inverse transform by, for example, the inverse transform processing unit 212 (and the corresponding inverse transform by the inverse transform processing unit 312 in, for example, the video decoder 30), and the corresponding scaling coefficient for the forward transform by the conversion processing unit 206 in the encoder 20 may be specified accordingly.
[0067] Embodiments of the video encoder 20 (each the conversion processing unit 206) may be configured to encode or compress conversion parameters, such as the type of one or more conversions, for example, directly or via the entropy encoding unit 270 and then output them, whereby, for example, the video decoder 30 may receive and use the conversion parameters for decoding.
[0068] [Quantization] The quantization unit 208 may be configured to quantize the transform coefficient 207 to obtain a quantized coefficient 209, for example, by applying scalar quantization or vector quantization. The quantized coefficient 209 may also be referred to as a quantized transform coefficient 209 or a quantized residual coefficient 209.
[0069] The quantization process may reduce the bit depth associated with some or all of the transform coefficients 207. For example, an n-bit transform coefficient may be rounded to an m-bit transform coefficient during quantization, where n is greater than m. The degree of quantization may be changed by adjusting the quantization parameter (QP). For example, in the case of scalar quantization, different scalings may be applied to achieve finer or coarser quantization. The smaller the quantization step size, the finer the quantization, while the larger the quantization step size, the coarser the quantization. The applicable quantization step size may be indicated by the quantization parameter (QP). The quantization parameter may be, for example, an index to a predefined set of applicable quantization step sizes. For example, a small quantization parameter may correspond to fine quantization (small quantization step size), and a large quantization parameter may correspond to coarse quantization (large quantization step size), or vice versa. Quantization may include division by the quantization step size. For example, the corresponding inverse quantization and / or dequantization by the inverse quantization unit 210 may include multiplication by the quantization step size. Some standards, such as embodiments according to HEVC, may be configured to use the quantization parameter to determine the quantization step size. Generally, the quantization step size may be calculated based on the quantization parameter using a fixed-point approximation of an equation that includes division. Additional scaling factors may be introduced for quantization and dequantization to restore the norm of the residual block that may be changed due to the scaling used in the fixed-point approximation of the equations for the quantization step size and quantization parameter. In one example implementation, the scaling for inverse transform and dequantization may be combined. Alternatively, a customized quantization table may be used, for example, signaled from the encoder to the decoder in the bitstream. Quantization is an irreversible operation, and the loss increases with an increase in the quantization step size.
[0070] Embodiments of the video encoder 20 (each quantization unit 208) may be configured to output, for example, after encoding via a quantization parameter (QP), such as directly or via the entropy encoding unit 270, whereby, for example, the video decoder 30 may receive and apply the quantization parameter for decoding.
[0071] [Inverse quantization] The inverse quantization unit 210 is configured to apply inverse quantization of the quantization unit 208 to the quantization coefficients, for example, by applying the inverse of the quantization scheme applied by the quantization unit 208 based on or using the same quantization step size as the quantization unit 208, to obtain the dequantization coefficients 211. The dequantization coefficients 211, which may also be referred to as dequantization residual coefficients 211, are typically not identical to the transform coefficients due to losses caused by quantization, but correspond to the transform coefficients 207.
[0072] [Inverse transform]
[0073] The inverse transform processing unit 212 is configured to apply an inverse transform of the transform applied by the transform processing unit 206, for example, an inverse discrete cosine transform (DCT) or an inverse discrete sine transform (DST) or other inverse transform, to obtain the reconstructed residual block 213 (or corresponding dequantization coefficients 213) in the sample region. The reconstructed residual block 213 may also be referred to as the transform block 213.
[0074] [Reconstruction] The reconstruction unit 214 (for example, an adder or accumulator 214) is configured to add the transform block 213 (i.e., the reconstructed residual block 213) to the prediction block 265, for example, by adding the sample values of the reconstructed residual block 213 and the sample values of the prediction block 265 for each sample, to obtain the reconstructed block 215 in the sample region.
[0075] [Filtering] The loop filter unit 220 (or, abbreviated as "loop filter" 220) is configured to filter the reconstructed block 215 to obtain a filtered block 221, or generally, to filter the reconstructed samples to obtain filtered samples. The loop filter unit is configured, for example, to smooth pixel transitions or otherwise improve video quality. The loop filter unit 220 may include a deblocking filter, a sample adaptive offset (SAO) filter, or one or more other filters, such as a bilateral filter, an adaptive loop filter (ALF), sharpening, a smoothing filter or a collaborative filter, or any combination thereof. Although the loop filter unit 220 is shown in FIG. 2 as being within the loop filter, in other configurations, the loop filter unit 220 may be implemented as a post-loop filter. The filtered block 221 may also be referred to as the filtered reconstructed block 221.
[0076] Embodiments of the video encoder 20 (each loop filter unit 220) may be configured to encode loop filter parameters (such as sample adaptive offset information, etc.) and then output them, for example, directly or via the entropy encoding unit 270, so that, for example, the decoder 30 may receive and apply the same loop filter parameters or respective loop filters for decoding.
[0077] [Decoded Image Buffer] The decoded picture buffer (DPB) 230 may be a memory that stores reference pictures for encoding video data by the video encoder 20, or generally reference picture data. The DPB 230 may be formed by any of various memory devices such as dynamic random access memory (DRAM), including synchronous DRAM (SDRAM), magnetoresistive RAM (MRAM), resistive RAM (RRAM (registered trademark)), or other types of memory devices. The decoded picture buffer (DPB) 230 may be configured to store one or more filtered blocks 221. The decoded picture buffer 230 may be further configured to store other previously filtered blocks, such as previously reconstructed and filtered blocks 221, of the same current picture or a different picture, e.g., a previously reconstructed picture, and may provide, for example, for inter prediction, a fully previously reconstructed, i.e., decoded, picture (and corresponding reference blocks and samples) and / or a partially reconstructed current picture (and corresponding reference blocks and samples). The decoded picture buffer (DPB) 230 may be configured to store, for example, one or more non-filtered reconstructed blocks 215, or generally non-filtered reconstructed samples, or any other further processed version of the reconstructed blocks or samples, if, for example, the reconstructed blocks 215 have not been filtered by the loop filter unit 220.
[0078] [Mode Selection (Segmentation and Prediction)] The mode selection unit 260 includes a segmentation unit 262, an inter prediction unit 244, and an intra prediction unit 254, and is configured to receive or obtain original image data, for example, the original block 203 (the current block 203 of the current image 17), and reconstructed image data, for example, filtered and / or unfiltered reconstructed samples or blocks from the same (current) image and / or from one or more previously decoded images, for example, from the decoded image buffer 230 or other buffers (e.g., a line buffer not shown). The reconstructed image data is used as reference image data for prediction, for example, inter prediction or intra prediction, to obtain the prediction block 265 or the prediction factor 265.
[0079] The mode selection unit 260 may be configured to determine or select a segmentation and prediction mode (e.g., an intra or inter prediction mode) for the current block prediction mode (without segmentation) and generate a corresponding prediction block 265, which is used for the calculation of the residual block 205 and the reconstruction of the reconstructed block 215.
[0080] An embodiment of the mode selection unit 260 may be configured to select a partitioning and prediction mode (e.g., from those supported by or available to the mode selection unit 260), thereby providing the best match, or in other words, the minimum residual (the minimum residual means a better compression rate for transmission or storage), or the minimum signaling overhead (the minimum signaling overhead means a better compression rate for transmission or storage), or a consideration or balance of both. The mode selection unit 260 may be configured to determine the partitioning and prediction modes based on rate-distortion optimization (RDO), i.e., to select a prediction mode that provides the minimum rate distortion. In this context, terms such as "best," "minimum," "optimal," etc. do not necessarily refer to the general "best," "minimum," "optimal," etc., but may refer to the achievement of an end or selection criterion such that the value exceeds or falls below a threshold or other constraint, potentially leading to a "quasi-optimal selection" but reducing complexity and processing time.
[0081] In other words, the partitioning unit 262 may be configured to repeatedly use, for example, quadtree partitioning (QT), binary tree partitioning (BT), or ternary tree partitioning (TT), or any combination thereof, to partition block 203 into smaller block partitions or sub-blocks (which also form blocks here), and also, for example, to perform prediction for each block partition or sub-block, where the mode selection includes the selection of the tree structure of the partitioned block 203, and the prediction mode is applied to each of the block partitions or sub-blocks.
[0082] The partitioning (e.g., by the partitioning unit 260) and prediction processing (by the inter prediction unit 244 and the intra prediction unit 254) performed by the exemplary video encoder 20 will be described in more detail below.
[0083] [Partitioning] The partitioning unit 262 may partition (or divide) the current block 203 into smaller partitions, for example, smaller blocks of square or rectangular size. These smaller blocks (which may also be referred to as sub-blocks) may be further partitioned into even smaller partitions. This is also referred to as tree partitioning or hierarchical tree partitioning, where, for example, a root block at root tree level 0 (hierarchical level 0, depth 0) may be recursively partitioned, for example, into two or more blocks of nodes at the next lower tree level, for example, tree level 1 (hierarchical level 1, depth 1), and these blocks may be further partitioned into two or more blocks at the next lower level, for example, tree level 2 (hierarchical level 2, depth 2), until the partitioning ends, for example, when an end criterion is achieved, for example, when the maximum tree depth or minimum block size is reached. Blocks that are not further partitioned are also referred to as leaf blocks or leaf nodes of the tree. A tree using partitioning into two partitions is referred to as a binary tree (BT), a tree using partitioning into three partitions is referred to as a ternary tree (TT), and a tree using partitioning into four partitions is referred to as a quaternary tree (QT).
[0084] As mentioned previously, the term "block" as used herein may be a portion of an image, particularly a square or rectangular portion. For example, referring to HEVC and VVC, a block may be or may correspond to a coding tree unit (CTU), a coding unit (CU), a prediction unit (PU), and a transform unit (TU), and / or may correspond to a corresponding block, for example, a coding tree block (CTB), a coding block (CB), a transform block (TB), or a prediction block (PB).
[0085] For example, a coding tree unit (CTU) may be or may include a CTB of luma samples of an image having three sample arrays, two corresponding CTBs of chroma samples, or a CTB of samples of an image coded using three separate color planes and syntax structures used to code a monochrome image or samples. Correspondingly, a coding tree block (CTB) can be samples of an N×N block for some value N such that the splitting into component CTBs is a partitioning. A coding unit (CU) may be or may include a coding block of luma samples, two corresponding coding blocks of chroma samples of an image having three sample arrays, or a coding block of samples of an image coded using three separate color planes and syntax structures used to code a monochrome image or samples. Correspondingly, a coding block (CB) can be samples of an M×N block for some values of M and N such that the splitting into coding blocks of a CTB is a partitioning.
[0086] For example, in an embodiment according to HEVC, a coding tree unit (CTU) may be split into CUs by using a quadtree structure shown as a coding tree. The decision of whether to use inter-picture (temporal) prediction or intra-picture (spatial) prediction to code an image area is made at the CU level. Each CU may be further split into one, two, or four PUs according to the PU split type. Inside one PU, the same prediction process is applied and the relevant information is sent to the decoder on a PU basis. After obtaining a residual block by applying a prediction process based on the PU split type, the CU can be partitioned into transform units (TUs) according to another quadtree structure similar to the coding tree for the CU.
[0087] For example, in an embodiment that complies with the latest video coding standard currently under development, referred to as versatile video coding (VVC), combined quadtree and binary tree (QTBT) partitioning is used, for example, to partition coding blocks. In the QTBT block structure, a CU can be either square or rectangular in shape. For example, a coding tree unit (CTU) is first partitioned by a quadtree structure. A quadtree leaf node is further partitioned by a binary or ternary tree structure. A leaf node of the partitioning tree is referred to as a coding unit (CU), and its segmentation is used for prediction and transform processing without any further partitioning. This means that the CU, PU, and TU have the same block size within the QTBT coding block structure. In parallel, multiple partitionings, for example, ternary partitioning, can be used together with the QTBT block structure.
[0088] In one example, the mode selection unit 260 of the video encoder 20 may be configured to perform any combination of the partitioning techniques described herein.
[0089] As described above, the video encoder 20 is configured to determine or select the best or optimal prediction mode from a set of prediction modes (e.g., pre-determined). The set of prediction modes may include, for example, an intra prediction mode and / or an inter prediction mode.
[0090] [Intra Prediction] The set of intra prediction modes may include 35 different intra prediction modes, such as non-directional modes like DC (or mean) mode and planar mode, or directional modes as defined, for example, in HEVC, or may include 67 different intra prediction modes, such as non-directional modes like DC (or mean) mode and planar mode, or directional modes as defined, for example, in VVC.
[0091] The intra prediction unit 254 is configured to generate an intra prediction block 265 according to an intra prediction mode among a set of intra prediction modes, using reconstructed samples of adjacent blocks of the same current image.
[0092] The intra prediction unit 254 (or generally the mode selection unit 260) is further configured to output an intra prediction parameter (or generally, information indicating the intra prediction mode selected for a block) in the form of a syntax element 266 to the entropy encoding unit 270 to be included in the encoded image data 21, whereby, for example, the video decoder 30 may receive and use the prediction parameter for decoding.
[0093] [Inter Prediction] A set of inter prediction modes (or possible inter prediction modes) depends on available reference images (i.e., for example, previously at least partially decoded images stored in the DPB 230) and other inter prediction parameters, for example, whether the entire reference image or only a part of the reference image, for example, a search window area around the area of the current block, was used for searching for the best matching reference block, and / or, for example, whether pixel interpolation, for example, half / semi pel and / or quarter pel interpolation, was applied.
[0094] In addition to the above prediction modes, a skip mode and / or a direct mode may be applied.
[0095] The inter prediction unit 244 may include a motion estimation (ME) unit and a motion compensation (MC) unit (both not shown in FIG. 2). The motion estimation unit may be configured to receive or obtain, for motion estimation, the image block 203 (the current image block 203 of the current image 17) and the decoded image 231, or at least one or a plurality of previously reconstructed blocks, for example, reconstructed blocks of one or a plurality of other / different previously decoded images 231. For example, a video sequence may include the current image and the previously decoded image 231, or in other words, the current image and the previously decoded image 231 are part of a series of images forming the video sequence or can form the sequence.
[0096] The encoder 20 may be configured to select, for example, reference blocks from a plurality of reference blocks of the same or different images among a plurality of other images, and provide the motion estimation unit with a reference image (or reference image index) and / or an offset (spatial offset) between the position (x, y coordinates) of the reference block and the position of the current block as inter prediction parameters. This offset is also called a motion vector (MV).
[0097] The motion compensation unit is configured to obtain, for example, receive inter prediction parameters, and perform inter prediction based on or using the inter prediction parameters to obtain an inter prediction block 265. The motion compensation performed by the motion compensation unit may involve fetching or generating a prediction block based on the motion / block vector determined by motion estimation and, optionally, performing interpolation up to sub-pixel accuracy. Interpolation filtering may generate additional pixel samples from known pixel samples, thus potentially increasing the number of candidate prediction blocks that can be used to code an image block. When receiving a motion vector for the PU of the current image block, the motion compensation unit may locate the prediction block indicated by the motion vector in one of the reference image lists.
[0098] The motion compensation unit can also generate the blocks used by the video decoder 30 and the syntax elements associated with the video slice when decoding the image blocks of the video slice. In addition to or instead of the slice and its respective syntax elements, tile groups and / or tiles, and their respective syntax elements can be generated or used.
[0099] [Entropy Coding] The entropy encoding unit 270 applies, for example, an entropy encoding algorithm or scheme (e.g., variable length coding (VLC) scheme, context adaptive VLC scheme (CAVLC), arithmetic coding scheme, binari zation, context adaptive binary arithmetic coding (CABAC), syntax based context adaptive binary arithmetic coding (SBAC), probability interval partitioning entropy (PIPE) coding, or another entropy encoding method or technique), or bypass (uncompressed), to the quantized coefficients 209, inter prediction parameters, intra prediction parameters, loop filter parameters, and / or other syntax elements, and is configured to obtain encoded image data 21 that can be output via output 272, for example, in the form of an encoded bitstream 21, whereby, for example, the video decoder 30 can receive and use the parameters for decoding. The encoded bitstream 21 can be transmitted to the video decoder 30 or stored in memory for later transmission or retrieval by the video decoder 30.
[0100] Other structural variations of the video encoder 20 can be used to encode a video stream. For example, the non-transform-based encoder 20 can directly quantize the residual signal without a transform processing unit 206 for a particular block or frame. In another implementation, the encoder 20 can have a quantization unit 208 and an inverse quantization unit 210 combined into a single unit.
[0101] [Decoder and Decoding Method] FIG. 3 shows an example of a video decoder 30 configured to implement the techniques of the present disclosure. The video decoder 30 is configured to receive encoded image data 21 (e.g., an encoded bitstream 21), encoded, for example, by the encoder 20, and obtain a decoded image 331. The encoded image data or bitstream includes information for decoding the encoded image data, e.g., data representing image blocks of an encoded video slice (and / or tile group or tile) and associated syntax elements.
[0102] In the example of FIG. 3, the decoder 30 includes an entropy decoding unit 304, an inverse quantization unit 310, an inverse transform processing unit 312, a reconstruction unit 314 (e.g., an adder 314), a loop filter 320, a decoded picture buffer (DPB) 330, a mode application unit 360, an inter prediction unit 344, and an intra prediction unit 354. The inter prediction unit 344 may be or include a motion compensation unit. In some examples, the video decoder 30 may generally execute a decoding coding path that is inverse to the encoding path described in relation to the video encoder 100 from FIG. 2.
[0103] As described with respect to the encoder 20, the inverse quantization unit 210, the inverse transform processing unit 212, the reconstruction unit 214, the loop filter 220, the decoded picture buffer (DPB) 230, the inter prediction unit 344, and the intra prediction unit 354 are also referred to as forming the "built-in decoder" of the video encoder 20. Accordingly, the inverse quantization unit 310 may have the same function as the inverse quantization unit 110, the inverse transform processing unit 312 may have the same function as the inverse transform processing unit 212, the reconstruction unit 314 may have the same function as the reconstruction unit 214, the loop filter 320 may have the same function as the loop filter 220, and the decoded picture buffer 330 may have the same function as the decoded picture buffer 230. Therefore, the description provided for each unit and function of the video 20 encoder is correspondingly applicable to each unit and function of the video decoder 30.
[0104] [Entropy Decoding] Entropy decoding unit 304 parses the bitstream 21 (or generally the encoded image data 21), and for example, performs entropy decoding on the encoded image data 21 to obtain, for example, quantization coefficients 309 and / or decoded coding parameters (not shown in FIG. 3), such as inter prediction parameters (e.g., reference image index and motion vector), intra prediction parameters (e.g., intra prediction mode or index), transform parameters, quantization parameters, loop filter parameters, and / or any or all of other syntax elements. Entropy decoding unit 304 may be configured to apply a decoding algorithm or scheme corresponding to the encoding scheme described for the entropy encoding unit 270 of encoder 20. Entropy decoding unit 304 may be further configured to provide inter prediction parameters, intra prediction parameters, and / or other syntax elements to mode application unit 360 and other parameters to other units of decoder 30. Video decoder 30 may receive syntax elements at the video slice level and / or at the video block level. In addition to or instead of slices and their respective syntax elements, tile groups and / or tiles, and their respective syntax elements may be received and / or used.
[0105] [Inverse quantization] The inverse quantization unit 310 receives the quantization parameter (QP) (or generally information related to inverse quantization) and the quantization coefficients from the encoded image data 21 (e.g., by the entropy decoding unit 304, e.g., by parsing and / or decoding), and is configured to apply inverse quantization to the decoded quantization coefficients 309 based on the quantization parameter to obtain the dequantization coefficients 311, which may also be referred to as transform coefficients 311. The inverse quantization process may include using the quantization parameter determined by the video encoder 20 for each video block within a video slice (or tile or tile group) to determine the degree of quantization and, similarly, the degree of inverse quantization to be applied.
[0106] [Inverse Transformation] The inverse transformation processing unit 312 receives the dequantization coefficients 311, which may also be referred to as transform coefficients 311, and is configured to apply a transformation to the dequantization coefficients 311 to obtain the reconstructed residual block 213 in the sample region. The reconstructed residual block 213 may also be referred to as the transform block 313. The transformation may be an inverse transformation, e.g., inverse DCT, inverse DST, inverse integer transformation, or a conceptually similar inverse transformation process. The inverse transformation processing unit 312 further receives the transformation parameter or corresponding information from the encoded image data 21 (e.g., by the entropy decoding unit 304, e.g., by parsing and / or decoding) and is further configured to determine the transformation to be applied to the dequantization coefficients 311.
[0107] [Reconstruction] The reconstruction unit 314 (e.g., adder or combiner 314) adds the reconstructed residual block 313 to the prediction block 365 to obtain the reconstructed block 315 in the sample region, e.g., by adding the sample values of the reconstructed residual block 313 and the sample values of the prediction block 365.
[0108] [Filtering] The loop filter unit 320 (which is either within the coding loop or after the coding loop) is configured to filter the reconstructed block 315, for example, to smooth pixel transitions or otherwise improve video quality, to obtain a filtered block 321. The loop filter unit 320 may include one or more loop filters, such as a deblocking filter, a sample adaptive offset (SAO) filter, or one or more other filters, for example, a bilateral filter, an adaptive loop filter (ALF), a sharpening, a smoothing filter, or a collaborative filter, or any combination thereof. Although the loop filter unit 320 is shown in FIG. 3 as being within the loop filter, in other configurations, the loop filter unit 320 may be implemented as a post-loop filter.
[0109] [Decoded Image Buffer] The decoded video block 321 of the image is then stored in the decoded image buffer 330, and the decoded image buffer 330 stores the decoded image 331 as a reference image for subsequent motion compensation of other images and / or for outputting each display.
[0110] The decoder 30 is configured to output the decoded image 331 for presentation or viewing by the user, for example, via output 332.
[0111] [Prediction] The inter prediction unit 344 may be identical to the inter prediction unit 244 (in particular, the motion compensation unit), the intra prediction unit 354 may be functionally identical to the intra prediction unit 254, and based on the partitioned and / or predicted parameters or respective information received (e.g., by the entropy decoding unit 304, e.g., by parsing and / or decoding) from the encoded image data 21, it performs the determination and prediction of partitioning or segmentation. The mode application unit 360 may be configured to perform prediction (intra or inter prediction) for each block based on the reconstructed image, block, or respective samples (filtered or unfiltered), and obtain the prediction block 365.
[0112] When a video slice is coded as an intra-coding (I) slice, the intra prediction unit 354 of the mode application unit 360 is configured to generate a prediction block 365 for an image block of the current video slice based on the signaled intra prediction mode and data from previously decoded blocks of the current picture. When the video picture is coded as an inter-coded (i.e., B or P) slice, the inter prediction unit 344 (e.g., motion compensation unit) of the mode application unit 360 is configured to generate a prediction block 365 for a video block of the current video slice based on motion vectors and other syntax elements received from the entropy decoding unit 304. In inter prediction, the prediction block may be generated from one of a plurality of reference pictures included within one of a plurality of reference picture lists. The video decoder 30 may construct the reference frame lists, list 0 and list 1, based on the reference pictures stored in the DPB 330 using default construction techniques. The same or similar may additionally or alternatively apply to embodiments that use tile groups (e.g., video tile groups) and / or tiles (e.g., video tiles) for slices (e.g., video slices). For example, the video may be coded using I, P, or B tile groups and / or tiles.
[0113] The mode application unit 360 is configured to determine prediction information for a video block of a current video slice by parsing motion vectors or related information and other syntax elements, and uses the prediction information to generate a prediction block for the currently decoded video block. For example, the mode application unit 360 uses some of the received syntax elements to determine the prediction mode (e.g., intra or inter prediction) used to code the video block of the video slice, the inter prediction slice type (e.g., B slice, P slice, or GPB slice), the construction information regarding one or more of the reference picture lists for the slice, the motion vector for each inter-coded video block of the slice, the inter prediction status for each inter-coded video block of the slice, and other information for decoding the video block within the current video slice. The same or similar applies to embodiments that additionally or alternatively use tile groups (e.g., video tile groups) and / or tiles (e.g., video tiles) for slices (e.g., video slices). For example, the video may be coded using I, P, or B tile groups and / or tiles.
[0114] The embodiment of the video decoder 30 shown in FIG. 3 may be configured to segment and / or decode an image by using slices (also referred to as video slices), where the image may be segmented or decoded using one or more slices (usually without overlap), and each slice may include one or more blocks (e.g., CTUs).
[0115] The embodiment of the video decoder 30 shown in FIG. 3 may be configured to segment and / or decode an image by using tile groups (also referred to as video tile groups) and / or tiles (also referred to as video tiles), the image may be segmented or decoded using one or more tile groups (usually non-overlapping), each tile group may include, for example, one or more blocks (e.g., CTUs) or one or more tiles, each tile may be, for example, rectangular in shape, and may include one or more blocks (e.g., CTUs), such as complete blocks or partial blocks.
[0116] Other variations of the video decoder 30 can be used to decode the encoded image data 21. For example, the decoder 30 can generate an output video stream without the loop filtering unit 320. For example, the non-transform-based decoder 30 can directly inverse quantize the residual signal without the inverse transform processing unit 312 for a particular block or frame. In another implementation, the video decoder 30 can have an inverse quantization unit 310 and an inverse transform processing unit 312 combined in a single unit.
[0117] It should be understood that in the encoder 20 and the decoder 30, the processing result of the current stage may be further processed and then output to the next stage. For example, after interpolation filtering, motion vector derivation, or loop filtering, further operations such as clipping or shifting may be performed on the processing result of interpolation filtering, motion vector derivation, or loop filtering.
[0118] It should be noted that further operations may be applied to the derived motion vectors of the current block (including, but not limited to, affine mode control point motion vectors, affine, planar, sub-block motion vectors in the ATMVP mode, temporal motion vectors, etc.). For example, the values of the motion vectors are restricted to a predefined range according to their representation bits. When the representation bits of the motion vector are bitDepth, the range is -2^(bitDepth - 1) to 2^(bitDepth - 1) - 1, where "^" means exponentiation. For example, when bitDepth is set equal to 16, the range is -32768 to 32767, and when bitDepth is set equal to 18, the range is -131072 to 131071. For example, the values of the derived motion vectors (e.g., the MVs of 4 four-by-four sub-blocks within one eight-by-eight block) are restricted such that the maximum difference between the integer parts of the 4 four-by-four sub-block MVs is N pixels or less, such as 1 pixel or less. Two methods for restricting the motion vectors according to bitDepth are provided below.
[0119] Method 1: Remove the overflow MSB (most significant bit) by the following operation.
Number
[0120] For example, after applying equations (1) and (2), if the value of mvx is -32769, the resulting value is 32767. In a computer system, decimal numbers are stored as two's complement.
[0121] The two's complement of -32769 is 1,0111,1111,1111,1111 (17 bits), and then the MSB is discarded, so the resulting two's complement is 0111,1111,1111,1111 (decimal 32767). This is the same as the output by applying equations (1) and (2).
[0122] [Number] The operation may be applied during the summation of mvp and mvd as shown in equations (5)-(8).
[0123] Method 2: Remove the overflow MSB by clipping the value. [Number] Here, vx is the horizontal component of the motion vector of an image block or sub-block, vy is the vertical component of the motion vector of an image block or sub-block, x, y, and z respectively correspond to the three input values of the MV clipping process, and the definition of the function Clip3 is as follows. [Number]
[0124] FIG. 4 is a schematic diagram of a video coding device 400 according to an embodiment of the present disclosure. The video coding device 400 is suitable for implementing the disclosed embodiments described herein. In an embodiment, the video coding device 400 can be a decoder such as the video decoder 30 of FIG. 1A, or an encoder such as the video encoder 20 of FIG. 1A.
[0125] The video coding device 400 includes an inlet port 410 (or input port 410) and a receiver unit (Rx) 420 for receiving data, a processor, logic unit, or central processing unit (CPU) 430 for processing data, a transmitter unit (Tx) 440 and an outlet port 450 (or output port 450) for transmitting data, and a memory 460 for storing data. The video coding device 400 may also include optical / electrical (OE) components and electrical / optical (EO) components connected to the inlet port 410, the receiver unit 420, the transmitter unit 440, and the outlet port 450 for the outlet or inlet of optical or electrical signals.
[0126] The processor 430 is implemented by hardware and software. The processor 430 may be implemented as one or more CPU chips, cores (e.g., multi-core processors), FPGAs, ASICs, and DSPs. The processor 430 communicates with the inlet port 410, the receiver unit 420, the transmitter unit 440, the outlet port 450, and the memory 460. The processor 430 includes a coding module 470. The coding module 470 implements the disclosed embodiments described above. For example, the coding module 470 implements, processes, prepares, or provides various coding operations. Thus, by including the coding module 470, a substantial improvement to the functionality of the video coding device 400 is provided, resulting in a conversion of the video coding device 400 to different states. Alternatively, the coding module 470 is implemented as instructions stored in the memory 460 and executed by the processor 430.
[0127] Memory 460 may include one or more disks, tape drives, and solid state drives, and may be used as an overflow data storage device to store the program related when the program is selected for execution and to store instructions and data read during the execution of the program. Memory 460 may be, for example, volatile and / or non-volatile, and may be read only memory (ROM), random access memory (RAM), ternary content addressable memory (TCAM), and / or static random access memory (SRAM).
[0128] FIG. 5 is a simplified block diagram of an apparatus 500 that may be used as either or both of the source device 12 and the destination device 14 from FIG. 1, according to an exemplary embodiment.
[0129] The processor 502 in the apparatus 500 can be a central processing unit. Alternatively, the processor 502 can be any other type of device or devices capable of operating on or processing information, existing currently or developed in the future. The disclosed implementation can be implemented using a single processor as shown, e.g., processor 502, but advantages in terms of speed and efficiency can be realized using more than one processor.
[0130] In one implementation, the memory 504 in the device 500 can be a read-only memory (ROM) device or a random access memory (RAM) device. Any other suitable type of storage device can be used as the memory 504. The memory 504 can include code and data 506 that are accessed by the processor 502 using the bus 512. The memory 504 can further include an operating system 508 and an application program 510, and the application program 510 includes at least one program that enables the processor 502 to execute the methods described herein. For example, the application program 510 can include applications 1 to N, and the applications 1 to N further include a video coding application that executes the methods described herein.
[0131] The device 500 can also include one or more output devices, such as a display 518. In one example, the display 518 can be a touch sensor type display that combines a display and a touch sensor element operable to detect touch input. The display 518 can be coupled to the processor 502 via the bus 512.
[0132] Although shown herein as a single bus, the bus 512 of the device 500 can be composed of multiple buses. Further, the secondary storage 514 can be directly connected to other components of the device 500 or accessed via a network, and can include a single integrated unit such as a memory card, or multiple units such as multiple memory cards. Therefore, the device 500 can be implemented in a variety of configurations.
[0133] [Parameter set] Parameter sets are basically similar and share the same basic design goals, namely bitrate efficiency, error resilience, and provision of a system layer interface. HEVC (H.265) has a hierarchy of parameter sets including Video Parameter Sets (VPS), Sequence Parameter Sets (SPS), and Picture Parameter Sets (PPS), which are similar to their counterparts in AVC and VVC. Each slice refers to a single active PPS, SPS, and VPS in order to access the information used to decode the slice. Since the PPS contains information applicable to all slices within a picture, all slices within a picture must refer to the same PPS. Slices in different pictures can also refer to the same PPS. Similarly, the SPS contains information applicable to all pictures in the same coded video sequence.
[0134] The PPS can be different for each individual picture, but it is common for many or all pictures in a coded video sequence to refer to the same PPS. Reusing parameter sets makes the bitrate efficient by avoiding the need to transmit shared information multiple times. It also makes it resilient to loss by enabling the content of the parameter sets to be carried over some more reliable external communication link or repeated frequently within the bitstream to ensure that it is not lost.
[0135] [Parameter Set] Parameter sets are basically similar and share the same basic design goals (i.e., bitrate efficiency, error resilience, and provision of system layer interfaces). HEVC (H.265) has a hierarchy of parameter sets including Video Parameter Sets (VPS), Sequence Parameter Sets (SPS), and Picture Parameter Sets (PPS), which are similar to their counterparts in AVC and VVC. Each slice refers to a single active PPS, SPS, and VPS to access the information used to decode the slice. Since the PPS contains information applicable to all slices within a picture, all slices within a picture must refer to the same PPS. Slices in different pictures can also refer to the same PPS. Similarly, the SPS contains information applicable to all pictures in the same coded video sequence.
[0136] The PPS can be different for each individual picture, but it is common for many or all pictures in a coded video sequence to refer to the same PPS. Reusing parameter sets makes the bitrate efficient by avoiding the need to transmit shared information multiple times. It also makes it resilient to loss as the content of the parameter sets can be carried by some more reliable external communication links or repeated frequently within the bitstream to ensure it is not lost.
[0137] [Sequence Parameter Set (SPS)] The SPS contains parameters applicable to one or more layers of a coded video sequence and does not change from picture to picture within the coded video sequence. Specifically, the SPS contains signaling information for sub - pictures.
[0138] Some parts of the following table show a snapshot of the sub - picture signaling of the SPS in ITU JVET - Q2001 - v11, and the download link is as follows. http: / / phenix.it-sudparis.eu / jvet / doc_end_user / documents / 17_Brussels / wg11 / JVET-Q2001-v11.zip. In the remainder of this application, this prior art document is referred to as VVC Draft8 for the sake of brevity.
[0139]
Table 1
Table 2
[0140] Some syntax elements in the SPS signal signal the position information and control flags of each sub - picture. The position information of the i - th sub - picture is · subpic_ctu_top_left_x[i] indicating the horizontal component of the top - left coordinate of sub - picture i within the image, or · subpic_ctu_top_left_y[i] indicating the vertical component of the top - left coordinate of sub - picture i within the image, or · subpic_width_minus1[i] indicating the width of sub - picture i within the image, or · subpic_height_minus1[i] indicating the height of sub - picture i within the image and includes.
[0141] Some syntax elements indicate the number of sub - pictures within the image, e.g., sps_num_subpics_minus1.
[0142] An image is divided into one or more rows of tiles and one or more columns of tiles. A tile is a series of CTUs that cover a rectangular region of the image. The CTUs within a tile are scanned in raster - scan order within that tile.
[0143] A slice is composed of an integer number of complete tiles, or an integer number of consecutive complete CTU rows within a tile of the image. Thus, each vertical slice boundary is always also a vertical tile boundary. It is also possible for the horizontal boundary of a slice to be composed of horizontal CTU boundaries within a tile but not a tile boundary: this occurs when a tile is divided into a plurality of rectangular slices, each of which is composed of an integer number of consecutive complete CTU rows within the tile.
[0144] Two modes of slices are supported, namely, the raster scan slice mode and the rectangular slice mode. In the raster scan slice mode, a slice contains a series of complete tiles in the tile raster scan of the image. In the rectangular slice mode, a slice contains either a number of complete tiles that collectively form a rectangular region of the image, or a number of consecutive complete CTU rows of one tile that collectively form a rectangular region of the image. The tiles within a rectangular slice are scanned in tile raster scan order within the rectangular region corresponding to that slice.
[0145] A sub-image contains one or more slices that collectively cover a rectangular region of the image. Thus, each sub-image boundary is always also a slice boundary, and each vertical sub-image boundary is always also a vertical tile boundary.
[0146] Assume that one or both of the following conditions are achieved for each sub-image and tile. - All CTUs within a sub-image belong to the same tile. - All CTUs within a tile belong to the same sub-image.
[0147] [Partitioning an image into CTUs, slices, tiles, and sub-images]
[0148] [Partitioning an image into CTUs] The picture is divided into a series of coding tree units (CTUs). The term CTB (Coding Tree Block) is sometimes used interchangeably. The concept of CTU is the same as that of HEVC. In the case of a picture with three sample arrays, a CTU is composed of an N×N block of luma samples together with two corresponding blocks of chroma samples. FIG. 6 shows an example of a picture divided into CTUs. The size of CTUs inside the frame must be the same, except at the picture boundary where there may be incomplete CTUs.
[0149] [Partitioning the picture into tiles] When tiles are enabled, the picture is divided into rectangular-shaped groups of CTUs separated by vertical and / or horizontal boundaries. The vertical and horizontal tile boundaries cross the picture from the bottom and at the bottom, and from the left picture boundary to the right picture boundary, respectively. The bitstream includes indications regarding the positions of the horizontal and vertical tile boundaries.
[0150] FIG. 7 illustrates partitioning the picture into nine tiles. In the example, the tile boundaries are shown as thick dashed lines. In other words, FIG. 7 shows the tile-based raster scan order of CTUs having nine tiles of different sizes within the picture. Note that the tile boundaries are shown as thick dashed lines.
[0151] If more than one tile exists inside the picture, the scan order of CTUs is changed. CTUs are scanned according to the following rules. 1. Tiles are scanned from left to right and from top to bottom in a raster scan order called the tile scan order in this disclosure. This means starting from the top-left tile, first all tiles in the same tile row are scanned from left to right. Then starting from the first tile in the second tile row (the tile row one below), all tiles are scanned from left to right in the second tile row. The process is repeated until all tiles are scanned. 2. Inside a tile, CTUs are scanned in raster scan order. Inside a CTU row, the CTUs are scanned from left to right, and the CTU rows are scanned from top to bottom. FIG. 7 illustrates the scanning order of CTUs when a tile exists, and the numbers inside the CTUs indicate the scanning order.
[0152] The concept of a tile provides a way to partition an image such that each tile is independently decodable from other tiles of the same image, where decoding refers to entropy, residual, and predictive decoding. Further, using tiles makes it possible to partition an image into regions having similar sizes. Thus, it is possible for the tiles of an image to be processed in parallel with each other, which is favorable for a multi-core processing environment where each processing core is identical to each other.
[0153] Terms such as processing order and scanning order are used in this disclosure as follows.
[0154] Processing refers to the encoding or decoding of CTUs performed within an encoder or decoder. Scanning order indicates the indexing of a particular partition within an image. The CTU scanning order in a tile means how the CTUs inside the tile are indexed, which may not be in the same order as the order in which the CTUs are processed.
[0155] [Partitioning an image into slices] The concept of a slice provides a partitioning of an image such that each slice is independently decodable from other slices of the same image, where decoding refers to entropy, residual, and predictive decoding. The difference from a tile is that a slice can have a more arbitrary shape (more flexible in terms of partitioning possibilities), and the purpose of slice partitioning is not parallel processing but packet size matching in a transmission environment and error resilience.
[0156] A slice can consist of a complete image as well as its parts. In HEVC, a slice contains a plurality of consecutive CTUs of an image in the processing order. A slice is identified by starting with the CTU address that is signaled in itself in the slice header or the picture parameter set or some other unit.
[0157] In Draft 8 of VVC, a slice contains an integer number of complete tiles, or an integer number of consecutive CTU rows within a tile of an image. Thus, each vertical slice boundary is always also a vertical tile boundary. It is also possible that the horizontal boundary of a slice is not a tile boundary but includes horizontal CTU boundaries within a tile: this occurs when a tile is divided into a plurality of rectangular slices, each of which contains an integer number of consecutive complete CTU rows within the tile.
[0158] In some examples, there are two slice modes, such as the raster scan slice mode and the rectangular slice mode. In the raster scan slice mode, a slice contains a series of tiles in the tile raster scan of an image. In the rectangular slice mode, a slice contains a number of tiles that collectively form a rectangular region of an image, or a slice contains a number of consecutive CTU rows of one tile that collectively form a rectangular region of an image. The tiles within a rectangular slice are scanned in tile raster scan order within the rectangular region corresponding to that slice.
[0159] All slices of an image collectively form the entire image, i.e., all CTUs of an image must be included in one of the plurality of slices of the image. Similar rules apply to tiles and sub-images.
[0160] [Dividing an image into sub-images] A sub - image is a rectangular partition of an image. A sub - image can be the whole image or a part of the image. A sub - image is a partition of an image such that each sub - image is independently decodable from other sub - images of the entire video sequence. In VVC Draft8, this is true when the sub - image is indicated in the bit - stream. That is, if the indication of subpic_treated_as_pic_flag[i] is true for sub - image i, then that sub - image i is independently decodable from other sub - images of the entire video sequence.
[0161] The difference between a sub - image and a tile or a slice is that a sub - image creates an independently decodable video sequence within the video sequence. On the other hand, in the case of tiles and slices, independent decoding is guaranteed only within a single image of the video sequence.
[0162] In VVC Draft8, a sub - image contains one or more slices that collectively cover the rectangular area of the image. Thus, each sub - image boundary is always a slice boundary, and each vertical sub - image boundary is always a vertical tile boundary.
[0163] Figure 8 provides examples of tiles, slices, and sub - images. In other words, Figure 8 shows an example of an image that contains four tiles, i.e., two tile columns and two tile rows, four rectangular slices, and three sub - images. Sub - image 1 contains two slices.
[0164] In the example shown in Figure 8, the image is partitioned into 216 CTUs, four tiles, four slices, and three sub - images. The value of sps_num_subpics_minus1 is 2, and the syntax elements related to the positions have the following values.
[0165] For sub - image 0, · subpic_ctu_top_left_x[0] is not signaled but is inferred to be 0. · subpic_ctu_top_left_y[0] is not signaled but is inferred to be 0. · The value of subpic_width_minus1[0] is 8. · The value of subpic_height_minus1[0] is 11.
[0166] For sub - picture 1, · The value of subpic_ctu_top_left_x[1] is 9. · The value of subpic_ctu_top_left_y[1] is 0. · The value of subpic_width_minus1[1] is 8. · The value of subpic_height_minus1[1] is 5.
[0167] For sub - picture 2, · The value of subpic_ctu_top_left_x[2] is 9. · The value of subpic_ctu_top_left_y[2] is 6. · subpic_width_minus1[2] is not signaled but is inferred to be 8. · subpic_height_minus1[2] is not signaled but is inferred to be 5.
[0168] [Tile Signaling] The following table illustrates the signaling of tile sizes and the coordinates of tiles within the picture (from the picture parameter set RBSP syntax table of VVC Draft8).
Table 3
[0169] The tile segmentation information (address and dimensions of each tile) is usually included in the parameter set. In the above example, first, the indication is included in the bitstream (no_pic_partition_flag), which indicates whether the image is segmented into slices and tiles. If this indication is true (meaning the image is not segmented into slices or tiles), it is inferred that the image is segmented into only one slice and one tile, and its boundary is aligned with the image boundary. Otherwise (when no_pic_partition_flag is false), the tile segmentation information is included in the bitstream.
[0170] The syntax element tile_column_width_minus1[i] indicates the width of the i'-th tile column. The syntax element tile_row_height_minus1[i] indicates the height of the i'-th tile row.
[0171] Both the height of the tile row and the width of the tile column can be either explicitly signaled or inferred within the bitstream. The syntax elements num_exp_tile_columns_minus1 and num_exp_tile_rows_minus1 indicate the number of tile columns and tile rows, respectively, whose width and height are explicitly signaled. The widths and heights of the remaining tile columns and rows are inferred according to a function.
[0172] The indexing of tiles follows the "tile scan order in the image". The tiles in the image are ordered (scanned) in raster scan order. The first tile at the upper left corner of the image is tile 0, and the index increases from left to right in each tile row. After the last tile in the tile row is scanned, it continues with the tile at the left end of the next tile row (one row below the current tile row).
[0173] [Slice Signaling] The following table illustrates the signaling of tile sizes and the coordinates of rectangular-shaped slices within an image (from the VVC Draft8 picture parameter set RBSP syntax table). [Table 4]
[0174] In VVC Draft8, the following relationship exists between slices and tiles. Either a slice contains one or more complete tiles, or a tile contains one or more complete slices. Therefore, the coordinates and sizes of slices are indicated w.r.t tile partitioning. In VVC Draft8, first, tile partitioning is signaled within the picture parameter set. The slice partitioning information is then signaled using the tile mapping information.
[0175] In the above table, the syntax element num_slices_in_pic_minus1 indicates the number of slices within the image. Tile_idx_delta[i] indicates the difference between the tile indices of the first tiles of the (i + 1)-th and i-th slices. For example, the index of the first tile of the first slice within the image is 0. If the tile index of the first tile of the second slice within the image is 5, then Tile_idx_delta[0] is equal to 5. In this context, the tile index is used as the address of the slice, i.e., the index of the first tile of the slice is the start address of the slice.
[0176] slice_width_in_tiles_minus1[i] and slice_height_in_tiles_minus1[i] indicate the width and height of the i-th slice within the image in terms of the number of tiles.
[0177] In the above table, when both slice_width_in_tiles_minus1[i] and slice_height_in_tiles_minus1[i] are equal to 0 (indicating that the maximum dimensions of the i-th slice are 1 tile in height and 1 tile in width), the syntax element num_exp_slices_in_tile[i] can be included in the bitstream. This syntax element indicates the number of slices within a tile.
[0178] As described above, according to VVC Draft8, a slice can contain multiple complete tiles, or a tile can contain multiple complete slices, and other alternatives are prohibited. According to the above syntax table, first, the number of tiles within a slice is indicated (by including slice_width_in_tiles_minus1[i] and slice_height_in_tiles_minus1[i]). In addition, when the number of tiles within a slice is equal to 1 according to the indication, the number of slices within the tile is indicated (by num_exp_slices_in_tile[i]). Therefore, when both slice_width_in_tiles_minus1[i] and slice_height_in_tiles_minus1[i] are equal to 1, the actual size of the slice may be equal to or smaller than 1 tile.
[0179] The syntax element single_slice_per_subpic_flag, when true, indicates the existence of a slice and that there is only one slice per subpicture for all subpictures of the slice (i.e., a subpicture cannot be divided into more than one slice).
[0180] According to one alternative signaling method, the slice map (slice start address and slice size) is shown in VVC Draft8 according to the following steps. 1. First, a tile partition map is indicated within the bitstream, where an index (which may be called tileIdx) is used to index all tiles within the image (in accordance with the tile scan order in the image). After this stage, the index, coordinates, and size of each tile are known. 2. The number of slices within the image is signaled. In one example, the number of slices can be indicated by the num_slices_in_pic_minus1 syntax element. 3. For the first slice within the image, only the width and height of the slice are indicated in terms of the number of tiles. The start address of the first slice is not explicitly signaled, but rather is inferred to be tileIdx 0 (where the first tile within the image is the first tile within the first slice of the image). 4. If the size of the first slice is equal to one tile in width and one tile in height, and if there are more than one CTU rows within the tile included in the first slice, the num_exp_slices_in_tile[0] syntax element is signaled, which indicates how many slices (called numSlicesInTile[0]) are included within the tile. 5. For each of the second slice to the last slice within the image (including the second slice but excluding the last slice), the width and height of the slice are explicitly indicated in terms of the number of tiles. The start address of the slice can be explicitly indicated by the tile_idx_delta[i] syntax element, where i is the index of the slice. If the start address is not explicitly signaled (e.g., if the slices are signaled in an order that allows the start position of the next slice to be inferred using the start position and the width and height of the current slice), then the start address of the slice is inferred via a function. 6. If the size of the n-th slice (where n is a number between 2 and the number of slices in the image minus 1) is equal to one tile in width and one tile in height, and there is more than one CTU row in the tile included in the first slice, the num_exp_slices_in_tile[n] syntax element is signaled, which indicates how many slices are included in the tile. 7. For the last slice in the image, the width and height of the slice are not explicitly signaled, but are inferred according to the number of tiles in the width of the image, the number of tiles in the height of the image, and the start address of the last slice. The start address of the last slice can be explicitly indicated or inferred. The inference of the width and height of the last slice in the image can be performed according to the following two equations, which are from Section 6.5.1 of VVC Draft8. [Number]
[0181] As can be seen from the steps described above, the width and height of the last slice are not signaled. Since it can be easily inferred when the start address of the slice is known, it is desirable not to include the width and height of the last slice in the bitstream. As a result, efficient compression is achieved by not including redundant information in the bitstream.
[0182] The variables tileX, tileY, NumTileColumns, and NumTileRows in the above equations will be described later.
[0183] [Section 6.5.1 of VVC Draft8] [6.5.1 Process of CTB raster scan, tile scan and sub-image scan] For rectangular slices, the list NumCtusInSlice[i] specifying the number of CTUs in the i-th slice for i in the range from 0 to num_slices_in_pic_minus1 (inclusive), the list SliceTopLeftTileIdx[i] specifying the index of the top-left tile of the slice for i in the range from 0 to num_slices_in_pic_minus1 (inclusive), and the matrix CtbAddrInSlice[i][j] specifying the picture raster scan address of the j-th CTB in the i-th slice for i in the range from 0 to num_slices_in_pic_minus1 (inclusive) and j in the range from 0 to NumCtusInSlice[i]-1 (inclusive) are derived as follows.
Number
Number
Number
[0184] Again, for completeness, Recommendation ITU-T H.266 (ISO / IEC 23090-3:2020), cited via http: / / handle.itu.int / 11.1002 / 1000 / 14336 on August 29, 2020, also cites substantially the same thing, the content of which is as follows.
[0185] When rect_slice_flag is equal to 1, the list NumCtusInSlice[i] specifying the number of CTUs in the i-th slice for i in the range from 0 to num_slices_in_pic_minus1 (inclusive), the list SliceTopLeftTileIdx[i] specifying the tile index of the tile containing the first CTU in the slice for i in the range from 0 to num_slices_in_pic_minus1 (inclusive), the matrix CtbAddrInSlice[i][j] specifying the picture raster scan address of the j-th CTB in the i-th slice for i in the range from 0 to num_slices_in_pic_minus1 (inclusive) and j in the range from 0 to NumCtusInSlice[i]-1 (inclusive), and the variable NumSlicesInTile[i] specifying the number of slices in the tile containing the i-th slice, are derived as follows.
Number
Number
Number
Number
Number
[0186] Here, refer to the text shown above in VVC Draft8.
[0187] The above step-by-step description of the signaling of the slice map inside the picture is an example of the signaling in VVC Draft8. More specifically, the description is for the case where rectangular-shaped slices are used, the number of slices per sub-picture is not shown to be equal to 1, there are more than one tile in the picture, and the number of CTU rows inside the tile is greater than 1. When some of the parameters are changed, other modes of the slice map signaling can be used. For example, when it is shown that there is only one slice per sub-picture, the width and height of the slice are not explicitly signaled in the bitstream, but rather are inferred to be equal to the width and height of the corresponding sub-picture.
[0188] Subclause 6.5.1 of VVC Draft8 specifies the scan order of the CTUs inside slice i, where i is the slice index. The output of this subclause, the matrix CtbAddrInSlice[i][n], specifies the scan order of the CTUs inside slice i, where n is the CTU index between 0 and the number of CTUs in slice i. The value of CtbAddrInSlice[i][n] specifies the address of the n-th CTU in slice i (in raster scan order in the picture).
[0189] Figure 9 shows, as an example, the raster scan order of the CTUs in the picture (the "CTU raster scan order in the picture") and one slice (slice 5, i.e., the fifth slice in the picture) inside the picture. In other words, Figure 9 shows the raster scan order of the CTUs inside the picture, where the picture is composed of one tile and one sub-picture.
[0190] According to this example, the values of CtbAddrInSlice are as follows. CtbAddrInSlice[4][0]=27 CtbAddrInSlice[4][1]=28 CtbAddrInSlice[4][2]=29 CtbAddrInSlice[4][3] = 30 CtbAddrInSlice[4][4] = 37 CtbAddrInSlice[4][5] = 38 CtbAddrInSlice[4][6] = 39 CtbAddrInSlice[4][7] = 40
[0191] [Terms used in this disclosure] · "Tile scan order in the image" as described in this disclosure · "CTU scan order inside the tile" as described in this disclosure · "CTU scan order inside the slice" as described in this disclosure · "CTU raster scan order in the image" as described in this disclosure · "Tile-based scan order of CTUs inside the image" · "Scanning order" refers to the indexing in Y of X in ascending order of the index. · "Processing" means decoding or encoding in an encoder or decoder. Thus, the processing order means the order in which X (e.g., CTU) is processed in the encoder or decoder.
[0192] In VVC Draft8, when there are more than one tile per image, the slice signaling is as follows.
[0193] 1. Determine the start tile address of the slice in terms of the number of tiles using explicit indication or inference. 2. For each slice except the last slice, signal how many tiles the slice contains. a. If it is determined that the slice contains only one tile, indicate how many slices are contained within the tile. 3. For the last slice in the image, infer the number of tiles within the slice if it is determined that the slice contains at least one complete tile.
[0194] In other words, in VVC Draft8, if the size of the last slice in the picture is greater than or equal to one tile in both width and height dimensions, the size of the last slice is inferred and not signaled.
[0195] This can be seen from Table 1, where slice_width_in_tiles_minus1[i] and slice_height_in_tiles_minus1[i] (indicating the width and height of the i-th slice in terms of the number of tiles respectively) are included in the bitstream if they are smaller than num_slices_in_pic_minus1 (due to the for-loop "for(i = 0; i < num_slices_in_pic_minus1; i++)"). Therefore, the width and height of the slice are not signaled when i is equal to num_slices_in_pic_minus1, i.e., for the last slice.
[0196] [Luma Mapping with Chroma Scaling (LMCS)] In VVC, a coding tool called Luma Mapping with Chroma Scaling (LMCS) is added as a new processing block before the loop filter. LMCS has two main components: 1) an in-loop mapping of the luma component based on an adaptive piecewise linear model, and 2) for the chroma component, luma-dependent chroma residual scaling is applied. Figure 11 shows the LMCS architecture from the perspective of the decoder. The cyan shaded blocks in Figure 11 indicate where processing is applied to the mapped domain, which includes inverse quantization, inverse transform, luma intra prediction, and adding the luma prediction to the luma residual. The unshaded blocks in Figure 11 indicate where processing is applied to the original (i.e., unmapped) domain, which includes deblocking, loop filters such as ALF and SAO, motion compensation prediction, chroma intra prediction, adding the chroma prediction to the chroma residual, and storing the decoded image as a reference image. The light yellow shaded blocks in Figure 11 are the new LMCS functional blocks, which include the forward and inverse mapping of the luma signal and the luma-dependent chroma scaling process. Similar to most other tools in VVC, LMCS can be enabled / disabled at the sequence level using the SPS flag.
[0197] Slice Header: A coded slice that contains data elements related to all tiles or CTU rows within a tile is represented in the slice.
[0198] [Slice Header]
Table 5
[0199] Table 3 exemplifies a part of the slice header syntax structure of VVC Draft8. The lines containing "…" indicate that some of the rows in the table are omitted.
[0200] In the slice header, the syntax elements are as follows: picture_header_in_slice_header_flag indicates whether the picture header syntax structure is present in the slice header. If the picture header syntax structure is not present in the slice header, it must be included in the picture header that must be included in the bitstream. slice_address indicates the tile index of the first tile of the slice. num_tiles_in_slice_minus1 indicates the number of tiles included in the slice.
[0201] Figure 10 illustrates an image partitioned into 12 tiles and 3 slices. Or, in other words, Figure 10 shows an image having an 18×12 luma CTU partitioned into 12 tiles and 3 raster scan slices.
[0202] In this example shown in Figure 10, the slice address and num_tiles_in_slice_minus1 syntax elements assume the following values for each slice of the image: ·Slice 1 ·slice_address = 0, and the slice start address is tile index 0. ·num_tiles_in_slice_minus1 = 1, and the slice is composed of 2 tiles. ·Slice 2 ·slice_address = 2, and the slice start address is tile index 2. ·num_tiles_in_slice_minus1 = 5, and the slice is composed of 5 tiles. ·Slice 3 ·slice_address = 7, and the slice start address is tile index 7. ·num_tiles_in_slice_minus1 = 4, and the slice is composed of 4 tiles.
[0203] A slice_lmcs_enabled_flag equal to 1 specifies that luma mapping using chroma scaling is enabled for the current slice. A slice_lmcs_enabled_flag equal to 0 specifies that luma mapping using chroma scaling is not enabled for the current slice. If the slice_lmcs_enabled_flag does not exist, it is inferred to be equal to 0.
[0204] The start tile of a slice (the address of the slice within the image) and the number of tiles within the image can be indicated in two ways. If rect_slice_flag is equal to 1 (indicating that the slice of the image has a rectangular shape), the signaling mechanism in Table 1 is used. Table 1 represents part of the picture parameter set. In this mechanism, the addresses and sizes of all slices of the image are signaled within the picture parameter set before the first slice of the image in the bitstream. Note that the bitstream has an order in which information (picture parameter set, slices of the image, and internal syntax elements such as syntax structure) is included in (or parsed from) the bitstream.
[0205] Otherwise, if rect_slice_flag is equal to 0 (indicating that the slice of the image need not have a rectangular shape), the slice_address and num_tiles_in_slice_minus1 syntax elements in the slice header indicate the address and size of the slice.
[0206] [Picture Header] [7.3.2.6 RBSP Syntax of Picture Header] [Table 6]
[0207] The above table presents the picture header syntax for VVC Draft8. It includes the picture header structure and rbsp_trailing_bits( ), which are filler bits that make the number of bits in the picture header equal to a multiple of 8.
[0208] [Picture header structure] [7.3.2.7 Syntax of picture header structure] [Table 7]
[0209] The picture header structure includes syntax elements applicable to all slices of the picture. Some of the syntax elements are included in the picture header structure presented in the above table. As an example, ph_lmcs_enabled_flag indicates whether the LMCS (Luma mapping using chroma scaling) coding tool is enabled for the slices of the picture. A ph_lmcs_enabled_flag equal to 1 specifies that the luma mapping using chroma scaling is enabled for all slices associated with PH. A ph_lmcs_enabled_flag equal to 0 specifies that the luma mapping using chroma scaling may be disabled for one or more or all slices associated with PH. If it does not exist, the value of ph_lmcs_enabled_flag is inferred to be equal to 0.
[0210] As can be understood from the above, the picture header structure can exist either in the slice header or in the picture header. According to VVC Draft8, the picture header structure must exist either in the slice header or in the picture header of the picture. When the picture header structure exists in the picture header, all slices of the picture referring to the picture header must not contain the picture header structure. The reverse is also true. When the picture header structure does not exist in the picture header, that is, when the picture header is not included in the bitstream of a specific picture, the picture header structure must exist in the slice header of the slice of the picture.
[0211] Furthermore, VVC Draft8 has another constraint. Here, when the picture header structure exists in the slice header, the picture must be composed of only one slice (that is, the picture cannot be divided into multiple slices).
[0212] The current VVC Draft8 is not efficient because slice_address and num_tiles_in_slice_minus1 are redundantly included in the bitstream in certain cases. Due to the redundant inclusion of slice_address and num_tiles_in_slice_minus1 in the bitstream, every slice header of the picture can contain this syntax element, thus reducing the compression efficiency and increasing the bitrate.
[0213] [Embodiment 1] According to the embodiment, the existence of the slice_address and num_tiles_in_slice_minus1 syntax elements in the slice header is controlled based on the existence of the picture header structure in the slice header.
Table 8
[0214] The present invention can be implemented as shown in the above table. According to the present invention, slice_address is included in the slice header when the condition on line 6 is true. In other words, slice_address is included in the slice header in the following cases. · The number of tiles in the image is more than 1 and non-rectangular slices are allowed (rect_slice_flag = 0) and the image header structure does not exist in the slice header. Or, · (rect_slice_flag = 1) and the number of slices in the current sub-image is more than 1.
[0215] Otherwise, slice_address is not included in the slice header, and its value can be inferred to be equal to 0.
[0216] In addition or alternatively, the presence of the num_tiles_in_slice_minus1 syntax element in the slice header can be controlled by the presence of the image header structure in the slice header. For example, num_tiles_in_slice_minus1 is not included in the slice header if there is an image header structure in the slice header.
[0217] Line 10 of the above table is · Indicates an implementation of the present invention where Rect_slice_flag is equal to 0, the number of tiles in the image is more than 1, and picture_header_in_slice_header_flag is equal to 0 and num_tiles_in_slice_minus1 is included in the slice header.
[0218] Otherwise, num_tiles_in_slice_minus1 is not included in the slice header, and its value can be inferred to be equal to the number obtained by subtracting 1 from the number of tiles in the image.
[0219] As described above, there is a bitstream compliance requirement in VVC Draft8 that restricts the inclusion of the image header structure in the slice header. According to VVC Draft8, the image header structure can be included in the slice header when there is one slice per image.
[0220] According to the present invention, the presence of the image header structure in the slice header is used to control the presence of slice_address and the number of tiles in the slice indication. This is because when there is a single slice in the image, the slice address must be equal to the first tile in the image, and the number of tiles in the slice must be equal to the number of tiles in the image.
[0221] [Embodiment 2] [Table 9]
[0222] In addition or alternatively, the presence of num_tiles_in_slice_minus1 in the slice header is controlled by the difference between the number of tiles in the image (e.g., NumTilesInPic in the above table) and slice_address.
[0223] More specifically, when the difference between the number of tiles in the image and slice_address is less than the threshold, num_tiles_in_slice_minus1 is not included in the slice header, and its value is inferred to be equal to a predefined number. For example, when the difference between NumTilesInPic and slice_address is less than or equal to 1, num_tiles_in_slice_minus1 is not included in the bitstream, and its value is inferred to be equal to 0 (indicating that there is one tile in the current slice).
[0224] slice_address specifies the slice address of the slice.
[0225] If it does not exist, the value of slice_address is inferred to be equal to 0.
[0226] When rect_slice_flag is equal to 0, the following applies.
[0227] The slice address is the raster scan tile index of the first tile within the slice.
[0228] The length of slice_address is Ceil(Log2(NumTilesInPic)) bits.
[0229] The value of slice_address shall be within the range from 0 to NumTilesInPic - 1 (inclusive).
[0230] Otherwise (rect_slice_flag is equal to 1), the following applies.
[0231] The slice address is the slice index at the sub - picture level of the current slice, i.e., SubpicLevelSliceIdx[j], where j is the slice index at the picture level of the current slice. The length of slice_address is Ceil(Log2(NumSlicesInSubpic[CurrSubpicIdx])) bits. The value of slice_address shall be within the range from 0 to NumSlicesInSubpic[CurrSubpicIdx] - 1 (inclusive).
[0232] It is a bit - stream conformity requirement to which the following constraints apply.
[0233] When rect_slice_flag is equal to 0 or sps_subpic_info_present_flag is equal to 0, the value of slice_address shall be different from the value of slice_address of any other coded slice NAL unit of the same coded picture.
[0234] Otherwise, the pair of values of subpic_id and slice_address shall be different from the pair of values of subpic_id and slice_address of any other coded slice NAL unit of the same coded picture.
[0235] The shape of a slice of a picture shall be such that each CTU, when decoded, has its entire left boundary and its entire upper boundary composed of the picture boundary or the boundary of a previously decoded CTU.
[0236] num_tiles_in_slice_minus1 + 1, if present, specifies the number of tiles in the slice. The value of num_tiles_in_slice_minus1 shall be in the range from 0 to NumTilesInPic - 1, inclusive. If not present, the value of num_tiles_in_slice_minus1 shall be inferred to be 0.
[0237] The variable NumCtusInCurrSlice that specifies the number of CTUs in the current slice, and the list CtbAddrInCurrSlice[i] that specifies the picture raster scan address of the i-th CTB in the slice for i in the range from 0 to NumCtusInCurrSlice - 1, inclusive, are derived as follows.
Number
[0238] The variables SubpicLeftBoundaryPos, SubpicTopBoundaryPos, SubpicRightBoundaryPos, and SubpicBotBoundaryPos are derived as follows.
Number
[0239] [Embodiment 3]
Table 10
[0240] In addition or alternatively, the presence of slice_lmcs_enabled_flag in the slice header is controlled based on the presence of the image header structure in the slice header. An exemplary implementation is included in row 15 of the above table.
[0241] More specifically, if the image header structure is included in the slice header, slice_lmcs_enabled_flag is not included in the slice header. In addition, when not included in the slice header, the value of slice_lmcs_enabled_flag can be inferred according to the following rules. · The value of slice_lmcs_enabled_flag is inferred to be equal to ph_lmcs_enabled_flag.
[0242] Alternatively or in addition, when the value of slice_lmcs_enabled_flag does not exist in the slice header, it can be inferred according to the following rules. · When the value of slice_lmcs_enabled_flag is inferred to be equal to ph_lmcs_enabled_flag if picture_header_in_slice_header_flag is equal to 1 (the image header structure is included in the slice header).
[0243] Alternatively or in addition, if the value of slice_lmcs_enabled_flag does not exist in the slice header, it can be inferred according to the following rules. · If the value of picture_header_in_slice_header_flag is equal to 0, the value of slice_lmcs_enabled_flag is inferred to be equal to 0.
[0244] The above embodiments can be implemented by replacing the conditions "!rect_slice_flag && NumTilesInPic > 1" in lines 6 and 10 with "!rect_slice_flag". In some exemplary implementations, if the value of rect_slice_flag is equal to 0 (indicating that the slices in the picture do not necessarily have to be rectangular), the value of the NumTilesInPic syntax element must be greater than 0 (e.g., the number of tiles in a slice must be greater than 1). In other words, the value of rect_slice_flag can only be equal to 0 when the number of tiles in the picture is greater than 1. In such an implementation, the conditions "!rect_slice_flag && NumTilesInPic > 1" and "!rect_slice_flag" will have the same result. Therefore, the condition (in lines 6 and 10 of all the above embodiments) that includes "!rect_slice_flag && NumTilesInPic > 1" as part of the condition can be replaced with "!rect_slice_flag".
[0245] The above embodiment can be implemented by replacing the conditions “!rect_slice_flag&&NumTilesInPic>1” in lines 6 and 10 with “!rect_slice_flag”. In some exemplary implementation manners, when the value of rect_slice_flag is equal to 0 (indicating that each slice in the image contains one or more tiles), and when picture_header_in_slice_header_flag is equal to 0 (indicating that the number of slices in the image is greater than 1), NumTilesInPic must be greater than 1 when picture_header_in_slice_header_flag is equal to 0 and rect_slice_flag is equal to 0.
[0246] The following is an explanation of the application of the encoding method and decoding method as shown in the above embodiment, and the system using them.
[0247] FIG. 14 is a diagram showing a flowchart of a method for decoding a video bitstream according to an embodiment of the present disclosure. The method shown in FIG. 14 is a method for decoding an image from a video bitstream implemented by a decoding device, the bitstream includes a slice header of a current slice and data representing the current slice, and the method includes a step of obtaining a parameter used to derive the number of tiles in the current slice from the slice header when a condition is satisfied, the condition includes that the slice address of the current slice is not the address of the last tile in the image where the current slice is located (step 1601), and a step of reconstructing the current slice using the number of tiles in the current slice and the data representing the current slice (step 1603).
[0248] FIG. 15 is a diagram showing a flowchart of another encoding method of a video bitstream according to an embodiment of the present disclosure. The method shown in FIG. 15 is a method of encoding a video bitstream implemented by an encoding device, the bitstream including a slice header of a current slice and data representing the current slice, the method including encoding a parameter used to derive the number of tiles in the current slice from the slice header when a condition is satisfied, the condition including that the slice address of the current slice is not the address of the last tile in the image where the current slice is located (step 1701), and reconstructing the current slice using the number of tiles in the current slice and the data representing the current slice (step 1703).
[0249] FIG. 16 is a diagram showing an apparatus for decoding a video bitstream according to an embodiment of the present disclosure, that is, a decoder (30). The apparatus shown in FIG. 16 is an apparatus for decoding an image from a video bitstream, the bitstream including a slice header of a current slice and data representing the current slice, the apparatus including an acquisition unit (3001) configured to acquire a parameter used to derive the number of tiles in the current slice from the slice header when a condition is satisfied, the condition including that the slice address of the current slice is not the address of the last tile in the image where the current slice is located, and a reconstruction unit (3003) configured to reconstruct the current slice using the number of tiles in the current slice and the data representing the current slice.
[0250] FIG. 17 is a diagram showing an apparatus for encoding a video bit stream according to an embodiment of the present disclosure, that is, an encoder (20). The apparatus shown in FIG. 17 is an apparatus for encoding a coded video bit stream, the bit stream including a slice header of a current slice and data representing the current slice, and the apparatus is an encoding unit configured to encode a parameter used to derive the number of tiles in the current slice from the slice header when a condition is satisfied, the condition including that the slice address of the current slice is not the address of the last tile in the image in which the current slice is located, an encoding unit (2001), and a reconstruction unit (2003) configured to reconstruct the current slice using the number of tiles in the current slice and the data representing the current slice, the apparatus (20).
[0251] The video decoding apparatus shown in FIG. 16 may be, or may be constituted by, the decoder 30 shown in FIGS. 1A, 1B and 3 and the video decoder 3206 shown in FIG. 13. Further, the decoding apparatus may be constituted by the video coding device 400 shown in FIG. 4, the apparatus 500 shown in FIG. 5, and the terminal device 3106 shown in FIG. 12. The encoding apparatus shown in FIG. 17 may be, or may be constituted by, the encoder 20 shown in FIGS. 1A, 1B and 3. Further, the encoding apparatus may be constituted by the video coding device 400 shown in FIG. 4, the apparatus 500 shown in FIG. 5, and the capture device 3102 shown in FIG. 12.
[0252] The present disclosure discloses the following further figures. FIG. 12 is a block diagram showing a content supply system 3100 for realizing a content delivery service. This content supply system 3100 includes a capture device 3102 and a terminal device 3106, and optionally includes a display 3126. The capture device 3102 communicates with the terminal device 3106 via a communication link 3104. The communication link may include the communication channel 13 described above. The communication link 3104 includes, but is not limited to, WIFI (registered trademark), Ethernet (registered trademark), cable, wireless (3G / 4G / 5G), USB, or any combination of these types.
[0253] The capture device 3102 can generate data and encode the data by the encoding method as shown in the above embodiments. Alternatively, the capture device 3102 may deliver the data to a streaming server (not shown), and the server encodes the data and transmits the encoded data to the terminal device 3106. The capture device 3102 includes, but is not limited to, a camera, a smartphone or a tablet, a computer or a laptop, a video conferencing system, a PDA, an in-vehicle device, or any combination thereof. For example, as described above, the capture device 3102 may include the source device 12. When the data includes video, the video encoder 20 included in the capture device 3102 may actually perform video encoding processing. When the data includes audio (i.e., voice), the audio encoder included in the capture device 3102 may actually perform audio encoding processing. For some actual scenarios, the capture device 3102 distributes the encoded video and audio data by multiplexing them together. For other actual scenarios, for example, in a video conferencing system, the encoded audio data and the encoded video data are not multiplexed. The capture device 3102 distributes the encoded audio data and the encoded video data separately to the terminal device 3106.
[0254] In the content supply system 3100, the terminal device 310 receives and reproduces the encoded data. The terminal device 3106 may be a device having a data reception and restoration function, for example, a smartphone or a pad 3108 capable of decoding the above-described encoded data, a computer or a laptop 3110, a network video recorder (NVR) / digital video recorder (DVR) 3112, a TV 3114, a set-top box (STB) 3116, a video conferencing system 3118, a video monitoring system 3120, a personal digital assistant (PDA) 3122, an in-vehicle device 3124, or any combination thereof. For example, as described above, the terminal device 3106 may include the destination device 14. When the encoded data includes video, the video decoder 30 included in the terminal device gives priority to performing video decoding. When the encoded data includes audio, the audio decoder included in the terminal device gives priority to performing audio decoding processing.
[0255] For a terminal device having the display, such as a smartphone or a Pad 3108, a computer or a laptop 3110, a network video recorder (NVR) / digital video recorder (DVR) 3112, a TV 3114, a personal digital assistant (PDA) 3122, or an in-vehicle device 3124, the terminal device can supply the decoded data to the display. For a terminal device not equipped with a display, such as an STB 3116, a video conferencing system 3118, or a video monitoring system 3120, an external display 3126 is internally contacted to receive and display the decoded data.
[0256] When each device in the system performs encoding or decoding, an image encoding device or an image decoding device as shown in the above-described embodiment can be used.
[0257] FIG. 13 is a diagram showing the structure of an example of the terminal device 3106. After the terminal device 3106 receives a stream from the capture device 3102, the protocol progress unit 3202 analyzes the transmission protocol of the stream. The protocol includes, but is not limited to, Real-Time Streaming Protocol (RTSP), Hypertext Transfer Protocol (HTTP), HTTP Live Streaming Protocol (HLS), MPEG-DASH, Real-Time Transport Protocol (RTP), Real-Time Messaging Protocol (RTMP), or any combination of these types.
[0258] After the protocol progress unit 3202 processes the stream, a stream file is generated. The file is output to the demultiplexing unit 3204. The demultiplexing unit 3204 can separate the multiplexed data into encoded audio data and encoded video data. As described above, in some actual scenarios, for example, in a video conferencing system, the encoded audio data and the encoded video data are not multiplexed. In this situation, the encoded data is sent to the video decoder 3206 and the audio decoder 3208 without passing through the demultiplexing unit 3204.
[0259] By this demultiplexing process, a video elementary stream (ES), an audio ES, and optionally subtitles are generated. The video decoder 3206 includes the video decoder 30 as described in the above embodiment, decodes the video ES by the decoding method shown in the above embodiment to generate video frames, and supplies this data to the synchronization unit 3212. The audio decoder 3208 decodes the audio ES to generate audio frames and supplies this data to the synchronization unit 3212. Alternatively, the video frames can be stored in a buffer (not shown in FIG. 13) before being supplied to the synchronization unit 3212. Similarly, the audio frames can be stored in a buffer (not shown in FIG. 13) before being supplied to the synchronization unit 3212.
[0260] The synchronization unit 3212 synchronizes video frames and audio frames and supplies video / audio to the video / audio display 3214. For example, the synchronization unit 3212 synchronizes the presentation of video and audio information. The information may be coded in syntax using time stamps related to the presentation of coded audio and visual data, and time stamps related to the delivery of the data stream itself.
[0261] When subtitles are included in the stream, the subtitle decoder 3210 decodes the subtitles and synchronizes them with the video frames and audio frames, and supplies video / audio / subtitles to the video / audio / subtitle display 3216.
[0262] The present invention is not limited to the above-described system, and either the image encoding device or the image decoding device in the above-described embodiment can be incorporated into other systems, for example, a car system.
[0263] [Mathematical operators] The mathematical operators used in this disclosure are the same as those used in the C programming language. However, the results of integer division and arithmetic shift operations are more precisely defined, and additional operations such as exponentiation and real-valued division are defined. The numbering and counting regulations generally start from 0. For example, "the first" is equivalent to the 0th, "the second" is equivalent to the 1st, and so on.
[0264] [Arithmetic operators] The following arithmetic operators are defined as follows. [Table 11]
[0265] [Logical operators] The following logical operators are defined as follows. x && y The Boolean logical "and" of x and y x || y The Boolean logical "or" of x and y ! The Boolean logical "not" x? y : z If x is TRUE or not equal to 0, the value of y is taken, otherwise the value of z is taken.
[0266] [Relational operators] The following relational operators are defined as follows. > Greater than >= Greater than or equal to < Less than <= Less than or equal to == Equal to != Not equal to
[0267] When a relational operator is applied to a syntax element or variable assigned the value "na" (not applicable), the value "na" is treated as a distinct value for that syntax element or variable. The value "na" is considered not equal to any other value.
[0268] [Bitwise operators] The following per-bit operators are defined as follows. & Bitwise "and". When performing an operation on an integer term, the operation is performed on the two's complement representation of the integer value. When performing an operation on a binary term with fewer bits than another term, the shorter term is extended by adding higher-order bits equal to 0. | Bitwise "or". When performing an operation on an integer term, the operation is performed on the two's complement representation of the integer value. When performing an operation on a binary term with fewer bits than another term, the shorter term is extended by adding higher-order bits equal to 0. ^ Bitwise "exclusive or". When performing an operation on an integer term, the operation is performed on the two's complement representation of the integer value. When performing an operation on a binary term with fewer bits than another term, the shorter term is extended by adding higher-order bits equal to 0. x >> y Arithmetic right shift of the two's complement integer representation of x by the number of bits in y. This function is defined only for non - negative integer values of y. The bit shifted into the most significant bit (MSB) as a result of the right shift has the same value as the MSB of x before the shift operation. x << y Arithmetic left shift of the two's complement integer representation of x by the number of bits in y. This function is defined only for non - negative integer values of y. The bit shifted into the least significant bit (LSB) as a result of the left shift has a value equal to 0.
[0269] [Assignment operator] The following arithmetic operators are defined as follows. = Assignment operator ++ Increment, i.e., x++ is equivalent to x = x + 1. When used as an array index, it takes the value of the variable before the increment operation. -- Decrement, i.e., x-- is equivalent to x = x - 1. When used as an array index, it takes the value of the variable before the decrement operation. += Increment by the specified amount, i.e., x += 3 is equivalent to x = x + 3 and x += (-3) is equivalent to x = x + (-3). -= Decrement by the specified amount, i.e., x -= 3 is equivalent to x = x - 3 and x -= (-3) is equivalent to x = x - (-3).
[0270] [Range notation] The following notations are used to specify a range of values. x = y..z x takes integer values from y to z (inclusive), where x, y, and z are integers and z is greater than y.
[0271] [Mathematical functions] The following mathematical functions are defined. [Number] Asin(x) The inverse sine function, performs the operation on an argument x in the range from - 1.0 to 1.0 (inclusive), The output value is in the range from -π÷2 to π÷2 (including both ends) in radians. Atan(x) is the inverse trigonometric tangent function, which performs an operation on the argument x, and the output value is in the range from -π÷2 to π÷2 (including both ends) in radians.
Number
Number
Number
Number
Number
Number
Number
[0272] [Order of operation precedence] If the order of precedence of an expression is not explicitly indicated using parentheses, the following rules apply. - Operations of higher precedence are evaluated before any operations of lower precedence. - Operations of the same precedence are evaluated sequentially from left to right.
[0273] The following table specifies the order of operation precedence from highest to lowest. A higher position in the table indicates a higher precedence.
[0274] For operators also used in the C programming language, the order of precedence used in this specification is the same as that used in the C programming language. Table: Order of operation precedence from highest (top of table) to lowest (bottom of table)
Table 12
[0275] [Text description of logical operations] In text, statements of logical operations are mathematically described in the following form.
Number
Number
[0276] For each statement in the text of "If...Otherwise,if...Otherwise,...", "as follows" or "the following applies" is immediately followed by "If...". The last condition of "If...Otherwise,if...Otherwise" is always "Otherwise,...". Interleaved "If...Otherwise,if...Otherwise,..." statements can be identified by matching "as follows" or "the following applies" that ends with "Otherwise,...".
[0277] In the text, statements of logical operations are mathematically described in the following form. [Number] can be described in the following way. [Number]
[0278] In the text, statements of logical operations are mathematically described in the following form. [Number] can be described in the following way. If condition 0, statement 0 If condition 1, statement 1
[0279] Embodiments of the present invention have been mainly described based on video coding. However, embodiments of the coding system 10, the encoder 20, and the decoder 30 (and the system 10 corresponding thereto) and other embodiments described herein can also be configured for still image processing or coding, i.e., processing or coding of individual images independent of any preceding or successive images such as video coding. It should be noted that generally, when image processing coding is limited to a single image 17, only the inter prediction units 244 (encoder) and 344 (decoder) may not be available. All other functions (also referred to as tools or techniques) of the video encoder 20 and the video decoder 30 can be equally used for still image processing, for example, residual calculation 204 / 304, transformation 206, quantization 208, inverse quantization 210 / 310, (inverse) transformation 212 / 312, segmentation 262 / 362, intra prediction 254 / 354, and / or loop filtering 220, 320, as well as entropy coding 270 and entropy decoding 304.
[0280] For example, the embodiments of the encoder 20 and the decoder 30, and the functions described herein with reference to, for example, the encoder 20 and the decoder 30 may be implemented in hardware, software, firmware, or any combination thereof. When implemented in software, the functions may be stored on a computer-readable medium or transmitted as one or more instructions or codes via a communication medium and executed by a hardware-based processing unit. The computer-readable medium may include a computer-readable storage medium corresponding to a tangible medium such as a data storage medium, or a communication medium including any medium that facilitates the movement of a computer program from one location to another, for example, in accordance with a communication protocol. Thus, the computer-readable medium may generally correspond to (1) a non-transitory tangible computer-readable storage medium, or (2) a communication medium such as a signal or a carrier wave. The data storage medium may be any available medium that can be accessed by one or more computers or one or more processors to obtain instructions, codes, and / or data structures for implementing the techniques described in the present disclosure. A computer program product may include a computer-readable medium.
[0281] By way of example and not limitation, such computer-readable storage media can include RAM, ROM, EEPROM, CD-ROM or other optical disk storage, magnetic disk storage, or other magnetic storage devices, flash memory, or any other medium that can be used to store the desired program code in the form of instructions or data structures and that can be accessed by a computer. Additionally, any connection is suitable to be called a computer-readable medium. For example, when instructions are transmitted from a website, server, or other remote source using coaxial cable, fiber optic cable, twisted pair, digital subscriber line (DSL), or wireless technologies such as infrared, radio, and microwave, coaxial cable, fiber optic cable, twisted pair, DSL, or wireless technologies such as infrared, radio, and microwave are included in the definition of the medium. However, computer-readable storage media and data storage media should be understood to relate to non-transitory, tangible storage media and do not include connections, carrier waves, signals, or other transient media. As used herein, disk and disc include compact disc (CD), laser disc, optical disc, digital versatile disc (DVD), floppy disk, and Blu-ray disc, where disk typically magnetically reproduces data and disc optically reproduces data by laser. The above combinations should also be included within the scope of computer-readable media.
[0282] The commands may be executed by one or more processors, such as one or more digital signal processors (DSPs), general-purpose microprocessors, application-specific integrated circuits (ASICs), field-programmable logic arrays (FPGAs), or other equivalent integrated or discrete logic circuitry. Thus, the term "processor" as used herein may refer to any of the foregoing structures, or any other structure suitable for implementation of the techniques described herein. Additionally, in some aspects, the functions described herein may be provided within dedicated hardware and / or software modules configured for encoding and decoding, or incorporated into a combined codec. Additionally, these techniques may be implemented entirely in one or more circuits or logic elements.
[0283] The techniques of the present disclosure may be implemented in a variety of devices or apparatuses, including wireless handsets, integrated circuits (ICs) or sets of ICs (e.g., chip sets). Although various components, modules, or units are described in the present disclosure to emphasize the functional aspects of devices configured to execute the disclosed techniques, implementation by different hardware units is not necessarily required. Rather, as described above, various units may be combined as codec hardware units in conjunction with suitable software and / or firmware, or provided by a set of interoperable hardware units including one or more processors, as described above.
[0284] The present disclosure discloses the following twenty-one further aspects.
[0285] (Item 1) An aspect of a method for decoding a video or image bitstream implemented by a decoding device, wherein the bitstream includes data representing a current slice, the method comprising: obtaining, conditional on an existence condition being satisfied, the slice address of the current slice from a slice header of the bitstream, wherein the existence condition includes that the picture header syntax structure does not exist in the slice header; and reconstructing the current slice based on the slice address of the current slice.
[0286] (Item 2) The non - existence of the picture header syntax structure in the slice header includes that a syntax element is equal to false, and the syntax element equal to false specifies that the picture header syntax structure does not exist in the slice header, according to the aspect of the method described in Aspect 1.
[0287] (Item 3) The value of the slice address of the current slice is inferred to be equal to zero when the existence condition is not satisfied, according to the aspect of the method described in Aspect 1 or 2.
[0288] (Item 4) An aspect of a method for decoding a video or image bitstream implemented by a decoding device, wherein the bitstream includes data representing a current slice, the method comprising: obtaining, conditional on an existence condition being satisfied, a parameter used to derive the number of tiles of the current slice from a slice header of the bitstream, wherein the existence condition includes that the picture header syntax structure does not exist in the slice header; and reconstructing the current slice based on the number of tiles within the current slice.
[0289] (Item 5) The fact that the image header syntax structure does not exist in the slice header includes that the syntax element is equal to false, and the false-equal syntax element is the aspect of the method according to aspect 4 that specifies that the image header syntax structure does not exist in the slice header.
[0290] (Item 6) The value of the parameter of the current slice is inferred to be equal to the total number of tiles in the image, which is the value obtained by subtracting 1 from the current slice when the existence condition is not satisfied, according to the aspect of the method described in aspect 4 or 5.
[0291] (Item 7) An aspect of a method for decoding a video or image bitstream implemented by a decoding device, the bitstream including data representing a current slice, the method comprising, on condition that an existence condition is satisfied, obtaining a parameter used to derive the number of tiles of the current slice from a slice header of the bitstream, the existence condition including that the slice address of the current slice is not the address of the last tile in the image in which the current slice is located, and reconstructing the current slice based on the number of tiles in the current slice.
[0292] (Item 8) The fact that the slice address of the current slice is the address of the last tile in the image includes that the value obtained by subtracting the slice address of the current slice from the number of tiles in the image is equal to 1, according to the aspect of the method described in aspect 7.
[0293] (Item 9) The value of the parameter of the current slice is inferred to be equal to a default value when the existence condition is not satisfied, according to the aspect of the method described in aspect 7 or 8.
[0294] (Item 10) The aspect of the method according to aspect 9, wherein the default value is equal to 0.
[0295] (Item 11) An aspect of a method for decoding a video or image bitstream implemented by a decoding device, wherein the bitstream includes data representing a current slice, and the method comprises: obtaining a parameter used to derive the number of tiles of the current slice from a slice header of the bitstream, provided that an existence condition is satisfied, wherein the existence condition includes that the slice address of the current slice is not the address of the last tile in the image where the current slice is located, and that the image header syntax structure does not exist in the slice header; and reconstructing the current slice based on the number of tiles in the current slice.
[0296] (Item 12) The aspect of the method according to aspect 11, wherein the value of the parameter is inferred to be equal to a first default value when the slice address of the current slice is the address of the last tile in the image, or is a second default value, and the image header syntax structure does not exist in the slice header.
[0297] (Item 13) An aspect of a method for decoding a video or image bitstream implemented by a decoding device, wherein the bitstream includes data representing a current slice, and the method comprises: obtaining a parameter (such as slice_lmcs_enabled_flag) used to specify whether chroma scaling-based luma mapping is enabled for the current slice from a slice header of the bitstream, provided that an existence condition is satisfied, wherein the existence condition includes that the image header syntax structure does not exist in the slice header; and reconstructing the current slice based on the number of tiles in the parameter.
[0298] (Item 14) The fact that the image header syntax structure does not exist in the slice header includes that the syntax element is equal to false, and the false-equal syntax element is the aspect of the method according to aspect 13 that specifies that the image header syntax structure does not exist in the slice header.
[0299] (Item 15) An aspect of a method for encoding a video or image into a bitstream implemented by an encoding device, wherein the bitstream includes data representing a current slice, and the method includes, on condition that an existence condition is satisfied, including the slice address of the current slice from a slice header of the bitstream into the bitstream, wherein the existence condition includes that the image header syntax structure does not exist in the slice header, and reconstructing the current slice based on the slice address of the current slice.
[0300] (Item 16) An aspect of a decoder (30) comprising a processing circuit for executing the method according to any one of aspects 1 to 15.
[0301] (Item 17) An aspect of a computer program product including program code for executing the method according to any one of the preceding aspects when executed on a computer or a processor.
[0302] (Item 18) An aspect of a decoder, wherein the decoder is One or more processors and a non-transitory computer-readable storage medium coupled to the processor and storing programming for execution by the processor, the programming, when executed by the processor, configuring the decoder to execute the method according to any one of the preceding modes 1 to 15, the mode comprising the non-transitory computer-readable storage medium.
[0303] (Item 19) A mode of a non-transitory computer-readable medium that holds program code that, when executed by a computer device, causes the computer device to execute the method according to any one of the preceding modes 1 to 15.
[0304] (Item 20)
[0305] A mode of an encoded bitstream for the video signal by including a plurality of syntax elements, the plurality of syntax elements including picture_header_in_slice_header_flag, and flags (such as slice_lmcs_enabled_flag) are conditionally signaled within the slice header based at least on the value of picture_header_in_slice_header_flag.
[0306] (Item 21) A mode of a non-transitory recording medium including an encoded bitstream decoded by an image decoding device, the bitstream being generated by dividing a frame of a video signal or an image signal into a plurality of blocks and including a plurality of syntax elements, the plurality of syntax elements including rect_slice_flag or sps_num_subpics_minus1, and flags (such as slice_lmcs_enabled_flag) are conditionally signaled within the slice header based at least on the value of picture_header_in_slice_header_flag. [Other possible items] [Item 1] A method for decoding an image from a video bitstream implemented by a decoding device, wherein the bitstream includes a slice header of a current slice and data representing the current slice, and the method includes: Obtaining a parameter used to derive the number of tiles in the current slice from the slice header when a condition is satisfied, the condition including that the slice address of the current slice is not the address of the last tile in the image where the current slice is located; Reconstructing the current slice using the number of tiles in the current slice and the data representing the current slice; A method comprising: [Item 2] The method according to item 1, wherein the slice address of the current slice being the address of the last tile in the image includes determining that a value obtained by subtracting the slice address of the current slice from the number of tiles in the image is equal to 1. [Item 3] The method according to item 1, wherein the slice address of the current slice not being the address of the last tile in the image includes determining that a value obtained by subtracting the slice address of the current slice from the number of tiles in the image is greater than 1. [Item 4] The method according to any one of items 1 to 3, wherein the value of the parameter of the current slice is inferred to be equal to a default value when the condition is not satisfied. [Item 5] The method according to item 4, wherein the default value is equal to 0. [Item 6] The method according to any one of items 1 to 5, wherein the slice address is in tile units. [Item 7] The method according to any one of items 1 to 6, wherein the condition further includes a step of determining that the current slice is in a raster scan mode. [Item 8] The step of reconstructing the current slice using the number of tiles in the current slice includes a step of determining a scan order of the coding tree units in the current slice using the number of tiles in the current slice, and a step of reconstructing the coding tree units in the current slice using the scan order. The method according to any one of items 1 to 7. [Item 9] A method for encoding a video bitstream implemented by an encoding device, the bitstream including a slice header of a current slice and data representing the current slice, the method comprising: Encoding a parameter used to derive the number of tiles in the current slice from the slice header when a condition is satisfied, the condition including that the slice address of the current slice is not the address of the last tile in the image where the current slice is located; Reconstructing the current slice using the number of tiles in the current slice and the data representing the current slice A method comprising. [Item 10] An apparatus for decoding an image from a video bitstream, the bitstream including a slice header of a current slice and data representing the current slice, the apparatus comprising: An acquisition unit configured to acquire a parameter used to derive the number of tiles in the current slice from the slice header when a condition is satisfied, the condition including that the slice address of the current slice is not the address of the last tile in the image where the current slice is located; A reconstruction unit configured to reconstruct the current slice using the number of tiles in the current slice and the data representing the current slice An apparatus comprising the same. [Item 11] An apparatus for encoding a coded video bitstream, the bitstream including a slice header of a current slice and data representing the current slice, the apparatus comprising: An encoding unit configured to encode a parameter used to derive the number of tiles in the current slice from the slice header when a condition is satisfied, the condition including that the slice address of the current slice is not the address of the last tile in the image where the current slice is located, the encoding unit; A reconstruction unit configured to reconstruct the current slice using the number of tiles in the current slice and the data representing the current slice An apparatus comprising the same. [Item 12] An encoder (20) comprising a processing circuit for performing the method according to item 9. [Item 13] A decoder (30) comprising a processing circuit for performing the method according to any one of items 1 to 8. [Item 14] A computer program product including program code for performing the method according to any one of items 1 to 9 when executed on a computer or a processor. [Item 15] A decoder, comprising: One or more processors; A non-transitory computer-readable storage medium coupled to the processor and storing programming for execution by the processor, the programming configuring the decoder to perform the method according to any one of items 1 to 8 when executed by the processor, the non-transitory computer-readable storage medium A decoder comprising [Item 16] An encoder, One or more processors, A non - transitory computer - readable storage medium coupled to the processor and storing programming for execution by the processor, the programming, when executed by the processor, configuring the encoder to execute the method according to item 9 An encoder comprising [Item 17] A non - transitory computer - readable medium holding program code that, when executed by a computer device, causes the computer device to execute the method according to any one of items 1 to 9 [Item 18] A non - transitory storage medium containing a video bitstream, the bitstream including a slice header of a current slice and data representing the current slice, the slice header including a slice address of the current slice When a condition is satisfied, the slice header further includes a parameter used to derive the number of tiles within the current slice from the slice header, the condition including that the slice address of the current slice is not the address of the last tile in the image in which the current slice is located
Claims
1. A method for decoding an image from a video bitstream, implemented by a decoding device, the video bitstream including a slice header for a current slice and data representing the current slice, the method comprising: obtaining a parameter used to derive a number of tiles in the current slice from the slice header if a condition is met, the condition being that a slice address of the current slice is not the address of the last tile in the image in which the current slice is located, the slice address of the current slice being a raster scan tile index of a first tile in the current slice; reconstructing the current slice using the number of tiles in the current slice and the data representing the current slice; A method comprising:
2. The method described in claim 1, wherein determining that the slice address of the current slice is the address of the last tile in the image includes determining that the number of tiles in the image minus the slice address of the current slice is equal to 1.
3. The method of claim 1, wherein determining that the slice address of the current slice is not the address of the last tile in the image includes determining that a value obtained by subtracting the slice address of the current slice from the number of tiles in the image is greater than 1.
4. A method according to any one of claims 1 to 3, wherein the value of the parameter of the current slice is inferred to be equal to 0 if the condition is not satisfied.
5. A method according to any one of claims 1 to 4, wherein the slice address is in units of tiles.
6. A method, implemented by an encoding device, for encoding a video bitstream, the video bitstream including a slice header for a current slice and data representing the current slice, the method comprising: encoding a parameter used to derive a number of tiles in the current slice from the slice header if a condition is met, the condition being that a slice address of the current slice is not the address of the last tile in the image in which the current slice is located, the slice address of the current slice being a raster scan tile index of a first tile in the current slice; A method comprising:
7. The method described in claim 6, comprising determining that the slice address of the current slice is not the address of the last tile in the image is greater than 1 when the number of tiles in the image minus the slice address of the current slice.
8. An encoder comprising a processing circuit for carrying out the method according to claim 6 or 7.
9. A decoder comprising processing circuitry for carrying out a method according to any one of claims 1 to 5.
10. A computer program causing a computer to carry out a method according to any one of claims 1 to 7.
11. A decoder comprising: one or more processors; a non-transitory computer readable storage medium coupled to the one or more processors and storing programming for execution by the one or more processors, the programming, when executed by the one or more processors, configuring the decoder to perform the method of any one of claims 1 to 5; A decoder comprising:
12. An encoder comprising: one or more processors; a non-transitory computer-readable storage medium coupled to the one or more processors and storing programming for execution by the one or more processors, the programming, when executed by the one or more processors, configuring the encoder to perform the method of claim 6 or 7; An encoder comprising:
13. A device for storing a bitstream, comprising at least one storage medium and at least one communication interface, comprising: the at least one communication interface is configured to receive or transmit the bitstream; the at least one storage medium is configured to store the bitstream; A device, wherein the bitstream includes a slice header for a current slice and data representing the current slice, the slice header having parameters used to derive the number of tiles in the current slice when a condition is met, the condition being that the slice address of the current slice is not the address of the last tile in the image in which the current slice is positioned.
14. The device of claim 13, wherein the slice address of the current slice is a raster scan tile index of a first tile in the current slice.
15. A device for transmitting a bitstream, comprising: at least one storage medium configured to store at least one bitstream, the bitstream including a slice header of a current slice and data describing the current slice, the slice header having parameters used to derive a number of tiles in the current slice if a condition is met, the condition being that a slice address of the current slice is not an address of a last tile in an image in which the current slice is located; at least one processor configured to obtain one or more bitstreams from one of the at least one storage medium and to transmit the one or more bitstreams; A device comprising:
16. The device of claim 15, wherein the slice address of the current slice is a raster scan tile index of a first tile in the current slice.