Method, apparatus, decoder, encoder, and program for cross-component linear modeling for intra prediction
The use of a cross-component linear model for intra predicting chroma samples in video coding optimizes the efficiency of lookup table value retrieval, enhancing compression efficiency and maintaining image quality.
Patent Information
- Application Number
- JP2025067286
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2018-12-31
- Filing Date
- 2025-04-16
- Publication Date
- 2025-07-30
- Estimated Expiration
- 2039-12-30
AI Technical Summary
Existing video coding technologies face challenges in achieving efficient compression of video data without significant loss in image quality, particularly in the context of limited network bandwidth and memory resources.
A method and apparatus for intra predicting chroma samples using a cross-component linear model, involving the determination of maximum and minimum luma sample values, calculating a difference, and fetching values from a lookup table using a bit set as an index to obtain linear model parameters for predicting chroma sample values.
This approach enhances the efficiency of fetching values from the lookup table, minimizing the size of multipliers and lookup table entries, thereby improving the compression ratio with minimal impact on image quality.
Smart Images

Figure 2025111541000001_ABST
Abstract
Description
Technical Field
[0001] [Cross - Reference to Related Applications] This application claims the benefit of U.S. Provisional Application No. 62 / 786,563, filed on December 31, 2018, entitled "Methods and Apparatus for Cross - Component Linear Modeling for Intra Prediction". The said application is incorporated herein by reference. Embodiments of the present application (disclosure) generally relate to the field of image processing, and more specifically, to intra prediction using cross - component linear modeling.
Background Art
[0002] Video coding (video encoding and video decoding) is used in a wide range of digital video applications, such as broadcast digital TV, video transmission via the Internet and mobile networks, real - time conversation applications such as video chat, video conferencing, DVDs and Blu - ray discs, video content acquisition and editing systems, and camcorders for security purposes.
[0003] Even to present relatively short videos, the amount of video data required can be quite substantial, and difficulties can arise when data is streamed or otherwise communicated over a communication network with limited bandwidth capacity. Thus, video data is generally compressed before being communicated over modern telecommunications networks. The size of a video can also be a problem when the video is stored on a storage device, because memory resources can be limited. Video compression devices often use software and / or hardware at the source to code the video data before transmission or storage, thereby reducing the amount of data required to represent a digital video image. The compressed data is then received at the destination by a video decompression device that decodes the video data. Given limited network resources and the continuing increase in demand for higher video quality, improved compression and decompression techniques that improve the compression ratio with little to no sacrifice in image quality are desirable. SUMMARY OF THE INVENTION
[0004] Embodiments of the present application provide an apparatus and method for encoding and decoding as recited in the independent claims.
[0005] The foregoing and other objects are realized by the subject matter of the independent claims. Further implementations become apparent from the dependent claims, the description, and the figures.
[0006] According to a first aspect, the present invention relates to a method for intra predicting chroma samples of a block by applying a cross-component linear model. The method includes obtaining reconstructed luma samples, determining a maximum luma sample value and a minimum luma sample value based on the reconstructed luma samples, obtaining a difference between the maximum luma sample value and the minimum luma sample value, and determining a position of a most significant bit of the difference between the maximum luma sample value and the minimum luma sample value. The method also includes fetching a value from a look-up table (LUT) by using a bit set as an index, where the bit set follows the position of the most significant bit of the difference between the maximum luma sample value and the minimum luma sample value, obtaining linear model parameters α and linear model parameter β based on the fetched value, and calculating a predicted chroma sample value by using the obtained linear model parameters α and linear model parameter β.
[0007] According to a first aspect of the present invention, the index of the LUT is calculated in an elegant way by extracting some bits in a binary representation. As a result, the efficiency of fetching a value from the LUT is increased.
[0008] In a possible implementation of the method according to such a first aspect, the method obtains the linear model parameters α and the linear model parameter β by multiplying the fetched value by a difference between a maximum value and a minimum value of the reconstructed chroma samples.
[0009] Since the efficiency of fetching a value from the LUT is increased, the size of a multiplier for obtaining the linear model parameters α and β is minimized.
[0010] In a possible implementation of the method according to such a first aspect, the LUT includes at least two adjacent values stored in the LUT corresponding to different stages of the obtained difference, and the values of this stage increase or are constant together with the difference value. The index of the LUT is calculated in an elaborate way of extracting some bits in the binary representation, and accordingly, the size of the entry in the LUT corresponding to the index is minimized. As a result, the size of the LUT is minimized.
[0011] An apparatus for intra predicting chroma samples of a block by applying a cross-component linear model is provided according to a second aspect of the present invention. The apparatus according to the second aspect of the present invention includes an acquisition unit, a determination unit, and a calculation unit. The acquisition unit is configured to acquire reconstructed luma samples. The determination unit is configured to determine a maximum luma sample value and a minimum luma sample value based on the reconstructed luma samples. The acquisition unit is further configured to acquire a difference between the maximum luma sample value and the minimum luma sample value. The determination unit is further configured to determine a position of the most significant bit of the difference between the maximum luma sample value and the minimum luma sample value. The calculation unit is configured to fetch a value from a look-up table (LUT) by using a bit set following the position of the most significant bit of the difference between the maximum luma sample value and the minimum luma sample value as an index, acquire linear model parameters α and β based on the fetched value, and calculate a predicted chroma sample value by using the acquired linear model parameters α and β.
[0012] According to the second aspect of the present invention, the apparatus calculates the index of the LUT in an elaborate way of extracting some bits in the binary representation. As a result, the efficiency of fetching a value from the LUT is improved.
[0013] According to a third aspect, the present invention relates to an apparatus for decoding a video stream, comprising a processor and a memory. The memory stores instructions for causing the processor to execute a method according to the first aspect or any possible embodiment of the first aspect.
[0014] According to a fourth aspect, the present invention relates to an apparatus for encoding a video stream, comprising a processor and a memory. The memory stores instructions for causing the processor to execute a method according to the first aspect or any possible embodiment of the first aspect.
[0015] According to a fifth aspect, there is proposed a computer-readable storage medium storing instructions which, when executed, cause one or more processors to be configured to code video data. The instructions cause the one or more processors to execute a method according to the first aspect or any possible embodiment of the first aspect.
[0016] According to a sixth aspect, the present invention relates to a computer program comprising program code for executing a method according to the first aspect or any possible embodiment of the first aspect when executed on a computer.
[0017] Details of one or more embodiments are set forth in the accompanying drawings and the description below. Other features, objects, and advantages will be apparent from the specification, drawings, and claims.
Brief Description of the Drawings
[0018] Embodiments of the present invention will be described in more detail below with reference to the accompanying figures and drawings.
Figure 1A
Figure 1B
Figure 2
Figure 3
Figure 4
Figure 5
Figure 6
Figure 7
Figure 8
Figure 9
Figure 10
Figure 11
Figure 12
Figure 13
Figure 14
DETAILED DESCRIPTION OF THE INVENTION
[0019] In the following description, reference is made to the accompanying drawings, which form a part hereof and which illustrate specific aspects of embodiments of the present disclosure or specific aspects in which embodiments of the present invention may be used. It is understood that embodiments of the present invention may be used in other aspects and may include structural or logical changes not shown in the drawings. Accordingly, the following detailed description should not be construed in a limiting sense, and the scope of the present invention is defined by the appended claims.
[0020] For example, it is understood that the disclosure related to the methods described may also apply to corresponding devices or systems configured to perform the methods, and vice versa. For example, when one or more specific method steps are described, the corresponding device may include one or more units (e.g., functional units) for performing the one or more described method steps (e.g., one unit for performing one or more steps, or multiple units each performing one or more of the multiple steps), even when such one or more units are not explicitly described or shown in the figures. On the other hand, for example, when a specific device is described based on one or more units (e.g., functional units), the corresponding method may include one step (e.g., one step for performing the functions of one or more units, or multiple steps each performing one or more of the functions of one or more of the multiple units) for performing the functions of the one or more units, even when such one or more steps are not explicitly described or shown in the figures. Further, it is understood that the features of the various exemplary embodiments and / or aspects described herein may be combined with each other, unless otherwise specifically stated.
[0021] Video coding typically refers to the processing of a series of images that form a video or video sequence. Instead of the term "image", the terms "frame" or "image" may be used as synonyms in the field of video coding. Video coding (or generally coding) includes two parts: video encoding and video decoding. Video encoding is performed on the source side and typically involves processing the original video images (e.g., by compression) to reduce the amount of data required to represent the video images (for more efficient storage and / or transmission). Video decoding is performed on the destination side and typically involves the reverse process compared to the encoder to reconstruct the video images. Embodiments that refer to the "coding" of video images (or generally images) are to be understood as relating to the "encoding" or "decoding" of video images or each video sequence. The combination of the encoding part and the decoding part is also referred to as a codec (coding and decoding).
[0022] In the case of lossless video coding, the original video images can be reconstructed (assuming no transmission loss or other data loss during storage or transmission). That is, the reconstructed video images have the same quality as the original video images. In the case of irreversible video coding, further compression is performed, for example by quantization, to reduce the amount of data representing video images that cannot be fully reconstructed at the decoder. That is, the quality of the reconstructed video images is lower or worse compared to the quality of the original video images.
[0023] Some video coding standards belong to the group of "irreversible hybrid video codecs" (i.e., they combine spatial and temporal prediction within the sample area with 2D transform coding for applying quantization within the transform area). Each image of a video sequence is typically partitioned into a set of non-overlapping blocks, and coding is typically performed at the block level. In other words, in the encoder, the video generates prediction blocks using, for example, spatial (intra-picture) prediction and / or temporal (inter-picture) prediction, subtracts the prediction blocks from the current block (the currently processed / future processed block) to obtain a residual block, transforms the residual block, quantizes the residual block within the transform area to reduce the amount of data to be transmitted (compression), and is typically processed, i.e., encoded, at the block (video block) level. On the other hand, in the decoder, the reverse process compared to the encoder is applied to the encoded or compressed block to reconstruct the current block for presentation. Further, the encoder repeats the decoder's processing loop so that both will generate the same prediction (e.g., intra prediction and inter prediction) and / or reconstruction for processing, i.e., coding, subsequent blocks.
[0024] In the following embodiments of the video coding system 10, the video encoder 20 and the video decoder 30 will be described with reference to FIGS. 1 to 3.
[0025] FIG. 1A is a schematic block diagram showing an exemplary coding system 10 that can utilize the technology of the present application, such as a video coding system 10 (or simply, coding system 10). The video encoder 20 (or simply encoder 20) and the video decoder 30 (or simply decoder 30) of the video coding system 10 represent examples of devices that can be configured to execute the technology according to various examples described in the present application.
[0026] As shown in FIG. 1A, the coding system 10 includes a source device 12 configured to provide encoded image data 21, for example, to a destination device 14 for decoding the encoded image data 21.
[0027] The source device 12 includes an encoder 20 and may additionally, i.e., optionally, include an image source 16, a preprocessor (or preprocessing unit) 18, for example, an image preprocessor 18, and a communication interface or communication unit 22.
[0028] The image source 16 may include any type of imaging device, for example, a camera for imaging real-world images, and / or any type of image generation device, for example, a computer graphics processor for generating computer-animated images, or any other type of device for acquiring and / or providing real-world images, computer-generated images (e.g., screen content, virtual reality (VR) images), and / or any combination thereof (e.g., augmented reality (AR) images). The image source may be any type of memory or storage for storing any of the above-described images.
[0029] Distinguished from the preprocessor 18 and the processing performed by the preprocessing unit 18, the image or image data 17 may also be referred to as raw image or raw image data 17.
[0030] The preprocessor 18 is configured to receive (raw) image data 17 and perform preprocessing on the image data 17 to obtain preprocessed image 19 or preprocessed image data 19. The preprocessing performed by the preprocessor 18 may include, for example, trimming, color format conversion (e.g., from RGB to YCbCr), color correction, or noise removal. It can be understood that the preprocessing unit 18 may be an optional component.
[0031] The video encoder 20 is configured to receive the preprocessed image data 19 and provide the encoded image data 21 (for example, based on FIG. 2, further details will be described below). The communication interface 22 of the source device 12 receives the encoded image data 21 and, for storage or direct reconstruction, transmits the encoded image data 21 (or any further processed version thereof) via the communication channel 13 to another device, such as the destination device 14 or any other device.
[0032] The destination device 14 includes a decoder 30 (for example, a video decoder 30) and additionally, that is, optionally, may include a communication interface or communication unit 28, a post-processor 32 (or post-processing unit 32), and a display device 34.
[0033] The communication interface 28 of the destination device 14 is configured to receive the encoded image data 21 (or any further processed version thereof), for example, directly from the source device 12 or from any other source, such as a storage device, for example, a storage device for the encoded image data, and provide the encoded image data 21 to the decoder 30.
[0034] The communication interface 22 and the communication interface 28 may be configured to transmit or receive the encoded image data 21 or the encoded data 13 via a direct communication link between the source device 12 and the destination device 14, such as a direct wired or wireless connection, or via any type of network, such as a wired network or a wireless network or any combination thereof, or any type of private network and public network or any combination thereof.
[0035] The communication interface 22 may be configured to process the encoded image data 21, for example, by packaging the encoded image data into a suitable format, such as packets, and / or using any kind of transmission encoding or processing for transmission over a communication link or communication network.
[0036] The communication interface 28 that forms the counterpart of the communication interface 22 may be configured to receive the transmitted data and process the transmitted data using any kind of corresponding transmission decoding or processing and / or unpackaging to obtain the encoded image data 21.
[0037] Both the communication interface 22 and the communication interface 28 may be configured as a unidirectional communication interface or a bidirectional communication interface as indicated by the arrow for the communication channel 13 from the source device 12 to the destination device 14 in FIG. 1A, for example, to send and receive messages, for example, to establish a connection and check and exchange any other information related to the communication link and / or data transmission, for example, the transmission of the encoded image data.
[0038] The decoder 30 is configured to receive the encoded image data 21 and provide the decoded image data 31 or the decoded image 31 (for example, based on FIG. 3 or FIG. 5, further details will be described below).
[0039] The post-processor 32 of the destination device 14 is configured to post-process the decoded image data 31 (also referred to as reconstructed image data), for example, the decoded image 31, to obtain post-processed image data 33, for example, the post-processed image 33. The post-processing executed by the post-processing unit 32 may include, for example, converting the decoded image data 31, for example, for the purpose of preparing it for display by the display device 34, for example, color format conversion (e.g., from YCbCr to RGB), color correction, trimming or resampling, or any other processing.
[0040] The display device 34 of the destination device 14 is configured to receive the post-processed image data 33 for displaying an image, for example, to a user or viewer. The display device 34 may be any type of display for representing the reconstructed image, for example, an integrated or external display or monitor, and may include this. The display may include, for example, a liquid crystal display (LCD), an organic light-emitting diode (OLED) display, a plasma display, a projector, a micro LED display, a liquid crystal on silicon (LCoS), a digital light processor (DLP), or any other type of display.
[0041] FIG. 1A shows the source device 12 and the destination device 14 as separate devices, but embodiments of the devices may also include both or both of their functions, that is, the source device 12 or the corresponding function and the destination device 14 or the corresponding function. In such embodiments, the source device 12 or the corresponding function and the destination device 14 or the corresponding function may be implemented using the same hardware and / or software, or by separate hardware and / or software, or any combination thereof.
[0042] As will be apparent to those of ordinary skill in the art based on the description, the functions of the different units or the presence and (exact) partitioning of the functions within the source device 12 and / or the destination device 14, as shown in FIG. 1A, may vary depending on the actual devices and applications.
[0043] The encoder 20 (e.g., a video encoder 20), decoder 30 (e.g., a video decoder 30), or both the encoder 20 and decoder 30 may be implemented via a processing circuit as shown in FIG. 1B, such as one or more microprocessors, digital signal processors (DSPs), application specific integrated circuits (ASICs), field programmable gate arrays (FPGAs), discrete logic, hardware, dedicated video coding, or any combination thereof. The encoder 20 may be implemented via the processing circuit 46 to embody the various modules discussed in relation to the encoder 20 of FIG. 2 and / or any other encoder system or encoder subsystem described herein. The decoder 30 may be implemented via the processing circuit 46 to embody the various modules discussed in relation to the decoder 30 of FIG. 3 and / or any other decoder system or decoder subsystem described herein. The processing circuit may be configured to perform various operations as discussed later. As shown in FIG. 5, when these techniques are implemented partially within software, the device may store instructions for the software in a suitable non-transitory computer-readable storage medium and execute the instructions in hardware using one or more processors to perform the techniques of the present disclosure. Either the video encoder 20 or the video decoder 30 may be integrated, for example, as part of a combined encoder / decoder (codec) within a single device, as shown in FIG. 1B.
[0044] The source device 12 and the destination device 14 may be any kind of handheld device or stationary device, such as a notebook computer or laptop computer, mobile phone, smartphone, tablet or tablet computer, camera, desktop computer, set-top box, television, display device, digital media player, video game console, video streaming device (such as a content service server or content delivery server), broadcast receiver device or broadcast transmitter device, etc., and may not use an operating system or may use any kind of operating system. In some cases, the source device 12 and the destination device 14 may be equipped for wireless communication. Accordingly, the source device 12 and the destination device 14 may be wireless communication devices.
[0045] In some cases, the video coding system 10 shown in FIG. 1A is merely an example, and the technology of the present application can be applied to a video coding setting (for example, video encoding or video decoding) that does not necessarily include any data communication between the encoding device and the decoding device. In other examples, the data is obtained from local memory, or streamed via a network, etc. The video encoding device may encode and store the data in memory, and / or the video decoding device may obtain and decode the data from memory. In some examples, encoding and decoding are performed by devices that do not communicate with each other, but simply encode data into memory and / or obtain and decode data from memory.
[0046] For the sake of convenience of explanation, embodiments of the present invention will be described herein with reference to reference software of High-Efficiency Video Coding (HEVC) or Versatile Video Coding (VVC), which is a next-generation video coding standard developed by, for example, the Joint Collaboration Team on Video Coding (JCT-VC) of ITU-T Video Coding Experts Group (VCEG) and ISO / IEC Motion Picture Experts Group (MPEG). Those skilled in the art will understand that the embodiments of the present invention are not limited to HEVC or VVC. [Encoder and Encoding Method]
[0047] FIG. 2 shows a schematic block diagram of an exemplary video encoder 20 configured to implement the technology of the present application. In the example of FIG. 2, the video encoder 20 includes an input 201 (or input interface 201), a residual calculation unit 204, a conversion processing unit 206, a quantization unit 208, an inverse quantization unit 210, an inverse conversion processing unit 212, a reconstruction unit 214, a loop filter unit 220, a decoded picture buffer (DPB) 230, a mode selection unit 260, an entropy encoding unit 270, and an output 272 (or output interface 272). The mode selection unit 260 may include an inter prediction unit 244, an intra prediction unit 254, and a partitioning unit 262. The inter prediction unit 244 may include a motion estimation unit and a motion compensation unit (not shown). The video encoder 20 as shown in FIG. 2 may also be referred to as a hybrid video encoder or a video encoder by a hybrid video codec.
[0048] The residual calculation unit 204, the conversion processing unit 206, the quantization unit 208, and the mode selection unit 260 may be referred to as forming the forward signal path of the encoder 20, while the inverse quantization unit 210, the inverse conversion processing unit 212, the reconstruction unit 214, the buffer 216, the loop filter 220, the buffer for decoded images (DPB) 230, the inter prediction unit 244, and the intra prediction unit 254 may be referred to as forming the reverse signal path of the video encoder 20. The reverse signal path of the video encoder 20 corresponds to the signal path of the decoder (see the video decoder 30 in FIG. 3). The inverse quantization unit 210, the inverse conversion processing unit 212, the reconstruction unit 214, the loop filter 220, the buffer for decoded images (DPB) 230, the inter prediction unit 244, and the intra prediction unit 254 are also referred to as forming the "built-in decoder" of the video encoder 20. [Image and Image Segmentation (Image and Block)]
[0049] The encoder 20 may be configured to receive, for example, via the input 201, an image 17 (or image data 17), for example, an image among a series of images forming a video or a video sequence. The received image or image data may be preprocessed image 19 (or preprocessed image data 19). For simplicity, in the following description, image 17 is referred to. (In particular, in video coding, to distinguish the current image from other images, for example, previously encoded and / or decoded images of the same video sequence, i.e., the video sequence including the current image,) image 17 may also be referred to as the current image or the image to be coded.
[0050] (Digital) images can be or are considered as two-dimensional arrays or matrices of samples having intensity values. Samples within the array can also be referred to as pixels (short for picture elements) or pels. The size and / or resolution of an image is determined by the number of samples in the horizontal and vertical directions (or axes) of the array or image. For color representation, typically three color components are used. That is, an image can be represented as or can include three sample arrays. In the RGB format or RGB color space, an image includes corresponding red, green, and blue sample arrays. However, in video coding, each pixel typically includes a luminance component represented by Y (in some cases, L may also be used instead) and two chrominance components represented by Cb and Cr, and is represented in YCbCr. The luminance (or simply, luma) component Y represents brightness or intensity of gray levels (such as in a grayscale image), while the two chrominance (or simply, chroma) components Cb and Cr represent chrominance components or color information components. Thus, an image in YCbCr format includes a luminance sample array of luminance sample values (Y) and two chrominance sample arrays of chrominance values (Cb and Cr). An image in RGB format may be converted or transformed into YCbCr format and vice versa, and this process is also known as color conversion or color transformation. If an image is monochrome, this image may include only a luminance sample array. Thus, an image can be, for example, an array of luma samples in monochrome format, or an array of luma samples and two corresponding arrays of chroma samples in color formats of 4:2:0, 4:2:2, and 4:4:4.
[0051] An embodiment of the video encoder 20 may include an image segmentation unit (not shown in FIG. 2) configured to segment the image 17 into a plurality of (typically non-overlapping) image blocks 203. These blocks may also be referred to as root blocks, macroblocks (H.264 / AVC), or coding tree blocks (CTB) or coding tree units (CTU) (H.265 / HEVC and VVC). The image segmentation unit may use the same block size for all images of the video sequence and the corresponding grid defining the block size, or may be configured to change the block size between images or subsets or groups of images to segment each image into corresponding blocks.
[0052] In a further embodiment, the video encoder may be configured to directly receive the blocks 203 of the image ..... For example, one, some, or all of the blocks forming the image 17. The image block 203 may also be referred to as the current image block or the image block to be coded.
[0053] Similar to the image 17, the image block 203 can also be considered or is also a two-dimensional array or matrix of samples having intensity values (sample values) but smaller in size than the image 17. In other words, the block 203 may include, for example, one sample array (e.g., a luma array in the case of a monochrome image 17, or a luma array or chroma array in the case of a color image), or three sample arrays (e.g., one luma array and two chroma arrays in the case of a color image 17), or any other number and / or type of arrays depending on the color format applied. The size of the block 203 is determined by the number of samples in the horizontal and vertical directions (or axes) of the block 203. Thus, the block may be, for example, an M×N (M columns × N rows) array of samples or an M×N array of transform coefficients.
[0054] An embodiment of the video encoder 20 as shown in FIG. 2 may be configured to encode the image 17 block by block. For example, encoding and prediction may be performed for each block 203. [Residual calculation]
[0055] The residual calculation unit 204 may be configured to calculate a residual block 205 (also referred to as a residual 205) based on the image block 203 and the prediction block 265 (further details about the prediction block 265 will be provided later), for example, by subtracting the sample values of the prediction block 265 from the sample values of the image block 203 sample by sample (pixel by pixel) within the sample region to obtain the residual block 205 within the sample region. [Transformation]
[0056] The transformation processing unit 206 may be configured to apply a transformation, such as a discrete cosine transform (DCT) or a discrete sine transform (DST), to the sample values of the residual block 205 to obtain the transformation coefficients 207 within the transformation region. The transformation coefficients 207 may also be referred to as transformation residual coefficients and represent the residual block 205 within the transformation region.
[0057] The transformation processing unit 206 may be configured to apply an integer approximation of DCT / DST, such as the transformation specified for H.265 / HEVC. Compared with the orthogonal DCT transform, such an integer approximation is typically scaled by a specific coefficient. To maintain the norm of the residual block processed by the forward and inverse transforms, an additional scaling coefficient is applied as part of the transformation process. The scaling coefficient is typically selected based on specific constraints such as a scaling coefficient that is a power of two for shift operations, the bit depth of the transformation coefficients, and the trade-off between accuracy and implementation cost. For example, a specific scaling coefficient may be specified for the inverse transform by the inverse transform processing unit 212 (and the corresponding inverse transform by the inverse transform processing unit 312 in the video decoder 30, for example), and the corresponding scaling coefficient for the forward transform by the transformation processing unit 206 in the encoder 20 may be specified accordingly.
[0058] Embodiments of the video encoder 20 (each of the conversion processing units 206) may be configured to output conversion parameters, such as the type of one conversion or multiple conversions, for example, after being encoded or compressed directly or via the entropy encoding unit 270, so that, for example, the video decoder 30 may receive and use the conversion parameters for decoding. [Quantization]
[0059] The quantization unit 208 may be configured to obtain the quantized coefficients 209 by quantizing the conversion coefficients 207, for example, by applying scalar quantization or vector quantization. The quantized coefficients 209 may also be referred to as quantized conversion coefficients 209 or quantized residual coefficients 209.
[0060] Through quantization processing, the bit depth associated with some or all of the transform coefficients 207 can be reduced. For example, an n-bit transform coefficient can have its fractional part truncated to an m-bit transform coefficient during quantization, where n is greater than m. The degree of quantization can be adjusted by adjusting the quantization parameter (QP). For example, in the case of scalar quantization, different scalings can be applied to achieve finer or coarser quantization. A smaller quantization step size corresponds to finer quantization, while a larger quantization step size corresponds to coarser quantization. The applicable quantization step size can be indicated by the quantization parameter (QP). The quantization parameter can be, for example, an index of a predefined set of applicable quantization step sizes. For example, a small quantization parameter can correspond to fine quantization (small quantization step size), a large quantization parameter can correspond to coarse quantization (large quantization step size), or vice versa. Quantization can include division by the quantization step size. For example, the corresponding and / or inverse dequantization by the inverse quantization unit 210 can include multiplication by the quantization step size. Embodiments according to some standards such as HEVC can be configured to determine the quantization step size using the quantization parameter. Generally, the quantization step size can be calculated based on the quantization parameter using a fixed-point approximation of an expression that includes division. Additional scaling factors can be introduced for quantization and dequantization to restore the norm of the residual block, which can be corrected due to the scaling used in the fixed-point approximation of the expressions for the quantization step size and the quantization parameter. In one exemplary implementation, the scaling for inverse transform and dequantization can be combined. Alternatively, a customized quantization table can be used and signaled, for example, from the encoder to the decoder in the bitstream. Quantization is an irreversible operation, and the loss increases as the quantization step size increases.
[0061] Embodiments of the video encoder 20 (each quantization unit 208) may be configured to output the quantization parameter (QP) after encoding it, for example, directly or via the entropy encoding unit 270. As a result, for example, the video decoder 30 may receive and apply the quantization parameter for decoding. [Inverse quantization]
[0062] The inverse quantization unit 210 is configured to apply inverse quantization of the quantization unit 208 to the quantized coefficients, for example, by applying the inverse of the quantization scheme applied by the quantization unit 208 based on or using the same quantization stage size as the quantization unit 208, to obtain the dequantized coefficients 211. The dequantized coefficients 211 may also be referred to as dequantized residual coefficients 211 and typically are not identical to the transform coefficients due to loss by quantization, but correspond to the transform coefficients 207. [Inverse transformation]
[0063] The inverse transform processing unit 212 is configured to apply an inverse transform of the transform applied by the transform processing unit 206, for example, an inverse discrete cosine transform (DCT) or inverse discrete sine transform (DST) or other inverse transform, to obtain a reconstructed residual block 213 (or corresponding dequantized coefficients 213) within the sample region. The reconstructed residual block 213 may also be referred to as the transform block 213. [Reconstruction]
[0064] The reconstruction unit 214 (for example, an adder or accumulator 214) is configured to add the transform block 213 (i.e., the reconstructed residual block 213) to the prediction block 265 by adding, for example, the sample values of the reconstructed residual block 213 and the sample values of the prediction block 265 on a sample-by-sample basis, to obtain a reconstructed block 215 within the sample region. [Filtering]
[0065] The loop filter unit 220 (or simply, the "loop filter" 220) is configured to filter the reconstructed block 215 to obtain a filtered block 221, or generally, to filter the reconstructed samples to obtain filtered samples. The loop filter unit is configured to, for example, smooth pixel transitions or otherwise improve video quality. The loop filter unit 220 may include one or more loop filters such as a deblocking filter, a sample adaptive offset (SAO) filter, or one or more other filters such as a bilateral filter, an adaptive loop filter (ALF), a sharpening filter, a smoothing filter or a collaborative filter, or any combination thereof. The loop filter unit 220 is shown as an in-loop filter in FIG. 2, but in other configurations, it may be implemented as a post-loop filter. The filtered block 221 may also be referred to as the filtered reconstructed block 221.
[0066] Embodiments of the video encoder 20 (each, the loop filter unit 220) may be configured to output loop filter parameters (such as sample adaptive offset information), for example, after being encoded directly or via the entropy encoding unit 270. As a result, for example, the decoder 30 may receive and apply the same loop filter parameters or respective loop filters for decoding. [Buffer for decoded image]
[0067] The decoded picture buffer (DPB) 230 may be a memory that stores reference pictures or generally reference picture data for encoding video data by the video encoder 20. The DPB 230 may be formed by any of various memory devices such as dynamic random access memory (DRAM) including synchronous DRAM (SDRAM), magnetoresistive RAM (MRAM), resistive RAM (RRAM), or other types of memory devices. The decoded picture buffer (DPB) 230 may be configured to store one or more filtered blocks 221. The decoded picture buffer 230 may further be configured to store other previously filtered blocks of the same current picture or a different picture, e.g., previously reconstructed pictures, e.g., previously reconstructed filtered blocks 221, and may provide a complete, previously reconstructed, i.e., decoded picture (corresponding reference blocks and reference samples) and / or a partially reconstructed current picture (corresponding reference blocks and reference samples) for, e.g., inter prediction. Also, the decoded picture buffer (DPB) 230 may be configured to store one or more unfiltered reconstructed blocks 215 or generally unfiltered reconstructed samples, or any other further processed version of the reconstructed blocks or samples, e.g., if the reconstructed block 215 has not been filtered by the loop filter unit 220. [Mode Selection (Segmentation and Prediction)]
[0068] The mode selection unit 260 includes a segmentation unit 262, an inter prediction unit 244, and an intra prediction unit 254, and is configured to receive or obtain original image data, for example, the original block 203 (the current block 203 of the current image 17), and reconstructed image data, for example, from the same (current) image and / or from one or more previously decoded images, for example, filtered and / or unfiltered reconstructed samples or blocks from the decoded image buffer 230 or other buffers (e.g., a line buffer not shown). The reconstructed image data is used as reference image data for prediction, for example, inter prediction or intra prediction, to obtain the prediction block 265 or the prediction factor 265.
[0069] The mode selection unit 260 may be configured to determine or select a segmentation for the current block prediction mode (without segmentation) and the prediction mode (e.g., the intra prediction mode or the inter prediction mode), and generate a corresponding prediction block 265 used for the calculation of the residual block 205 and the reconstruction of the reconstructed block 215.
[0070] Embodiments of the mode selection unit 260 may be configured to select a partitioning and prediction mode (e.g., from those supported by or available for the mode selection unit 260). Thereby, the best matching, or in other words, the minimum residual (the minimum residual means a better compression rate for transmission or storage) or the minimum signaling overhead (the minimum signaling overhead means a better compression rate for transmission or storage) is provided, or both are considered, or a balance between both is taken. The mode selection unit 260 may be configured to determine the partitioning and prediction mode based on rate distortion optimization (RDO), i.e., to select a prediction mode that provides the minimum rate distortion. Terms such as "best", "minimum", "optimal", etc. in this context do not necessarily refer to an overall "best", "minimum", "optimal", etc., but may also refer to achieving the criteria for termination or selection, such as a value exceeding or falling below a threshold, or other constraints that potentially lead to a "quasi-optimal selection" but reduce complexity and processing time.
[0071] In other words, the partitioning unit 262 may be configured to repeatedly partition the block 203 into smaller block partitions or sub-blocks (re-forming the blocks) using, for example, quad tree partitioning (QT), binary partitioning (BT), or triple tree partitioning (TT), or any combination thereof, and to perform a prediction for each of the block partitions or sub-blocks, for example. The mode selection includes the selection of the tree structure of the partitioned block 203, and the prediction mode is applied to each of the block partitions or sub-blocks.
[0072] The partitioning and prediction processing (e.g., by the partitioning unit 262) and prediction processing (by the inter prediction unit 244 and the intra prediction unit 254) performed by the exemplary video encoder 20 will be described in more detail below. [Partitioning]
[0073] The partitioning unit 262 may partition (or divide) the current block 203 into smaller partitions, for example, smaller blocks of square or rectangular size. These smaller blocks (which may also be referred to as sub-blocks) may be further partitioned into even smaller partitions. This is also referred to as tree partitioning or hierarchical tree partitioning. The root block, for example, root tree level 0 (hierarchical level 0, depth 0) may be recursively partitioned, for example, into two or more blocks of nodes at the next lower tree level, for example, tree level 1 (hierarchical level 1, depth 1). These blocks may be further partitioned into two or more blocks at the next lower level, for example, tree level 2 (hierarchical level 2, depth 2), etc., until the partitioning is terminated, for example, when an end criterion is met, for example, when the maximum tree depth or minimum block size is reached. Blocks that are not further partitioned are also referred to as leaf blocks or leaf nodes of the tree. A tree using partitioning into two partitions is referred to as a binary tree (BT), a tree using partitioning into three partitions is referred to as a ternary tree (TT), and a tree using partitioning into four partitions is referred to as a quadtree (QT).
[0074] As previously mentioned, the term "block" as used herein may be a portion of an image, particularly a square or rectangular portion. Referring to, for example, HEVC and VVC, a block may be a coding tree unit (CTU), a coding unit (CU), a prediction unit (PU) or a transform unit (TU) and / or a corresponding block, for example, a coding tree block (CTB), a coding block (CB), a transform block (TB) or a prediction block (PB), and may correspond thereto.
[0075] For example, a coding tree unit (CTU) may be a CTB of luma samples of an image having three sample arrays and two corresponding CTBs of chroma samples, or a CTB of samples of a monochrome image or an image coded using three separate color planes and syntax structures used for coding samples, and may include them. Accordingly, a coding tree block (CTB) may be an N×N block of samples of a certain value N such that the division of components into the CTB is hierarchical. A coding unit (CU) may be a coding block of luma samples of an image having three sample arrays and two corresponding coding blocks of chroma samples, or a coding block of samples of a monochrome image or an image coded using three separate color planes and syntax structures used for coding samples, and may include them. Accordingly, a coding block (CB) may be an M×N block of samples of certain values M and N such that the division of CTBs into the coding block is hierarchical.
[0076] For example, in some embodiments according to HEVC, a coding tree unit (CTU) may be divided into CUs by using a quad-tree structure represented as a coding tree. The decision of whether to use inter-picture (temporal) prediction or intra-picture (spatial) prediction to code an image area is made at the CU level. Each CU may be further divided into one, two, or four PUs according to the partition type of the PU. Inside one PU, the same prediction process is applied and related information is sent to the decoder on a PU basis. After obtaining a residual block by applying a prediction process based on the partition type of the PU, the CU may be divided into transform units (TUs) according to another quad-tree structure similar to the coding tree of the CU.
[0077] For example, in an embodiment according to the latest video coding standard currently under development, called Versatile Video Coding (VVC), quad tree and binary tree (QTBT) partitioning is used to partition coding blocks. In the QTBT block structure, a CU can have either a square or rectangular shape. For example, a coding tree unit (CTU) is first partitioned by a quad tree structure. The quad tree leaf nodes are further partitioned by a binary tree or ternary (or triple) tree structure. The partitioned tree leaf nodes are called coding units (CUs), and their segmentation is used for prediction and transformation processing without any further partitioning. This means that CUs, PUs, and TUs have the same block size within the QTBT coding block structure. In parallel, it has also been proposed to use multiple levels of partitioning, for example, triple tree partitioning, together with the QTBT block structure.
[0078] In one example, the mode selection unit 260 of the video encoder 20 can be configured to perform any combination of the partitioning techniques described herein.
[0079] As described above, the video encoder 20 is configured to determine or select the best or optimal prediction mode from a set of (predetermined) prediction modes. The set of prediction modes can include, for example, intra prediction modes and / or inter prediction modes. [Intra Prediction]
[0080] The set of intra prediction modes may include 35 different intra prediction modes, such as non - directional modes like DC (or average) mode and planar mode, or directional modes defined, for example, in HEVC, or may include 67 different intra prediction modes, such as non - directional modes like DC (or average) mode and planar mode, or directional modes defined for example for VVC.
[0081] The intra prediction unit 254 is configured to generate an intra prediction block 265 according to a certain intra prediction mode from a set of intra prediction modes using the reconstructed samples of adjacent blocks of the same current image.
[0082] The intra prediction unit 254 (or generally, the mode selection unit 260) may further be configured to output an intra prediction parameter (or generally, information indicating the intra prediction mode selected for a block) in the form of a syntax element 266 to the entropy encoding unit 270 for inclusion in the encoded image data 21. As a result, for example, the video decoder 30 may receive and use the prediction parameters for decoding. [Inter Prediction]
[0083] The set of inter prediction modes (or possible inter prediction modes) depends on the available reference images (i.e., for example, previously at least partially decoded images stored in the DPB 230) and other inter prediction parameters, such as whether only the whole or a part of the reference image, for example, the search window area around the area of the current block of the reference image, is used to search for the reference block with the best matching, and / or, for example, whether pixel interpolation, such as half interpolation / semi-pel interpolation and / or quarter-pel interpolation, is applied. In addition to the above prediction modes, a skip mode and / or a direct mode may be applied.
[0084] The inter-prediction unit 244 may include a motion estimation (ME) unit and a motion compensation (MC) unit (both not shown in FIG. 2). The motion estimation unit may be configured to receive or obtain, for motion estimation, the image block 203 (the current image block 203 of the current image 17) and the decoded image 231, or at least one or more previously reconstructed blocks such as reconstructed blocks of one or more other / different previously decoded images 231. For example, the video sequence may include the current image and the previously decoded image 231, or in other words, the current image and the previously decoded image 231 may be part of or form a series of images that form the video sequence.
[0085] The encoder 20 may be configured to select a reference block from among a plurality of reference blocks of the same or different images of a plurality of other images, and provide the reference image (or reference image index), and / or the offset (spatial offset) between the position (x coordinate, y coordinate) of the reference block and the position of the current block to the motion estimation unit as inter-prediction parameters. This offset is also called a motion vector (MV).
[0086] The motion compensation unit is configured to obtain, for example, receive the inter-prediction parameters, and perform inter-prediction based on or using the inter-prediction parameters to obtain the inter-prediction block 265. The motion compensation performed by the motion compensation unit may involve fetching or generating a prediction block based on the motion / block vector determined by motion estimation, and in some cases, performing interpolation with sub-pixel accuracy. By interpolation filtering, additional pixel samples may be generated from known pixel samples. Thus, the number of candidate prediction blocks that can be used to code an image block potentially increases. When receiving the motion vector for the PU of the current image block, the motion compensation unit may identify the position of the prediction block indicated by the motion vector in one of the reference image lists.
[0087] Also, the motion compensation unit may generate blocks and syntax elements related to the video slice that are used by the video decoder 30 in decoding the image blocks of the video slice. [Entropy Coding]
[0088] The entropy encoding unit 270 is configured to apply, for example, an entropy encoding algorithm or an entropy encoding scheme (e.g., variable length coding (VLC) scheme, context adaptive VLC scheme (CAVLC), arithmetic coding scheme, binning, context adaptive binary arithmetic coding (CABAC), syntax-based context adaptive binary arithmetic coding (SBAC), probability interval partitioning entropy (PIPE) coding or another entropy encoding method or entropy encoding technique) or bypass (uncompressed) to the quantized coefficients 209, inter prediction parameters, intra prediction parameters, loop filter parameters and / or other syntax elements to obtain encoded image data 21 that can be output via output 272, for example, in the form of an encoded bitstream 21. As a result, for example, the video decoder 30 can receive and use the parameters for decoding. The encoded bitstream 21 can be transmitted to the video decoder 30 or stored in a memory for later transmission or acquisition by the video decoder 30.
[0089] Other structural variations of the video encoder 20 can be used to encode the video stream. For example, the non-transform based encoder 20 can directly quantize the residual signal without a transform processing unit 206 for a particular block or frame. In another implementation, the encoder 20 may have a quantization unit 208 and an inverse quantization unit 210 combined into a single unit. [Decoder and Decoding Method]
[0090] Figure 3 shows an example of a video decoder 30 configured to implement the technology of the present application. The video decoder 30 is configured to receive, for example, encoded image data 21 (e.g., an encoded bitstream 21) encoded by an encoder 20 and obtain a decoded image 331. The encoded image data or bitstream includes information for decoding the encoded image data, for example, data representing image blocks of an encoded video slice and associated syntax elements.
[0091] In the example of Figure 3, the decoder 30 includes an entropy decoding unit 304, an inverse quantization unit 310, an inverse transform processing unit 312, a reconstruction unit 314 (e.g., an adder 314), a loop filter 320, a buffer for decoded images (DPB) 330, an inter prediction unit 344, and an intra prediction unit 354. The inter prediction unit 344 may be or include a motion compensation unit. The video decoder 30 may execute a decoding path that is generally inverse to the encoding path described in relation to the video encoder 20 of Figure 2 in some examples.
[0092] As described in relation to the encoder 20, the inverse quantization unit 210, the inverse transform processing unit 212, the reconstruction unit 214, the loop filter 220, the decoded picture buffer (DPB) 230, the inter prediction unit 244, and the intra prediction unit 254 are also referred to as forming the "built-in decoder" of the video encoder 20. Thus, the inverse quantization unit 310 may be functionally identical to the inverse quantization unit 210, the inverse transform processing unit 312 may be functionally identical to the inverse transform processing unit 212, the reconstruction unit 314 may be functionally identical to the reconstruction unit 214, the loop filter 320 may be functionally identical to the loop filter 220, and the decoded picture buffer 330 may be functionally identical to the decoded picture buffer 230. Thus, the description provided for each unit and function of the video encoder 20 applies to correspond to each unit and function of the video decoder 30. [Entropy Decoding]
[0093] Entropy decoding unit 304 analyzes the bitstream 21 (or generally, the encoded image data 21), and for example, performs entropy decoding on the encoded image data 21 to obtain, for example, the quantized coefficients 309 and / or the decoded coding parameters (not shown in FIG. 3), for example, the inter prediction parameters (e.g., reference image index and motion vector), the intra prediction parameters (e.g., intra prediction mode or intra prediction index), the transform parameters, the quantization parameters, the loop filter parameters, and / or any or all of other syntax elements. The entropy decoding unit 304 may be configured to apply a decoding algorithm or a decoding scheme corresponding to the encoding scheme as described in relation to the entropy encoding unit 270 of the encoder 20. The entropy decoding unit 304 may further be configured to provide the inter prediction parameters, the intra prediction parameters, and / or other syntax elements to the mode selection unit 360, and other parameters to other units of the decoder 30. The video decoder 30 may receive the syntax elements at the video slice level and / or the video block level. [Inverse quantization]
[0094] Inverse quantization unit 310 receives the quantization parameter (QP) (or generally, information related to inverse quantization) and the quantized coefficients from the encoded image data 21 (e.g., by the entropy decoding unit 304, for example, through analysis and / or decoding), and is configured to apply inverse quantization to the decoded quantized coefficients 309 based on the quantization parameter to obtain the dequantized coefficients 311, which may also be referred to as transform coefficients 311. The inverse quantization process may include the use of the quantization parameter determined by the video encoder 20 for each video block within the video slice to determine the degree of quantization and, similarly, the degree of inverse quantization to be applied. [Inverse transform]
[0095] The inverse transformation processing unit 312 may be configured to receive the dequantized coefficient 311, also referred to as the transformation coefficient 311, and apply a transformation to the dequantized coefficient 311 to obtain the reconstructed residual block 313 within the sample region. The reconstructed residual block 313 may also be referred to as the transformation block 313. This transformation may be an inverse transformation, such as an inverse DCT, inverse DST, inverse integer transformation, or a conceptually similar inverse transformation process. The inverse transformation processing unit 312 may further receive transformation parameters or corresponding information from the encoded image data 21 (e.g., by the entropy decoding unit 304, e.g., by analysis and / or decoding) to determine the transformation applied to the dequantized coefficient 311. [Reconstruction]
[0096] The reconstruction unit 314 (e.g., adder or summer 314) may be configured to add the reconstructed residual block 313 to the prediction block 365 by adding, for example, the sample values of the reconstructed residual block 313 and the sample values of the prediction block 365 to obtain the reconstructed block 315 within the sample region. [Filtering]
[0097] (Either within or after the coding loop), the loop filter unit 320 is configured to filter the reconstructed block 315, for example, to smooth pixel transitions or otherwise improve video quality, to obtain a filtered block 321. The loop filter unit 320 may include one or more loop filters, such as a deblocking filter, a sample adaptive offset (SAO) filter, or one or more other filters, such as a bilateral filter, an adaptive loop filter (ALF), a sharpening filter, a smoothing filter, or a collaborative filter, or any combination thereof. The loop filter unit 320 is shown as an in-loop filter in FIG. 3, but in other configurations, it may be implemented as a post-loop filter. [Buffer for decoded image]
[0098] The decoded video block 321 of the image is then stored in a decoded image buffer 330 that stores the decoded image 331 as a reference image for subsequent motion compensation of other images and / or for output or display, respectively.
[0099] The decoder 30 is configured to output the decoded image 311, for example, via the output 312, for presentation to or viewing by the user. [Prediction]
[0100] The inter prediction unit 344 may be identical to the inter prediction unit 244 (in particular, the motion compensation unit), the intra prediction unit 354 may be functionally identical to the intra prediction unit 254, and based on the partition parameters and / or prediction parameters or respective information received from the encoded image data 21 (e.g., by the entropy decoding unit 304, e.g., by parsing and / or decoding), it performs the determination and prediction of partitioning or segmentation. The mode selection unit 360 may be configured to perform prediction (intra prediction or inter prediction) for each block based on the reconstructed image, block, or respective samples (filtered or unfiltered) to obtain the prediction block 365.
[0101] When the video slice is coded as an intra coding (I) slice, the intra prediction unit 354 of the mode selection unit 360 is configured to generate a prediction block 365 for the image block of the current video slice based on the signaling intra prediction mode and the data from the previously decoded blocks of the current image. When the video image is coded as an inter coded (i.e., B or P) slice, the inter prediction unit 344 (e.g., the motion compensation unit) of the mode selection unit 360 is configured to generate a prediction block 365 for the video block of the current video slice based on the motion vector and other syntax elements received from the entropy decoding unit 304. In the case of inter prediction, the prediction block may be generated from one of the reference images included in one of the reference image lists. The video decoder 30 may construct the reference frame lists of list 0 and list 1 based on the reference images stored in the DPB 330 using the default construction technique.
[0102] The mode selection unit 360 is configured to determine prediction information about the video blocks of the current video slice by analyzing motion vectors and other syntax elements, and generate a prediction block for the currently decoded video block using the prediction information. For example, the mode selection unit 360 uses some of the received syntax elements to determine a prediction mode (e.g., intra prediction or inter prediction) used to code the video blocks of the video slice, an inter prediction slice type (e.g., B slice, P slice, or GPB slice), construction information regarding one or more of the reference picture lists for the slice, the motion vector of each of the inter-encoded video blocks of the slice, the inter prediction status of each of the inter-coded video blocks of the slice, and other information for decoding the video blocks within the current video slice.
[0103] For decoding the encoded image data 21, other variations of the video decoder 30 may be used. For example, the decoder 30 can generate an output video stream without the loop filtering unit 320. For example, the non-transform-based decoder 30 can directly inverse quantize the residual signal without the inverse transform processing unit 312 for a particular block or frame. In another implementation, the video decoder 30 may have an inverse quantization unit 310 and an inverse transform processing unit 312 combined into a single unit.
[0104] It should be understood that in the encoder 20 and the decoder 30, the processing results of the current stage can be further processed and then output to the next stage. For example, after interpolation filtering, motion vector derivation, or loop filtering, further operations such as Clip or shift can be performed on the processing results of interpolation filtering, motion vector derivation, or loop filtering.
[0105] Note that further operations may be applied to the derived motion vectors of the current block (including, but not limited to, the affine mode control point motion vectors, affine mode, plane mode, sub-block motion vectors in ATMVP mode, and temporal motion vectors, etc.). For example, the value of the motion vector is restricted to a predefined range according to its representation bit number. When the representation bit of the motion vector is bitDepth, the range is -2^(bitDepth-1) to 2^(bitDepth-1)-1. "^" means exponentiation. For example, when bitDepth is set equal to 16, the range is -32768 to 32767, and when bitDepth is set equal to 18, the range is -131072 to 131071. For example, the value of the derived motion vector (e.g., the MV of 4 4×4 sub-blocks within one 8×8 block) is restricted such that the maximum difference between the integer parts of the 4 4×4 sub-block MVs is N or fewer pixels, such as 1 or fewer pixels. Here, two methods are provided to constrain the motion vector according to bitDepth.
[0106] Method 1: Remove the overflow MSB (Most Significant Bit) by the following operation.
Number
Number
Number
Number
[0107] For example, when the value of mvx is -32769, after applying equations (1) and (2), the resulting value is 32767. In a computer system, decimal numbers are stored as two's complements. The two's complement of -32769 is 1,0111,1111,1111,1111 (17 bits), and then, since the MSB is discarded, the resulting two's complement is the same as the output by applying equations (1) and (2), which is 0111,1111,1111,1111 (decimal 32767).
Number
Number
Number
Number
[0108] Method 2: Remove the overflow MSB by clipping the value.
Number
Number
Number
[0109] FIG. 4 is a schematic diagram of a video coding device 400 according to an embodiment of the present disclosure. The video coding device 400 is suitable for implementing the embodiments of the present disclosure described herein. In an embodiment, the video coding device 400 may be a decoder such as the video decoder 30 of FIG. 1A, or an encoder such as the video encoder 20 of FIG. 1A.
[0110] The video coding device 400 includes an inlet port 410 (or input port 410) and a receiver unit (Rx) 420 for receiving data, a processor, logic unit, or central processing unit (CPU) 430 for processing data, a transmitter unit (Tx) 440 and an outlet port 450 (or output port 450) for transmitting data, and a memory 460 for storing data. The video coding device 400 may also include optical / electrical (OE) components and electrical / optical (EO) components connected to the inlet port 410, the receiver unit 420, the transmitter unit 440, and the outlet port 450 for the outlet or inlet of optical or electrical signals.
[0111] Processor 430 is implemented by hardware and software. Processor 430 can be implemented as one or more CPU chips, cores (e.g., as a multi-core processor), FPGAs, ASICs, and DSPs. Processor 430 communicates with an input port 410, a receiver unit 420, a transmitter unit 440, an output port 450, and a memory 460. Processor 430 has a coding module 470. Coding module 470 implements the disclosed embodiments described above. For example, coding module 470 implements, processes, prepares, or provides various coding operations. Thus, by including coding module 470, a substantial improvement in the functionality of video coding device 400 is provided, resulting in the conversion of video coding device 400 to different states. Alternatively, coding module 470 is implemented as instructions stored in memory 460 and executed by processor 430.
[0112] Memory 460 may comprise one or more disks, tape drives, and solid state drives and may be used as an overflow data storage device for storing such programs when selected for execution and for storing instructions and data read during program execution. Memory 460 may be, for example, volatile and / or non-volatile and may be read only memory (ROM), random access memory (RAM), ternary content addressable memory (TCAM), and / or static random access memory (SRAM).
[0113] FIG. 5 is a simplified block diagram of an apparatus 500 that can be used as either or both of the source device 12 and the destination device 14 of FIG. 1A according to an exemplary embodiment.
[0114] The processor 502 within the device 500 may be a central processing unit. Alternatively, the processor 502 may be any other type of device or devices, existing or to be developed in the future, capable of manipulating or processing information. The disclosed implementation may be carried out using a single processor, such as processor 502 as shown, but speed and efficiency advantages may be realized using more than one processor.
[0115] In an implementation, the memory 504 within the device 500 may be a read-only memory (ROM) device or a random access memory (RAM) device. Any other suitable type of storage device may be used as the memory 504. The memory 504 may include code and data 506 that is accessed by the processor 502 using the bus 512. The memory 504 may further include an operating system 508 and an application program 510, and the application program 510 includes at least one program that enables the processor 502 to execute the methods described herein. For example, the application program 510 may include applications 1 through N, and applications 1 through N further include a video coding application that executes the methods described herein.
[0116] The device 500 may also include one or more output devices, such as a display 518. The display 518 may, in one example, be a touch sensor display that combines a display with a touch sensor element operable to detect touch input. The display 518 may be coupled to the processor 502 via the bus 512.
[0117] Although shown here as a single bus, the bus 512 of the apparatus 500 may be composed of a plurality of buses. Further, the secondary storage 514 may be directly connected to other components of the apparatus 500, may be accessed via a network, and may include a single integrated unit such as a memory card or a plurality of units such as a plurality of memory cards. Thus, the apparatus 500 can be implemented in a wide variety of configurations. Intra prediction of chroma samples can be performed using samples of the reconstructed luma block.
[0118] During the development of HEVC, the cross-component linear model (CCLM) chroma intra prediction [J. Kim, S.-W. Park, J.-Y. Park, and B.-M. Jeon, Intra Chroma Prediction Using Inter Channel Correlation, document JCTVC-B021, Jul. 2010] was proposed. CCLM uses the linear correlation between chroma samples and luma samples at corresponding positions within a coding block. When a chroma block is coded using CCLM, a linear model is derived by linear regression from the reconstructed adjacent luma samples and chroma samples. The chroma samples within the current block can then be predicted by the reconstructed luma samples within the current block using the derived linear model (see FIG. 6).
Number
Number
Number
[0119] If the encoded or decoded image has a format that specifies a different number of samples for the luma and chroma components (e.g., 4:2:0 YCbCr format), the luma samples are downsampled before modeling and prediction.
[0120] The method is adopted for use in VTM 2.0. Specifically, the derivation of parameters is performed as follows.
Number
Number
[0121] [G. Laroche, J. Taquet, C. Gisquet, P. Onno (Canon), "CE3: Cross-component linear model simplification (Test 5.1)", Input document to 12 th JVET Meeting in Macao, China, Oct. 2018], different methods for deriving α and β were proposed (see Fig. 7). In particular, the linear model parameters α and β are obtained according to the following equations.
Number
Number
Number
Number
[0122] It has also been proposed to implement a division operation using multiplication by a number stored in a lookup table (LUT) specified in Table 1. This replacement is possible by using the following method.
Number
[0123] Table 1 provides a matching of the list of values stored in the LUT with the LUT index range (given in the first row of this table). Each list corresponds to its index range.
Number
Number
[0124] Using the LUT defined in Table 1 (or equivalently calculated using the above formula), the calculation of α is performed as follows.
Number
[0125] The shift parameter S can be decomposed into several parts, that is,
Number
Number
[0126] In this case, the linear model coefficient α has a fixed - point integer representation of a fractional value, and the precision of α is used in obtaining the value of the chroma prediction sample [Number] is determined by the value of [Number]
[0127] The size of the LUT is quite important in the hardware implementation of an encoder or decoder. The most straightforward way to solve this problem is to periodically subsample the LUT by keeping only every Nth (where N is the subsampling ratio) element of the initial LUT as specified, i.e., as in Table 1
[0128] After normal subsampling at a subsampling ratio that is a power of 2 of N, the fetches from the LUT are different, i.e., [Number] instead of [Number] is defined as
[0129] In the case of natural images, [Number] The probability has a small value and is found to be larger than the probability that this difference becomes large. In other words, the occurrence probability of the values in Table 1 decreases from the left column to the right column, and within each column, this probability decreases from the first element belonging to that column to the last element.
Number
[0130] Therefore, it is far from optimal to maintain only every Nth element of the initial LUT. This is because it corresponds to an equal probability distribution of its arguments that do not apply.
[0131] By considering this distribution, it is possible to achieve a better trade-off between the size of the LUT and the accuracy of the calculation than that provided by normal subsampling.
[0132] Specifically, it is proposed to define the LUT using non-linear indexing such that two adjacent LUT entries
Number
[0133] One of the computationally efficient solutions is
Number
Number
[0134] Specific embodiments are shown in FIGS. 9 and 10. FIG. 9 shows a flowchart of LUT value calculation, and FIG. 10 shows
Number
[0135] By using the steps shown in FIG. 9, it is possible to obtain the values stored in the LUT for further use in CCLM modeling. In FIG. 9, the "ctr" variable
Number
[0136] As shown in FIG. 9, the flowchart shows an exemplary look-up table generation process according to an embodiment of the present invention. In step 902, the video coding device starts the look-up table generation process. The video coding device may be a decoder such as the video decoder 30 of FIGS. 1A, 1B, and 3, or an encoder such as the video encoder 20 of FIGS. 1A, 1B, and 2, or the video coding device 400 of FIG. 4, or the device 500 of FIG. 5.
[0137] In step 904, set ctr = 1 and lut_shift = 1. In step 906, it is determined whether the index lut_shift < lut_shiftmax. If lut_shift < lut_shiftmax, the step is calculated in step 908 as 1<<max(0, lut_shift) and col = 0; otherwise, the generation process ends in step 922. In step 910, the start offset is provided by the step "ctr = ctr+(step>>1)". Then, it is determined in step 912 whether col < colmax. If col < colmax, in step 914, LUT[ctr]=(1<<S) / ctr, where one row of the LUT is generated by the pre-calculated LUT values defined from step 912 to step 918; otherwise, in step 920, set the index lut_shift = lut_shift + 1. In step 916, the value of ctr corresponding to the start point of each sub-range is set as ctrl + step. In step 918, the generation process moves to the next column, and then this process returns to step 912.
[0138] An exemplary LUT generated using the flowchart shown in FIG. 9 is given in Table 2. The rows of this table correspond to sub-ranges having the "lut_shift" index. This table is obtained using "lut_shift" equal to 6 and "col" equal to 3, and thus 48 entries are brought about in the LUT. The values of ctr corresponding to the start points of each sub-range (zero "col" values) are 0, 8, 17, 35, 71, 143, 287. max ", and "col" equal to 3 max ", and thus 48 entries are brought about in the LUT. The values of ctr corresponding to the start points of each sub-range (zero "col" values) are 0, 8, 17, 35, 71, 143, 287.
[0139] These values are not always powers of two because the subsampling within each subrange is performed relative to the center of the subrange. The corresponding start offset is provided by the "ctr=ctr+(step>>1)" step shown in FIG. 9. The value of "step" for "lut_shift" values not greater than 0 is set to be equal to 1. Table 2 Exemplary LUT Generated Using the Flowchart Shown in FIG. 9 [Table 2]
[0140] In FIG. 10(A), to determine the position of the corresponding entry in the LUT shown in Table 2, the binary representation 1001 of the input value (e.g., difference [Number] ) is being processed. The most significant non-zero bit of the input value is shown as 1002. The "msb" value is determined by the position of this bit. In fact, the msb value is log2() of the input value 1001. Subtracting 1 from "msb" gives the "lut_shift" value for selecting the row in Table 2. Otherwise, the step is calculated as "1<<lut_shift".
[0141] Selection of the column is performed by putting the "col max " bit next to 1002. The value of "col" in Table 2 is obtained as follows. The value of "high_bits" is obtained by selecting the "col max +1" bits following the most significant bit 1020, and col is set to be equal to the value stored in "high_bits" decremented by 1.
[0142] If the alignment step "ctr = ctr+(step >> 1)" is not executed in FIG. 9, the derivation of the "lut_shift" value is the same and the derivation of the "col" value is even simpler (FIG. 10(B)). The value of "col" is obtained by selecting the "col" max bits following the most significant bit 1020. This index derivation method corresponds to Table 3. The values of "ctr" (FIG. 9) corresponding to the start points (zero "col" values) of each sub-range are 0, 8, 16, 32, 64, 128, 256. Table 3 Another exemplary LUT generated using the flowchart shown in FIG. 9 when "ctr = ctr+(step >> 1)" is skipped
Table 3
[0143] When deriving the value of "col", the msb may be less than or equal to "col" max ". In this case, the value of "col" is set to be equal to the "col" max " least significant bit of the input difference 1001.
[0144] In a practical implementation, the LUT index is one-dimensional. The LUTs shown in Table 2 and Table 3 are
Number
[0145] Both of the LUTs shown in Tables 2 and 3 store values of very different magnitudes. Therefore, it is reasonable for all the values stored in the LUT to have the same number of bits. The value fetched from the LUT can be further left-shifted according to the value of lut_shift. The only exception to this rule is the first four values, which have a different precision from the last four values in this row. However, this problem can be solved by an additional LUT that stores this additional shift for the first four values. In this embodiment, the value of the multiplier is restored from the value fetched from the LUT as follows. [Number] Here, [Number] That is. The value of δ is set to be equal to 3, 2, 1, 1 respectively for "idx" values less than or equal to 4. The look-up table of this embodiment is given in Table 4. Table 4 Another exemplary LUT generated using the flowchart shown in FIG. 9 when the precision is equal within a plurality of ranges [Table 4]
[0146] It can be seen that the last rows of Table 4 are very similar to each other. Therefore, it is possible to reduce the size of the LUT by storing only one row for some settings of the sub-range. In a particular embodiment, when the value of lut_shift is greater than a particular threshold, the value of lut_shift is set to be equal to this threshold, and the value of δ is reduced by the difference between the initial value "lut_shift" and the threshold.
[0147] FIG. 11 is a flowchart showing an exemplary intra prediction of chroma samples of a block by applying a cross-component linear model. At step 1102, a video coding device obtains reconstructed luma samples. The video coding device may be a decoder such as video decoder 30 of FIGS. 1A, 1B, 3, or an encoder such as video encoder 20 of FIGS. 1A, 1B, 2, or video coding device 400 of FIG. 4, or device 500 of FIG. 5.
[0148] At step 1104, the video coding device determines the positions of the maximum and minimum reconstructed sample values within the reconstructed luma samples. For example, the reconstructed luma samples are adjacent reconstructed luma samples corresponding to chroma samples.
[0149] At step 1106, the video coding device obtains the value of the difference between the maximum and minimum values of the reconstructed luma samples.
[0150] At step 1108, the video coding device calculates an index of a look-up table (LUT) and fetches a value of a multiplier corresponding to the value of the difference between the maximum and minimum values of the reconstructed luma samples.
[0151] For example, the video coding device determines the position of the most significant bit of the difference between the maximum luma sample value and the minimum luma sample value, and fetches a value using a bit set following the position of the most significant bit of the difference between the maximum luma sample value and the minimum luma sample value as an index of the LUT. The position of the most significant bit of the difference between the maximum luma sample value and the minimum luma sample value may be obtained as the logarithm to the base 2 of the difference. The video coding device determines a bit set following the position of the most significant bit of the difference. As a possible result, the bit set includes four bits.
[0152] The LUT is generated regardless of the presence or absence of an alignment stage. The LUT may include at least two adjacent values stored in the LUT corresponding to different stages of the obtained difference, and the values of this stage may increase together with the difference value or be constant.
[0153] As disclosed in exemplary Tables 1 to 4, the LUT may include a sub-range of values. The stage of the difference value between the maximum value and the minimum value of the reconstructed luma sample is constant within a certain sub-range, and the stages of different sub-ranges are different. As an example, the stage of the difference value between the maximum value and the minimum value of the reconstructed luma sample increases as the sub-index increases. For example, the stage of the difference value between the maximum value and the minimum value of the reconstructed luma sample may be a power of 2 of the sub-index.
[0154] The LUT includes at least three values, namely a first value, a second value, and a third value. Among the three values, the first value and the second value are two adjacent values, and the second value and the third value are two adjacent values. The stage (i.e., precision or difference) between the first value and the second value may be equal to or different from the stage between the second value and the third value. When the first value is indexed by a first bit set and the second value is indexed by a second bit set, if the value of the first bit set is greater than the value of the second bit set, the first value is smaller than the second value, or if the value of the first bit set is smaller than the value of the second bit set, the first value is larger than the second value.
[0155] The LUT is divided into sub-ranges. The sub-index is determined using the position of the most significant non-zero bit of the difference between the maximum value and the minimum value of the reconstructed luma sample. As an example, the size of the sub-range is set to 8 and the number of sub-ranges is 6. As another example, different adjacent sub-ranges have different value increases for the same stage.
[0156] The LUT may include a non-linear index. Two adjacent LUT entries correspond to different levels of L(B)-L(A). L(B) represents the maximum value of the reconstructed luma sample, and L(A) represents the minimum value of the reconstructed luma sample. The level value of this entry may increase with the index of this entry.
[0157] When the LUT uses some of the most significant bits of L(B)-L(A), the position of the most significant bits of L(B)-L(A) defines the accuracy based on the position of the most significant bits (i.e., the level between two adjacent entries of the LUT). The larger the value of the position of the most significant bits, the lower the accuracy and the larger the level value may correspond.
[0158] In stage 1110, the video coding device obtains the linear model parameters α and β by multiplying the fetched value by the difference between the maximum value and the minimum value of the reconstructed chroma sample.
[0159] In stage 1112, the video coding device calculates the predicted chroma sample value using the obtained linear model parameters α and β.
[0160] FIG. 12 is a block diagram showing an exemplary structure of an apparatus 1200 for intra prediction of chroma samples of a block by applying a cross-component linear model. The apparatus 1200 is configured to execute the above method, an acquisition unit 1210 configured to acquire the reconstructed luma samples, a determination unit 1220 configured to determine a maximum luma sample value and a minimum luma sample value based on the reconstructed luma samples, and may include, The acquisition unit 1210 is further configured to acquire the difference between the maximum luma sample value and the minimum luma sample value. The decision unit 1220 is further configured to determine the position of the most significant bit of the difference between the maximum luma sample value and the minimum luma sample value. As an example, the position of the most significant bit of the difference between the maximum luma sample value and the minimum luma sample value is the logarithm to the base 2 of the difference. In an implementation, the most significant bit is the first non-zero bit.
[0161] The apparatus 1200 further includes a calculation unit 1230 configured to fetch a value from a look-up table (LUT) by using, as an index, a bit set following the position of the most significant bit of the difference between the maximum luma sample value and the minimum luma sample value, obtain linear model parameters α and β based on the fetched value, and calculate a predicted chroma sample value by using the obtained linear model parameters α and β. For example, the bit set includes four bits.
[0162] The calculation unit may obtain the linear model parameters α and β based on the fetched value and the difference between the maximum value and the minimum value of the reconstructed chroma samples. For example, the calculation unit obtains the linear model parameters α and β by multiplying the fetched value by the difference between the maximum value and the minimum value of the reconstructed chroma samples.
[0163] [Advantages of Embodiments of the Present Invention] 1. The index of the LUT is calculated in an elegant way that extracts some bits within the binary representation. As a result, the efficiency of fetching a value from the LUT is increased. 2. Since the efficiency of fetching a value from the LUT is increased, the size of the multiplier for obtaining the linear model parameters α and β is minimized. 3. The size of the LUT is minimized. The curve of the division function (f(x)=1 / x, known as a hyperbola) in the embodiments of the present invention is approximated in the following way. 1) The size of the LUT table may be 16. i. (Minimum number of entries for approximation of the 1 / x curve having a derivative that varies from 0 to infinity) 2) The elements of the LUT have a non-linear dependence on the entry index. i. (For approximating the 1 / x curve) 3) The multiplier (element of the LUT) is a 3-bit unsigned integer (0…7). i. (Minimum precision for approximating the 1 / x curve having a derivative that varies from 0 to infinity)
[0164] The following is an explanation of an encoding method, a decoding method, and applications thereof, and a system using them as shown in the embodiments mentioned above.
[0165] FIG. 13 is a block diagram showing a content supply system 3100 for realizing a content delivery service. This content supply system 3100 includes an imaging device 3102, a terminal device 3106, and optionally includes a display 3126. The imaging device 3102 communicates with the terminal device 3106 via a communication link 3104. The communication link may include the communication channel 13 described above. The communication link 3104 includes, but is not limited to, WiFi, Ethernet (registered trademark), cable, wireless (3G / 4G / 5G), or USB, or any combination of these types.
[0166] The imaging device 3102 can generate data and encode the data by an encoding method as shown in the above embodiment. Alternatively, the imaging device 3102 may deliver the data to a streaming server (not shown in the figure), and the server encodes the data and transmits the encoded data to the terminal device 3106. The imaging device 3102 includes, but is not limited to, a camera, a smartphone or a tablet, a computer or a laptop, a video conferencing system, a PDA, an in-vehicle device, or any combination thereof. For example, the imaging device 3102 may include the source device 12 as described above. When the data includes video, the video encoder 20 included in the imaging device 3102 may actually execute video encoding processing. When the data includes audio (i.e., voice), the audio encoder included in the imaging device 3102 may actually execute audio encoding processing. In some actual scenarios, the imaging device 3102 distributes the encoded video data and audio data by multiplexing them together. In other actual scenarios, for example, in a video conferencing system, the encoded audio data and the encoded video data are not multiplexed. The imaging device 3102 distributes the encoded audio data and the encoded video data to the terminal device 3106 separately.
[0167] In the content supply system 3100, the terminal device 310 receives and reproduces the encoded data. The terminal device 3106 may be a device having a data receiving and recovery function, for example, a smartphone or a pad 3108 capable of decoding the encoded data mentioned above, a computer or a laptop 3110, a network video recorder (NVR) / digital video recorder (DVR) 3112, a TV 3114, a set-top box (STB) 3116, a video conferencing system 3118, a video surveillance system 3120, a personal digital assistant (PDA) 3122, an in-vehicle device 3124, or any combination thereof. For example, the terminal device 3106 may include the destination device 14 as described above. When the encoded data includes video, the video decoder 30 included in the terminal device gives priority to performing video decoding. When the encoded data includes audio, the audio decoder included in the terminal device gives priority to performing audio decoding processing.
[0168] For a terminal device having a display, such as a smartphone or a pad 3108, a computer or a laptop 3110, a network video recorder (NVR) / digital video recorder (DVR) 3112, a TV 3114, a personal digital assistant (PDA) 3122 or an in-vehicle device 3124, the terminal device can supply the decoded data to its display. For a terminal device not equipped with a display, such as an STB 3116, a video conferencing system 3118 or a video surveillance system 3120, an external display 3126 is internally contacted to receive and display the decoded data.
[0169] When each device in this system performs encoding or decoding, an image encoding device or an image decoding device can be used as shown in the above-described embodiment.
[0170] FIG. 14 is a diagram showing an example structure of the terminal device 3106. After the terminal device 3106 receives a stream from the imaging device 3102, the protocol processing unit 3202 analyzes the transmission protocol of the stream. The protocol includes, but is not limited to, Real-Time Streaming Protocol (RTSP), Hypertext Transfer Protocol (HTTP), HTTP Live Streaming Protocol (HLS), MPEG-DASH, Real-Time Transport Protocol (RTP), Real-Time Messaging Protocol (RTMP), or any kind of combination thereof, etc.
[0171] After the protocol processing unit 3202 processes the stream, a stream file is generated. This file is output to the demultiplexing unit 3204. The demultiplexing unit 3204 can separate the multiplexed data into encoded audio data and encoded video data. As described above, in some actual scenarios, for example, in a video conferencing system, the encoded audio data and the encoded video data are not multiplexed. In this situation, the encoded data is sent to the video decoder 3206 and the audio decoder 3208 without passing through the demultiplexing unit 3204.
[0172] Through inverse multiplexing processing, a video elementary stream (ES), an audio ES, and optionally subtitles are generated. A video decoder 3206 including a video decoder 30 as described in the above-mentioned embodiments decodes the video ES by a decoding method as shown in the above-mentioned embodiments to generate video frames, and supplies this data to a synchronization unit 3212. An audio decoder 3208 decodes the audio ES to generate audio frames, and supplies this data to the synchronization unit 3212. Alternatively, the video frames can be stored in a buffer (not shown in FIG. 14) before being supplied to the synchronization unit 3212. Similarly, the audio frames can be stored in a buffer (not shown in FIG. 14) before being supplied to the synchronization unit 3212.
[0173] The synchronization unit 3212 synchronizes the video frames and the audio frames, and supplies the video / audio to a video / audio display 3214. For example, the synchronization unit 3212 synchronizes the presentation of video information and audio information. The information can be coded in syntax using time stamps related to the presentation of the coded audio data and visual data, and time stamps related to the delivery of the data stream itself.
[0174] If the stream includes subtitles, a subtitle decoder 3210 decodes the subtitles, synchronizes them with the video frames and the audio frames, and supplies the video / audio / subtitle to a video / audio / subtitle display 3216.
[0175] The present invention is not limited to the above-mentioned system, and any of the image encoding device or the image decoding device in the above-mentioned embodiments can be incorporated into other systems, such as an automotive system.
[0176] Although embodiments of the present invention have been mainly described based on video coding, embodiments of the coding system 10, the encoder 20, and the decoder 30 (and thus the system 10 accordingly), as well as other embodiments described herein, may be configured for the processing or coding of still images, i.e., individual images independent of any previous or consecutive images as in video coding. It should be noted that generally, when image processing coding is limited to a single image 17, only the inter prediction units 244 (encoder) and 344 (decoder) may not be available. All other functions (also referred to as tools or techniques) of the video encoder 20 and the video decoder 30 can be equally used for still image processing, e.g., residual calculation 204 / 304, transformation 206, quantization 208, inverse quantization 210 / 310, (inverse) transformation 212 / 312, segmentation 262 / 362, intra prediction 254 / 354, and / or loop filtering 220, 320 as well as entropy coding 270 and entropy decoding 304.
[0177] For example, the embodiments of the encoder 20 and the decoder 30, and the functions described herein with reference to, for example, the encoder 20 and the decoder 30, may be implemented in hardware, software, firmware, or any combination thereof. When implemented in software, these functions may be stored on a computer-readable medium, transmitted as one or more instructions or codes via a communication medium, and executed by a hardware-based processing unit. The computer-readable medium may include a computer-readable storage medium corresponding to a tangible medium such as a data storage medium, or a communication medium including any medium that facilitates transfer of a computer program from one location to another, for example, in accordance with a communication protocol. Thus, the computer-readable medium may generally correspond to (1) a non-transitory tangible computer-readable storage medium, or (2) a communication medium such as a signal or a carrier wave. The data storage medium may be any available medium that can be accessed by one or more computers or one or more processors to obtain instructions, codes, and / or data structures for implementation of the techniques described in the present disclosure. A computer program product may include a computer-readable medium.
[0178] By way of example and not limitation, such computer-readable storage media can include RAM, ROM, EEPROM, CD-ROM or other optical disk storage, magnetic disk storage or other magnetic storage devices, flash memory, or any other medium that can be used to store the desired program code in the form of instructions or data structures and that can be accessed by a computer. Also, any connection can be properly termed a computer-readable medium. For example, when instructions are transmitted from a website, server, or other remote source using coaxial cable, fiber optic cable, twisted pair, digital subscriber line (DSL), or wireless technologies such as infrared, radio, and microwave, the coaxial cable, fiber optic cable, twisted pair, DSL, or wireless technologies such as infrared, radio, and microwave are included in the definition of the medium. However, it should be understood that computer-readable storage media and data storage media do not include connections, carrier waves, signals, or other transient media, but rather are directed to non-transient tangible storage media. As used herein, disks (disk and disc) include compact discs (CDs), laser discs, optical discs, digital versatile discs (DVDs), floppy disks, and Blu-ray discs, where disk typically reproduces data magnetically, while disc reproduces data optically with a laser. Combinations of the above should also be included within the scope of computer-readable media.
[0179] The commands can be executed by one or more processors, such as one or more digital signal processors (DSPs), general-purpose microprocessors, application-specific integrated circuits (ASICs), field-programmable logic arrays (FPGAs), or other equivalent integrated or discrete logic circuits. Accordingly, the term "processor" as used herein may refer to any of the foregoing structures, or any other structure suitable for implementation of the techniques described herein. Additionally, in some aspects, the functions described herein may be provided within dedicated hardware modules and / or software modules configured for encoding and decoding, and / or incorporated within a combined codec. Further, these techniques may be implemented entirely in one or more circuits or logic elements.
[0180] The techniques of the present disclosure may be implemented in a variety of devices or apparatuses, including wireless handsets, integrated circuits (ICs) or IC sets (e.g., chip sets). Although various components, modules, or units are described in the present disclosure to emphasize functional aspects of devices configured to perform the disclosed techniques, implementation by different hardware units is not necessarily required. Rather, as described above, various units may be combined with a codec hardware unit along with suitable software and / or firmware, or may be provided by a set of interoperable hardware units including one or more processors as described above. [Other possible items] (Item 1) A method for intra predicting chroma samples of a block by applying a cross-component linear model, comprising: obtaining reconstructed luma samples; determining a maximum luma sample value and a minimum luma sample value based on the reconstructed luma samples; obtaining a difference between the maximum luma sample value and the minimum luma sample value; Determining a position of a most significant bit of the difference between the maximum luma sample value and the minimum luma sample value; Fetching a value from a look-up table (LUT) by using a bit set as an index, wherein the bit set follows the position of the most significant bit of the difference between the maximum luma sample value and the minimum luma sample value; Obtaining a linear model parameter α and a linear model parameter β based on the fetched value; Calculating a predicted chroma sample value by using the obtained linear model parameter α and the linear model parameter β A method comprising the steps above. (Item 2) The method according to Item 1, wherein the position of the most significant bit of the difference between the maximum luma sample value and the minimum luma sample value is obtained as a logarithm to the base 2 of the difference. (Item 3) The method according to Item 1 or 2, wherein determining the bit set following the position of the most significant bit of the difference results in a bit set including four bits. (Item 4) The method according to any one of Items 1 to 3, wherein the most significant bit is the first non-zero bit. (Item 5) Obtaining the linear model parameter α and the linear model parameter β based on the fetched value and a difference between a maximum value and a minimum value of the reconstructed chroma sample; The method according to any one of Items 1 to 4, comprising the steps above. (Item 6) Obtaining the linear model parameter α and the linear model parameter β by multiplying the fetched value by the difference between the maximum value and the minimum value of the reconstructed chroma sample; The method according to Item 5, comprising the steps above. (Item 7) The LUT includes at least three values, namely a first value, a second value, and a third value. Among the three values, the first value and the second value are two adjacent values, and the second value and the third value are two adjacent values. The method according to any one of items 1 to 6. (Item 8) The method according to item 7, wherein a step between the first value and the second value is equal to a step between the second value and the third value. (Item 9) The method according to item 7, wherein a step between the first value and the second value is different from a step between the second value and the third value. (Item 10) The first value is indexed by a first bit set, and the second value is indexed by a second bit set. If a value of the first bit set is greater than a value of the second bit set, the first value is smaller than the second value, or If a value of the first bit set is smaller than a value of the second bit set, the first value is greater than the second value. The method according to items 7 to 9. (Item 11) The LUT includes a plurality of sub-ranges of a plurality of values, and a step between any two adjacent values is constant within one sub-range. The method according to items 1 to 10. (Item 12) An apparatus for intra-predicting chroma samples of a block by applying a cross-component linear model, the apparatus being an encoder or a decoder, the apparatus comprising: An acquisition unit configured to acquire a reconstructed luma sample, the acquisition unit being further configured to acquire a difference between a maximum luma sample value and a minimum luma sample value. A determination unit configured to determine the maximum luma sample value and the minimum luma sample value based on the reconstructed luma sample, wherein the determination unit is further configured to determine the position of the most significant bit of the difference between the maximum luma sample value and the minimum luma sample value, and a determination unit By using a bit set following the position of the most significant bit of the difference between the maximum luma sample value and the minimum luma sample value as an index, a value is fetched from a look-up table (LUT), and based on the fetched value, a linear model parameter α and a linear model parameter β are obtained, and by using the obtained linear model parameter α and the linear model parameter β, a calculation unit configured to calculate a predicted chroma sample value An apparatus comprising: (Item 13) The apparatus according to item 12, wherein the position of the most significant bit of the difference between the maximum luma sample value and the minimum luma sample value is the base-2 logarithm of the difference. (Item 14) The apparatus according to item 12 or 13, wherein the bit set includes four bits. (Item 15) The apparatus according to any one of items 12 to 14, wherein the most significant bit is the first non-zero bit. (Item 16) The apparatus according to any one of items 12 to 15, wherein the calculation unit is configured to obtain the linear model parameter α and the linear model parameter β based on the fetched value and the difference between the maximum value and the minimum value of the reconstructed chroma sample. (Item 17) The calculation unit is configured to obtain the linear model parameter α and the linear model parameter β by multiplying the fetched value by the difference between the maximum value and the minimum value of the reconstructed chroma sample. The apparatus according to item 16. (Item 18) An encoder comprising a processing circuit for performing the method according to any one of items 1 to 11. (Item 19) A decoder comprising a processing circuit for performing the method according to any one of items 1 to 11. (Item 20) A computer program product comprising program code for performing the method according to any one of items 1 to 11. (Item 21) A decoder, one or more processors, a non-transitory computer-readable storage medium coupled to the processor and storing programming for execution by the processor, the programming configuring the decoder to perform the method according to any one of items 1 to 11 when executed by the processor, the non-transitory computer-readable storage medium comprising a decoder. (Item 22) An encoder, one or more processors, a non-transitory computer-readable storage medium coupled to the processor and storing programming for execution by the processor, the programming configuring the encoder to perform the method according to any one of items 1 to 11 when executed by the processor, the non-transitory computer-readable storage medium comprising an encoder. (Item 23) A non-transitory recording medium comprising a predicted encoded bitstream decoded by a device, the bitstream being generated according to any one of items 1 to 11.
Claims
**Claim 1** A method for generating, by an encoder, a bitstream in which chroma samples of blocks are intra-predicted, comprising: encoding the bitstream by applying a cross-component linear model; The encoding of the bitstream comprises: obtaining a reconstructed luma sample; determining a maximum luma sample value and a minimum luma sample value based on the reconstructed luma sample; obtaining a difference between the maximum luma sample value and the minimum luma sample value; obtaining an index according to log2() of the difference between the maximum luma sample value and the minimum luma sample value; fetching a value from a look-up table (LUT) by using the index; obtaining a linear model parameter α and a linear model parameter β based on the fetched value; calculating a predicted chroma sample value by using the obtained linear model parameter α and the linear model parameter β; The predicted chroma sample value satisfies C(x, y)=α×L(x, y)+β, where C(x, y) represents the predicted chroma sample value, and L(x, y) represents a luma sample value corresponding to the predicted chroma sample value. A method. **Claim 2** The method according to claim 1, wherein the index is determined by using some most significant bits of the difference between the maximum luma sample value and the minimum luma sample value. **Claim 3** The method according to claim 1 or 2, wherein the LUT comprises at least two adjacent values corresponding to different levels of the obtained difference, and the values of the levels increase or are constant together with the difference values. **Claim 4** The method according to any one of claims 1 to 3, further comprising obtaining the linear model parameter α and the linear model parameter β based on the fetched value and a difference between a maximum value and a minimum value of the reconstructed chroma sample. **Claim 5** The method according to claim 4, further comprising obtaining the linear model parameter α and the linear model parameter β by multiplying the fetched value by the difference between the maximum value and the minimum value of the reconstructed chroma sample. **Claim 6** The method according to any one of claims 1 to 5, wherein the LUT includes a plurality of sub-ranges of a plurality of values, and the step between two adjacent values is constant within one sub-range.
7. A method for intra-predicting chroma samples of a block by a decoder, comprising: analyzing a bitstream by applying a cross-component linear model; comprising: wherein the step of analyzing the bitstream obtaining reconstructed luma samples; determining a maximum luma sample value and a minimum luma sample value based on the reconstructed luma samples; obtaining a difference between the maximum luma sample value and the minimum luma sample value; obtaining an index according to log2() of the difference between the maximum luma sample value and the minimum luma sample value; fetching a value from a look-up table (LUT) by using the index; obtaining linear model parameter α and linear model parameter β based on the fetched value; calculating a predicted chroma sample value by using the obtained linear model parameter α and the linear model parameter β. comprising: wherein the predicted chroma sample value satisfies C(x, y)=α×L(x, y)+β, C(x, y) represents the predicted chroma sample value, and L(x, y) represents a luma sample value corresponding to the predicted chroma sample value. Method.
8. The method according to claim 7, wherein the index is determined by using some most significant bits of the difference between the maximum luma sample value and the minimum luma sample value.
9. The method according to claim 7 or 8, wherein the LUT includes at least two adjacent values corresponding to different steps of the obtained difference, and the values of the steps increase or are constant together with the difference values.
10. obtaining the linear model parameter α and the linear model parameter β based on the fetched value and a difference between a maximum value and a minimum value of the reconstructed chroma samples. The method according to any one of claims 7 to 9, comprising:
11. obtaining the linear model parameter α and the linear model parameter β by multiplying the fetched value by the difference between the maximum value and the minimum value of the reconstructed chroma samples. The method according to claim 10, comprising
12. The method according to any one of claims 7 to 11, wherein the LUT comprises a plurality of sub-ranges of a plurality of values, and the step between two adjacent values is constant within one sub-range.
13. An encoder comprising a processing circuit for performing the method according to any one of claims 1 to 6.
14. A decoder comprising a processing circuit for performing the method according to any one of claims 7 to 12.
15. A program for causing a processor to perform the method according to any one of claims 1 to 12.
16. An encoder, comprising one or more processors, and a non-transitory computer-readable storage medium coupled to the one or more processors and storing programming for execution by the one or more processors, the programming configuring the encoder to perform the method according to any one of claims 1 to 6 when executed by the one or more processors An encoder comprising
17. A decoder, comprising one or more processors, and a non-transitory computer-readable storage medium coupled to the one or more processors and storing programming for execution by the one or more processors, the programming configuring the decoder to perform the method according to any one of claims 7 to 12 when executed by the one or more processors A decoder comprising
Citation Information
Patent Citations
Method and device for predicting color difference components of an image using the luminance component of the image
JP2014523698A
Complexity reduction in parameter derivation for intra prediction
WO2020094059A1
Method and apparatus for video encoding or decoding
WO2020131512A1