Image encoding / decoding method and apparatus, and recording medium storing bit stream
By generating virtual samples and deriving compensation parameters, the intra-frame prediction and inter-frame prediction in image compression technology are optimized, solving the problem of insufficient prediction accuracy for high-resolution and high-quality images and achieving more efficient coding results.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- LG ELECTRONICS INC
- Filing Date
- 2024-09-23
- Publication Date
- 2026-04-24
Smart Images

Figure CN121925845A_ABST
Abstract
Description
Technical Field
[0001] This disclosure relates to an image encoding / decoding method and apparatus, as well as a recording medium for storing bit streams. Background Technology
[0002] Recently, the demand for high-resolution and high-quality images, such as HD (high-definition) and UHD (ultra-high-definition) images, has been increasing in various application areas, and therefore, efficient image compression technologies are being discussed.
[0003] Various techniques exist, such as inter-frame prediction techniques that use video compression technology to predict pixel values included in the current frame from frames before or after the current frame, intra-frame prediction techniques that use pixel information in the current frame to predict pixel values included in the current frame, and entropy coding techniques that assign short symbols to values that occur frequently and long symbols to values that occur infrequently. These image compression techniques can be used to effectively compress image data and send or store it. Summary of the Invention
[0004] Technical issues
[0005] This disclosure provides a method and apparatus for refining prediction samples obtained through intra-frame prediction or inter-frame prediction.
[0006] This disclosure provides a model-based compensation method and apparatus for predicted samples.
[0007] This disclosure provides a method and apparatus for generating virtual samples for model-based compensation.
[0008] This disclosure provides a method and apparatus for deriving parameters for model-based compensation.
[0009] This disclosure provides a method and apparatus for signaling information about model-based compensation.
[0010] Technical solution
[0011] The image decoding method and apparatus according to the present invention can obtain a predicted sample of the current block, generate a virtual sample of the current block, derive model-based compensation parameters for the current block based on the predicted sample of the current block and the generated virtual sample, obtain a modified predicted sample by applying the parameters to the predicted sample of the current block, and reconstruct the current block based on the modified predicted sample of the current block and the residual sample of the current block.
[0012] In the image decoding method and apparatus according to this disclosure, the virtual samples can be generated based on a planar pattern for intra-frame prediction.
[0013] In the image decoding method and apparatus according to this disclosure, the virtual samples can be generated based on DC patterns for intra-frame prediction.
[0014] In the image decoding method and apparatus according to this disclosure, the virtual sample can be generated by a weighted sum of a sample generated by applying a predetermined prediction method to the current block and a reference sample of the current block. Here, the reference sample may include at least one of the upper neighbor sample or the left neighbor sample adjacent to the current block.
[0015] In the image decoding method and apparatus according to the present disclosure, the virtual sample can be generated based on a predetermined intra-frame prediction mode, and the predetermined intra-frame prediction mode can be derived based on decoder-side intra-frame mode derivation (DIMD) or template-based intra-frame mode derivation (TIMD).
[0016] In the image decoding method and apparatus according to this disclosure, the virtual samples can be generated based on intra-template matching prediction (intra-TMP).
[0017] In the image decoding method and apparatus according to the present disclosure, the virtual samples can be generated based on matrix-based intra-frame prediction (MIP).
[0018] In the image decoding method and apparatus according to this disclosure, the virtual sample can be generated for the sample position of a portion of the current block.
[0019] In the image decoding method and apparatus according to this disclosure, the parameters may be parameters of a linear model or parameters of a convolutional model.
[0020] In the image decoding method and apparatus according to this disclosure, the parameter can be derived as a parameter that minimizes a predetermined cost. Here, the cost can be defined for each sample position within the current block based on the difference between the modified predicted sample and the virtual sample.
[0021] In the image decoding method and apparatus according to this disclosure, a modified prediction sample can be obtained by applying the parameters to the prediction sample and neighboring samples adjacent to the prediction sample. The neighboring samples may include at least one of the top neighboring sample, bottom neighboring sample, left neighboring sample, right neighboring sample, top-left neighboring sample, top-right neighboring sample, bottom-left neighboring sample, or bottom-right neighboring sample of the prediction sample.
[0022] According to the image coding method and apparatus of this disclosure, a predicted sample of the current block can be obtained, a virtual sample of the current block can be generated, parameters for model-based compensation for the current block can be derived based on the predicted sample of the current block and the generated virtual sample, a modified predicted sample can be obtained by applying the parameters to the predicted sample of the current block, a residual sample of the current block can be generated based on the modified predicted sample of the current block, and the residual sample of the current block can be encoded.
[0023] A computer-readable digital storage medium is provided for storing encoded video / image information, which enables a decoding device according to the present disclosure to perform an image decoding method.
[0024] A computer-readable digital storage medium is provided for storing video / image information generated based on the image encoding method of this disclosure.
[0025] A method and apparatus are provided for transmitting video / image information generated according to the image encoding method of this disclosure.
[0026] Beneficial effects
[0027] According to this disclosure, prediction accuracy can be improved by refining the prediction samples via intra-frame prediction or inter-frame prediction.
[0028] According to this disclosure, prediction accuracy can be improved by model-based compensation for the predicted samples.
[0029] According to this disclosure, coding efficiency can be improved by proposing a method for generating virtual samples and a method for deriving parameters for model-based compensation.
[0030] According to this disclosure, information about model-based compensation can be effectively communicated using signals. Attached Figure Description
[0031] Figure 1 A video / image encoding system according to this disclosure is shown.
[0032] Figure 2 A schematic block diagram of an encoding apparatus to which embodiments of the present disclosure are applicable and which performs encoding of video / image signals is shown.
[0033] Figure 3 A schematic block diagram of a decoding apparatus to which embodiments of the present disclosure are applicable and which performs decoding of video / image signals is shown.
[0034] Figure 4 A model-based compensation method performed by a decoding device 300 as an embodiment of the present disclosure is illustrated.
[0035] Figure 5 An illustrative configuration of a decoding device 300 that performs a model-based compensation method according to the present disclosure is shown.
[0036] Figure 6 An example is given of a model-based compensation method performed by an encoding device 200 as an embodiment of the present disclosure.
[0037] Figure 7 An illustrative configuration of an encoding device 200 that performs a model-based compensation method according to the present disclosure is shown.
[0038] Figure 8 Examples of content streaming systems to which embodiments of the present disclosure can be applied are shown. Detailed Implementation
[0039] Because this disclosure can be modified in various ways and has multiple embodiments, specific embodiments will be shown in the accompanying drawings and described in detail in the specific embodiments. However, this disclosure is not intended to be limited to the specific embodiments, but should be understood to include all changes, equivalents, and substitutions included within the spirit and scope of this disclosure. Similar reference numerals are used for similar components in the description of the various figures.
[0040] Terms such as "first," "second," etc., may be used to describe various components, but components should not be limited by these terms. Terms are used only to distinguish one component from others. For example, without departing from the scope of this disclosure, a first component may be referred to as a second component, and similarly, a second component may be referred to as a first component. Terms include any combination of one or more of the associated terms.
[0041] When a component is described as "connected" or "linked" to another component, it should be understood that it can be directly connected or linked to the other component, but the other component can exist in between. On the other hand, when a component is described as "directly connected" or "directly linked" to another component, it should be understood that there is no other component in between.
[0042] The terminology used in this application is for describing particular embodiments only and is not intended to limit this disclosure. Singular expressions include plural expressions unless the context clearly indicates otherwise. In this application, it should be understood that terms such as “comprising” or “having” are intended to specify the presence of the features, quantities, steps, operations, components, portions, or combinations thereof described in this specification, but do not preclude the possibility of the presence or addition of one or more other features, quantities, steps, operations, components, portions, or combinations thereof.
[0043] This disclosure relates to video / image coding. For example, the methods / implementations disclosed herein can be applied to methods disclosed in the Multifunctional Video Coding (VVC) standard. Additionally, the methods / implementations disclosed herein can be applied to methods disclosed in the Basic Video Coding (EVC) standard, the AOMedia Video 1 (AV1) standard, the Audio Video Coding 2 (AVS2) standard, or next-generation video / image coding standards (e.g., H.267 or H.268).
[0044] This specification sets forth various implementations of video / image encoding, and unless otherwise specified, these implementations may be combined with each other.
[0045] In this article, video can refer to a collection of images over time. A frame typically refers to a unit representing an image within a specific time period, and a slice / tile is a unit that forms part of a frame during encoding. A slice / tile can include at least one Code Tree Unit (CTU). A frame can consist of at least one slice / tile. A tile is a rectangular area composed of multiple CTUs within a specific tile column and a specific tile row of a frame. A tile column is a rectangular area of CTUs with a height equal to the height of the frame and a width specified by the syntax requirements of the frame parameter set. A tile row is a rectangular area of CTUs with a height specified by the frame parameter set and a width equal to the width of the frame. CTUs within a tile can be arranged continuously according to CTU raster scans, and tiles within a frame can be arranged continuously according to tile raster scans. A slice can include an integer number of complete tiles of a frame that can be exclusively included in a single NAL unit, or an integer number of consecutive complete CTU rows within a tile. Furthermore, a frame can be divided into at least two sub-frames. A sub-frame can be a rectangular area of at least one slice within a frame.
[0046] A pixel, or pelin, can represent the smallest unit that makes up a frame (or image). Additionally, "sample" can be used as the term corresponding to a pixel. A sample can typically represent a pixel or a pixel value, and can represent only the pixel / pixel value of the luminance component or only the pixel / pixel value of the chrominance component.
[0047] A unit can represent the basic unit of image processing. A unit may include a specific region of the image and at least one of the information associated with that region. A unit may include a luminance block and two chrominance (e.g., cb, cr) blocks. In some cases, the term "unit" may be used interchangeably with terms such as "block" or "region". In general, an M×N block may include a set (or array) of transform coefficients or samples (or sample arrays) consisting of M columns and N rows.
[0048] In this document, “A or B” can mean “A only”, “B only”, or “both A and B”. In other words, “A or B” can be interpreted as “A and / or B”. For example, “A, B or C” can mean “A only”, “B only”, “C only”, or “any combination of A, B and C”.
[0049] The forward slash ( / ) or comma used in this article can indicate "and / or". For example, "A / B" can mean "A and / or B". Therefore, "A / B" can mean "A only", "B only", or "both A and B". For example, "A, B, C" can mean "A, B, or C".
[0050] In this document, "at least one of A and B" can mean "only A", "only B" or "both A and B". Furthermore, in this document, expressions such as "at least one of A or B" or "at least one of A and / or B" can be interpreted in the same way as "at least one of A and B".
[0051] Additionally, in this document, "at least one of A, B, and C" can mean "A only", "B only", "C only" or "any combination of A, B, and C". Furthermore, "at least one of A, B, or C" or "at least one of A, B, and / or C" can mean "at least one of A, B, and C".
[0052] Additionally, the parentheses used in this document can indicate "for example". Specifically, when the indication is "prediction (intra-frame prediction)", "intra-frame prediction" can be cited as an example of "prediction". In other words, "prediction" in this document is not limited to "intra-frame prediction", and "intra-frame prediction" can be cited as an example of "prediction". Furthermore, even when the indication is "prediction (i.e., intra-frame prediction)", "intra-frame prediction" can be cited as an example of "prediction".
[0053] In this article, the technical features described individually in a single diagram can be implemented individually or simultaneously.
[0054] Figure 1 A video / image encoding system according to this disclosure is shown.
[0055] Reference Figure 1 A video / image encoding system may include a first device (source device) and a second device (receiving device).
[0056] A source device can transmit encoded video / image information or data to a receiving device in the form of a file or stream via a digital storage medium or network. The source device may include a video source, an encoding device, and a transmitting unit. The receiving device may include a receiving unit, a decoding device, and a renderer. The encoding device may be referred to as a video / image encoding device, and the decoding device may be referred to as a video / image decoding device. The transmitter may be included in the encoding device. The receiver may be included in the decoding device. The renderer may include a display unit, and the display unit may consist of a separate device or external components.
[0057] A video source can acquire video / images through processes that capture, synthesize, or generate video / images. A video source may include means for capturing video / images and means for generating video / images. Means for capturing video / images may include at least one camera, a video / image archive containing previously captured video / images, etc. Means for generating video / images may include a computer, tablet computer, smartphone, etc., and can generate video / images (electronically). For example, virtual video / images can be generated by a computer, etc., and in this case, the process of capturing video / images can be replaced by a process of generating related data.
[0058] Encoding devices can encode input video / images. They can perform a series of processes such as prediction, transformation, and quantization for compression and encoding efficiency. The encoded data (encoded video / image information) can be output as a bitstream.
[0059] The transmitting unit can send encoded video / image information or data, output in bitstream form, to the receiving unit of the receiving device in the form of a file or stream via a digital storage medium or network. The digital storage medium can include various storage media such as USB, SD, CD, DVD, Blu-ray, HDD, SSD, etc. The transmitting unit can include elements for generating media files according to a predetermined file format, and may include elements for transmission via a broadcast / communication network. The receiving unit can receive / extract the bitstream and send it to a decoding device.
[0060] Decoding devices can decode video / images by performing a series of processes, such as dequantization, inverse transform, and prediction, that correspond to the operations of encoding devices.
[0061] The renderer can render decoded video / images. The rendered video / images can be displayed through a display unit.
[0062] Figure 2 A rough block diagram of an encoding apparatus that can be applied to embodiments of the present disclosure and perform encoding of video / image signals is shown.
[0063] Reference Figure 2The encoding device 200 may consist of an image segmenter 210, a predictor 220, a residual processor 230, an entropy encoder 240, an adder 250, a filter 260, and a memory 270. The predictor 220 may include an inter-frame predictor 221 and an intra-frame predictor 222. The residual processor 230 may include a transformer 232, a quantizer 233, a dequantizer 234, and an inverse transformer 235. The residual processor 230 may also include a subtractor 231. The adder 250 may be referred to as a reconstructor or a reconstruction block generator. According to embodiments, the image segmenter 210, predictor 220, residual processor 230, entropy encoder 240, adder 250, and filter 260 may be configured by at least one hardware component (e.g., an encoder chipset or processor). Additionally, the memory 270 may include a decoded picture buffer (DPB) and may be configured by a digital storage medium. The hardware component may also include the memory 270 as an internal / external component.
[0064] Image segmenter 210 can segment an input image (or picture or frame) input to encoding device 200 into at least one processing unit. As an example, a processing unit can be referred to as a coding unit (CU). In this case, the coding unit can be recursively segmented from coding tree unit (CTU) or maximum coding unit (LCU) according to a quadtree-binary-tritree (QTBTTT) structure.
[0065] For example, a coding unit can be segmented into multiple deeper coding units based on a quadtree, binary tree, and / or ternary tree structure. In this case, for example, a quadtree structure can be applied first, followed by a binary tree and / or ternary tree structure. Alternatively, a binary tree structure can be applied before the quadtree structure. The coding process according to this specification can be performed based on the final coding unit that is no longer segmented. In this case, based on image characteristics, coding efficiency, etc., the largest coding unit can be directly used as the final coding unit, or, if necessary, the coding unit can be recursively segmented into deeper coding units, and the coding unit with the optimal size can be used as the final coding unit. Here, the coding process can include processes such as prediction, transformation, and reconstruction, as described later.
[0066] As another example, the processing unit may also include a prediction unit (PU) or a transform unit (TU). In this case, the prediction unit and the transform unit can be divided or segmented from the aforementioned final encoding unit, respectively. The prediction unit may be a unit for predicting samples, and the transform unit may be a unit for deriving transform coefficients and / or a unit for deriving residual signals from transform coefficients.
[0067] In some cases, a unit can be used interchangeably with terms such as block or region. Generally, an M×N block can represent a set of transform coefficients or samples consisting of M columns and N rows. Samples can typically represent pixels or pixel values, and can represent only the pixel / pixel value of the luminance component, or only the pixel / pixel value of the chrominance component. Samples can be used as a term to form a frame (or image) corresponding to a pixel or cell.
[0068] Encoding device 200 can subtract the prediction signal (prediction block, prediction sample array) output from inter-frame predictor 221 or intra-frame predictor 222 from the input image signal (original block, original sample array) to generate a residual signal (residual signal, residual sample array), and the generated residual signal is sent to converter 232. In this case, the unit in encoding device 200 that subtracts the prediction signal (prediction block, prediction sample array) from the input image signal (original block, original sample array) can be called subtractor 231.
[0069] Predictor 220 can perform prediction on the block to be processed (hereinafter referred to as the current block) and generate a prediction block that includes prediction samples of the current block. Predictor 220 can determine whether to apply intra-frame prediction or inter-frame prediction on a per-block or per-unit basis. Predictor 220 can generate various information about the prediction (e.g., prediction mode information) and send it to entropy encoder 240, as described later in the description of the various prediction modes. The information about the prediction can be encoded in entropy encoder 240 and output as a bitstream.
[0070] Intra-predictor 222 can predict the current block by referencing samples within the current frame. Depending on the prediction mode, the referenced samples can be located near the current block or positioned at a specific distance away from the current block. In intra-prediction, the prediction mode can include at least one non-directional mode and multiple directional modes. The non-directional mode can include at least one of a DC mode or a planar mode. Depending on the level of detail of the prediction direction, the directional modes can include 33 or 65 directional modes. However, this is just an example; more or fewer directional modes can be used depending on the configuration. Intra-predictor 222 can determine the prediction mode applied to the current block by using prediction modes applied to neighboring blocks.
[0071] Inter-frame predictor 221 can deduce the predicted block of the current block based on a reference block (reference sample array) specified by motion vectors on a reference frame. In this case, to reduce the amount of motion information transmitted in inter-frame prediction mode, motion information can be predicted on a block, sub-block, or sample basis based on the correlation between motion information between neighboring blocks and the current block. Motion information may include motion vectors and reference frame indices. Motion information may also include inter-frame prediction direction information (L0 prediction, L1 prediction, Bi prediction, etc.). For inter-frame prediction, neighboring blocks may include spatially neighboring blocks existing in the current frame and temporally neighboring blocks existing in the reference frame. The reference frame including the reference block and the reference frame including the temporally neighboring block may be the same or different. The temporally neighboring block may be referred to as a co-located reference block, a co-located CU (colCU), etc., and the reference frame including the temporally neighboring block may be referred to as a co-located frame (colPic). For example, inter-frame predictor 221 can configure a motion information candidate list based on neighboring blocks and generate information indicating which candidate to use to deduce the motion vector and / or reference frame index of the current block. Inter-frame prediction can be performed based on various prediction modes. For example, in skip mode and merge mode, the inter-frame predictor 221 can use motion information from neighboring blocks as motion information for the current block. In skip mode, unlike merge mode, residual signals may not be sent. In motion vector prediction (MVP) mode, motion vectors from surrounding blocks are used as motion vector predictors, and the motion vector difference is signaled to indicate the motion vector of the current block.
[0072] Predictor 220 can generate a prediction signal based on various prediction methods described later. For example, the predictor can not only apply intra-frame prediction or inter-frame prediction to predict a block, but can also apply both intra-frame and inter-frame prediction simultaneously. This can be referred to as the Inter-intra-frame Combined Prediction (CIIP) mode. Alternatively, the predictor can predict blocks based on the Intra-Block Copy (IBC) prediction mode, or it can predict blocks based on a palette mode. The IBC prediction mode or palette mode can be used for content image / video coding such as Screen Content Coding (SCC) in games, etc. IBC essentially performs prediction within the current frame, but it can be performed similarly to inter-frame prediction because it derives a reference block within the current frame. In other words, IBC can use at least one of the inter-frame prediction techniques described herein. The palette mode can be considered an example of intra-frame coding or intra-frame prediction. When applying a palette mode, sample values within the frame can be signaled based on information about the palette table and palette index. The prediction signal generated by predictor 220 can be used to generate a reconstructed signal or a residual signal.
[0073] Transformer 232 can generate transform coefficients by applying transform techniques to the residual signal. For example, the transform techniques may include at least one of Discrete Cosine Transform (DCT), Discrete Sine Transform (DST), Karhunen–Loève Transform (KLT), Graph-Based Transform (GBT), or Conditional Nonlinear Transform (CNT). Here, GBT represents the transform obtained from a graph when the relationship information between pixels is represented as a graph. CNT represents the transform obtained based on generating a prediction signal using all previously reconstructed pixels. Furthermore, the transform processing can be applied to square pixel blocks of the same size, or it can be applied to non-square blocks of variable size.
[0074] Quantizer 233 can quantize the transform coefficients and send them to entropy encoder 240, which can encode the quantized signal (information about the quantized transform coefficients) and output it as a bitstream. This information about the quantized transform coefficients can be called residual information. Quantizer 233 can rearrange the block-form quantized transform coefficients into a one-dimensional vector based on the coefficient scan order, and can generate information about the quantized transform coefficients based on this one-dimensional vector form.
[0075] The entropy encoder 240 can perform various encoding methods such as Golomb, context-adaptive variable-length coding (CAVLC), and context-adaptive binary arithmetic coding (CABAC). The entropy encoder 240 can encode information required for video / image reconstruction other than quantization transform coefficients (e.g., values of syntax elements, etc.) together or separately.
[0076] Encoded information (e.g., encoded video / image information) can be transmitted or stored in bitstream form at the Network Abstraction Layer (NAL) unit level. The video / image information may also include information about various parameter sets such as Adaptive Parameter Set (APS), Picture Parameter Set (PPS), Sequence Parameter Set (SPS), or Video Parameter Set (VPS). Additionally, the video / image information may include general constraint information. Information and / or syntax elements transmitted from the encoding device to / signaled to the decoding device may be included in the video / image information. The video / image information can be encoded and included in the bitstream through the encoding process described above. The bitstream can be transmitted over a network or stored in a digital storage medium. Here, the network may include broadcast networks and / or communication networks, and the digital storage medium may include various storage media such as USB, SD, CD, DVD, Blu-ray, HDD, SSD, etc. A transmitting unit (not shown) for transmitting the signal output from the entropy encoder 240 and / or a storage unit (not shown) for storing the signal may be configured as internal / external components of the encoding device 200, or the transmitting unit may also be included in the entropy encoder 240.
[0077] The quantized transform coefficients output from quantizer 233 can be used to generate a prediction signal. For example, the residual signal (residual block or residual sample) can be reconstructed by applying dequantization and inverse transform to the quantized transform coefficients via dequantizer 234 and inverse transformer 235. Adder 250 can add the reconstructed residual signal to the prediction signal output from inter-frame predictor 221 or intra-frame predictor 222 to generate a reconstructed signal (reconstructed frame, reconstructed block, reconstructed sample array). When there is no residual for the block to be processed (similar to when a skip mode is applied), the prediction block can be used as a reconstructed block. Adder 250 can be referred to as a reconstructor or reconstructed block generator. The generated reconstructed signal can be used for intra-frame prediction of the next block to be processed in the current frame, and can also be used for inter-frame prediction of the next frame by filtering, as described later. Furthermore, luminance mapping and chroma scaling (LMCS) can be applied in frame encoding and / or reconstruction processing.
[0078] Filter 260 can improve subjective / objective image quality by applying filtering to the reconstructed signal. For example, filter 260 can generate a modified reconstructed image by applying various filtering methods to the reconstructed image, and the modified reconstructed image can be stored in memory 270 (specifically, the DPB of memory 270). Various filtering methods can include deblocking filtering, sample adaptive offsetting, adaptive loop filtering, bilateral filtering, etc. Filter 260 can generate various information about the filtering and send it to entropy encoder 240. The information about the filtering can be encoded in entropy encoder 240 and output as a bitstream.
[0079] The modified reconstructed frame sent to memory 270 can be used as a reference frame in inter-frame predictor 221. When inter-frame prediction is applied through it, the encoding device can avoid prediction mismatch between the encoding device 200 and the decoding device, and can also improve encoding efficiency.
[0080] The DPB of memory 270 can store modified reconstructed frames for use as reference frames in inter-frame predictor 221. Memory 270 can store motion information of blocks in the current frame from which motion information is derived (or encoded) and / or of blocks in previously reconstructed frames. The stored motion information can be sent to inter-frame predictor 221 as motion information for spatially or temporally neighboring blocks. Memory 270 can store reconstructed samples of reconstructed blocks in the current frame and send them to intra-frame predictor 222.
[0081] Figure 3 A rough block diagram of a decoding device that can be implemented using embodiments of the present disclosure and perform decoding of video / image signals is shown.
[0082] Reference Figure 3 The decoding device 300 can be configured to include an entropy decoder 310, a residual processor 320, a predictor 330, an adder 340, a filter 350, and a memory 360. The predictor 330 may include an inter-frame predictor 332 and an intra-frame predictor 331. The residual processor 320 may include a dequantizer 321 and an inverse transformer 322.
[0083] According to the implementation, the entropy decoder 310, residual processor 320, predictor 330, adder 340, and filter 350 described above can be configured by a single hardware component (e.g., a decoder chipset or processor). Additionally, the memory 360 may include a decoded screen buffer (DPB) and can be configured by a digital storage medium. The hardware component may also include the memory 360 as an internal / external component.
[0084] When the input includes a bitstream containing video / image information, the decoding device 300 can respond to... Figure 2 The encoding device processes video / image information to reconstruct the image. For example, the decoding device 300 can deduce units / blocks based on block segmentation information obtained from the bitstream. The decoding device 300 can perform decoding using processing units applied in the encoding device. Therefore, the decoding processing unit can be an encoding unit, and the encoding unit can be segmented from the encoding tree unit or a larger encoding unit according to a quadtree structure, binary tree structure, and / or ternary tree structure. At least one transform unit can be derived from the encoding unit. Furthermore, the reconstructed image signal decoded and output by the decoding device 300 can be played back by a playback device.
[0085] Decoding device 300 can receive data in bitstream form from... Figure 2The signal output by the encoding device can be decoded by the entropy decoder 310. For example, the entropy decoder 310 can parse the bitstream to derive the information (e.g., video / image information) required for image reconstruction (or picture reconstruction). The video / image information may also include information about various parameter sets such as Adaptive Parameter Set (APS), Picture Parameter Set (PPS), Sequence Parameter Set (SPS), or Video Parameter Set (VPS). In addition, the video / image information may also include general constraint information. The decoding device can also decode the picture based on the information about the parameter sets and / or general constraint information. The information and / or syntax elements that are signaled / received, as described later herein, can be decoded and obtained from the bitstream through the decoding process. For example, the entropy decoder 310 can decode the information in the bitstream based on encoding methods such as Exponential Golomb coding, CAVLC, CABAC, etc., and output the values of the syntax elements required for image reconstruction and the quantized values of the transform coefficients of the residuals. More specifically, the CABAC entropy decoding method can receive bins corresponding to each syntax element from the bitstream, determine a context model using information about the syntax element to be decoded, decoding information of surrounding blocks and the block to be decoded, or information about symbols / bins decoded in previous steps, perform arithmetic decoding of bins by predicting the occurrence probability of bins based on the determined context model, and generate symbols corresponding to the values of each syntax element. In this case, after determining the context model, the CABAC entropy decoding method can update the context model by using information about the decoded symbols / bins for the context model of the next symbol / bin. Among the information decoded in the entropy decoder 310, information about prediction is provided to the predictors (inter-frame predictor 332 and intra-frame predictor 331), and the residual values (i.e., quantization transform coefficients and related parameter information) from which entropy decoding has been performed in the entropy decoder 310 can be input to the residual processor 320. The residual processor 320 can derive residual signals (residual blocks, residual samples, residual sample arrays). In addition, information about filtering from the information decoded in the entropy decoder 310 can be provided to the filter 350. Furthermore, the receiving unit (not shown) that receives the signal output from the encoding device can be further configured as an internal / external component of the decoding device 300, or the receiving unit can be a component of the entropy decoder 310.
[0086] Furthermore, the decoding device according to this specification may be referred to as a video / image / screen decoding device, and the decoding device may be divided into an information decoder (video / image / screen information decoder) and a sample decoder (video / image / screen sample decoder). The information decoder may include an entropy decoder 310, and the sample decoder may include at least one of a dequantizer 321, an inverse transformer 322, an adder 340, a filter 350, a memory 360, an inter-frame predictor 332, and an intra-frame predictor 331.
[0087] Dequantizer 321 can dequantize the quantized transform coefficients and output the transform coefficients. Dequantizer 321 can rearrange the quantized transform coefficients into two-dimensional blocks. In this case, the rearrangement can be performed based on the coefficient scan order performed in the encoding device. Dequantizer 321 can perform dequantization on the quantized transform coefficients using quantization parameters (e.g., quantization step size information) and obtain the transform coefficients.
[0088] The inverse transformer 322 performs an inverse transformation on the transformation coefficients to obtain the residual signal (residual block, residual sample array).
[0089] Predictor 320 can perform prediction on the current block and generate a prediction block that includes prediction samples of the current block. Predictor 320 can determine whether to apply intra-frame prediction or inter-frame prediction to the current block based on the prediction information output from entropy decoder 310, and determine the specific intra-frame / inter-frame prediction mode.
[0090] Predictor 320 can generate prediction signals based on various prediction methods described later. For example, predictor 320 can not only apply intra-frame prediction or inter-frame prediction to predict a block, but can also apply intra-frame prediction and inter-frame prediction simultaneously. This can be referred to as the Inter-Frame Intra-Frame Combined Prediction (CIIP) mode. Alternatively, the predictor can predict blocks based on the Intra-Frame Block Copy (IBC) prediction mode, or it can predict blocks based on a palette mode. The IBC prediction mode or palette mode can be used for content image / video coding such as Screen Content Coding (SCC) in games, etc. IBC essentially performs prediction within the current frame, but it can be performed similarly to inter-frame prediction because it derives a reference block within the current frame. In other words, IBC can use at least one of the inter-frame prediction techniques described herein. The palette mode can be considered an example of intra-frame coding or intra-frame prediction. When the palette mode is applied, information about the palette table and palette index can be included in the video / image information and signaled.
[0091] Intra-predictor 331 can predict the current block by referencing samples within the current frame. Depending on the prediction mode, the referenced samples can be located near the current block or at a specific distance away. In intra-prediction, the prediction mode can include at least one non-directional mode and multiple directional modes. Intra-predictor 331 can determine the prediction mode applied to the current block by using prediction modes applied to neighboring blocks.
[0092] Inter-frame predictor 332 can deduce the predicted block of the current block based on a reference block (reference sample array) specified by motion vectors on a reference frame. In this case, to reduce the amount of motion information transmitted in inter-frame prediction mode, motion information can be predicted on a block, sub-block, or sample basis based on the correlation of motion information between neighboring blocks and the current block. Motion information may include motion vectors and reference frame indices. Motion information may also include inter-frame prediction direction information (L0 prediction, L1 prediction, Bi prediction, etc.). For inter-frame prediction, neighboring blocks may include spatially neighboring blocks existing in the current frame and temporally neighboring blocks existing in the reference frame. For example, inter-frame predictor 332 can configure a motion information candidate list based on neighboring blocks and deduce the motion vector and / or reference frame index of the current block based on the received candidate selection information. Inter-frame prediction can be performed based on various prediction modes, and the information about the prediction may include information indicating the inter-frame prediction mode of the current block.
[0093] Adder 340 can add the obtained residual signal to the prediction signal (prediction block, prediction sample array) output from the predictor (including inter-frame predictor 332 and / or intra-frame predictor 331) to generate a reconstruction signal (reconstructed frame, reconstruction block, reconstruction sample array). When there is no residual for the block to be processed (similar to when a skip mode is applied), the prediction block can be used as a reconstruction block.
[0094] Adder 340 can be referred to as a reconstructor or reconstruction block generator. The generated reconstructed signal can be used for intra-frame prediction of the next block to be processed in the current frame, can be output through filtering as described later, or can be used for inter-frame prediction of the next frame. In addition, luminance mapping and chroma scaling (LMCS) can be applied in the frame decoding process.
[0095] Filter 350 can improve subjective / objective image quality by applying filtering to the reconstructed signal. For example, filter 350 can generate a modified reconstructed image by applying various filtering methods to the reconstructed image and send the modified reconstructed image to memory 360 (specifically, the DPB of memory 360). Various filtering methods may include deblocking filtering, sample adaptive offset, adaptive loop filter, bilateral filter, etc.
[0096] The (modified) reconstructed frame stored in the DPB of memory 360 can be used as a reference frame in inter-frame predictor 332. Memory 360 can store motion information of blocks in the current frame from which motion information is derived (or decoded) and / or motion information of blocks in previously reconstructed frames. The stored motion information can be sent to inter-frame predictor 332 as motion information of spatially or temporally neighboring blocks. Memory 360 can store reconstructed samples of reconstructed blocks in the current frame and send them to intra-frame predictor 331.
[0097] The embodiments described in this document in the filter 260, inter-frame predictor 221 and intra-frame predictor 222 of the encoding device 200 can also be applied equivalently or correspondingly to the filter 350, inter-frame predictor 332 and intra-frame predictor 331 of the decoding device 300, respectively.
[0098] This disclosure relates to a method for obtaining predicted samples. Specifically, according to this disclosure, a model can be derived to compensate for the difference in sample values between a virtual sample and a predicted sample of the current block, said model being applied to the predicted sample of the current block to obtain a modified predicted sample. Hereinafter, it will be referred to as model-based compensation or model-based sample modification. This disclosure can be distinguished from existing methods for using neighboring samples adjacent to both the predicted block and the current block.
[0099] Figure 4 A model-based compensation method performed by a decoding device 300 as an embodiment of the present disclosure is illustrated.
[0100] Reference Figure 4 This allows us to obtain the prediction sample S400 for the current block.
[0101] Predicted samples for the current block can be obtained based on intra-frame prediction.
[0102] As an example, a predicted sample for the current block can be obtained based on intra-frame template matching prediction (intra-frame TMP). In other words, a region that matches or is most similar to the template of the current block can be searched within a pre-reconstructed region in the current frame, and the block using the corresponding region as a template can be determined as a reference block. A predicted sample for the current block can be obtained based on the reconstructed sample of the corresponding reference block. The search can be performed within a predefined search range in the pre-reconstructed region. The cost of using the template of the current block can be calculated by traversing within the predefined search range. The region with the minimum cost among the calculated costs can be selected as the block with the template cost. Here, the cost can be calculated based on the sum of absolute differences (SAD), the sum of absolute transform differences (SATD), or the sum of squared differences (SSE). The predefined search range can include at least one of the current coding tree unit (CTU) to which the current block belongs or the neighboring CTUs adjacent to the current CTU. Here, the neighboring CTUs can include at least one of the top-left neighboring CTU, the top neighboring CTU, or the left neighboring CTU. Alternatively, the current block can be a block encoded in intra-block copy (IBC) mode. In this scenario, a reference block can be determined based on the block vector of the current block, and a predicted sample for the current block can be obtained based on the reconstructed sample of the corresponding reference block. The block vector can specify the position of the reference block within a pre-reconstructed region of the current frame. Similarly, the reference block can belong to the same current frame as the current block.
[0103] Alternatively, prediction samples for the current block can be obtained based on inter-frame prediction.
[0104] As an example, the predicted sample for the current block can be obtained based on a motion compensation method using motion vectors (e.g., merge mode, AMVP mode). Alternatively, the predicted sample for the current block can be obtained based on combined inter-frame and intra-frame prediction (CIIP). Alternatively, the predicted sample for the current block can be obtained based on a geometric segmentation mode. When a geometric segmentation mode is applied to the current block, the current block can be divided into two partitions (i.e., a first partition and a second partition), and the predicted sample for the current block can be obtained based on a weighted sum of a first predicted sample for the first partition and a second predicted sample for the second partition. Here, both the first and second predicted samples can be generated based on intra-frame prediction. Alternatively, both the first and second predicted samples can be generated based on inter-frame prediction. Alternatively, one of the first or second predicted samples can be generated based on intra-frame prediction, and the other can be generated based on inter-frame prediction. Alternatively, the predicted sample for the current block can be obtained based on a multiple hypothesis prediction (MHP) mode. The MHP mode can be a method for obtaining a predicted sample based on a weighted sum of an initial predicted sample and one or more additional predicted samples. Here, at least one of the initial prediction sample or the additional prediction sample can be obtained based on the motion compensation method described above.
[0105] Reference Figure 4 It can generate a virtual sample S410 for the current block.
[0106] Virtual samples according to this disclosure can be generated to represent the sample values of the current block. Virtual samples of the current block can be generated based on reconstructed samples surrounding the current block. Here, the virtual sample can be a predicted sample or a reconstructed sample generated by the method described below. The virtual sample can correspond to the same sample position as the predicted sample of the current block. The generated virtual sample can be used as input to a compensation model for deriving the predicted sample. To generate virtual samples with the same values in the encoding and decoding devices, adjacent reconstructed samples that can be used in the encoding and decoding steps of the current block can be used. Virtual samples of the current block can be generated by using samples in regions encoded / decoded prior to the current block. Virtual samples of the current block can be generated by one method described below, or by a combination of two or more methods.
[0107] Method 1
[0108] Virtual samples for the current block can be generated based on the planar pattern used in intra-frame prediction. Virtual samples for the current block can also be generated based on the weighted sum of at least two neighboring samples adjacent to the current block. Here, at least two neighboring samples can include at least two of the top neighbor, bottom-left neighbor, left neighbor, or top-right neighbor samples. As an example, virtual samples for the current block can be generated as shown in Equation 1 below.
[0109] [Equation 1]
[0110] In Equation 1, `genSamples[x][y]` represents the virtual sample generated at position (x, y) for the current block. Here, the value of x can range from 0 to (nTbW - 1), and the value of y can range from 0 to (nTbH - 1). `nTbW` represents the width of the current block, and `nTbH` represents the height of the current block. `ref[x][y]` can represent the reference sample used to generate the virtual sample for the current block. Specifically, `ref[x][-1]` is the upper neighbor sample of the current block, which can have the same x-coordinate as the virtual sample to be generated. `ref[-1][nTbH]` can be the upper right neighbor sample of the current block. `ref[-1][y]` is the left neighbor sample of the current block, which can have the same y-coordinate as the virtual sample to be generated. `ref[nTbW][-1]` can be the lower left neighbor sample of the current block.
[0111] Method 2
[0112] Virtual samples for the current block can be generated based on the DC pattern used in intra-frame prediction. Virtual samples for the current block can also be generated based on the average value of reconstructed samples belonging to the neighboring regions of the current block. Here, the neighboring regions can include at least one of the upper neighboring region, left neighboring region, upper-left neighboring region, lower-left neighboring region, or upper-right neighboring region.
[0113] As an example, the virtual sample of the current block can be generated as shown in Equation 2 below.
[0114] [Equation 2]
[0115] In Equation 2, the value of dcVal can be calculated based on the width (nTbW) and height (nTbH) of the current block as follows. Specifically, when nTbW and nTbH are the same, the value of dcVal can be calculated as in Equation 3.
[0116] [Equation 3]
[0117] Alternatively, when nTbW is greater than nTbH, the value of dcVal can be calculated as in Equation 4.
[0118] [Equation 4]
[0119] Alternatively, when nTbW is less than nTbH, the value of dcVal can be calculated as in Equation 5.
[0120] [Equation 5]
[0121] Alternatively, regardless of the width and height of the current block, the value of dcVal can be calculated as in Equation 6.
[0122] [Equation 6]
[0123] In Equation 2, `genSamples[x][y]` represents the virtual sample generated for the current block at position (x, y). Here, the value of x can range from 0 to (nTbW - 1), and the value of y can range from 0 to (nTbH - 1). In Equations 3 through 6, `nTbW` represents the width of the current block, and `nTbH` represents the height of the current block. `ref[x][y]` can represent the reference sample used to generate the virtual sample for the current block. Specifically, `ref[x'][-1]` can be the upper neighbor sample adjacent to the current block. `ref[-1][y']` can be the left neighbor sample adjacent to the current block.
[0124] Method 3
[0125] A virtual sample for the current block can be generated by weighting a sample generated by applying a predetermined prediction method to the current block with a reference sample of the current block. The reference sample can be a neighboring sample adjacent to the current block, and this neighboring sample can include at least one of the upper neighbor or left neighbor of the current block. Here, the upper neighbor sample can be a sample with the same x-coordinate as the virtual sample, while the left neighbor sample can be a sample with the same y-coordinate as the virtual sample.
[0126] Here, the prediction method can be the intra-template matching prediction or intra-block copy (IBC) mode described above. Alternatively, the prediction method can be a motion compensation method using motion vectors. Alternatively, the prediction method can be a planar mode or a DC mode.
[0127] Alternatively, the samples generated by applying a predetermined prediction method to the current block can be derived as the average of the reference samples of the current block. The reference samples of the current block may include the upper neighbor and left neighbor samples that are adjacent to the current block.
[0128] As an example, the virtual sample of the current block can be generated as shown in Equation 7 below.
[0129] [Equation 7]
[0130] In Equation 7, `genSamples[x][y]` represents the virtual sample generated at position (x, y) for the current block. Here, the value of x can range from 0 to (nTbW - 1), and the value of y can range from 0 to (nTbH - 1). `refL[x][y]` and `refT[x][y]` can represent the reference samples used to generate the virtual sample at position (x, y), respectively. For example, `refL[x][y]` can represent the left neighbor sample adjacent to the current block, and `refT[x][y]` can represent the top neighbor sample adjacent to the current block. `wL[x]` can represent the weight applied to the reference sample `refL[x][y]`, which can be determined based on the x-coordinate of the virtual sample and / or the width and height of the current block. `wT[y]` can represent the weight applied to the reference sample `refT[x][y]`, which can be determined based on the y-coordinate of the virtual sample and / or the width and height of the current block. `nTbW` represents the width of the current block, and `nTbH` represents the height of the current block. p[x][y] can represent a sample at position (x, y) generated by applying a predetermined prediction method to the current block.
[0131] Method 4
[0132] The intra-prediction mode of the current block can be derived based on the decoder-side intra-mode derivation (DIMD) method, and virtual samples of the current block can be generated based on the corresponding intra-prediction mode.
[0133] Specifically, gradients can be calculated based on at least two samples belonging to neighboring regions of the current block. Here, the gradient can include at least one of a horizontal gradient or a vertical gradient. The intra-prediction mode for the current block can be derived based on at least one of the calculated gradient or gradient magnitude. Here, the gradient magnitude can be determined based on the sum of the horizontal and vertical gradients. This derivation method can derive one intra-prediction mode, or two or more intra-prediction modes for the current block.
[0134] Gradients can be calculated in units of windows of a predetermined size. An angle representing the orientation of samples within the corresponding window can be calculated based on the calculated gradient. The calculated angle can correspond to any of the multiple predefined intra-prediction modes described above. The magnitude of the gradient can be stored / updated for the intra-prediction mode corresponding to the calculated angle. In this process, an intra-prediction mode corresponding to the calculated gradient can be determined for each window, and the magnitude of the gradient can be stored / updated for the determined intra-prediction mode. The top T intra-prediction modes with the largest magnitudes among the stored gradients can be selected, and the selected intra-prediction mode can be set as the intra-prediction mode for the current block. Here, T can be an integer of 1, 2, 3, or greater.
[0135] The neighboring region used to compute the gradient is a pre-reconstructed region preceding the current block, which may include at least one of the left, top, top-left, bottom-left, or top-right regions adjacent to the current block. The neighboring region may include at least one of the adjacent sample lines adjacent to the current block, a first non-adjacent sample line separated from the current block by 1 sample, or a second non-adjacent sample line separated from the current block by 2 samples. However, it is not limited to this and may further include non-adjacent sample lines separated from the current block by N samples, where N can be an integer greater than or equal to 3.
[0136] Neighboring regions can be regions predefined by both the encoding and decoding devices for calculating gradients. Alternatively, neighboring regions can be variably determined based on information about the location of specified neighboring regions. In this case, the location information of specified neighboring regions can be signaled via the bitstream. Alternatively, the location of neighboring regions can be determined based on at least one of the following: whether the current block is located at the boundary of a coding tree unit, the size of the current block (e.g., width, height, width-to-height ratio, width-to-height product), the segmentation type of the current block, the prediction mode of neighboring regions, or the availability of neighboring regions.
[0137] Method 5
[0138] The intra-prediction mode of the current block can be derived using the Template-Based Intra-Mode Derivation (TIMD) method, and virtual samples of the current block can be generated based on the corresponding intra-prediction mode.
[0139] Specifically, the cost of each predetermined candidate mode can be calculated. A predetermined candidate mode can refer to multiple intra-prediction modes predefined equivalently for both the encoding and decoding devices. Alternatively, for template region-based derivation, a candidate list consisting of candidate modes can be generated, and the cost can be calculated for each candidate mode belonging to the candidate list. Alternatively, the cost can be calculated only for the first N candidate modes in the generated candidate list. Here, N can be a value predefined equivalently for both the encoding and decoding devices. As an example, N can be an integer of 2, 3, 4, 5, or greater.
[0140] The cost can be calculated as the sum of the absolute differences (SAD) between the reconstructed samples and the predicted samples within the template region. Alternatively, the cost can also be calculated as the sum of the absolute transform differences (SATD) between the reconstructed samples and the predicted samples within the template region. Here, SATD can refer to the SAD transformed to the frequency domain. As an example of a transform, the Hadamard transform can be used, but it is not limited to this. Predicted samples for the template region can be generated based on the candidate patterns described above.
[0141] The template region used to calculate the cost can be a pre-reconstructed region adjacent to the current block. As an example, the template region can include at least one of the current block's upper adjacent region, left adjacent region, upper-left adjacent region, lower-left adjacent region, or upper-right adjacent region.
[0142] The template region can be a region predefined equally for both the encoding and decoding devices to calculate the cost. Alternatively, the template region can be variably determined based on information about the location of the specified template region. In this case, the location information of the specified template region can be signaled via the bitstream. Alternatively, the location of the template region can be determined based on at least one of the following: whether the current block is located at the boundary of a coding tree unit, the size of the current block (e.g., width, height, width-to-height ratio, width-to-height product), the segmentation type of the current block, the prediction pattern of adjacent regions, or the availability of adjacent regions.
[0143] You can select the candidate pattern with the lowest cost among those calculated for each candidate pattern. As an example, you can calculate the cost for each of the five candidate patterns in the candidate list. The five candidate patterns in the candidate list can be reordered in ascending order of their calculated costs. You can then select the first candidate pattern from the five reordered patterns.
[0144] Alternatively, at least two candidate patterns can be selected that have the minimum cost among the costs calculated for each candidate pattern. As an example, the cost can be calculated separately for each of the five candidate patterns in the candidate list. The five candidate patterns in the candidate list can be reordered in ascending order of their calculated costs. The first two candidate patterns can be selected from the five reordered candidate patterns.
[0145] One or more candidate modes selected through the above process can be set as the intra-prediction mode for the current block.
[0146] Alternatively, when at least two candidate modes are selected through the above process, the intra-prediction mode of the current block can be derived based on comparisons between the selected candidate modes and / or comparisons between at least one of the selected candidate modes and a threshold. As an example, the intra-prediction mode of the current block can be derived based on whether the selected candidate modes satisfy the following conditions.
[0147] [condition]
[0148] Under this condition, costMode1 can refer to the cost calculated based on any of the selected candidate modes, and costMode2 can refer to the cost calculated based on the other of the selected candidate modes. For example, costMode1 can refer to the cost calculated based on the candidate mode with the smaller cost among the selected candidate modes, while costMode2 can refer to the cost calculated based on the candidate mode with the larger cost among the selected candidate modes. In this case, K represents a predetermined comparison factor, which can be a predefined value for both the encoding and decoding devices. For example, K can be an integer of 1, 2, or larger, or it can refer to a real number such as 1 / 2 or 1 / 4.
[0149] When the conditions are met, the selected candidate mode can be set as the intra-prediction mode for the current block. On the other hand, when the conditions are not met, the candidate mode with a cost of costMode1 can be set as the intra-prediction mode for the current block, while the candidate mode with a cost of costMode2 may not be used as the intra-prediction mode for the current block.
[0150] Method 6
[0151] Virtual samples for the current block can be generated based on intra-frame template matching prediction (intra-frame TMP).
[0152] Specifically, a region matching or most similar to the template of the current block can be searched within a pre-reconstructed area in the current frame, and the block using the corresponding region as a template can be identified as a reference block. A virtual sample of the current block can be generated based on the reconstruction sample of the corresponding reference block.
[0153] A search can be performed within a predefined search range within a pre-reconstructed region. The cost of using the template of the current block can be calculated by traversing within the predefined search range. A block with the minimum cost among the calculated costs can be selected as the template. Here, the cost can be calculated based on the sum of absolute differences (SAD), the sum of absolute transform differences (SATD), or the sum of squared differences (SSE). The predefined search range can include at least one of the current coding tree unit (CTU) to which the current block belongs or the neighboring CTUs adjacent to the current CTU. Here, the neighboring CTUs can include at least one of the top-left neighboring CTU, the top neighboring CTU, or the left neighboring CTU.
[0154] The template of the current block can be composed of adjacent regions. Here, adjacent regions can include at least one of the left adjacent region, the top adjacent region, or the top-left adjacent region. The top adjacent region can consist of N horizontal sample rows, while the left adjacent region can consist of M vertical sample rows. n and M are integers greater than or equal to 1. N and M can be the same or different from each other. The size and / or position of the template of the current block can be variably determined based on the size and / or position of the current block.
[0155] Method 7
[0156] Virtual samples for the current block can be generated based on matrix-based intra-frame prediction (MIP). Any one of several matrices predefined equivalently for both the encoding and decoding devices can be selectively used. Any one of the matrices can be specified based on the size of the current block. The coefficients of each matrix can be predefined equivalently for both the encoding and decoding devices.
[0157] Method 8
[0158] Virtual samples of the current block can be generated through intra-frame prediction based on multiple reference rows (MRL).
[0159] As an example, a predicted sample for the current block can be obtained by using at least one sample belonging to a reference row adjacent to the current block (e.g., the first reference row) as a reference sample. In this case, a virtual sample for the current block can be generated by applying the same method as used to obtain the predicted sample, but using at least one sample belonging to a reference row not adjacent to the current block (e.g., the second, third, or fourth reference row) as a reference sample. Alternatively, a virtual sample for the current block can be generated by applying any of methods 1 to 7 described above, but using at least one sample belonging to a reference row not adjacent to the current block (e.g., the second, third, or fourth reference row) as a reference sample.
[0160] Alternatively, a predicted sample for the current block can be obtained by using at least one sample belonging to a reference row adjacent to the current block as a reference sample. In this case, a virtual sample for the current block can be generated by applying the same method as used to obtain the predicted sample, but using at least one sample belonging to a reference row adjacent to the current block as a reference sample. Alternatively, a virtual sample for the current block can be generated by applying any of methods 1 to 7 described above, but using at least one sample belonging to a reference row adjacent to the current block as a reference sample.
[0161] According to this disclosure, the size / range of the region in which virtual samples are generated within the current block can be variable.
[0162] As an example, virtual samples can be generated for all sample locations within the current block. In the current block... When generating a block, the number of virtual samples generated for the current block can be ( ).
[0163] Virtual samples can also be generated only for sample positions belonging to a portion of the current block. For example, a portion of the block can consist of N vertical sample rows from the left and M horizontal sample rows from the top. n can be an integer greater than or equal to 0 and less than the width (W) of the current block. m can be an integer greater than or equal to 0 and less than the height (H) of the current block. Alternatively, the size / range of the portion of the block can be variably determined based on the size of the current block. For example, the size / range of the portion of the block can be determined according to the size of the current block as shown in Equation 8 below.
[0164] [Equation 8]
[0165] In Equation 8, the value of i can be any one of 1, 2, 3, 4, or 5. The value of j can be any one of 1, 2, 3, 4, or 5. i and j can be predefined values equivalent to those for both the encoding and decoding devices. i and j can be defined the same or different.
[0166] A portion of the current block can be defined as the remaining area excluding the predetermined bottom-right region of the current block. Here, the bottom-right region can be defined as... The block. N can be an integer greater than or equal to 0 and less than the width (W) of the current block, and M can be an integer greater than or equal to 0 and less than the height (H) of the current block. Alternatively, the lower right region can be variably determined based on the size of the current block. For example, the lower right region can be determined as in Equation 8 above.
[0167] Virtual samples can be generated only for sample locations located on the block boundaries of the current block. Here, the block boundary can include at least one of the top, left, right, or bottom boundaries of the current block.
[0168] In this way, when virtual samples are not generated for all sample locations within the current block, a model for refining the prediction samples can be derived using only the generated virtual samples.
[0169] The regions described above are merely examples, and it goes without saying that any region within the current block can be predefined, and virtual samples can be generated only for the corresponding regions.
[0170] A virtual sample can be generated for each sample location within the current block. Alternatively, two or more virtual samples can be generated for each sample location. One of the two or more virtual samples can be generated based on any of methods 1 to 8 described above, while the other can be generated based on a different method. In this case, the final virtual sample can be generated by a weighted sum of the two or more virtual samples.
[0171] Alternatively, an index can be defined specifying any one of a plurality of predefined methods for generating virtual samples of the current block. Any of the plurality of methods can be selected based on the index, and virtual samples of the current block can be generated based on the selected method. The plurality of predefined methods are candidate methods for generating virtual samples, and may include at least one of methods 1 to 8 described above. The index can be signaled via a bitstream or derived based on the properties of the current block. The properties of the current block may include at least one of the following: size, prediction method used to obtain predicted samples, component type, or whether geometric segmentation is applied.
[0172] As an example, when multiple methods for generating virtual samples are defined for a sequence or image, a flag or index for selecting any of the multiple methods can be signaled on a per-picture or per-slice basis. Table 1 below shows an example of signaling ph_gen_samples_idx in the per-picture header.
[0173] [Table 1]
[0174] In Table 1, ph_gen_samples_idx can be an index for selecting one of several methods used to generate virtual samples in the current frame. Table 2 below is an example showing the mapping between the values of ph_gen_samples_idx and the various methods mentioned above.
[0175] [Table 2]
[0176] Alternatively, a flag or index for selecting any of the various methods for generating virtual samples can be signaled in the coding unit syntax for the current block. Alternatively, a flag or index for selecting any of the various methods for generating virtual samples can be signaled at the sequence level, and a separate signal can be sent indicating whether the corresponding flag or index should be included in the lower-level syntax.
[0177] Reference Figure 4 The parameters S420 used for model-based compensation can be derived based on the predicted samples and virtual samples of the current block.
[0178] Specifically, one or more models can be defined based on the predicted samples and virtual samples of the current block, and the parameters of each model can be derived.
[0179] As an example, the model according to this disclosure can have the form of a first-order linear model, as shown in Equation 9 below.
[0180] [Equation 9]
[0181] In the model of Equation 9, P[x][y] can represent the predicted sample of the current block, and P'[x][y] can represent the modified predicted sample based on the model of this disclosure. When the width and height of the current block are nTbW and nTbH respectively, the value of x in P[x][y] can range from 0 to (nTbW - 1), and the value of y can range from 0 to (nTbH - 1). a0 and a1 are parameters of the linear model and can be derived according to Equation 10 below. a0 and a1 can correspond to the weight parameter and offset parameter in the parameters of the linear model, respectively.
[0182] [Equation 10]
[0183] According to Equation 10, for each sample location within the current block, the square of the difference between the virtual sample and the modified predicted sample can be calculated, and ErrCost can be derived as the sum of the calculated values. Here, the modified predicted sample can be derived based on the parameters of the linear model, as shown in Equation 9. In this case, a0 and a1 can be derived as values that minimize ErrCost for the current block. Clearly, when virtual samples are generated only for sample locations belonging to a portion of the current block, the model parameters can be derived solely based on the corresponding virtual sample and its corresponding predicted sample.
[0184] Alternatively, the model according to this disclosure can have a convolutional form. In this case, the current sample (C) and one or more neighboring samples can be used. The neighboring samples can include at least one of the following: left sample (L), right sample (R), top sample (A), bottom sample (B), top-left sample (AL), top-right sample (AR), bottom-left sample (BL), or bottom-right sample (BR). Here, the current sample (C) is the predicted sample obtained in S400 and can refer to the sample to be modified.
[0185] The convolutional model according to this disclosure can be defined as shown in Equation 11 below.
[0186] [Equation 11]
[0187] According to Equation 11, a model can be defined using the current sample (C) and the samples located to its left, right, top, and bottom (L, R, A, B). a0 to a4 are the parameters of the convolutional model and can be derived as shown in Equation 12 below. a0 to a4 can correspond to the weight parameters of the convolutional model.
[0188] [Equation 12]
[0189] According to Equation 12, for each sample position within the current block, the square of the difference between the virtual sample and the modified predicted sample can be calculated, and ErrCost can be derived as the sum of the calculated values. Here, the modified predicted sample can be derived based on the parameters of the convolutional model, as shown in Equation 11. In this case, a0 to a4 can be derived as values that minimize the ErrCost of the current block.
[0190] Equation 12 above can also be expressed as equation 13 below.
[0191] [Equation 13]
[0192] Convolutional models can be defined in various forms based on the number of neighboring samples used in addition to the current sample (C). As an example, all eight neighboring samples (L, R, A, B, AL, AR, BL, BR) of the current sample (C) can be used. Alternatively, five neighboring samples of the current sample (C) can be used. Here, the five neighboring samples can consist of {L, R, A, AL, AR} or {L, R, B, BL, BR}. Alternatively, two neighboring samples of the current sample (C) can be used. Here, the two neighboring samples can consist of {L, R}, {A, B}, {AL, AR} or {BL, BR}.
[0193] The parameters of a convolutional model can include at least one of an offset parameter or a nonlinear parameter. The offset parameter can be defined as the median value of the bit depth of the input image, or it can be defined as the average value of the input samples. Here, the input samples can refer to the current sample (C) input to the convolutional model and its neighboring samples. The nonlinear parameter can refer to a parameter with nonlinear characteristics, such as the squared value of the current sample. The nonlinear parameter can be a weight parameter applied to the squared values of the input samples (at least one of the current sample or neighboring samples) in the convolutional model.
[0194] For ease of explanation, the example above uses neighboring samples of size 3×3 with the current sample (C) as the center sample, but is not limited to this. For example, all or part of neighboring samples of size, such as 5×5 or 13×13, with the current sample (C) as the center sample can be used.
[0195] The number of neighboring samples used in a convolutional model, whether to use offset parameters, and nonlinear parameters can be variably determined by considering complexity and model performance.
[0196] As described above, the parameters of the model according to this disclosure can be derived based on the predicted sample region of the current block. However, it is not limited to this, and the parameters of the model can also be derived based on the difference signal region (or residual sample region) of the current block. In this case, the input samples of the linear model or convolutional model can be the corresponding residual samples, rather than the predicted samples.
[0197] Meanwhile, neighboring samples used for the convolution model may not be available. For example, neighboring samples used for the convolution model might be located outside the current block. To address this, a padding process can be performed on the outer region adjacent to the boundary of the current block. The number of padding samples can be determined based on the number of neighboring samples used for convolution.
[0198] The padding process can be a process of replacing unavailable samples among neighboring samples input to the convolutional model with available samples from the current block. Available samples can be samples located at the boundaries of the current block. For example, if the upper sample (A, AL, AR) of the current sample is unavailable, it can be replaced with the current sample or a lower sample adjacent to the corresponding sample. If the left sample (L, AL, BL) of the current sample is unavailable, it can be replaced with the current sample or a right sample adjacent to the corresponding left sample. If the lower sample (B, BL, BR) of the current sample is unavailable, it can be replaced with the current sample or an upper sample adjacent to the corresponding lower sample. If the right sample (R, AR, BR) of the current sample is unavailable, it can be replaced with the current sample or a left sample adjacent to the corresponding right sample. Alternatively, padding can be performed in a mirror form centered on the boundaries of the current block.
[0199] The parameters of either a predefined linear or convolutional model can be derived for the current block. Alternatively, an index can be defined for any of a plurality of models. Any of the multiple models can be selected based on the corresponding index, and the parameters of the selected model can be derived. Here, the index can be signaled via a bitstream, or it can be derived based on the attributes of the current block. The attributes of the current block may include at least one of the following: size, prediction method used to obtain prediction samples, component type, or whether geometric segmentation is applied. Alternatively, the parameters of each of the multiple models can be derived for the current block.
[0200] As an example, when multiple models are defined for a sequence or image, a flag or index for selecting any of the multiple models can be signaled on a picture or slice basis. Table 3 below is an example where ph_mbc_model_idx is signaled in the image header.
[0201] [Table 3]
[0202] In Table 3, ph_mbc_model_idx can be an index for selecting any of several models used for model-based compensation in the current frame. Table 4 below is an example showing the mapping between the values of ph_mbc_model_idx and the aforementioned models.
[0203] [Table 4]
[0204] The index for selecting any of the multiple models used for model-based compensation can be signaled in the coding unit syntax for the current block. This index can be represented as cu_mbc_model_idx, and it can be signaled as shown in Table 5 below.
[0205] [Table 5]
[0206] According to Table 5, when the flag (mbc_flag) indicating whether model-based compensation is applied in the current coding unit (or current block) is 1, cu_mbc_model_idx can be notified by signaling. On the other hand, when mbc_flag is 0, cu_mbc_model_idx may not be notified by signaling.
[0207] When cu_mbc_model_idx is 0, it indicates that a linear model is selected as the model for model-based compensation. When cu_mbc_model_idx is 1, it indicates that a convolutional model is selected as the model for model-based compensation. However, this is just an example, and multiple models can be additionally included besides the linear and convolutional models.
[0208] Reference Figure 4 Modified prediction samples S430 can be obtained based on the prediction samples of the current block and the parameters of the derived model.
[0209] The parameters of the derived linear or convolutional model can be applied to the predicted samples of the current block to modify the predicted samples of the current block.
[0210] When deriving the parameters of each of the multiple models for the current block, the parameters of each model can be applied to the predicted samples of the current block. This yields multiple modified predicted samples corresponding to the multiple models, and the final predicted sample can be obtained based on the weighted sum of these modified predicted samples.
[0211] The derived model aims to improve coding efficiency by compensating for the difference in sample values between the virtual block and the predicted block in the current block.
[0212] At the same time, information about the aforementioned model-based compensation can be communicated via signals.
[0213] Specifically, a flag indicating whether model-based compensation is applied to a specific unit at a corresponding high level can be signaled from the high-level syntax. Alternatively, the flag can indicate from the high-level syntax whether model-based compensation is available for a specific unit at a corresponding high level. Here, the high-level syntax may include at least one of a sequence parameter set (SPS), a picture parameter set (PPS), a picture header (PH), or a slice header (SH).
[0214] As an example, the flag can be represented as sps_model_based_compensation_enabled_flag, which can be signaled as shown in Table 6 below.
[0215] [Table 6]
[0216] According to Table 6, `sps_model_based_compensation_enabled_flag` can be a syntax signaled in SPS. `sps_model_based_compensation_enabled_flag` can indicate the presence of a flag (`mbc_flag`) indicating whether model-based compensation is applied to the current block. For example, when `sps_model_based_compensation_enabled_flag` is 1, it indicates that `mbc_flag` can exist in the coding unit syntax for the current block. On the other hand, when `sps_model_based_compensation_enabled_flag` is 0, it indicates that `mbc_flag` does not exist in the coding unit syntax for the current block. When `sps_model_based_compensation_enabled_flag` does not exist (i.e., when `sps_model_based_compensation_enabled_flag` is not signaled), the value of `sps_model_based_compensation_enabled_flag` can be deduced to be 0.
[0217] A flag (mbc_flag) indicating whether model-based compensation should be applied to the current block can be signaled. As an example, this flag can be signaled as shown in Table 7 below.
[0218] [Table 7]
[0219] According to Table 7, the mbc_flag can be signaled in the coding unit syntax for the current block. mbc_flag indicates whether model-based compensation is applied in the current coding unit (or current block). When mbc_flag is 1, it indicates that model-based compensation is applied to the current coding unit. Conversely, when mbc_flag is 0, it indicates that model-based compensation is not applied to the current coding unit. When mbc_flag is not present (i.e., when mbc_flag is not signaled), its value can be deduced to be 0.
[0220] Regardless of the current block size, `mbc_flag` can be notified by a signal. Alternatively, `mbc_flag` can be notified only when the current block size is greater than or equal to a predefined threshold (T1). Alternatively, `mbc_flag` can be notified only when the current block size is less than or equal to a predefined threshold (T2). Alternatively, `mbc_flag` can be notified only when the current block size is greater than or equal to T1 and less than or equal to T2. Here, the size of the current block can be defined by the number of samples belonging to the current block (or the product of the width and height of the current block), either the width or height of the current block, the ratio of the width and height of the current block, or the maximum or minimum value of the width and height of the current block.
[0221] As an example, `mbc_flag` can be signaled only when the number of samples belonging to the current block is less than 1024. Alternatively, `mbc_flag` can be signaled only when the number of samples belonging to the current block is greater than 32 or 64. Since model-based prediction accuracy may be low for blocks that are too small or too large, the application of model-based compensation can be limited according to the size of the current block, thereby reducing the signaling burden of the flag and improving coding efficiency.
[0222] The mbc_flag flag can be adaptively signaled based on the prediction mode of the current block. The mbc_flag flag can be signaled when the prediction mode of the current block corresponds to a prediction mode that allows model-based compensation. Prediction modes that allow model-based compensation can include at least one of the following: AMVP mode, merge mode, affine mode, sub-block-based merge mode, geometric segmentation mode, CIIP mode, or MHP mode.
[0223] The mbc_flag can be adaptively notified with signals based on sps_model_based_compensation_enabled_flag. As an example, when sps_model_based_compensation_enabled_flag is 1, mbc_flag can be notified with signals, while when sps_model_based_compensation_enabled_flag is 0, mbc_flag can be notified without signals.
[0224] It can be determined whether to apply model-based compensation to the current block. Based on information about the model-based compensation, it can be determined whether to apply it to the current block. When it is determined that model-based compensation should be applied to the current block (e.g., when mbc_flag is 1), S410 to S430 described above can be executed for the current block. Otherwise (e.g., when mbc_flag is 0), the predicted sample obtained in S400 can be set as the final predicted sample for the current block.
[0225] Figure 5 An illustrative configuration of a decoding device 300 that performs a model-based compensation method according to the present disclosure is shown.
[0226] Reference Figure 5 The decoding device 300 may include at least one of a prediction sample obtainr 500, a virtual sample generator 510, a model parameter derivator 520, or a prediction sample modifier 530.
[0227] The prediction sample acquirer 500 can obtain the prediction sample of the current block. The prediction sample acquirer 500 can obtain the prediction sample of the current block based on intra-frame prediction or inter-frame prediction, which is different from the reference... Figure 4 The description is the same.
[0228] Virtual sample generator 510 can generate virtual samples for the current block. Virtual sample generator 510 can generate virtual samples for the current block based on reconstructed samples surrounding the current block, and the method for generating virtual samples is consistent with reference to... Figure 4 The methods described are the same.
[0229] The model parameter derivator 520 can derive parameters for model-based compensation based on predicted samples and virtual samples of the current block. The model parameter derivator 520 can derive parameters based on a model predefined in the decoding device (e.g., a linear or convolutional model), and can derive parameters by selectively using any of a plurality of models predefined in the decoding device. This is in contrast to the reference... Figure 4 The descriptions are the same, and overlapping descriptions will be omitted here.
[0230] The prediction sample modifier 530 can obtain modified prediction samples based on the prediction samples of the current block and the parameters used for model-based compensation.
[0231] The decoding device 300 may also include a model-based compensation determiner (not shown). The model-based compensation determiner can determine whether model-based compensation is applied to the current block. The model-based compensation determiner can determine whether to apply model-based compensation to the current block based on the aforementioned information about model-based compensation. The information about model-based compensation can be decoded by the entropy decoder 310. When it is determined that model-based compensation should be applied to the current block, modified prediction samples can be obtained by the virtual sample generator 510, the model parameter derivator 520, and the prediction sample modifier 530.
[0232] Figure 6 An example is given of a model-based compensation method performed by an encoding device 200 as an embodiment of the present disclosure.
[0233] Reference Figure 6 This allows obtaining the prediction sample S600 for the current block. The prediction sample for the current block can be obtained based on intra-frame prediction or inter-frame prediction, which is different from the reference... Figure 4 The description is the same.
[0234] Reference Figure 6 This allows for the generation of virtual samples S610 for the current block. Virtual samples for the current block can be generated based on reconstructed samples surrounding it, and methods for generating virtual samples are provided in reference. Figure 4 The methods described are the same.
[0235] Reference Figure 6 The parameters S620 for model-based compensation can be derived based on the predicted samples and virtual samples of the current block. The parameters can be derived based on a model predefined in the encoding device (e.g., a linear model or a convolutional model), and can be derived by selectively using any one of a plurality of models predefined in the encoding device. For this purpose, an index specifying any one of the plurality of models can be defined. In other words, the model for model-based compensation for the current block can be determined among a plurality of models, and the index indicating the determined model can be encoded into the bitstream. A method for signaling the index is provided, along with a reference. Figure 4 The description is the same. Alternatively, the index can be derived based on the attributes of the current block. Methods and references for deriving parameters for model-based compensation are provided. Figure 4 The descriptions are the same, and overlapping descriptions will be omitted here.
[0236] Reference Figure 6Modified prediction samples S630 can be obtained based on the prediction samples of the current block and the parameters for model-based compensation.
[0237] It can be determined whether to apply model-based compensation to the current block. When it is determined that model-based compensation should be applied to the current block, steps S610 to S630 described above can be performed on the current block. Otherwise, the prediction sample obtained in S600 can be set as the final prediction sample for the current block. Based on the determination of whether to apply model-based compensation to the current block, the information regarding model-based compensation can be encoded into a bitstream. A method and reference for signaling information regarding model-based compensation are provided. Figure 4 The description is the same.
[0238] Figure 7 An illustrative configuration of an encoding device 200 that performs a model-based compensation method according to the present disclosure is shown.
[0239] Reference Figure 7 The encoding device 200 may include at least one of a prediction sample obtainr 700, a virtual sample generator 710, a model parameter derivator 720, or a prediction sample modifier 730.
[0240] The prediction sample acquirer 700 can acquire the prediction sample of the current block. The prediction sample acquirer 700 can acquire the prediction sample of the current block based on intra-frame prediction or inter-frame prediction, which is different from the reference... Figure 4 The description is the same.
[0241] The virtual sample generator 710 can generate virtual samples for the current block. The virtual sample generator 710 can generate virtual samples for the current block based on reconstructed samples surrounding the current block, and the method for generating virtual samples is referenced. Figure 4 The methods described are the same.
[0242] The model parameter derivator 720 can derive parameters for model-based compensation based on predicted samples and virtual samples of the current block. The model parameter derivator 720 can derive parameters based on a predefined model in the encoding device (e.g., a linear or convolutional model), and can derive parameters by selectively using any of a plurality of predefined models in the encoding device. This is different from the reference... Figure 4 and Figure 6 The descriptions are the same, and overlapping descriptions will be omitted here.
[0243] The prediction sample modifier 730 can obtain modified prediction samples based on the prediction samples of the current block and the parameters for model-based compensation.
[0244] The encoding device 200 may further include a model-based compensation determiner (not shown). The model-based compensation determiner can determine whether model-based compensation is applied to the current block. When it is determined that model-based compensation will be applied to the current block, modified prediction samples can be obtained through the virtual sample generator 710, the model parameter derivator 720, and the prediction sample modifier 730. Based on the determination of whether model-based compensation is applied to the current block, the aforementioned information regarding model-based compensation can be encoded into a bitstream. Methods and references for signaling information regarding model-based compensation are also provided. Figure 4 The description is the same.
[0245] In the above embodiments, the method is described based on a flowchart as a series of steps or blocks. However, the corresponding embodiments are not limited to this order of steps. Some steps may occur simultaneously with other steps or in a different order, as described above. In addition, those skilled in the art will understand that the steps shown in the flowchart are not exclusive. Other steps may be included, or one or more steps in the flowchart may be deleted, without affecting the scope of the embodiments of this disclosure.
[0246] The methods described above according to embodiments of the present disclosure can be implemented in software, and the encoding and / or decoding devices according to the present disclosure can be included in an apparatus for performing image processing, such as a TV, computer, smartphone, set-top box, display device, etc.
[0247] In this disclosure, when the implementation is implemented as software, the above-described method can be implemented as a module (process, function, etc.) performing the above-described functions. The module can be stored in memory and can be executed by a processor. The memory can be internal or external to the processor and can be connected to the processor by various well-known means. The processor may include an application-specific integrated circuit (ASIC), another chipset, logic circuitry, and / or data processing devices. The memory may include read-only memory (ROM), random access memory (RAM), flash memory, memory cards, storage media, and / or other storage devices. In other words, the implementations described herein can be executed by implementation on a processor, microprocessor, controller, or chip. For example, the functional units shown in the various figures can be executed by implementation on a computer, processor, microprocessor, controller, or chip. In this case, information for implementation (e.g., information about instructions) or algorithms can be stored in a digital storage medium.
[0248] Furthermore, decoding and encoding devices employing embodiments of this disclosure can be included in multimedia broadcasting transmitting and receiving devices, mobile communication terminals, home theater video devices, digital cinema video devices, surveillance cameras, video conferencing devices, real-time communication devices similar to video communication, mobile streaming devices, storage media, cameras, devices for providing video-on-demand (VOD) services, OTT (over-the-top) video devices, devices for providing internet streaming services, three-dimensional (3D) video devices, virtual reality (VR) devices, augmented reality (AR) devices, video telephony devices, transportation terminals (e.g., vehicle (including autonomous vehicle) terminals, aircraft terminals, ship terminals, etc.), and medical video devices, and can be used to process video signals or data signals. For example, OTT (over-the-top) video devices can include game consoles, Blu-ray players, internet-connected televisions, home theater systems, smartphones, tablet PCs, digital video recorders (DVRs), etc.
[0249] Furthermore, the processing methods applying the embodiments of this disclosure can be generated in the form of a computer-executable program and can be stored in a computer-readable recording medium. Multimedia data with data structures according to the embodiments of this disclosure can also be stored in a computer-readable recording medium. Computer-readable recording media include all types of storage devices and distributed storage devices that store computer-readable data. Computer-readable recording media can include, for example, Blu-ray discs (BD), Universal Serial Bus (USB), ROM, PROM, EPROM, EEPROM, RAM, CD-ROM, magnetic tape, floppy disks, and optical media storage devices. Additionally, computer-readable recording media include media implemented in the form of carrier waves (e.g., transmission via the Internet). Furthermore, bitstreams generated by encoding methods can be stored in computer-readable recording media or transmitted via wired / wireless communication networks.
[0250] Furthermore, the embodiments of this disclosure can be implemented by a computer program product using program code, and the program code can be executed on a computer using the embodiments of this disclosure. The program code can be stored on a computer-readable medium.
[0251] Figure 8 Examples of content streaming systems to which embodiments of this disclosure can be applied are shown.
[0252] Reference Figure 8 A content streaming system that applies embodiments of the present disclosure may generally include an encoding server, a streaming server, a network server, a media storage device, a user device, and a multimedia input device.
[0253] An encoding server generates a bitstream by compressing content input from multimedia input devices such as smartphones, cameras, and camcorders into digital data and then sends it to a streaming server. As another example, when multimedia input devices such as smartphones, cameras, and camcorders directly generate bitstreams, the encoding server can be omitted.
[0254] A bitstream can be generated by an encoding method or bitstream generation method that applies the embodiments of this disclosure, and the streaming server can temporarily store the bitstream during the sending or receiving of the bitstream.
[0255] A streaming server sends multimedia data to a user device via a web server based on a user request, and the web server acts as a medium to inform the user what services are available. When a user requests a desired service from the web server, the web server transmits the request to the streaming server, and the streaming server sends the multimedia data to the user. In this scenario, the content streaming system may include a separate control server, which controls the commands / responses between the various devices in the content streaming system.
[0256] A streaming server can receive content from media storage and / or encoding servers. For example, when receiving content from an encoding server, content can be received in real time. In this case, to provide a smooth streaming service, the streaming server can store the bitstream for a specific time period.
[0257] Examples of user devices may include mobile phones, smartphones, laptops, digital broadcasting terminals, personal digital assistants (PDAs), portable multimedia players (PMPs), navigators, touchscreen PCs, tablet PCs, ultrabooks, wearable devices (e.g., smartwatches, smart glasses, head-mounted displays (HMDs)), digital TVs, desktop computers, digital signage, etc.
[0258] In a content streaming system, each server can operate as a distributed server, and in this case, the data received from each server can be distributed and processed.
[0259] The claims set forth herein can be combined in various ways. For example, the technical features of the method claims of this disclosure can be combined and implemented as an apparatus, and the technical features of the apparatus claims of this disclosure can be combined and implemented as a method. Furthermore, the technical features of the method claims and the apparatus claims of this disclosure can be combined and implemented as an apparatus, and the technical features of the method claims and the apparatus claims of this disclosure can be combined and implemented as a method.
Claims
1. A method, the method comprising: Obtain the prediction sample for the current block; Generate a virtual sample for the current block; Based on the predicted samples and generated virtual samples of the current block, derive the parameters for model-based compensation for the current block; A modified prediction sample is obtained by applying the parameters to the prediction sample of the current block; as well as The current block is reconstructed based on the modified prediction sample and the residual sample of the current block.
2. The method according to claim 1, wherein, The virtual samples are generated based on planar patterns for intra-frame prediction.
3. The method according to claim 1, wherein, The virtual samples are generated based on DC patterns for intra-frame prediction.
4. The method according to claim 1, wherein, The virtual sample is generated by weighting a sample generated by applying a predetermined prediction method to the current block with a reference sample of the current block, and The reference sample includes at least one of the upper neighbor sample or the left neighbor sample that is adjacent to the current block.
5. The method according to claim 1, wherein, The virtual samples are generated based on a predetermined intra-frame prediction mode, and The predetermined intra-frame prediction mode is derived based on decoder-side intra-frame mode derivation DIMD or template-based intra-frame mode derivation TIMD.
6. The method according to claim 1, wherein, The virtual samples are generated based on intra-template matching prediction (intra-TMP).
7. The method according to claim 1, wherein, The virtual samples are generated based on matrix-based intra-frame prediction (MIP).
8. The method according to claim 1, wherein, The virtual sample is generated based on the sample location of a portion of the current block.
9. The method according to claim 1, wherein, The parameters are either parameters of a linear model or parameters of a convolutional model.
10. The method according to claim 1, wherein, The parameters are derived to minimize a predetermined cost, and The cost is defined based on the difference between the modified predicted sample and the virtual sample for each sample position within the current block.
11. The method according to claim 1, wherein, The modified predicted sample is obtained by applying the parameters to the predicted sample and its neighboring samples. The adjacent samples include at least one of the upper adjacent sample, lower adjacent sample, left adjacent sample, right adjacent sample, upper left adjacent sample, upper right adjacent sample, lower left adjacent sample, or lower right adjacent sample of the predicted sample.
12. A method, the method comprising: Obtain the prediction sample for the current block; Generate a virtual sample for the current block; Based on the predicted samples and generated virtual samples of the current block, derive the parameters for model-based compensation for the current block; A modified prediction sample is obtained by applying the parameters to the prediction sample of the current block; Generate residual samples for the current block based on the modified prediction samples for the current block; as well as The residual sample of the current block is encoded.
13. A computer-readable storage medium storing a bit stream generated by the method according to claim 12.
14. A method, the method comprising: Obtaining a bitstream of image information, wherein the bitstream is generated based on the following steps: obtaining a prediction sample for the current block; generating a virtual sample for the current block; deriving model-based compensation parameters for the current block based on the prediction sample and the generated virtual sample; obtaining a modified prediction sample by applying the parameters to the prediction sample for the current block; generating a residual sample for the current block based on the modified prediction sample; and encoding the residual sample for the current block; and Send data including the bit stream.