Image decoding method and apparatus therefor
By using RR-IBC modes predicted between or within frames, the block vector and flip type are derived, solving the problem of efficient compression of high-resolution, high-quality image/video data and improving coding efficiency and prediction performance.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- LX SEMICON CO LTD
- Filing Date
- 2025-01-03
- Publication Date
- 2026-07-31
AI Technical Summary
As images/videos reach higher resolutions and higher quality, their data size increases, leading to higher transmission and storage costs. Existing technologies struggle to efficiently compress and transmit high-resolution, high-quality image/video data, especially in immersive media and multi-domain applications.
The video/image coding method based on inter-frame prediction or intra-frame prediction is adopted. By reconstructing and reordering intra-block copying (RR-IBC) mode, the block vector and flip type are derived, the reference block is modified to improve the prediction performance, and the flip and rotation type information is signaled during the coding process.
It enhances video/image compression efficiency and prediction performance by performing IBC prediction considering various flip types of objects, thereby improving coding efficiency and quality.
Smart Images

Figure CN122498145A_ABST
Abstract
Description
Technical Field
[0001] This disclosure relates to image / video coding methods and image / video coding devices. Background Technology
[0002] Image / video encoding is used in a variety of applications, such as digital storage media, television broadcasting, video streaming services, and real-time communications. Recently, the demand for high-resolution and high-quality images / videos has been growing across various sectors.
[0003] As images / videos reach higher resolutions and higher quality, their data size increases, leading to a relative increase in the amount of information or bits transmitted. Therefore, the costs of transmitting and storing image data rise when using conventional media (such as wired and wireless broadband lines) or existing storage media.
[0004] Recently, there has been an increasing interest in and demand for immersive media, such as virtual reality (VR), augmented reality (AR), mixed reality (MR) content, and holograms. Furthermore, there is a growing trend of using immersive media to deliver immersive experiences in gaming, education, healthcare, real estate, and marketing.
[0005] Therefore, efficient image / video compression technology is needed to efficiently compress, send, store, and play high-resolution and high-quality image / video information with various characteristics. Summary of the Invention
[0006] Technical solution
[0007] According to embodiments of this disclosure, methods and apparatus for enhancing video / image coding efficiency are provided.
[0008] According to embodiments of this disclosure, a video / image coding method and a video / image coding apparatus based on inter-frame prediction or intra-frame prediction are provided.
[0009] According to embodiments of this disclosure, an image decoding method performed by a decoding device is provided. The method includes the following steps: deriving a prediction mode of the current block as a Reconstruction-Reordering Intra-Block Copy (RR-IBC) mode based on prediction-related information; deriving the block vector (BV) and flip type of the current block based on flip type information included in the prediction-related information; deriving a reference block of the current block based on the block vector; deriving a modified reference block based on the reference block and the flip type; and deriving a prediction sample of the current block based on the modified reference block.
[0010] According to embodiments of this disclosure, an image encoding method performed by an encoding device is provided. The method includes the following steps: deriving a prediction mode of the current block as a Reconstruction-Reordering Intra-Block Copy (RR-IBC) mode; deriving a modified reference block based on the flip type of the current block; deriving a prediction sample of the current block based on the modified reference block; and encoding image information including prediction-related information of the current block, wherein the prediction-related information includes a block vector (BV) for the modified reference block and flip type information indicating the flip type.
[0011] According to embodiments of this disclosure, a decoding apparatus for image decoding is provided. The decoding apparatus includes a memory and at least one processor connected to the memory, wherein the at least one processor is configured to perform the following operations: derive a prediction mode of a current block as a Reconstruction-Reordering Intra-Block Copy (RR-IBC) mode based on prediction-related information; derive a block vector (BV) and a flip type of the current block based on flip type information included in the prediction-related information; derive a reference block of the current block based on the block vector; derive a modified reference block based on the reference block and the flip type; and derive a prediction sample of the current block based on the modified reference block.
[0012] According to embodiments of this disclosure, an encoding apparatus for image encoding is provided. The encoding apparatus includes a memory and at least one processor connected to the memory, wherein the at least one processor is configured to perform the following operations: derive a prediction mode of a current block as a Reconstruction-Reordering Intra-Block Copy (RR-IBC) mode; derive a modified reference block based on the flip type of the current block; derive a prediction sample of the current block based on the modified reference block; and encode image information including prediction-related information of the current block, wherein the prediction-related information includes a block vector (BV) for the modified reference block and flip type information indicating the flip type.
[0013] According to embodiments of the present disclosure, a method for transmitting video / image data is provided, the video / image data comprising a bitstream generated by a video / image encoding method according to at least one of embodiments of the present disclosure.
[0014] According to embodiments of the present disclosure, an apparatus for transmitting video / image data is provided, the video / image data comprising a bitstream generated by a video / image encoding method according to at least one of embodiments of the present disclosure.
[0015] According to embodiments of the present disclosure, a computer-readable storage medium is provided that stores a program for performing a method according to at least one of the embodiments of the present disclosure.
[0016] According to embodiments of the present disclosure, a computer-readable digital storage medium is provided that stores encoded video / image information generated by a video / image encoding method according to at least one of the embodiments of the present disclosure.
[0017] According to embodiments of the present disclosure, a computer-readable digital storage medium is provided that stores encoded information or encoded video / image information such that a decoding device performs a video / image decoding method according to at least one of the embodiments of the present disclosure.
[0018] Beneficial effects
[0019] According to the embodiments of this disclosure, the overall video / image compression efficiency can be enhanced.
[0020] According to embodiments of this disclosure, prediction performance for the current block can be enhanced.
[0021] According to embodiments of this disclosure, in RR-IBC mode, information indicating flip types, including flips and rotations, can be signaled, enabling IBC prediction to be performed by taking into account various flip types of objects within the reference block. Attached Figure Description
[0022] Figure 1 Examples of video / image coding systems to which embodiments of the present disclosure may be applied are illustrated schematically.
[0023] Figure 2 This is a diagram that schematically illustrates the configuration of a video / image encoding device to which embodiments of the present disclosure may be applied.
[0024] Figure 3 This is a diagram that schematically illustrates the configuration of a video / image decoding device to which embodiments of the present disclosure may be applied.
[0025] Figure 4 The inter-frame prediction process is illustrated.
[0026] Figure 5 An example of a video / image coding method based on inter-frame prediction is shown.
[0027] Figure 6 An example of a video / image decoding method based on inter-frame prediction is shown.
[0028] Figure 7 An example of a partition shape supported by GPM is shown.
[0029] Figure 8 An example is shown of intra-prediction modes that can be used as IPM candidates.
[0030] Figure 9A reference template for deriving the TM cost, which is the MVD prediction cost of the MV candidate, is illustrated.
[0031] Figure 10 A reference template for deriving the TM cost is shown, which is the MVD prediction cost of MV candidates in the affine AMVP pattern or the affine MMVD pattern.
[0032] Figure 11 Examples of MMVD candidates are shown, which are combinations of available signs and magnitudes.
[0033] Figure 12 An example is shown of the flip type of the RR-IBC mode.
[0034] Figure 13 An implementation of aligning the current block and the reference block based on the flip type of the current block is illustrated.
[0035] Figure 14 A video / image coding method according to an embodiment of the present disclosure is illustrated schematically.
[0036] Figure 15 A video / image decoding method according to an embodiment of the present disclosure is illustrated schematically. Detailed Implementation
[0037] Because this disclosure can have various modifications and implementations, specific implementations are illustrated in the accompanying drawings and will be described in detail. However, it should be understood that there is no intention to limit the implementations of this disclosure to that specific implementation. The terminology used herein is for the purpose of describing specific implementations only and is not intended to limit the technical spirit of this disclosure. As used herein, the singular form is intended to include the plural form unless the context clearly indicates otherwise. As used herein, the term "and / or" includes any and all combinations of the associated listed items. As used herein, the terms "comprising," "including," and "having" specify the presence of the stated features, quantities, operations, elements, components, and / or combinations thereof, but do not exclude the presence or addition of one or more other features, quantities, operations, elements, components, and / or combinations thereof. In this disclosure, the use of the term "may" (e.g., regarding what the example or implementation may include or implement) in conjunction with an example or implementation indicates the existence of at least one example or implementation that includes or implements that feature, but all examples are not limited thereto, and the corresponding feature or configuration may be omitted.
[0038] Each component in the accompanying drawings described in this disclosure is shown independently to facilitate the explanation of its different features and functions, which does not imply that each component is implemented as separate hardware or separate software. For example, two or more of these components may be combined to form a single component, or a single component may be divided into multiple components. Embodiments in which components are integrated and / or separated are also included within the scope of this disclosure without departing from its spirit.
[0039] In this disclosure, “A or B” can mean “A only”, “B only”, or “both A and B”. In other words, “A or B” can be interpreted in this disclosure as “A and / or B”. For example, “A, B or C” in this disclosure can mean “A only”, “B only”, “C only”, or “any and all combinations of A, B and C”.
[0040] As used in this article, a forward slash ( / ) or a comma can mean "and / or". For example, "A / B" can mean "A and / or B". Therefore, "A / B" can mean "A only", "B only", or "both A and B". For example, "A, B, C" can mean "A, B, or C".
[0041] In this disclosure, "at least one of A and B" can mean "only A", "only B" or "both A and B". Furthermore, the expressions "at least one of A or B" or "at least one of A and / or B" can be interpreted in the same way as "at least one of A and B".
[0042] In this disclosure, "at least one of A, B, and C" can mean "A only", "B only", "C only", or "any and all combinations of A, B, and C". Furthermore, "at least one of A, B, or C" or "at least one of A, B, and / or C" can mean "at least one of A, B, and C".
[0043] The brackets used in this disclosure may mean "for example". Specifically, when indicated as "prediction (intra-frame prediction)", "intra-frame prediction" can be presented as an example of "prediction". In other words, "prediction" in this disclosure is not limited to "intra-frame prediction", and "intra-frame prediction" can be presented as an example of "prediction". Furthermore, even when indicated as "prediction (i.e., intra-frame prediction)", "intra-frame prediction" can be presented as an example of "prediction".
[0044] In this disclosure, the technical features described in a single figure may be implemented independently or simultaneously.
[0045] This disclosure relates to video / image coding. For example, the methods / implementations described in this disclosure can be applied to enhanced compression models or methods disclosed in the H.267 standard. Furthermore, the methods / implementations disclosed herein can be applied to methods disclosed in the AOMedia Video 2 (AV2) standard or next-generation video / image coding standards (e.g., H.268 and H.269).
[0046] In this disclosure, encoding may include encoding and / or decoding. In this disclosure, image encoding may be used interchangeably with video encoding.
[0047] In this disclosure, video can refer to a collection of images over time. An image typically refers to a unit representing a single image at a particular time, and a slice / tile refers to a unit that forms part of an image in encoding. A slice / tile may include one or more coding tree units (CTUs). A single image may include one or more slices / tiles. A tile may represent a rectangular area of a CTU within a specific tile row and a specific tile column of an image.
[0048] A single image can be divided into two or more sub-images. A sub-image can be a rectangular region of one or more slices of the image.
[0049] A pixel or cell can refer to the smallest unit that makes up a picture (or image). A "sample" can be used as the term corresponding to a pixel. A sample can generally represent a pixel or pixel value, and can represent only the pixel / pixel value of the luminance component or only the pixel / pixel value of the chrominance component.
[0050] A unit can represent the basic unit of image processing. A unit may include a specific region of an image and at least one piece of information associated with that region. A single unit may include a luminance block and two chrominance (e.g., Cb and Cr) blocks. The term "unit" may be used interchangeably with terms such as "block" or "region" in some cases. Typically, an M×N block may include a set (or array) of samples or transform coefficients in M columns and N rows.
[0051] Hereinafter, embodiments of the present disclosure will be described in detail with reference to the accompanying drawings. Furthermore, the same reference numerals may be used throughout the drawings to indicate the same elements, and redundant descriptions of the same elements may be omitted.
[0052] Figure 1 Examples of video / image coding systems to which embodiments of the present disclosure may be applied are illustrated schematically.
[0053] refer to Figure 1A video / image encoding system may include a first device (encoding device) and a second device (decoding device). The first device may deliver encoded video / image information or data to the second device in the form of a file or stream via a digital storage medium or network.
[0054] A video / image encoding system may also include a video / image acquisition device and a video / image renderer. The video / image acquisition device may be included in the encoding device, or it may be configured as a separate device or an external component. The video / image renderer may be included in the decoding device, or it may be configured as a separate device or an external component.
[0055] The first device may include a transmitter as an internal component, or a transmitter as a separate device or an external component.
[0056] The second device may include a receiver as an internal component, or as a receiver as a separate device or an external component.
[0057] An encoding device can be called an encoder, and a decoding device can be called a decoder. A transmitter can be included in the encoding device. A receiver can be included in the decoding device. A renderer can include a display, and the display can be configured as a separate device or an external component.
[0058] The decoding and encoding devices applied in one or more embodiments of this disclosure can be included in multimedia broadcasting transmitters / receivers, mobile communication terminals, home theater video devices, digital cinema video devices, surveillance cameras, video conferencing devices, real-time communication devices such as video communication, mobile streaming devices, storage media, cameras, video-on-demand (VoD) service providers, over-the-air (OTT) video devices, internet streaming service providers, three-dimensional (3D) video devices, virtual reality (VR) devices, augmented reality (AR) devices, video telephony devices, transportation terminals (e.g., vehicle terminals (including autonomous vehicle terminals), aircraft terminals, and ship terminals), and medical video devices, and can be used to process video signals or data signals. For example, over-the-air (OTT) video devices can include game consoles, Blu-ray players, internet-connected TVs, home theater systems, smartphones, tablet PCs, and digital video recorders (DVRs).
[0059] A video / image acquisition device can acquire video / image sources. The video / image acquisition device can acquire video / images through processes of capturing, compositing, or generating video / images. The video / image acquisition device may include a video / image capture device and / or a video / image generation device. The video / image capture device may include, for example, one or more cameras and a video / image archive including previously captured video / images. The video / image generation device may include, for example, a camera, a computer, a tablet PC, and a smartphone, and can (electronically) generate video / images. For example, virtual video / images can be generated by a computer, in which case the video / image capture process can be replaced by a process of generating relevant data. The video / image source can perform a video / image preprocessing process to input the optimized video / image into the encoder.
[0060] Encoding devices can encode input video / images. They can perform a series of processes, such as prediction, transformation, and quantization, to achieve compression and encoding efficiency. The encoded data (encoded video / image information) can be output as a bitstream.
[0061] The transmitter can send encoded images / image information or data, output as a bitstream, to a receiver of a receiving device via a digital storage medium or network, either as a file or a stream. The encoded images / image information or data output as a bitstream can be sent to the receiver via a streaming server. The digital storage medium can include various storage media such as USB, SD, CD, DVD, Blu-ray, HDD, and SSD. The transmitter can include elements for generating media files according to a predetermined file format and may include elements for transmission via a broadcast / communication network. The receiver can receive / extract the bitstream and send the received bitstream to a decoding device.
[0062] A streaming server can temporarily store bitstreams during the sending or receiving of bitstreams. Based on user requests via a web server, the streaming server sends multimedia data to the user's device, with the web server acting as an intermediary to notify the user of available services. When a user requests a desired service from the web server, the web server forwards the request to the streaming server, and the streaming server sends the multimedia data to the user. The content streaming system may include a separate control server, in which case the control server manages the commands / responses between devices within the content streaming system.
[0063] A streaming server can receive content from media storage devices and / or encoding devices. For example, when receiving content from an encoding device, the content can be received in real time. In this case, the streaming server can store the bitstream for a certain period of time to provide a smooth streaming service.
[0064] Decoding devices can decode video / images by performing a series of processes such as dequantization, inverse transform, and prediction, which correspond to the operations of encoding devices.
[0065] The renderer can render the decoded video / images. The rendered video / images can then be displayed on the monitor.
[0066] Figure 2 This diagram schematically illustrates the configuration of a video / image encoding apparatus to which embodiments of the present disclosure may be applied. Hereinafter, the encoding apparatus may include image encoding apparatus and / or video encoding apparatus.
[0067] refer to Figure 2 The encoding device 200 may include an image partitioner 210, a predictor 220, a residual processor 230, an entropy encoder 240, an adder 250, a filter 260, and a memory 270. The predictor 220 may include an inter-frame predictor and an intra-frame predictor. The residual processor 230 may include a transformer 232, a quantizer 233, a dequantizer 234, and an inverse transformer 235. The residual processor 230 may also include a subtractor 231. The adder 250 may be referred to as a reconstructor or a reconstruction block generator. According to embodiments, the image partitioner 210, predictor 220, residual processor 230, entropy encoder 240, adder 250, and filter 260 may be configured as at least one hardware component (e.g., an encoder chipset or processor). The memory 270 may include a decoded picture buffer (DPB) or may be configured as a digital storage medium. The hardware component may also include the memory 270 as an internal / external component.
[0068] Image partitioner 210 can partition an input image (or picture or frame) input to encoding device 200 into one or more processing units. For example, a processing unit may be referred to as a coding unit (CU). In this case, the coding unit can be recursively partitioned from a coding tree unit (CTU) or a maximum coding unit (LCU) according to a quadtree-binary-tritree (QTBTTT) structure. For example, a single coding unit can be partitioned into multiple coding units of greater depth based on a quadtree structure, a binary tree structure, and / or a ternary tree structure. In this case, for example, a quadtree structure can be applied first, and a binary tree structure and / or a ternary tree structure can be applied later. Alternatively, a binary tree structure can be applied first. The encoding process according to this disclosure can be performed based on the final coding unit that is no longer partitioned. In this case, based on the encoding efficiency according to the image characteristics, the maximum coding unit can be used as the final coding unit, or, if necessary, the coding unit can be recursively partitioned into deeper coding units, thereby using the coding unit with the optimal size as the final coding unit. Here, the encoding process may include prediction, transformation, and reconstruction processes, which will be described below. In another example, the processing unit may also include a prediction unit (PU) or a transformation unit (TU). In this case, the prediction unit and the transformation unit can be split or partitioned from the aforementioned final encoding unit. The prediction unit may be a unit for sample prediction, and the transformation unit may be a unit for deriving the transform coefficients and / or a unit for deriving the residual signal from the transform coefficients.
[0069] The term "unit" can be used interchangeably with the terms "block" or "region" depending on the context. Typically, an M×N block can represent an array of samples or transform coefficients arranged in M rows and N columns. Samples can typically represent pixels or pixel values, and can represent only the pixel / pixel value of the luminance component or only the pixel / pixel value of the chrominance component. "Sample" can be used as a term corresponding to pixels or cells in a single picture (or image).
[0070] Encoding device 200 generates a residual signal (residual signal, residual block, or residual sample array) by subtracting the prediction signal (prediction block or prediction sample array) output from the predictor from the input image signal (original block or original sample array), and the generated residual signal is sent to converter 232. In this case, as illustrated, the component in encoder 200 used to subtract the prediction signal (prediction block or prediction sample array) from the input image signal (original block or original sample array) can be referred to as subtractor 231. The predictor can perform prediction on the processing target block (hereinafter referred to as the current block) and can generate a prediction block including prediction samples for the current block. The predictor can determine whether intra-frame prediction or inter-frame prediction is applied based on the current block or CU. The predictor can generate various information about the prediction, such as prediction mode information, and can send the generated information to entropy encoder 240, as described below in the description of each prediction mode. The information about the prediction can be encoded by entropy encoder 240 and output as a bitstream.
[0071] An intra-frame predictor can refer to samples within the current image to predict the current block. The referenced samples can be located adjacent to (near) the current block, or, depending on the prediction mode, located far from the current block. In intra-frame prediction, prediction modes can include multiple non-directional modes and multiple directional modes. Non-directional modes can include, for example, DC modes and planar modes. Directional modes can include, for example, 33 or 65 directional prediction modes based on the granularity of the prediction direction. However, this example is for illustration only, and more or fewer directional prediction modes can be used depending on the configuration. The intra-frame predictor can determine the prediction mode to be applied to the current block based on the prediction modes applied to neighboring blocks.
[0072] Inter-frame predictors can derive predicted blocks for the current block based on reference blocks (reference sample arrays) specified by motion vectors on a reference image. Here, to reduce the amount of motion information transmitted in inter-frame prediction mode, motion information can be predicted based on the correlation between the motion information of neighboring blocks and the current block, on a block, sub-block, or sample basis. Motion information can include motion vectors and reference image indices. Motion information can also include inter-frame prediction direction (L0 prediction, L1 prediction, and bidirectional prediction) information. In inter-frame prediction, neighboring blocks can include spatially adjacent blocks existing within the current image and temporally adjacent blocks existing in the reference image. The reference image including the reference block and the reference image including the temporally adjacent block can be the same or different. Temporally adjacent blocks can be referred to as collinear reference blocks or collinear CUs (colCUs), and the reference image including temporally adjacent blocks can also be referred to as collinear images (colPics). For example, the inter-frame predictor can configure a motion information candidate list based on neighboring blocks and can generate information indicating candidates for deriving the motion vectors and / or reference image indices for the current block. Inter-frame prediction can be performed based on various prediction modes. For example, in skip and merge modes, the inter-frame predictor can use motion information about neighboring blocks as motion information about the current block. In skip mode, unlike merge mode, the residual signal may not be sent. In motion vector prediction (MVP) mode, the motion vector of the current block can be indicated by using the motion vectors of neighboring blocks as the motion vector predictor and signaling the motion vector difference.
[0073] Predictor 220 can generate prediction signals based on various prediction methods described below. For example, the predictor can not only apply intra-frame prediction or inter-frame prediction to predict a block, but can also apply intra-frame prediction and inter-frame prediction simultaneously, which can be referred to as combined inter-frame and intra-frame prediction (CIIP). Furthermore, the predictor can predict blocks based on an intra-block copy (IBC) prediction mode or a palette mode. The IBC prediction mode or palette mode can be used, for example, for screen content coding (SCC). IBC essentially performs prediction within the current frame, but can be performed similarly to inter-frame prediction in that it derives a reference block within the current frame based on a block vector. That is, IBC can utilize at least one of the inter-frame prediction techniques described in this disclosure.
[0074] The predicted signal generated by predictor 220 can be used to generate a reconstructed signal or a residual signal. Transformer 232 can generate transform coefficients by applying transform techniques to the residual signal. For example, the transform techniques may include at least one of Discrete Cosine Transform (DCT), Discrete Sine Transform (DST), Karhunen-Loève Transform (KLT), Graph-Based Transform (GBT), or Conditional Nonlinear Transform (CNT).
[0075] Quantizer 233 can quantize the transform coefficients and send the quantized transform coefficients to entropy encoder 240, which can encode the quantized signal (information about the quantized transform coefficients) and output the encoded signal as a bitstream. The information about the quantized transform coefficients can be referred to as residual information. Quantizer 233 can rearrange the block-form quantized transform coefficients into a one-dimensional vector form based on the coefficient scan order, and can generate information about the transform coefficients based on the one-dimensional vector form of the quantized transform coefficients. Entropy encoder 240 can perform various encoding methods, such as Exponential Golomb, Context Adaptive Variable Length Coding (CAVLC), and Context Adaptive Binary Arithmetic Coding (CABAC). Entropy encoder 240 can encode information necessary for video / image reconstruction (e.g., values of syntax elements) other than the quantized transform coefficients, either together or separately. The encoded information (e.g., encoded video / image information) can be sent or stored as a bitstream on a Network Abstraction Layer (NAL) basis. The video / image information may also include information about various parameter sets, such as Adaptive Parameter Set (APS), Picture Parameter Set (PPS), Sequence Parameter Set (SPS), or Video Parameter Set (VPS). Furthermore, the video / image information may also include general constraint information. In this disclosure, information and / or syntax elements transmitted / signed from the encoding device to the decoding device can be included in the video / image information. The video / image information can be encoded by the aforementioned encoding process and included in a bitstream. The bitstream can be transmitted over a network or stored in a digital storage medium. The network may include broadcast networks and / or communication networks, and the digital storage medium may include various storage media such as USB, SD, CD, DVD, Blu-ray, HDD, and SSD. A transmitter (not shown) and / or storage device (not shown) for transmitting and / or storing signals output from the entropy encoder 240 may be configured as an internal / external element of the encoding device 200, or the transmitter may be included in the entropy encoder 240.
[0076] The quantized transform coefficients output from quantizer 233 can be used to generate a prediction signal. For example, the residual signal (residual block or residual sample) can be reconstructed by applying dequantization and inverse transform to the quantized transform coefficients using dequantizer 234 and inverse transform unit 235. Adder 250 can add the reconstructed residual signal to the prediction signal output from the predictor to generate a reconstructed signal (reconstructed image, reconstructed block, or array of reconstructed samples). When there is no residual for the processing target block, such as when a skip mode is applied, the prediction block can be used as a reconstructed block. Adder 250 can be referred to as a reconstructor or reconstructed block generator. The generated reconstructed signal can be used for intra-frame prediction of the next processing target block in the current image, or it can be used for inter-frame prediction of the next image after being filtered as follows.
[0077] Luminance mapping with chroma scaling (LMCS) can be applied in image encoding and / or reconstruction processes.
[0078] Filter 260 can improve subjective / objective image quality by applying filtering to the reconstructed signal. For example, filter 260 can generate a modified reconstructed image by applying various filtering methods to the reconstructed image, and the modified reconstructed image can be stored in memory 270, specifically in the DPB of memory 270. Various filtering methods can include, for example, deblocking filtering, sample adaptive offset, adaptive loop filtering, and bilateral filtering. Filter 260 can generate filtering-related information and can send the generated information to entropy encoder 240. The filtering-related information can be encoded by entropy encoder 240 and output as a bitstream.
[0079] The modified reconstructed image sent to memory 270 can be used as a reference image in the inter-frame predictor. When inter-frame prediction is applied via the modified reconstructed image, the encoding device can avoid prediction mismatch between the encoding device 200 and the decoding device, and can improve encoding efficiency.
[0080] The DPB of memory 270 can store modified reconstructed images for use as reference images in the inter-frame predictor. Memory 270 can store motion information about blocks in the current image from which motion information is derived (or encoded) and / or about blocks in the reconstructed image. The stored motion information can be sent to the inter-frame predictor to be used as motion information about spatially adjacent blocks or about temporally adjacent blocks. Memory 270 can store reconstructed samples of reconstructed blocks in the current image and can send these reconstructed samples to the intra-frame predictor.
[0081] Figure 3 This diagram schematically illustrates the configuration of a video / image decoding device to which embodiments of the present disclosure can be applied. Hereinafter, the decoding device may include an image decoding device and / or a video decoding device.
[0082] refer to Figure 3The decoding device 300 may include an entropy decoder 310, a residual processor 320, a predictor 330, an adder 340, a filter 350, and a memory 360. The predictor 330 may include an inter-frame predictor and an intra-frame predictor. The residual processor 320 may include a dequantizer 321 and an inverse transformer 322. According to embodiments, the entropy decoder 310, residual processor 320, predictor 330, adder 340, and filter 350 may be configured as a single hardware component (e.g., a decoder chipset or processor). The memory 360 may include a decoded picture buffer (DPB) and may be configured as a digital storage medium. The hardware component may also include the memory 360 as an internal / external component.
[0083] When a bitstream including video / image information is input, the decoding device 300 can, based on the... Figure 2 The decoding device 300 reconstructs an image by processing video / image information in an encoding device. For example, the decoding device 300 can derive units / blocks based on information related to block partitions obtained from the bitstream. The decoding device 300 can perform decoding using processing units applied to the encoding device. Therefore, the processing unit used for decoding can be, for example, an encoding unit, and the encoding unit can be partitioned from encoding tree units or maximum encoding units according to a quadtree structure, binary tree structure, and / or ternary tree structure. One or more transform units can be derived from the encoding unit. The reconstructed image signal decoded and output by the decoding device 300 can be reproduced via a reproduction device.
[0084] Decoding device 300 can receive signals output from encoding device in the form of a bitstream, and the received signals can be decoded by entropy decoder 310. For example, entropy decoder 310 can parse the bitstream to derive information (e.g., video / image information) necessary for image reconstruction (or picture reconstruction). The video / image information may also include information about various parameter sets, such as adaptive parameter sets (APS), picture parameter sets (PPS), sequence parameter sets (SPS), or video parameter sets (VPS). Furthermore, the video / image information may also include general constraint information. Decoding device can further decode the picture based on the information about the parameter sets and / or general constraint information. In this disclosure, the information and / or syntax elements that are signaled / received, as described below, can be decoded through a decoding process and can be obtained from the bitstream. For example, entropy decoder 310 can decode information in the bitstream based on encoding methods such as exponential Golomb coding, CAVLC, or CABAC, and can output the values of the syntax elements required for image reconstruction and the quantized values of the transform coefficients for the residuals. More specifically, the CABAC entropy decoding method can receive bins corresponding to each syntax element in the bitstream, determine a context model using information about the target syntax element and decoding information about adjacent and target blocks or information about symbols / bins decoded in previous stages, and generate symbols corresponding to the values of each syntax element by predicting the occurrence probability of bins based on the determined context model and performing arithmetic decoding on the bins. Here, after determining the context model, the CABAC entropy decoding method can update the context model using information about the decoded symbols / bins for use in the context model of the next symbol / bin. Prediction-related information from the information decoded by the entropy decoder 310 can be provided to the predictor 330, and the residual values (i.e., quantized transform coefficients and related parameter information) obtained by entropy decoding in the entropy decoder 310 can be input to the residual processor 320. The residual processor 320 can derive residual signals (residual blocks, residual samples, or arrays of residual samples). Filtering-related information from the information decoded by the entropy decoder 310 can be provided to the filter 350. A receiver (not shown) for receiving signals output from the encoding device may be further configured as an internal / external element of the decoding device 300, or the receiver may be a component of the entropy decoder 310. The decoding device according to this disclosure may be referred to as a video / image / picture decoding device and may be divided into an information decoder (video / image / picture information decoder) and a sample decoder (video / image / picture sample decoder). The information decoder may include the entropy decoder 310, and the sample decoder may include at least one of a dequantizer 321, an inverse transformer 322, an adder 340, a filter 350, a memory 360, and a predictor 330.
[0085] Dequantizer 321 can dequantize the quantized transform coefficients to output transform coefficients. Dequantizer 321 can rearrange the quantized transform coefficients into two-dimensional blocks. In this case, the rearrangement can be performed based on the coefficient scan order performed in the encoding device. Dequantizer 321 can perform dequantization on the quantized transform coefficients using quantization parameters (e.g., quantization step size information) and obtain the transform coefficients.
[0086] The inverse transformer 322 performs an inverse transformation on the transformation coefficients to obtain the residual signal (residual block or residual sample array).
[0087] The predictor can perform prediction for the current block and generate a prediction block that includes prediction samples of the current block. The predictor can determine whether to apply intra-frame prediction or inter-frame prediction to the current block based on prediction-related information output from the entropy decoder 310, and determine the specific intra-frame / inter-frame prediction mode.
[0088] Predictor 330 can generate a prediction signal based on various prediction methods described below. For example, the predictor can not only apply intra-frame prediction or inter-frame prediction to predict a block, but can also apply intra-frame prediction and inter-frame prediction simultaneously, which can be referred to as combined inter-frame and intra-frame prediction (CIIP). Furthermore, the predictor can predict blocks based on an intra-block copy (IBC) prediction mode or a palette mode. The IBC prediction mode or palette mode can be used, for example, for screen content coding (SCC). IBC essentially performs prediction within the current frame, but can be performed similarly to inter-frame prediction in that it derives a reference block within the current frame based on a block vector. That is, IBC can utilize at least one of the inter-frame prediction techniques described in this disclosure.
[0089] An intra-frame predictor can refer to samples within the current image to predict the current block. The referenced samples can be located adjacent to (near) the current block, or, depending on the prediction mode, located far from the current block. In intra-frame prediction, the prediction mode can include multiple non-directional modes and multiple directional modes. The intra-frame predictor can determine the prediction mode to be applied to the current block based on the prediction modes applied to neighboring blocks.
[0090] An inter-frame predictor can deduce a predicted block for the current block based on a reference block (reference sample array) specified by motion vectors on a reference image. Here, to reduce the amount of motion information transmitted in the inter-frame prediction mode, motion information can be predicted on a block, sub-block, or sample basis based on the correlation between the motion information of neighboring blocks and the current block. Motion information may include motion vectors and reference image indices. Motion information may also include inter-frame prediction direction (L0 prediction, L1 prediction, and bidirectional prediction) information. In inter-frame prediction, neighboring blocks may include spatially adjacent blocks existing within the current image and temporally adjacent blocks existing in the reference image. For example, inter-frame predictor 332 can configure a motion information candidate list based on neighboring blocks and can deduce the motion vector and / or reference image index of the current block based on the received candidate selection information. Inter-frame prediction can be performed based on various prediction modes, and the information about the prediction may include information indicating the inter-frame prediction mode for the current block.
[0091] Adder 340 can add the obtained residual signal to the prediction signal (prediction block or prediction sample array) output from predictor 330 to generate a reconstruction signal (reconstructed image, reconstruction block, or reconstruction sample array). When there is no residual for the processing target block, such as when a skip mode is applied, the prediction block can be used as a reconstruction block.
[0092] Adder 340 can be referred to as a reconstructor or reconstructed block generator. The generated reconstructed signal can be used for intra-frame prediction of the next processing target block in the current image, or it can be output after being filtered as described below, or used for inter-frame prediction of the next image.
[0093] Luminance mapping with chroma scaling (LMCS) can be applied in the image decoding process.
[0094] Filter 350 can improve subjective / objective image quality by applying filtering to the reconstructed signal. For example, filter 350 can generate a modified reconstructed image by applying various filtering methods to the reconstructed image, and the modified reconstructed image can be sent to memory 360, specifically to the DPB in memory 360. Various filtering methods can include, for example, unblocking filtering, sample adaptive shifting, adaptive loop filtering, and bilateral filtering.
[0095] The (modified) reconstructed image stored in the DPB of memory 360 can be used as a reference image in the inter-frame predictor. Memory 360 can store motion information about blocks in the current image from which motion information is derived (or decoded) and / or about blocks in the reconstructed image. The stored motion information can be sent to the inter-frame predictor to be used as motion information about spatially adjacent blocks or about temporally adjacent blocks. Memory 360 can store reconstructed samples of reconstructed blocks in the current image and can send the reconstructed samples to the intra-frame predictor.
[0096] The embodiments described in this specification for the filter 260 and predictor 220 of the encoding device 200 can also be applied equally or correspondingly to the filter 350 and predictor 330 of the decoding device 300.
[0097] As described above, during video encoding, prediction is performed to increase compression efficiency. Prediction generates prediction blocks that include prediction samples for the current block, which is the target block for encoding. Prediction blocks include prediction samples in the spatial domain (or pixel domain). Prediction blocks are derived identically in both the encoding and decoding devices, and the encoding device can enhance image coding efficiency by signaling information about the residuals between the original block and the prediction blocks (residual information) to the decoding device, rather than the original sample values of the original block. The decoding device can derive residual blocks including residual samples based on the residual information, generate reconstructed blocks including reconstructed samples by combining the residual blocks with the prediction blocks, and generate a reconstructed image including the reconstructed blocks.
[0098] Residual information can be generated through transformation and quantization processes. For example, an encoding device can derive a residual block between the original block and the prediction block, perform a transformation process on the residual samples (residual sample array) included in the residual block to derive transform coefficients, perform a quantization process on the transform coefficients to derive quantized transform coefficients, and signal the relevant residual information to a decoding device (via a bitstream). Residual information may include information such as the values and locations of the quantized transform coefficients, the transform technique, the transform kernel, and the quantization parameters. The decoding device can perform dequantization / inverse transform processes based on the residual information to derive residual samples (or residual blocks). The decoding device can generate a reconstructed image based on the prediction block and the residual block. The encoding device can perform dequantization / inverse transform on the quantized transform coefficients to derive residual blocks used for reference in inter-frame prediction of subsequent images, and can generate a reconstructed image based on the residual blocks.
[0099] In this disclosure, at least one of quantization / dequantization and / or transform / inverse transform may be omitted. When quantization / dequantization is omitted, the quantized transform coefficients may be referred to as transform coefficients. When transform / inverse transform is omitted, the transform coefficients may be referred to as coefficients or residual coefficients, or, for the sake of consistency, may still be referred to as transform coefficients.
[0100] Furthermore, in this disclosure, quantized transform coefficients and transform coefficients can be referred to as transform coefficients and scaled transform coefficients, respectively. In this case, residual information may include information about one or more transform coefficients, and the information about one or more transform coefficients may be signaled via residual coding syntax. Transform coefficients can be derived based on residual information (or information about one or more transform coefficients), and scaled transform coefficients can be derived by inverse transforming (scaling) the transform coefficients. Residual samples can be derived based on inverse transforming (scaling) the scaled transform coefficients. These details can be equally applied to other parts of this disclosure or described in other parts of this disclosure.
[0101] As described above, the predictor of the encoding / decoding device can perform inter-frame prediction on a block-by-block basis to derive prediction samples. Inter-frame prediction can refer to predictions derived in a manner that depends on data elements (i.e., sample values, motion information, etc.) of images other than the current image. When inter-frame prediction is applied to the current block, the prediction block (prediction sample array) of the current block can be derived based on the reference block (reference sample array) on the reference image indicated by the reference image index, which is specified by motion vectors. Here, in order to reduce the amount of motion information transmitted in inter-frame prediction mode, the motion information of the current block can be predicted on a block-by-block, sub-block, or sample-by-sample basis based on the correlation between the motion information of neighboring blocks and the current block. Motion information can include motion vectors and reference image indices. Motion information can also include inter-frame prediction type (L0 prediction, L1 prediction, bidirectional prediction, etc.) information. When inter-frame prediction is applied, neighboring blocks can include spatially adjacent blocks in the current image and temporally adjacent blocks in the reference image. The reference image including the reference block and the reference image including the temporally adjacent block can be the same or different. Temporally adjacent blocks can be referred to as co-located reference blocks, co-located CUs (colCUs), etc., and reference pictures including temporally adjacent blocks can be referred to as co-located pictures (colPic). For example, a candidate list of motion information can be constructed based on the neighboring blocks of the current block, and a flag or index information indicating which candidate to select (use) to derive the motion vector and / or reference picture index of the current block can be signaled. Inter-frame prediction can be performed based on various prediction modes. For example, in skip mode and merge mode, the motion information of the current block can be the same as the motion information of the selected neighboring blocks. In skip mode, unlike merge mode, residual signals may not be sent. In motion vector prediction (MVP) mode, the motion vectors of the selected neighboring blocks can be used as motion vector predictors, and the motion vector difference can be signaled. In this case, the motion vector of the current block can be derived using the sum of the motion vector predictor and the motion vector difference.
[0102] Motion information can include L0 motion information and / or L1 motion information, depending on the inter-frame prediction type (L0 prediction, L1 prediction, bidirectional prediction, etc.). Motion vectors in the L0 direction can be referred to as L0 motion vectors or MVL0, and motion vectors in the L1 direction can be referred to as L1 motion vectors or MVL1. Prediction based on L0 motion vectors can be called L0 prediction, prediction based on L1 motion vectors can be called L1 prediction, and prediction based on both L0 and L1 motion vectors can be called bidirectional prediction. Here, L0 motion vectors can represent motion vectors associated with a reference image list L0 (L0), and L1 motion vectors can represent motion vectors associated with a reference image list L1 (L1). The reference image list L0 can include images that are earlier than the current image in the output order as reference images, and the reference image list L1 can include images that are later than the current image in the output order. Earlier images can be called forward (reference) images, and later images can be called backward (reference) images. The reference image list L0 can also include images that are later than the current image in the output order as reference images. In this case, within the reference image list L0, earlier images can be indexed first, followed by later images. The reference image list L1 can also include images that are earlier than the current image in the output order as reference images. In this case, within the reference image list L1, later images can be indexed first, followed by earlier images. Here, the output order can correspond to the Image Order Count (POC) order.
[0103] The video / image coding process and predictors in coding devices based on inter-frame prediction can typically perform the following operations, for example, for inter-frame prediction.
[0104] Figure 4 The inter-frame prediction process is illustrated.
[0105] refer to Figure 4 As described above, the inter-frame prediction process may include an inter-frame prediction mode / type determination step, a motion vector derivation / refinement step, and an inter-frame prediction execution (prediction sample generation) step. The inter-frame prediction process may be performed in an encoding device and a decoding device as described above. In this document, the encoding apparatus may include an encoding apparatus and / or a decoding device.
[0106] The coding device determines the inter-frame prediction mode / type (S400).
[0107] The encoding device can determine the inter-frame prediction mode / type to be applied to the current block from the various inter-frame prediction modes / types described in this disclosure, and can generate prediction-related information. The prediction-related information may include inter-frame prediction mode information indicating the inter-frame prediction mode applied to the current block and / or inter-frame prediction type information indicating the inter-frame prediction type applied to the current block. The decoding device can determine the inter-frame prediction mode / type to be applied to the current block based on the prediction-related information.
[0108] The coding device derives / refines the motion vector of the current block (S410). The coding device can derive / refine the motion vector of the current block based on the determined inter-frame prediction mode / type. Here, motion information of neighboring blocks of the current block can be used to derive / refine the motion vector.
[0109] For example, when a skip mode or merge mode is applied to the current block, the encoding device can construct a merge candidate list and select one merge candidate from the merge candidates included in the list. Information indicating the selected merge candidate (e.g., merge index) can be included in the prediction-related information.
[0110] As another example, when the (A)MVP mode is applied to the current block, the decoding device can construct an (A)MVP candidate list, and the motion vector of the MVP candidate selected from the MVP (Motion Vector Predictor) candidates included in the (A)MVP candidate list can be used as the MVP for the current block. The selection can be indicated based on selection information (MVP flag or MVP index). In this case, not only the selection information but also information about MVD can be included in the prediction-related information.
[0111] Furthermore, as described below, the motion information of the current block can be derived without constructing a candidate list, and in this case, the motion information of the current block can be derived according to the process disclosed in the prediction pattern / type described below. In this case, the candidate list construction as described above can be omitted.
[0112] The encoding device predicts (generates prediction samples) the current block based on the derived / refined motion vectors (S420). The encoding device can derive prediction samples for the current block using samples of a reference block indicated by motion vectors on a reference image.
[0113] The encoding process based on inter-frame prediction can typically include, for example, the following operations.
[0114] Figure 5 An example of a video / image coding method based on inter-frame prediction is shown.
[0115] refer to Figure 5S500 can be executed by the predictor of the encoding device, S505 can be executed by the residual processor of the encoding device, and S510 or S515 can be executed by the entropy encoder of the encoding device. Specifically, prediction-related information can be derived by the predictor and encoded by the entropy encoder. Residual information can be derived by the residual processor and encoded by the entropy encoder. Residual information is information about residual samples. Residual information may include information about the quantized transform coefficients for the residual samples. As mentioned above, residual samples can be derived into transform coefficients by the transformer of the encoding device, and transform coefficients can be derived into quantized transform coefficients by the quantizer. Information about the quantized transform coefficients can be encoded by the entropy encoder via the residual encoding process.
[0116] The encoding device performs inter-frame prediction on the current block (S500). The encoding device can deduce the inter-frame prediction mode / type and motion information of the current block, and can generate prediction samples for the current block. Here, the inter-frame prediction mode / type determination, motion information deduction, and prediction sample generation processes can be performed simultaneously, or one process can be performed before the other. For example, the inter-frame predictor of the encoding device can search for blocks similar to the current block within a specific region (search region) of the reference image through motion estimation, and can deduce reference blocks with a minimum difference from the current block, or a smaller difference based on a specific criterion. Based on this, a reference image index indicating the location of the reference block in the reference image can be derived, and a motion vector can be derived based on the positional difference between the reference block and the current block. The encoding device can determine the mode to be applied to the current block from various prediction modes. The encoding device can compare the RD costs of various prediction modes and determine the optimal prediction mode for the current block.
[0117] For example, when a skip mode or merge mode is applied to the current block, the encoding device can construct a merge candidate list and deduce a reference block from among the reference blocks indicated by the merge candidates included in the merge candidate list that has the minimum difference with the current block, or a reference block with a specific criterion or smaller. In this case, the merge candidate associated with the deduced reference block can be selected, and merge index information indicating the selected merge candidate can be generated and signaled to the decoding device. The motion information of the current block can be deduced using the motion information of the selected merge candidate.
[0118] As another example, when the (A)MVP mode is applied to the current block, the encoding device can construct the (A)MVP candidate list described below, and can use the motion vector of the MVP candidate selected from the MVP (Motion Vector Predictor) candidates included in the (A)MVP candidate list as the MVP of the current block. In this case, for example, the motion vector of the reference block derived through the motion estimation described above can be used as the motion vector of the current block, and among the MVP candidates, the MVP candidate with the motion vector having the smallest difference from the motion vector of the current block can be the selected MVP candidate. The motion vector difference (MVD) can be derived, that is, the difference obtained by subtracting the MVP from the motion vector of the current block. In this case, information about the MVD can be signaled to the decoding device. Furthermore, when the (A)MVP mode is applied, the value of the reference picture index can be configured as reference picture index information and signaled separately to the decoding device.
[0119] The encoding device can perform residual processing based on the predicted samples (S505). The encoding device can derive residual samples based on the predicted samples. The encoding device can derive residual samples by comparing the original samples of the current block with the predicted samples. Residual information can be generated based on the residual samples. The residual information may include information about the quantized transform coefficients as described above.
[0120] The encoding device encodes image information, including prediction-related information and / or residual information (S510 or S515). The encoding device can output the encoded image information as a bitstream. Prediction-related information is information related to the prediction process and may include prediction mode information (e.g., skip flag, merge flag, mode index, etc.) and information about motion information. Information about motion information may include candidate selection information (e.g., merge index, MVP flag, or MVP index), which is information used to derive motion vectors. Furthermore, information about motion information may include information about the aforementioned MVD and / or reference image index information. Additionally, information about motion information may include information indicating whether L0 prediction, L1 prediction, or bidirectional prediction is applied. Residual information is information about residual samples. Residual information may include information about the quantized transform coefficients for the residual samples.
[0121] The output bitstream can be stored in a (digital) storage medium and delivered to a decoding device, or it can be delivered to a decoding device via a network.
[0122] Furthermore, as mentioned above, the encoding device can generate a reconstructed image (including reconstructed samples and reconstructed blocks) based on reference samples and residual samples. This is to derive the same prediction result in the encoding device as the prediction result performed in the decoding device, thereby improving encoding efficiency. Therefore, the encoding device can store the reconstructed image (or reconstructed samples, reconstructed blocks) in memory and use it as a reference image for inter-frame prediction. Loop filtering processes, etc., can be further applied to the reconstructed image as described above.
[0123] Decoding devices can perform operations corresponding to those performed in encoding devices. A video / image decoding process based on inter-frame prediction may, for example, include the following operations.
[0124] Figure 6 An example of a video / image decoding method based on inter-frame prediction is shown.
[0125] refer to Figure 6 S600 can be executed by the entropy decoder of the decoding device, S610 can be executed by the predictor of the decoding device, S615 can be executed by the residual processor of the decoding device, and S620 can be executed by the adder or reconstructor of the decoding device.
[0126] Specifically, the decoding device obtains image / video information from the bitstream (S600). The image / video information may include prediction-related information and / or residual information.
[0127] The decoding device performs inter-frame prediction based on prediction-related information (S610). The decoding device can deduce the inter-frame prediction mode / type of the current block, deduce / refine the motion information of the current block, and generate prediction samples in the current block based on the inter-frame prediction mode / type and / or motion information. In this case, the decoding device can perform a prediction sample filtering process. Prediction sample filtering can be referred to as post-filtering. Some or all of the prediction samples can be filtered through the prediction sample filtering process. In some cases, the prediction sample filtering process can be omitted.
[0128] The decoding device performs residual processing based on the residual information (S615). The decoding device can derive the residual samples of the current block based on the residual information. Specifically, the dequantizer of the residual processor can perform dequantization based on the quantized transform coefficients derived from the residual information to derive the transform coefficients, and the inverse transformer of the residual processor can perform an inverse transform on the transform coefficients to derive the residual samples of the current block.
[0129] The decoding device generates a reconstructed block / image (S620). The decoding device can generate reconstructed samples for the current block based on predicted samples and / or residual samples, and can derive a reconstructed block including the reconstructed samples. A reconstructed image of the current image can be generated based on the reconstructed block. Loop filtering processes, etc., can be further applied to the reconstructed image as described above.
[0130] Prediction-related information can be encoded / decoded using the binarization and encoding methods described in this document. For example, prediction-related information can be binarized using fixed-length binarization, truncated Rice binarization, truncated unary binarization, etc. For example, prediction-related information can be encoded / decoded using entropy coding (e.g., CABAC, CAVLC).
[0131] Furthermore, for example, according to this document, geometric partitioning mode (GPM) can be applied as an implementation of inter-frame prediction. GPM can be considered as one of the prediction techniques or prediction types. When GPM is applied to the current block, the current block can be partitioned into two partitions, motion information of each of the two partitions can be derived, and prediction samples for the current block can be derived by performing inter-frame prediction on each partition based on the motion information of each of the two partitions.
[0132] Figure 7 An example of a partition shape supported by GPM is illustrated.
[0133] refer to Figure 7 GPM can support 64 partition shapes from combinations of 20 angles and 4 distances. For example, the 64 partition shapes can include 32 partition shapes as combinations of 8 angles and 4 distances, 24 partition shapes as combinations of 8 angles and 3 distances, and 8 partition shapes as combinations of 4 angles and 2 distances.
[0134] For example, a signal can be used to indicate the GPM partition index, which specifies the partition shape of the current block. For instance, the partition shape of the current block can be derived from the partition shape indicated by the GPM partition index. The following table can represent the partition shapes indicated by the GPM partition index.
[0135] [Table 1]
[0136] Here, gpm_partition_idx can represent the GPM partition index, angleIdx can represent the angle index, and distanceIdx can represent the distance index. The partition shape of the current block can be derived based on the GPM partition index notified by signals, and the current block can be partitioned into multiple partitions based on the derived partition shape.
[0137] Furthermore, GPM can also be applied to non-square blocks. For example, Figure 7 (b) can represent the partition shape of a GPM applied to a non-square block with a W / H ratio of 2. Here, W can represent the width of the block, and H can represent the height of the block. For example, when the current block is a non-square block with a W / H ratio of 2, the GPM partition index notified by the signal can represent the partition shape shown in the table below.
[0138] [Table 2]
[0139] For example, when the current block is a non-square block with a W / H of 2, the partition shape of the current block can be derived based on Table 2 and the GPM partition index instead of Table 1.
[0140] In addition, for example, Figure 7 (c) can represent the partition shape of a GPM applied to a non-square block with a W / H of 4. For example, when the current block is a non-square block with a W / H of 4, the GPM partition index notified by the signal can represent the partition shape shown in the table below.
[0141] [Table 3]
[0142] For example, when the current block is a non-square block with a W / H of 4, the partition shape of the current block can be derived based on Table 3 and the GPM partition index.
[0143] For example, the table used for partition shape can be derived based on the size of the current block, and the partition shape of the current block can be derived based on the derived table and the GPM partition index. Alternatively, for example, the table used for partition shape can be derived based on the width and / or height of the current block, and the partition shape of the current block can be derived based on the derived table and the GPM partition index. Alternatively, for example, the table used for partition shape can be derived based on the ratio of the width to the height of the current block, and the partition shape of the current block can be derived based on the derived table and the GPM partition index.
[0144] Specifically, for example, when the W / H of the current block is 1, Table 1 can be derived as a table for the partition shape of the current block, and the partition shape of the current block can be derived based on Table 1 and the GPM partition index of the current block. Alternatively, for example, when the W / H of the current block is 2, Table 2 can be derived as a table for the partition shape of the current block, and the partition shape of the current block can be derived based on Table 2 and the GPM partition index of the current block. Alternatively, for example, when the W / H of the current block is 4, Table 3 can be derived as a table for the partition shape of the current block, and the partition shape of the current block can be derived based on Table 3 and the GPM partition index of the current block.
[0145] Furthermore, for example, there may be constraints on the size of the block to which GPM is applied. That is, whether or not GPM is applied can be determined based on the size of the current block. For example, the minimum width or minimum height of the block to which GPM is applied can be 8 or 4. That is, when the width or height of the current block is less than 8 or 4, GPM may not be applied. Furthermore, for example, the maximum width or maximum height of the block to which GPM is applied can be 64 or 128. That is, when the width or height of the current block is greater than 64 or 128, GPM may not be applied. Furthermore, for example, the maximum width-to-height ratio of the block to which GPM is applied can be 1:4 or 4:1. That is, when the maximum width-to-height ratio of the current block is greater than 1:4 or 4:1, GPM may not be applied. Alternatively, for example, the maximum width-to-height ratio of the block to which GPM is applied can be 1:8 or 8:1. That is, when the maximum width-to-height ratio of the current block is greater than 1:8 or 8:1, GPM may not be applied.
[0146] Furthermore, for example, according to this document, a GPM with inter-frame and intra-frame prediction can be applied. A GPM with inter-frame and intra-frame prediction can be considered one of the prediction techniques or prediction types. A GPM with inter-frame and intra-frame prediction can be referred to as an inter-intra-frame prediction GPM. When a GPM with inter-frame and intra-frame prediction is applied to the current block, the current block can be partitioned into two partitions, and inter-frame prediction can be applied to one of the two partitions, while intra-frame prediction can be applied to the other partition.
[0147] For example, in a GPM with inter-frame and intra-frame prediction, the final prediction sample can be generated by applying weights to the inter-frame prediction sample and the intra-frame prediction sample for each partition. That is, the weights of each of the inter-frame prediction sample (i.e., the prediction sample of the partition to which inter-frame prediction is applied) and the intra-frame prediction sample (i.e., the prediction sample of the partition to which intra-frame prediction is applied) can be derived, and the final prediction sample can be generated based on a weighted sum of the inter-frame and intra-frame prediction samples. The inter-frame prediction sample can be derived based on the inter-frame GPM, and the intra-frame prediction sample can be derived based on the intra-frame prediction mode (IPM) candidate list and an index signaled from the coding device. The inter-frame prediction mode applied to one partition of the current block can be derived from the inter-frame GPM, and the intra-frame prediction mode applied to another partition of the current block can be derived from the IPM candidates in the constructed IPM candidate list, indicated by the index signaled. For example, the size of the IPM candidate list can be predefined as 3.
[0148] Figure 8 Intra-prediction modes that can be used as IPM candidates are illustrated by example.
[0149] refer to Figure 8The intra-prediction modes that can be used as IPM candidates (i.e., available IPM candidates) can be parallel angular modes (parallel modes) about GPM block boundaries, vertical angular modes (vertical modes) about GPM block boundaries, and / or planar modes. Figure 8 (a) can represent the parallel mode. Figure 8 (b) can represent the vertical mode, and Figure 8 (c) can represent a planar pattern.
[0150] Furthermore, for example, during the construction of the IPM candidate list, intra-prediction modes derived from the decoder-side intra-mode derivation (DIMD) method and / or neighboring block derivation can be derived as IPM candidates. When intra-prediction modes derived from DIMD and / or neighboring blocks are derived as IPM candidates, the aforementioned parallel modes can be derived as IPM candidates first. For example, when the size of the IPM candidate list is 3, after the parallel modes are derived as IPM candidates, intra-prediction modes derived from DIMD and / or neighboring blocks can be derived as IPM candidates. Therefore, when there are no identical IPM candidates in the IPM candidate list, up to two IPM candidates can be derived from DIMD and / or neighboring blocks. In the case of deriving neighboring intra-prediction modes (i.e., intra-prediction modes derived from neighboring blocks), the number of available neighboring block positions is at most 5, but this can be limited by the GPM block boundary angle as shown in the table below.
[0151] [Table 4]
[0152] Here, the GPM angle can represent the GPM block boundary angle indicated by the index, the first partition can represent the position of the available neighboring blocks of the first partition, and the second partition can represent the position of the available neighboring blocks of the second partition. For example, when the current block is partitioned by the GPM block boundary angle of index 0, the available neighboring blocks of the first partition may include the upper neighboring block, and the available neighboring blocks of the second partition may include the left neighboring block and the upper neighboring block. Alternatively, for example, when the current block is partitioned by the GPM block boundary angle of index 5, the available neighboring blocks of the first partition may include the left neighboring block and the upper neighboring block, and the available neighboring blocks of the second partition may include the left neighboring block.
[0153] For example, in GPM with inter-frame and intra-frame prediction, motion information for a partition applying inter-frame prediction can be derived based on a regular GPM MV candidate list. For example, the regular GPM MV candidate list can be a merge candidate list derived based on neighboring blocks of the partition applying inter-frame prediction. For example, the merge candidate list can be derived based on neighboring blocks of the partition applying inter-frame prediction, and motion information for the partition applying inter-frame prediction can be derived based on merge candidates in the merge candidate list indicated by the merge index for the partition applying inter-frame prediction.
[0154] Furthermore, GPM with inter-frame and intra-frame prediction can be combined with GPM-MMVD (GPM with motion vector difference merging). That is, in GPM with inter-frame and intra-frame prediction, GPM-MMVD can be applied to partitions where inter-frame prediction is applied.
[0155] For example, MMVD information for a partition using inter-frame prediction can be signaled, the MMVD of the partition can be derived from the MMVD information, and the motion information of the partition can be derived from the merge candidate and MMVD of the partition. The MMVD information can include an MMVD distance index and an MMVD direction index. The MMVD distance index represents the distance to the MMVD, and the MMVD direction index represents the direction of the MMVD.
[0156] Furthermore, regression-based GPM can be applied to GPM with inter-frame and intra-frame prediction, for example. For instance, a pairing list can be constructed for partitions of the current block applying GPM with inter-frame and intra-frame prediction, and the pairing index indicating the selected pairing candidate can be signaled. Pairing candidates can include intra-prediction mode candidates for partitions of the current block applying intra-frame prediction and MV candidates for partitions of the current block applying inter-frame prediction.
[0157] For example, a signal can be used to indicate whether a regression-based GPM with inter-frame and intra-frame prediction is applied. This signal can be used at the CU level. Here, the regression-based GPM with inter-frame and intra-frame prediction can be referred to as a regression-based inter-intra-frame prediction GPM, and the signal can be referred to as the regression-based inter-intra-frame GPM signal. Furthermore, for example, a signal indicating whether a regression-based GPM with inter-frame and intra-frame prediction is available can be used at a higher-level syntax (e.g., SPS, PPS). When the availability signal is 0, the signaling indicating whether a regression-based GPM with inter-frame and intra-frame prediction is applied can be omitted. The availability signal can be referred to as the regression-based inter-intra-frame prediction GPM availability signal.
[0158] When a regression-based GPM with inter-frame and intra-frame prediction is applied to the current block, a pairing list can be constructed for the partitions of the current block. For example, an intra-frame prediction mode candidate can be one of the top six intra-frame prediction modes in the MPM list, and an MV candidate can be one of the regular GPM MV candidates. The pairing of intra-frame prediction mode candidates and MV candidates can be derived based on two integer mixing matrices derived from the regression model for the template of the current block. Here, the integer mixing matrices can be as follows.
[0159] [Formula 1]
[0160] Here, parameters a, b, and c can be derived as parameters that minimize the mean squared error (MSE) of the template for the current block.
[0161] A signal can be used to indicate the pairing index for one of the pairing candidates in the pairing list, and the pairing candidate for the current block can be derived based on the pairing index. Motion information for the current block's partitions, applied through inter-frame prediction, can be derived based on the MV candidates of the derived pairing candidates.
[0162] Furthermore, to further improve coding performance, in GPMs with inter-frame and intra-frame prediction, TIMD can be used as an IPM candidate for partitions applying intra-frame prediction. For example, in the IPM candidate list for partitions applying intra-frame prediction, the parallel mode can be derived as an IPM candidate first, and then TIMD, DIMD, and neighboring blocks can be derived as IPM candidates in this order. That is, for example, when constructing the IPM candidate list, IPM candidates can be derived in the order of parallel mode, TIMD, DIMD, and intra-frame prediction modes derived from neighboring blocks.
[0163] Alternatively, for example, in a GPM with inter-frame and intra-frame prediction, the IPM candidate for the partition to which intra-frame prediction is applied may include at least one of vertical intra-frame prediction mode, horizontal intra-frame prediction mode, planar mode, DC mode, or directional planar mode.
[0164] Furthermore, for example, signaling for GPM with inter-frame and intra-frame prediction can be performed as follows.
[0165] [Table 5]
[0166] Referring to Table 5, signals can be used to notify the GPM partition index, inter-frame-intra-frame prediction GPM flag, and / or inter-frame-intra-frame prediction GPM mode flag for the current block.
[0167] For example, `gpm_partition_idx` can represent the GPM partition index. For example, the GPM partition index can represent the partition shape of the current block to which GPM is applied. For example, the block boundary angle and the distance between the midpoint of the current block and the block boundary can be derived from `gpm_partition_idx`.
[0168] The `gpm_partition_interintra_flag` indicates whether inter-frame intra-frame prediction (GPM) is applied. For example, when `gpm_partition_interintra_flag` is 1, it indicates that inter-frame intra-frame prediction GPM is applied to the current block, and when it is 0, it indicates that inter-frame intra-frame prediction GPM is not applied to the current block. Here, `gpm_partition_interintra_flag` can be referred to as the inter-frame intra-frame prediction GPM flag.
[0169] Furthermore, for example, when `gpm_partition_interintra_flag` indicates that inter-frame-intra prediction (GPM) is applied to the current block, `gpm_prediction_interintra_mode_flag` can be signaled. `gpm_prediction_interintra_mode_flag` can indicate the partitions for which inter-frame prediction is applied and the partitions for which intra-frame prediction is applied. Here, the partitions for which inter-frame prediction is applied can be called inter-partitions, and the partitions for which intra-frame prediction is applied can be called intra-partitions. Additionally, `gpm_prediction_interintra_mode_flag` can be referred to as the inter-frame-intra prediction GPM mode flag.
[0170] For example, when the value of `gpm_prediction_interintra_mode_flag` is 1, it indicates that inter-frame prediction is applied to the first partition of the current block and intra-frame prediction is applied to the second partition. Conversely, when the value of `gpm_prediction_interintra_mode_flag` is 0, it indicates that intra-frame prediction is applied to the first partition of the current block and inter-frame prediction is applied to the second partition. In other words, when the value of `gpm_prediction_interintra_mode_flag` is 1, it indicates that the first partition of the current block is an inter-frame partition and the second partition of the current block is an intra-frame partition; and when the value of `gpm_prediction_interintra_mode_flag` is 0, it indicates that the first partition of the current block is an intra-frame partition and the second partition of the current block is an inter-frame partition. Alternatively, for example, when the value of gpm_prediction_interintra_mode_flag is 0, gpm_prediction_interintra_mode_flag can indicate that inter-frame prediction is applied to the first partition of the current block and intra-frame prediction is applied to the second partition; and when the value of gpm_prediction_interintra_mode_flag is 1, gpm_prediction_interintra_mode_flag can indicate that intra-frame prediction is applied to the first partition of the current block and inter-frame prediction is applied to the second partition. That is, when the value of gpm_prediction_interintra_mode_flag is 0, gpm_prediction_interintra_mode_flag can indicate that the first partition of the current block is an inter-frame partition and the second partition of the current block is an intra-frame partition; and when the value of gpm_prediction_interintra_mode_flag is 1, gpm_prediction_interintra_mode_flag can indicate that the first partition of the current block is an intra-frame partition and the second partition of the current block is an inter-frame partition. Here, gpm_prediction_interintra_mode_flag can be referred to as the inter-frame-intra prediction GPM mode flag.
[0171] Furthermore, for the partitions of the current block, the indexes of motion information candidates for inter-frame prediction and / or the indexes of IPM candidates can be indicated using signaling. For example, when the first partition of the current block is an inter-frame partition and the second partition is an intra-frame partition, the index of the motion information candidate for the first partition, gpm_inter_idx0, can be indicated using signaling, and the index of the IPM candidate for the second partition, gpm_intra_idx1, can be indicated using signaling. Similarly, when the first partition of the current block is an intra-frame partition and the second partition is an inter-frame partition, the index of the IPM candidate for the first partition, gpm_intra_idx0, can be indicated using signaling, and the index of the motion information candidate for the second partition, gpm_inter_idx1, can be indicated using signaling.
[0172] Alternatively, for example, signaling for GPM with inter-frame and intra-frame prediction can be performed as follows.
[0173] [Table 6]
[0174] Referring to Table 6, the GPM partition index and inter-frame / intra-frame prediction GPM indicator for the current block can be signaled.
[0175] For example, `gpm_partition_interintra_idc` can indicate whether the GPM for the first partition of the current block is an inter-frame partition and the second partition of the current block is an intra-frame partition, the GPM for the first partition of the current block is an intra-frame partition and the second partition of the current block is an inter-frame partition, or whether a GPM with both inter-frame and intra-frame predictions is not applied to the current block. `gpm_partition_interintra_idc` can be binarized based on truncated Ricean (or truncated unary). Here, `gpm_partition_interintra_idc` can represent an inter-frame-intra-frame prediction GPM indicator.
[0176] For example, the inter-frame-intra prediction GPM mode and binarization indicated by gpm_partition_interintra_idc can be as follows.
[0177] [Table 7]
[0178] Referring to Table 7, when the value of gpm_partition_interintra_idc is 0, gpm_partition_interintra_idc can indicate that the inter-frame-intra prediction GPM is not applied. When the value of gpm_partition_interintra_idc is 1, it can indicate that the GPM for the first partition being an inter-frame partition and the second partition of the current block being an intra-frame partition is applied. And when the value of gpm_partition_interintra_idc is 2, it can indicate that the GPM for the first partition being an intra-frame partition and the second partition of the current block being an inter-frame partition is applied.
[0179] Furthermore, for the partitions of the current block, the indexes of motion information candidates for inter-frame prediction and / or the indexes of IPM candidates can be indicated using signaling. For example, when the first partition of the current block is an inter-frame partition and the second partition is an intra-frame partition, the index of the motion information candidate for the first partition, gpm_inter_idx0, can be indicated using signaling, and the index of the IPM candidate for the second partition, gpm_intra_idx1, can be indicated using signaling. Similarly, when the first partition of the current block is an intra-frame partition and the second partition is an inter-frame partition, the index of the IPM candidate for the first partition, gpm_intra_idx0, can be indicated using signaling, and the index of the motion information candidate for the second partition, gpm_inter_idx1, can be indicated using signaling.
[0180] Furthermore, for example, the partitions applying inter-frame prediction and the partitions applying intra-frame prediction can be determined based on the partition type (i.e., partition shape) of the current block. That is, for example, the partitions applying inter-frame prediction and the partitions applying intra-frame prediction can be determined based on the partition shape of the current block. For example, when one of the partitions of the current block is partitioned as a partition that is not adjacent to neighboring samples, the partition that is not adjacent to neighboring samples can be determined as the partition applying inter-frame prediction, and the other partition can be determined as the partition applying intra-frame prediction. In this case, the syntax indicating the partitions applying inter-frame prediction and the partitions applying intra-frame prediction can be omitted (e.g., the gpm_prediction_interintra_mode_flag mentioned above). Neighboring samples can include the upper neighboring sample and the left neighboring sample. Partitions that are not adjacent to neighboring samples may have low dependence on neighboring blocks and are therefore very likely not to apply intra-frame prediction. Therefore, determining the partitions applying inter-frame prediction and the partitions applying intra-frame prediction based on the partition shape can improve coding efficiency and reduce the number of bits in the inter-frame-intra-frame prediction GPM.
[0181] Furthermore, as described above, when the (A)MVP mode is applied to the current block, the decoding device can construct an (A)MVP candidate list, derive the motion vector of the MVP candidate selected from the MVP (motion vector predictor) candidates included in the (A)MVP candidate list as the MVP of the current block, derive the MVD of the current block based on the information about MVD notified by the signal, and derive the motion information of the current block based on the MVP and MVD.
[0182] Furthermore, according to this document, MVD prediction can be applied. Based on MVD prediction, the syntax indicating the MVD prediction index and the syntax indicating the MVD magnitude can be parsed. MVD candidates can be derived by combining possible symbols and MVD magnitudes derived based on the syntax. The MVs combined with the derived MVD candidates can be reordered using template matching (TM) cost. The candidate indicated by the MVD prediction index among the reordered candidates can be derived as the MVD of the current block.
[0183] Specifically, based on MVD prediction, the decoding device can derive MVD as follows.
[0184] For example, 1) the decoding device can parse the amplitude of the MVD component. That is, the MVD amplitude syntax, which indicates the size of the MVD component, can be signaled. Then, 2) the decoding device can parse the context-coded MVD prediction index. The MVD prediction index can indicate one of the MVD candidates. 3) the decoding device can generate MVD candidates, which are combinations of available symbols and available amplitudes derived based on the MVD amplitude syntax, and can construct MV candidates by adding them to the MV predictor (MVP) of the current block. 4) the decoding device can derive the MVD prediction cost of each MV candidate and can sort the MV candidates based on their MVD prediction costs. Here, the MVD prediction cost of an MV candidate can be the template matching (TM) cost of the MV candidate. 5) the decoding device can select the true MVD from the sorted MV candidates based on the MVD prediction index signaled.
[0185] Furthermore, for example, bilateral filters can be used to generate reference templates for deriving the MVD prediction cost of MV candidates.
[0186] Figure 9 A reference template for deriving the TM cost, which is the MVD prediction cost of the MV candidate, is illustrated.
[0187] like Figure 9As illustrated, the TM cost of an MV candidate can be derived as the sum of absolute differences (SAD) between the current template, which includes the neighboring samples of the current block, and the reference template, which includes the neighboring reference samples of the reference block indicated by the MV candidate. For example, the TM cost of an MV candidate can be derived based on the following formula.
[0188] [Equation 2]
[0189] Here, i and j represent the positions (i, j) of the samples within the template, and Cost distortion Indicates cost, Temp ref This represents the sample value of the reference template of the reference block indicated by the MV candidate, and Temp cur This represents the sample value of the current template in the current block. The difference between corresponding samples between the reference template and the current template can be accumulated, and the accumulated difference can be used as a cost function to rank the MV candidates of the current block.
[0190] Furthermore, for example, MVD prediction can be applied not only to AMVP mode, but also to affine AMVP mode, MMVD mode, and affine MMVD mode. Additionally, when orbital motion compensation is available, orbital offset can be considered to prune MV candidates.
[0191] For example, when MVD prediction is applied to affine AMVP or affine MMVD patterns, sub-block-based templates can be used. For instance, the template matching cost of each sub-block of the current block can be accumulated to derive the final cost of the MV candidates for the current block.
[0192] Figure 10 A reference template for deriving the TM cost is shown, which is the MVD prediction cost of MV candidates in the affine AMVP pattern or the affine MMVD pattern.
[0193] like Figure 10 As illustrated, in the affine AMVP mode or the affine MMVD mode, the TM cost of each sub-block of the current block can be derived, and the TM costs of the corresponding sub-blocks can be accumulated to derive the final cost of the MV candidate of the current block.
[0194] Furthermore, for example, when encoding the MVD magnitude syntax, the first six valid suffixes of the MVD magnitude syntax (bin) can be context-encoded. Valid suffixes (bin) can include symbol bins.
[0195] Furthermore, for example, the number of valid suffix bins in the context-encoded MVD magnitude syntax can be derived based on the block size. For example, for blocks with a width and height greater than N, up to 6 valid suffix bins, including the symbol bin, can be context-encoded, and for blocks with a width or height equal to or less than N, up to 2 valid suffix bins can be encoded. For example, N can be 4. Alternatively, for example, for blocks with a width and height greater than N, up to 6 valid suffix bins, including the symbol bin, can be context-encoded, and for blocks with a width or height equal to or less than N, up to 4 valid suffix bins can be encoded. For example, N can be 4.
[0196] Furthermore, for example, the number of combinations of signs and amplitudes in MMVD can be 16.
[0197] Figure 11 Examples of MMVD candidates are shown, which are combinations of available signs and magnitudes.
[0198] like Figure 11 As illustrated, there can be 16 MMVD candidates. Furthermore, for example, when MVD prediction is applied to 16 MMVDs, a TM cost can be derived for each MMVD candidate, and the MMVD candidates can be sorted according to the derived TM costs, such that only the first 8 candidates can be derived as MMVD candidates for the current block. The TM cost for each MMVD candidate can be derived as the sum of absolute differences (SAD) between the reference template of the reference block indicated by the MV candidates derived from the current block's MVP and MMVD candidates and the template of the current block.
[0199] Furthermore, this document proposes Reconstruction-Reordering IBC (RR-IBC) as a prediction method. For example, RR-IBC can be applied according to this document. RR-IBC can be considered as a prediction technique or prediction type. RR-IBC patterns can be applied to IBC-coded blocks. However, this is merely an example, and RR-IBC patterns can also be applied to inter-frame prediction blocks, in which case RR-IBC can be referred to as Reconstruction-Reordering (Inter-Frame) Prediction. In this case, the block vector of RR-IBC can be replaced with motion vectors.
[0200] On the encoding device side, for example, the original block of the current block can be flipped before motion information search and residual calculation, and the predicted block can be derived without being flipped. Furthermore, on the decoding device side, for example, when applying RR-IBC mode, the block vector (BV) for the current block can be derived, the reconstructed block (i.e., the reference block) indicated by the BV can be flipped according to the flip type, and the predicted block of the current block can be derived based on the flipped reconstructed block.
[0201] Figure 12 An example is shown of the flip type of the RR-IBC mode.
[0202] For example, refer to Figure 12 The flip type can include vertical flip and horizontal flip. Alternatively, for example, see reference. Figure 12 The flip type can include vertical flip, horizontal flip, vertical-horizontal flip, 90-degree clockwise rotation and / or 90-degree counterclockwise rotation.
[0203] Figure 12 (a) can represent a reference block. Figure 12 (b) can represent a reference block that has been horizontally flipped (flipped along the horizontal axis). Figure 12 (c) can represent a reference block that has been vertically flipped (flipped along the vertical axis), and Figure 12 The (d) can represent a reference block that has been flipped vertically and horizontally (flipped along two axes). Furthermore, Figure 12 (e) can represent a reference block rotated 90 degrees clockwise, and Figure 12 (f) can represent a reference block rotated 90 degrees counterclockwise.
[0204] Furthermore, for example, a 90-degree clockwise rotation and / or a 90-degree counterclockwise rotation can be applied only if the current block is a square block. For example, when the current block is a square block, the flip type of the current block can be derived as one of vertical flip, horizontal flip, vertical-horizontal flip, 90-degree clockwise rotation, and 90-degree counterclockwise rotation, and when the current block is not a square block, the flip type of the current block can be derived as one of vertical flip, horizontal flip, and vertical-horizontal flip.
[0205] Alternatively, in addition to or instead of 90-degree clockwise and / or 90-degree counterclockwise rotation, 45-degree clockwise and / or 45-degree counterclockwise rotation may be included. Furthermore, when applying a 45-degree clockwise / counterclockwise rotation, the block vector (or motion vector) can be modified based on the rotation direction and angle. When applying a 45-degree clockwise / counterclockwise rotation, a reference block larger than the current block can be derived, and reference samples based on the rotation of the current block can be derived within the larger reference block. Predicted samples of the current block can be derived based on the reference samples.
[0206] Furthermore, for example, signals can be used to inform the syntax of the RR-IBC pattern. That is, for example, signals can be used to inform the prediction-related information for the RR-IBC pattern.
[0207] First, it can be determined whether the current block applies IBC AMVP mode or IBC merge mode. For example, it can be determined based on the current block's IBC prediction flag and / or merge flag.
[0208] For example, a signal can be used to indicate whether the IBC prediction flag is applied to the current block. For instance, when the IBC prediction flag is 1, it indicates that the IBC prediction flag is applied to the current block, and when the IBC prediction flag is 0, it indicates that the IBC prediction flag is not applied to the current block.
[0209] Furthermore, when the IBC mode is applied to the current block (i.e., when the IBC prediction flag indicates that the IBC mode is applied to the current block), a signal can be used to notify the merge flag indicating whether the merge mode is applied to the current block. For example, when the merge flag indicates that the merge mode is applied to the current block, the prediction mode of the current block can be deduced as the IBC merge mode, and when the merge flag indicates that the merge mode is not applied to the current block, the prediction mode of the current block can be deduced as the IBC AMVP mode. That is, for example, when the merge flag indicates that the merge mode is applied to the current block, the IBC merge mode can be applied to the current block, and when the merge flag indicates that the merge mode is not applied to the current block, the IBC AMVP mode can be applied to the current block.
[0210] For example, when an IBC merge pattern is applied to the current block, a flip (and / or rotation) presence flag indicating whether the RR-IBC pattern is applied to the current block can be signaled. And when the flip (and / or rotation) presence flag indicates that the RR-IBC pattern is applied to the current block, a flip type flag / index does not need to be signaled. The flip (and / or rotation) presence flag can also be called the RR-IBC pattern flag. When an IBC merge pattern is applied to the current block, prediction-related information for the RR-IBC pattern can include the flip presence flag. For example, the flip type of a neighboring block indicated by the merge index can be inherited as the flip type of the current block. That is, for example, when an IBC merge pattern is applied to the current block, the merge index of the current block can be signaled, and a merge candidate list can be constructed based on the neighboring blocks of the current block. For example, merge candidates including the flip types and BVs of the neighboring blocks of the current block can be derived to construct a merge candidate list. Furthermore, the merge index can indicate one of the merge candidates in the merge candidate list. For example, the flip type and BV of the current block can be derived based on the merge candidate indicated by the merge index.
[0211] Alternatively, for example, when the IBC AMVP mode is applied to the current block, a signal can be used to indicate whether the RR-IBC mode is applied to the current block's flip (and / or rotation) presence flag, and when the flip presence flag indicates that the RR-IBC mode is applied to the current block, a signal can be used to indicate the flip type flag / index indicating the flip type of the current block. The flip (and / or rotation) presence flag can also be called the RR-IBC mode flag. The flip type flag / index can indicate the flip type of the current block. When the IBC AMVP mode is applied to the current block, prediction-related information for the RR-IBC mode can include the flip presence flag and / or the flip type flag / index.
[0212] As an example, a type index can indicate one of a whole set of flip type candidates, including flip and rotation. For example, the flip type indicated by the type index could be as follows.
[0213] [Table 8]
[0214] For example, referring to Table 8, the type index can indicate a horizontal flip, a vertical flip, a vertical-horizontal flip, a 90-degree clockwise rotation, and / or a 90-degree counterclockwise rotation. The flip type indicated by the type index can be deduced as the flip type of the current block.
[0215] Alternatively, as another example, a combination of flipping and / or rotating can be configured, and one of the combinations can be indicated by a type index. For example, the flipping type indicated by the type index can be as follows.
[0216] [Table 9]
[0217] For example, referring to Table 9, the type index can indicate a horizontal flip, a combination of a horizontal flip and a 90-degree clockwise rotation, a combination of a horizontal flip and a 90-degree counterclockwise rotation, a vertical flip, a combination of a vertical flip and a 90-degree clockwise rotation, a combination of a vertical flip and a 90-degree counterclockwise rotation, a vertical-horizontal flip, and / or a vertical-horizontal flip and a 90-degree clockwise rotation. The flip type indicated by the type index can be deduced as the flip type of the current block.
[0218] Alternatively, as another example, multiple indices can be used to signal the flip / rotation type. These multiple indices can include a flip type index and a rotation type index. For example, the flip type indicated by the flip type index and the rotation type indicated by the rotation type index could be as follows.
[0219] [Table 10]
[0220] [Table 11]
[0221] For example, referring to Table 10, the flip type index can indicate no flip, horizontal flip, horizontal flip, and / or vertical-horizontal flip. Furthermore, for example, the rotation type index can indicate a 90-degree clockwise rotation, a 90-degree counter-clockwise rotation, a 45-degree clockwise rotation, and / or a 45-degree counter-clockwise rotation. The flip type indicated by the flip type index and the rotation type index can be deduced as the flip type of the current block.
[0222] Furthermore, for example, when the IBC AMVP mode is applied to the current block, the prediction-related information for the RR-IBC mode can include the block vector candidate index and / or block vector difference (BVD) related information for the current block. For example, a list of block vector candidates can be constructed based on the neighboring blocks of the current block. For example, block vector candidates including the block vectors of the neighboring blocks of the current block can be derived to construct the block vector candidate list. Furthermore, the block vector candidate index can indicate one of the block vector candidates in the block vector candidate list. For example, the block vector predictor for the current block can be derived based on the block vector candidate indicated by the block vector candidate index. The BVD of the current block can be derived based on the BVD related information, and the block vector of the current block can be derived based on the block vector predictor and the BVD. Furthermore, for example, the MVD prediction described above can be applied to the BVD.
[0223] Furthermore, when applying the RR-IBC mode, by taking into account horizontal or vertical symmetry, the current block and the reference block can be aligned horizontally or vertically depending on the flip type.
[0224] Figure 13 An implementation of aligning the current block and the reference block based on the flip type of the current block is illustrated.
[0225] refer to Figure 13 To better utilize symmetry properties, the block vector candidate can be refined by applying a flip-aware BV adjustment method. Figure 13 In the middle, (x nbr y nbr (x) can represent the coordinates of the center sample of an adjacent block. cur y cur () can represent the coordinates of the center sample of the current block, BV nbr It can represent the BV of adjacent blocks, and BV cur It can represent the BV of the current block.
[0226] For example, when encoding adjacent blocks using horizontal flipping, instead of directly inheriting BV from adjacent blocks, such as... Figure 13 As shown in (a), the motion shift can be achieved by relating it to BV. nbrThe horizontal component (denoted as BV) nbr h ) Add them together to calculate BV cur The horizontal component, that is, BV cur h For example, when encoding adjacent blocks using horizontal flipping, the BV of the current block can be calculated as shown in the following formula.
[0227] [Formula 3]
[0228] Furthermore, for example, when a horizontal flip is applied to the current block, the vertical component of the BV may not be signaled and may be inferred to be equal to 0.
[0229] Furthermore, for example, when encoding adjacent blocks using vertical flipping, instead of directly inheriting BV from adjacent blocks, such as... Figure 13 As shown in (b), the motion shift can be achieved by relating it to BV. nbr The vertical component (denoted as BV) nbr v ) Add them together to calculate BV cur The vertical component, that is, BV cur v For example, when encoding adjacent blocks using vertical flipping, the BV of the current block can be calculated as shown in the following formula.
[0230] [Formula 4]
[0231] Furthermore, for example, when a vertical flip is applied to the current block, the horizontal component of BV may not be signaled and may be inferred to be equal to 0.
[0232] Furthermore, for example, in RR-IBC mode, the BV of the current block can be modified based on the TM cost. For instance, the TM cost between the template of the current block and the template of the flipped reference block can be calculated, and the modified BV indicating the reference block with the minimum TM cost within the search region can be derived. The search region can be derived based on the position indicated by the initial BV of the current block. Additionally, the template of the current block and the template of the reference block can be derived based on the flip type of the current block. For example, when the flip type of the current block is horizontal flip, the template of the current block can include the upper adjacent sample (i.e., the upper template), and when the flip type of the current block is vertical flip, the template of the current block can include the left adjacent sample (i.e., the left template).
[0233] Figure 14 A video / image coding method according to an embodiment of the present disclosure is illustrated schematically. Figure 14 The method disclosed in the article can be derived from Figure 2 The encoding device disclosed in the document executes the code. Specifically, for example, Figure 14 S1400 to S1420 can be performed by the predictor 220 of the encoding device 200, and Figure 14 S1430 can be executed by the entropy encoder 240 of the encoding device 200. Figure 14 The methods disclosed herein may include the embodiments described above.
[0234] refer to Figure 14 The encoding device derives the prediction mode of the current block as the RR-IBC mode (S1400). The encoding device can derive the RR-IBC mode as the prediction mode applied to the current block among various prediction modes / types.
[0235] The encoding device derives a modified reference block based on the flip type of the current block (S1410). The encoding device can determine the flip type of the current block from among multiple flip types.
[0236] As an example, multiple flip types can include horizontal flip, vertical flip, vertical-horizontal flip, 90-degree clockwise rotation, and / or 90-degree counterclockwise rotation.
[0237] Alternatively, as another example, multiple flip types may include combinations of flip types and rotation types. For example, multiple flip types may include horizontal flip, a combination of horizontal flip and 90-degree clockwise rotation, a combination of horizontal flip and 90-degree counterclockwise rotation, vertical flip, a combination of vertical flip and 90-degree clockwise rotation, a combination of vertical flip and 90-degree counterclockwise rotation, vertical-horizontal flip, and / or a combination of vertical-horizontal flip and 90-degree clockwise rotation.
[0238] Alternatively, as another example, the multiple flip types may include a flip type and / or a rotation type. For example, the flip types included in the multiple flip types may include a horizontal flip, a vertical flip, and / or a vertical-horizontal flip, and the rotation types included in the multiple flip types may include a 90-degree clockwise rotation, a 90-degree counterclockwise rotation, a 45-degree clockwise rotation, and / or a 45-degree counterclockwise rotation.
[0239] The encoding device can search for blocks similar to the current block within the current image and can deduce a reference block whose difference from the current block is the smallest, equal to, or less than a certain criterion. The encoding device can deduce the block vector based on the positional difference between the reference block and the current block.
[0240] The encoding device can derive a modified reference block based on a reference block at a location indicated by a block vector and a flip type. For example, the modified reference block can be derived by flipping and / or rotating the reference block according to the flip type.
[0241] The encoding device derives the prediction samples for the current block based on the modified reference block (S1420). The prediction block for the current block can be derived based on the modified reference block. In this case, as described above, a prediction sample filtering process for all or some of the prediction samples of the current block can be further performed as needed.
[0242] The encoding device encodes image information including prediction-related information for the current block (S1430). The encoding device can generate prediction-related information based on the derived block vector and flip type. The prediction-related information may include the block vector for the modified reference block and flip type information indicating the flip type.
[0243] For example, prediction-related information could include an RR-IBC mode flag indicating whether the RR-IBC mode is applied to the current block. For instance, when the IBC merge mode or the IBC AMVP mode is applied to the current block, the RR-IBC mode flag indicating whether the RR-IBC mode is applied to the current block could be signaled.
[0244] Furthermore, prediction-related information may include, for example, IBC prediction flags and / or merge flags. For instance, the IBC prediction flag may indicate whether the IBC mode is applied to the current block, and the merge flag may indicate whether the merge mode is applied to the current block. For example, when the IBC prediction flag is 1, it indicates that the IBC mode is applied to the current block, and when the IBC prediction flag is 0, it indicates that the IBC mode is not applied to the current block.
[0245] Furthermore, for example, when the IBC mode is applied to the current block (i.e., when the IBC prediction flag indicates that the IBC mode is applied to the current block), the merge flag indicating whether the merge mode is applied to the current block can be signaled. For example, when the merge flag indicates that the merge mode is applied to the current block, the prediction mode of the current block can be deduced as the IBC merge mode, and when the merge flag indicates that the merge mode is not applied to the current block, the prediction mode of the current block can be deduced as the IBC AMVP mode. That is, for example, when the merge flag indicates that the merge mode is applied to the current block, the IBC merge mode can be applied to the current block, and when the merge flag indicates that the merge mode is not applied to the current block, the IBC AMVP mode can be applied to the current block.
[0246] In addition, for example, the prediction-related information may include the block vector for the modified reference block and flip type information indicating the flip type.
[0247] As an example, flip type information can include a type index. The flip type indicated by the type index can be deduced as the flip type of the current block. For example, the type index can indicate a horizontal flip, a vertical flip, a vertical-horizontal flip, a 90-degree clockwise rotation, and / or a 90-degree counterclockwise rotation.
[0248] Alternatively, as another example, the flip type information may include a type index, and the flip type of the current block may be inferred based on the combination of the flip type and rotation type indicated by the type index. For example, the type index may indicate a horizontal flip, a combination of a horizontal flip and a 90-degree clockwise rotation, a combination of a horizontal flip and a 90-degree counterclockwise rotation, a vertical flip, a combination of a vertical flip and a 90-degree clockwise rotation, a combination of a vertical flip and a 90-degree counterclockwise rotation, a vertical-horizontal flip, and / or a combination of a vertical-horizontal flip and a 90-degree clockwise rotation.
[0249] Alternatively, as another example, the flip type information may include a flip type index and a rotation type index, and the flip type of the current block may be inferred based on the flip type indicated by the flip type index and the rotation type indicated by the rotation type index. For example, the flip type index may indicate a horizontal flip, a vertical flip, or a vertical-horizontal flip, and the rotation type index may indicate a 90-degree clockwise rotation, a 90-degree counterclockwise rotation, a 45-degree clockwise rotation, or a 45-degree counterclockwise rotation.
[0250] Furthermore, for example, when the IBC AMVP mode is applied to the current block, the flip type information may include the block vector candidate index and / or block vector difference (BVD) related information of the current block. For example, a block vector candidate list can be constructed based on the neighboring blocks of the current block. For example, block vector candidates including the block vectors of the neighboring blocks of the current block can be derived to construct the block vector candidate list. Furthermore, the block vector candidate index may indicate one of the block vector candidates in the block vector candidate list. For example, the block vector predictor of the current block can be derived based on the block vector candidate indicated by the block vector candidate index. The BVD of the current block can be derived based on BVD related information, and the block vector of the current block can be derived based on the block vector predictor and BVD. Furthermore, for example, the MVD prediction described above can be applied to the BVD.
[0251] Alternatively, as another example, the flip type information may include a merge index for the current block. For example, when an IBC merge mode is applied to the current block, the flip type information may include a merge index for the current block. For example, a merge candidate list may be constructed based on the current block's neighboring blocks. For example, a merge candidate list including the flip types and block vectors of the current block's neighboring blocks may be derived to construct the merge candidate list. Furthermore, the merge index may indicate one of the merge candidates in the merge candidate list. For example, the block vector and flip type of the current block may be derived based on the merge candidate indicated by the merge index.
[0252] For example, prediction-related information could include the CU syntax for the current block. Image information could be referred to as video information.
[0253] Furthermore, according to embodiments of this disclosure, image information may include various types of information. For example, image information may include information disclosed in at least one of the tables above.
[0254] In addition, image information may include residual information. Residual information is information about the residual samples. Residual information may include information about the quantized transform coefficients for the residual samples.
[0255] Encoded image information can be output as a bitstream. The bitstream can be sent to a decoding device via a network or storage medium.
[0256] Furthermore, as mentioned above, the encoding device can generate a reconstructed image (including reconstructed samples and reconstructed blocks) based on reference samples and residual samples. This is to derive the same prediction result at the encoding device as the prediction result performed at the decoding device, thereby improving encoding efficiency. Therefore, the encoding device can store the reconstructed image (or reconstructed samples and reconstructed blocks) in memory and use it as a reference image for inter-frame prediction. As mentioned above, loop filtering processes, etc., can be further applied to the reconstructed image.
[0257] According to the above implementation, in IBC prediction using a reference block within the current image, IBC prediction of the current block can be performed by considering the flipping and / or rotation flipping type of objects within the reference block, thereby improving the prediction accuracy of the IBC mode and increasing coding efficiency.
[0258] Furthermore, in the process of deriving the block vector for IBC prediction, a template for cost derivation can be set by considering the shape of the current block, and thus, template matching costs can be derived using highly correlated neighboring samples, allowing for more accurate modification of the block vector, which can improve the prediction accuracy of IBC patterns and increase coding efficiency.
[0259] Figure 15 A video / image decoding method according to an embodiment of the present disclosure is illustrated schematically. Figure 15 The method disclosed in the article can be derived from Figure 3 The decoding device disclosed in the document performs the operation. Specifically, for example, Figure 15 S1500 to S1540 can be performed by the predictor 330 of the decoding device 300. Figure 15 The methods disclosed herein may include the embodiments described above.
[0260] refer to Figure 15 The decoding device derives the prediction mode of the current block into RR-IBC mode (S1500) based on prediction-related information.
[0261] The decoding device can deduce the prediction mode of the current block as RR-IBC mode based on prediction-related information.
[0262] For example, a decoding device can obtain image information, including prediction-related information, from a bitstream. As mentioned above, the image information can also include residual information.
[0263] For example, prediction-related information could include an RR-IBC mode flag indicating whether the RR-IBC mode is applied to the current block. For instance, when the IBC merge mode or the IBC AMVP mode is applied to the current block, the RR-IBC mode flag indicating whether the RR-IBC mode is applied to the current block could be signaled.
[0264] Furthermore, prediction-related information may include, for example, IBC prediction flags and / or merge flags. For instance, the IBC prediction flag may indicate whether the IBC mode is applied to the current block, and the merge flag may indicate whether the merge mode is applied to the current block. For example, when the IBC prediction flag is 1, it indicates that the IBC mode is applied to the current block, and when the IBC prediction flag is 0, it indicates that the IBC mode is not applied to the current block.
[0265] Furthermore, for example, when the IBC mode is applied to the current block (i.e., when the IBC prediction flag indicates that the IBC mode is applied to the current block), the merge flag indicating whether the merge mode is applied to the current block can be signaled. For example, when the merge flag indicates that the merge mode is applied to the current block, the prediction mode of the current block can be deduced as the IBC merge mode, and when the merge flag indicates that the merge mode is not applied to the current block, the prediction mode of the current block can be deduced as the IBC AMVP mode. That is, for example, when the merge flag indicates that the merge mode is applied to the current block, the IBC merge mode can be applied to the current block, and when the merge flag indicates that the merge mode is not applied to the current block, the IBC AMVP mode can be applied to the current block.
[0266] The decoding device derives the block vector (BV) and flip type of the current block based on the flip type information included in the prediction-related information (S1510).
[0267] As an example, flip type information can include a type index. The flip type indicated by the type index can be deduced as the flip type of the current block. For example, the type index can indicate a horizontal flip, a vertical flip, a vertical-horizontal flip, a 90-degree clockwise rotation, and / or a 90-degree counterclockwise rotation.
[0268] Alternatively, as another example, the flip type information may include a type index, and the flip type of the current block may be inferred based on the combination of the flip type and rotation type indicated by the type index. For example, the type index may indicate a horizontal flip, a combination of a horizontal flip and a 90-degree clockwise rotation, a combination of a horizontal flip and a 90-degree counterclockwise rotation, a vertical flip, a combination of a vertical flip and a 90-degree clockwise rotation, a combination of a vertical flip and a 90-degree counterclockwise rotation, a vertical-horizontal flip, and / or a combination of a vertical-horizontal flip and a 90-degree clockwise rotation.
[0269] Alternatively, as another example, the flip type information may include a flip type index and a rotation type index, and the flip type of the current block may be inferred based on the flip type indicated by the flip type index and the rotation type indicated by the rotation type index. For example, the flip type index may indicate a horizontal flip, a vertical flip, or a vertical-horizontal flip, and the rotation type index may indicate a 90-degree clockwise rotation, a 90-degree counterclockwise rotation, a 45-degree clockwise rotation, or a 45-degree counterclockwise rotation.
[0270] Furthermore, for example, when the IBC AMVP mode is applied to the current block, the flip type information may include the block vector candidate index and / or block vector difference (BVD) related information of the current block. For example, a block vector candidate list can be constructed based on the neighboring blocks of the current block. For example, block vector candidates including the block vectors of the neighboring blocks of the current block can be derived to construct the block vector candidate list. Furthermore, the block vector candidate index may indicate one of the block vector candidates in the block vector candidate list. For example, the block vector predictor of the current block can be derived based on the block vector candidate indicated by the block vector candidate index. The BVD of the current block can be derived based on BVD related information, and the block vector of the current block can be derived based on the block vector predictor and BVD. Furthermore, for example, the MVD prediction described above can be applied to the BVD.
[0271] Alternatively, as another example, the flip type information may include a merge index for the current block. For example, when an IBC merge mode is applied to the current block, the flip type information may include a merge index for the current block. For example, a merge candidate list may be constructed based on the current block's neighboring blocks. For example, a merge candidate list including the flip types and block vectors of the current block's neighboring blocks may be derived to construct the merge candidate list. Furthermore, the merge index may indicate one of the merge candidates in the merge candidate list. For example, the block vector and flip type of the current block may be derived based on the merge candidate indicated by the merge index.
[0272] Furthermore, for example, the block vector of the current block can be modified based on template matching. That is, the modified block vector can be derived based on the template of the current block.
[0273] For example, the decoding device can derive the search region based on the block vector, derive the template matching (TM) cost of the reference block within the search region based on the template of the current block, and derive a modified block vector indicating the reference block with the minimum TM cost among the reference blocks. Here, the TM cost of the reference block can be the sum of absolute differences (SAD) or the sum of absolute transformation differences (SATD) between the template of the current block and the template of the reference block. When deriving the modified block vector based on the template, the reference block of the current block can be derived based on the modified block vector. Furthermore, for example, when the flip type of the current block is horizontal flip, the template of the current block can be derived as the upper template including the upper adjacent sample, and when the flip type of the current block is vertical flip, the template of the current block can be derived as the left template including the left adjacent sample.
[0274] In addition, for example, when a 45-degree clockwise rotation or a 45-degree counterclockwise rotation is applied as the flip type of the current block, the block vector of the current block can be modified based on the rotation direction and angle of the 45-degree clockwise rotation or the 45-degree counterclockwise rotation.
[0275] The decoding device derives the reference block of the current block based on the block vector (S1520). The decoding device can derive the reconstructed block at the position indicated by the block vector as the reference block of the current block.
[0276] The decoding device derives a modified reference block based on the reference block and the flip type (S1530). The decoding device can derive a modified reference block based on the reference block and the flip type.
[0277] For example, a modified reference block can be derived by flipping and / or rotating a reference block according to a flip type. For example, a decoding device can derive a modified reference block by flipping and / or rotating a reference block according to a flip type.
[0278] The decoding device derives the prediction samples for the current block based on the modified reference block (S1540). The prediction block for the current block can be derived based on the modified reference block. In this case, as described above, a prediction sample filtering process for all or some of the prediction samples of the current block can be further performed as needed.
[0279] The decoding device can generate reconstructed samples based on the predicted samples of the current block. For example, the decoding device can generate reconstructed samples for the current block based on the residual samples and predicted samples for the current block. Residual samples for the current block can be generated based on the received residual information. Furthermore, as an example, the decoding device can generate a reconstructed image including the reconstructed samples. Thereafter, as described above, loop filtering processes, etc., can be further applied to the reconstructed image.
[0280] According to the above implementation, in IBC prediction using a reference block within the current image, IBC prediction of the current block can be performed by considering the flipping and / or rotation flipping type of objects within the reference block, thereby improving the prediction accuracy of the IBC mode and increasing coding efficiency.
[0281] Furthermore, in the process of deriving the block vector for IBC prediction, a template for cost derivation can be set by considering the shape of the current block, and thus, template matching costs can be derived using highly correlated neighboring samples, allowing for more accurate modification of the block vector, which can improve the prediction accuracy of IBC patterns and increase coding efficiency.
[0282] Although the method is described as a series of steps or blocks based on the flowchart in the above embodiments, the embodiments are not limited to the order of the steps, and a step may occur in a different order than another step described above or simultaneously with another step described above. Furthermore, those skilled in the art will understand that the steps shown in the flowchart are not exclusive, and other steps may be included or one or more steps in the flowchart may be deleted without affecting the scope of the embodiments of this disclosure.
[0283] The methods described above according to embodiments of this disclosure can be implemented in software, and the encoding and / or decoding devices according to this disclosure can be included in devices for performing image processing, such as televisions, computers, smartphones, set-top boxes, and display devices.
[0284] The embodiments described above can be implemented in the form of a recording medium including computer-executable (program) instructions, such as a program module executed by a computer. The module can be stored in memory and executed by a processor. The memory can be located inside or outside the processor and can be connected to the processor by various known means. The computer-readable medium can be any available medium accessible to a computer and can include both volatile and non-volatile media, as well as removable and non-removable media. Furthermore, the computer-readable medium can include both computer storage media and communication media. Computer storage media can include both volatile and non-volatile media, as well as removable and non-removable media, implemented using any method or technology for storing information (such as computer-readable instructions, data structures, program modules, or other data). Communication media typically include computer-readable instructions, data structures, program modules, other data in modulated data signals (such as carrier waves), or other transmission mechanisms, and include any information delivery medium.
[0285] Furthermore, the embodiments described above in this disclosure can be implemented as a computer program (or computer program product) including computer-executable instructions. The computer program may include programmable machine instructions processed by a processor and may be implemented in a high-level programming language, an object-oriented programming language, assembly language, or machine language. Additionally, the computer program may be recorded on a tangible computer-readable recording medium (e.g., memory, hard disk, magnetic / optical media, or solid-state drive (SSD)).
[0286] Therefore, embodiments of this disclosure can be implemented by executing the computer program described above using a computing device. The computing device may include at least some of a processor, memory, storage devices, high-speed interfaces connected to the memory and high-speed expansion ports, and low-speed interfaces connected to low-speed buses and storage devices. These components may be interconnected via various buses and may be mounted on a common motherboard or otherwise suitable.
[0287] A processor can process instructions within a computing device. These instructions may include those stored in memory or storage devices to display graphical information on an external input / output device (such as a display) connected to a high-speed interface, providing a graphical user interface (GUI). In another embodiment, multiple processors and / or multiple buses may be appropriately utilized along with multiple memories and memory types. Furthermore, the processor may be implemented as a chipset comprising multiple independent analog and / or digital processors.
[0288] Memory stores information within a computing device. For example, memory may include volatile memory cells or a collection of volatile memory cells. In another example, memory may include non-volatile memory cells or a collection of non-volatile memory cells. Memory may also be another form of computer-readable medium, such as a magnetic disk or optical disk.
[0289] Storage devices can provide large-capacity storage space for computing devices. Storage devices can be computer-readable media or components that include computer-readable media. For example, storage devices can include devices or other components within a storage area network (SAN) and can be floppy disk devices, hard disk devices, optical disk devices, magnetic tape devices, flash memory, other similar semiconductor storage devices, or device arrays.
[0290] The network can be implemented as a wired network, such as a local area network (LAN), a wide area network (WAN), or a value-added network (VAN), or various types of wireless networks, such as mobile radio communication networks or satellite communication networks.
[0291] Although this disclosure has been described with reference to embodiments illustrated in the accompanying drawings, these embodiments are merely exemplary. Those skilled in the art will understand that various modifications and variations are possible. That is, the scope of this disclosure is not limited to the described embodiments, and various modifications and alterations made by those skilled in the art based on the basic concepts defined in the appended claims also fall within the scope of the claims. Therefore, the true technical scope of this disclosure should be determined by the technical spirit of the appended claims.
Claims
1. An image decoding method performed by a decoding device, the image decoding method comprising the following steps: Based on prediction-related information, the prediction mode of the current block is derived as the Reconstruction-Reordering Intra-Block Copy (RR-IBC) mode. Based on the flip type information included in the prediction-related information, the block vector BV and flip type of the current block are derived; The reference block of the current block is derived based on the block vector; The modified reference block is derived based on the reference block and the flip type; as well as The predicted sample of the current block is derived based on the modified reference block.
2. The image decoding method of claim 1, wherein, The flip type information includes a type index, and The flip type indicated by the type index is deduced as the flip type of the current block.
3. The image decoding method according to claim 2, wherein, The type index indicates horizontal flip, vertical flip, vertical-horizontal flip, 90-degree clockwise rotation, or 90-degree counterclockwise rotation.
4. The image decoding method according to claim 1, wherein, The flip type information includes a type index, and The flip type of the current block is derived based on a combination of flip type and rotation type indicated by the type index.
5. The image decoding method according to claim 4, wherein, The type index indicates a horizontal flip, a combination of a horizontal flip and a 90-degree clockwise rotation, a combination of a horizontal flip and a 90-degree counterclockwise rotation, a vertical flip, a combination of a vertical flip and a 90-degree clockwise rotation, a combination of a vertical flip and a 90-degree counterclockwise rotation, a vertical-horizontal flip, or a combination of a vertical-horizontal flip and a 90-degree clockwise rotation.
6. The image decoding method according to claim 1, wherein, The flip type information includes a flip type index and a rotation type index, and The flip type of the current block is derived based on the flip type indicated by the flip type index and the rotation type indicated by the rotation type index.
7. The image decoding method according to claim 6, wherein, The flip type index indicates a horizontal flip, a vertical flip, or a vertical-horizontal flip.
8. The image decoding method according to claim 6, wherein, The rotation type index indicates a 90-degree clockwise rotation, a 90-degree counterclockwise rotation, a 45-degree clockwise rotation, or a 45-degree counterclockwise rotation.
9. The image decoding method according to claim 1, wherein, The prediction-related information includes an RR-IBC mode flag indicating whether the RR-IBC mode is applied to the current block.
10. The image decoding method according to claim 1, wherein, The steps of deriving the block vector and the flip type of the current block include: The search region is derived based on the block vector; Based on the template of the current block, derive the template matching TM cost of the reference blocks within the search area; and The modified block vector of the reference block with the minimum TM cost is derived, and The reference block of the current block is derived based on the modified block vector.
11. The image decoding method according to claim 10, wherein, The template of the current block is derived based on the flip type of the current block.
12. The image decoding method according to claim 11, wherein, When the flip type of the current block is horizontal flip, the template of the current block is derived to include the upper template of the upper adjacent sample, and Wherein, when the flip type of the current block is vertical flip, the template of the current block is derived as a left template including the left adjacent sample.
13. An image encoding method performed by an encoding device, the image encoding method comprising the following steps: The prediction mode of the current block is derived as the Reconstruction-Reordering Intra-Block Copy (RR-IBC) mode; The modified reference block is derived based on the flip type of the current block; The predicted sample of the current block is derived based on the modified reference block; as well as The image information, including prediction-related information for the current block, is encoded. The prediction-related information includes the block vector BV for the modified reference block and flip type information indicating the flip type.
14. The image encoding method according to claim 13, wherein, The flip type information includes a type index, and The type index indicates the flip type of the current block.
15. The image encoding method according to claim 14, wherein, The type index indicates horizontal flip, vertical flip, vertical-horizontal flip, 90-degree clockwise rotation, or 90-degree counterclockwise rotation.
16. The image encoding method according to claim 13, wherein, The flip type information includes a type index, and The flip type of the current block is derived based on a combination of flip type and rotation type indicated by the type index.
17. The image encoding method according to claim 16, wherein, The type index indicates a horizontal flip, a combination of a horizontal flip and a 90-degree clockwise rotation, a combination of a horizontal flip and a 90-degree counterclockwise rotation, a vertical flip, a combination of a vertical flip and a 90-degree clockwise rotation, a combination of a vertical flip and a 90-degree counterclockwise rotation, a vertical-horizontal flip, or a combination of a vertical-horizontal flip and a 90-degree clockwise rotation.
18. The image encoding method according to claim 13, wherein, The flip type information includes a flip type index and a rotation type index, and The flip type of the current block is derived based on the flip type indicated by the flip type index and the rotation type indicated by the rotation type index.
19. The image encoding method according to claim 6, wherein, The flip type index indicates a horizontal flip, a vertical flip, or a vertical-horizontal flip.
20. A method for transmitting image data, the method comprising the following steps: Obtain a bitstream generated by an image coding method, wherein the image coding method includes the following steps: deriving the prediction mode of the current block as a Reconstruction-Reordering Intra-Block Copy (RR-IBC) mode; deriving a modified reference block based on the flip type of the current block; deriving prediction samples of the current block based on the modified reference block; and encoding image information including prediction-related information of the current block; and Send the image data including the bitstream. The prediction-related information includes the block vector BV for the modified reference block and flip type information indicating the flip type.