METHODS AND TOOLS FOR SUBIMAGE-BASED IMAGE ENCODING / DECODING, AND METHODS FOR TRANSMITTING BIT STREAM

IDP000106420BActive Publication Date: 2026-07-13GUANGDONG OPPO MOBILE TELECOMMUNICATIONS CORP LTD

Patent Information

Authority / Receiving Office
ID · ID
Patent Type
Patents
Current Assignee / Owner
GUANGDONG OPPO MOBILE TELECOMMUNICATIONS CORP LTD
Filing Date
2020-09-23
Publication Date
2026-07-13

AI Technical Summary

Technical Problem

The increasing demand for high resolution and high quality images has led to a rise in transmission and storage costs due to the increased amount of information or bits required, necessitating highly efficient image compression technologies.

Method used

The implementation of image encoding/decoding methods that utilize subimages and techniques like Bi-Directional Optical Flow (BDOF) or Prediction Refinement with Optical Flow (PROF) to improve encoding/decoding efficiency, along with methods for transmitting bit streams generated by these processes.

Benefits of technology

These methods enhance encoding/decoding efficiency, allowing for more effective transmission and storage of high resolution and high quality images while reducing costs.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0_ABST
    Figure 0_ABST
Patent Text Reader

Abstract

The present invention provides an image encoding / decoding method and apparatus. An image decoding method according to the present disclosure is performed by an image decoding apparatus. The image decoding method includes determining whether bi-directional optical flow (BDOF) or predictive smoothing with optical flow (PROF) is applied to the current block, based on the BDOF or PROF applied to the current block, sampling a prediction of the current block from a reference image of the current block based on motion information of the current block, and deriving a refined prediction sample for the current block, by applying the BDOF or PROF to the current block based on the retrieved prediction sample.
Need to check novelty before this filing date? Find Prior Art

Description

Description METHODS AND TOOLS FOR SUBIMAGE-BASED IMAGE ENCODING / DECODING, AND METHODS FOR TRANSMITTING BIT STREAM Invention Engineering Field The present disclosure relates to image encoding / decoding methods and apparatus and methods of transmitting bit streams, and, more particularly, to image encoding / decoding methods and apparatus for performing subimage encoding / decoding and methods of transmitting bit streams generated by the image encoding methods / apparatus of the present disclosure. Background of the Invention Recently, the demand for high-resolution and high-quality images, such as high-definition (HD) and ultra-high-definition (UHD) images, has increased in various fields. As the resolution and quality of image data increase, the amount of information, or bits, transmitted increases relative to existing image data. This increase in the amount of information or bits transmitted leads to increased transmission and storage costs. Thus, there is a need for highly efficient image compression technology to effectively transmit, store and reproduce information about high-resolution and high-quality images. Brief Description of the Invention Technical Issues The object of the present disclosure is to provide methods and apparatus for image encoding / decoding with improved encoding / decoding efficiency. Another object of the present disclosure is to provide image encoding / decoding methods and apparatus for encoding / decoding images based on subimages. Another object of the present disclosure is to provide image encoding / decoding methods and apparatus for performing BDOF or PROF based on determining whether the current sub-image is treated as an image. Another object of the present disclosure is to provide a method of transmitting a bit stream generated by an image encoding method or apparatus according to the present disclosure. Another object of the present disclosure is to provide a recording medium that stores a bit stream generated by an image encoding method or apparatus according to the present disclosure. Another object of the present disclosure is to provide a recording medium that stores a bit stream that is received, decoded and used to reconstruct an image by an image decoding apparatus according to the present disclosure. The technical issues addressed by this disclosure are not limited to the above technical issues and other technical issues not described herein will become apparent to one skilled in the art from the following description. Technical Solutions A method of decoding an image according to an aspect of the present disclosure may include determining whether bidirectional optical flow (BDOF) or prediction refinement with optical flow (PROF) is applied to the current block, based on the BDOF or PROF applied to the current block, sampling predictions of the current block from a reference image of the current block based on motion information of the current block, and deriving refined prediction samples for the current block, by applying BDOF or PROF to the current block based on the retrieved prediction samples. In the image decoding method of the present disclosure, predictive sampling of the current block may be performed based on whether the current sub-image includes the current block treated as an image. In the image decoding method of the present disclosure, whether the current sub-image is treated as an image can be determined based on the marker information signaled through the bit stream. In the image decoding method of the present disclosure, marker information may be signaled via a sequence parameter set (SPS). In the image decoding method of the present disclosure, the prediction sampling of the current block may be performed based on the position of the prediction sample to be taken, and wherein the position of the prediction sample may be cropped within a predetermined range. In the image decoding method of the present disclosure, based on the current sub-image being treated as an image, a predetermined range can be specified by the boundary position of the current sub-image. In the image decoding method of the present disclosure, the position of the prediction sample to be retrieved may include an x ​​coordinate and a y coordinate, the x coordinate may be truncated within the range of the left boundary position and the right boundary position of the current sub-image, and the y coordinate may be truncated within the range of the upper boundary position and the lower boundary position of the current sub-image. In the image decoding method of the present disclosure, the left boundary position of the current sub-block can be derived as the product of position information of a predetermined unit specifying the left position of the current sub-image and the width of the predetermined unit, the right boundary position of the current sub-image can be derived by performing a -1 operation on the product of position information of a predetermined unit specifying the right position of the current sub-image and the width of the predetermined unit, the upper boundary position of the current sub-image can be derived as the product of position information of a predetermined unit specifying the top position of the current sub-image and the height of the predetermined unit,and the lower boundary position of the current sub-image can be derived by performing a -1 operation on the position information product of the specified unit that specifies the lower position of the current sub-image and the height of the specified unit., In the image decoding method of the present disclosure, the predetermined units may be grids or CTUs. In the image decoding method of the present disclosure, based on the current sub-image not being treated as an image, the predetermined range may be a range of the current image that includes the current block. An image decoding apparatus according to another aspect of the disclosure of the present invention may include a memory and at least one processor. The at least one processor may determine whether bi-directional optical flow (BDOF) or predictive smoothing with optical flow (PROF) is applied to the current block, based on the BDOF or PROF applied to the current block, derive prediction samples of the current block from a reference image of the current block based on motion information of the current block, and derive smoothed prediction samples for the current block, by applying BDOF or PROF to the current block based on the derived prediction samples. In the image decoding apparatus of the present disclosure, the at least one processor may sample a prediction from the current block based on whether the current sub-image includes the current block treated as an image. Image encoding methods according to other aspects of the present disclosure may include determining whether bidirectional optical flow (BDOF) or predictive smoothing with optical flow (PROF) is applied to the current block, based on the BDOF or PROF applied to the current block, sampling predictions of the current block from a reference image of the current block based on motion information of the current block, and deriving smoothed prediction samples for the current block, by applying BDOF or PROF to the current block based on the retrieved prediction samples. In the image encoding method of the present disclosure, predictive sampling of the current block may be performed based on whether the current sub-image includes the current block treated as an image. Transmission methods according to other aspects of the present disclosure may transmit a bit stream generated by the image encoding method and / or image encoding apparatus of the present disclosure to an image decoding apparatus. In addition, a computer-readable recording medium according to another aspect of the present disclosure may store a bit stream generated by an image encoding apparatus or image encoding method of the present disclosure. The features briefly summarized above relating to this disclosure are only exemplary aspects of the complete description of the disclosure of the invention below, and do not limit the scope of this disclosure. Superior Effect of Invention According to the present disclosure, it is possible to provide image encoding / decoding methods and apparatus with improved encoding / decoding efficiency. Also, according to the present disclosure, it is possible to provide image encoding / decoding methods and apparatus for encoding / decoding images based on subimages. Also, according to the present disclosure, it is possible to provide image encoding / decoding methods and apparatus for performing BDOF or PROF based on determining whether the current subimage is treated as an image. Also, according to the present disclosure, it is possible to provide a method of transmitting a bit stream generated by an image encoding method or apparatus according to the present disclosure. In addition, according to the present disclosure, it is possible to provide a recording medium that stores a bit stream generated by an image encoding apparatus or method according to the present disclosure. In addition, according to the present disclosure, it is possible to provide a recording medium that stores a bit stream that is received, decoded, and used to reconstruct an image by an image decoding apparatus according to the present disclosure. It will be understood by those skilled in the art that the effects that may be achieved by this disclosure are not limited to those specifically described above and other advantages of this disclosure will be more clearly understood from the full description. Short Description of Image 1 is a view schematically illustrating a video coding system, to which an embodiment of the present disclosure can be applied. Figure 2 is a view schematically illustrating an image encoding apparatus, to which embodiments of the present disclosure may be applied. Figure 3 is a view schematically illustrating an image decoding apparatus, to which an embodiment of the present disclosure may be applied. Figure 4 is a flowchart illustrating intermediate prediction based on video / image encoding methods. Figure 5 is a view illustrating the configuration of the intermediate prediction unit (180) according to the present disclosure. Figure 6 is a flowchart illustrating intermediate prediction based on video / image decoding methods. Figure 7 is a view illustrating the configuration of the intermediate prediction unit (260) according to the present disclosure. Figure 8 is a view illustrating the motion that can be expressed in affine mode. Figure 9 is a view illustrating the parameter model of the affine mode. Figure 10 is a view illustrating the method for generating an affine merge candidate list. Figure 11 is a view illustrating the CPMV derived from adjacent blocks. Figure 12 is a view illustrating the adjacent blocks for deriving inherited affine join candidates. Figure 13 is a view illustrating the adjacent blocks for deriving the constructed affine union candidates. Figure 14 is a view illustrating the method for generating a list of affine MVP candidates. Figure 15 is a view illustrating the adjacent blocks of the sub-block based TMVP mode. Figure 16 is a view illustrating the method of deriving the motion vector field according to the subblock-based TMVP mode. Figure 17 is a view illustrating a CU extended to perform BDOF. Figure 18 is a view illustrating the relationship between Δν(ί, j), v(i, j) and the subblock motion vector. Figure 19 is a view illustrating a syntactic embodiment for signaling subimage syntactic elements in SPS. Figure 20 is a view illustrating the implementation of the algorithm for deriving a predefined variable such as SubPicTop. Figure 21 is a view illustrating a method of encoding an image using a subimage by an encoding apparatus according to an embodiment. Figure 22 is a view illustrating a method of decoding an image using a subimage by a decoding apparatus according to an embodiment. Figure 23 is a screen illustrating the process of deriving a prediction sample from the current block by applying BDOF. Figure 24 is a view illustrating the input and output of a BDOF process according to an embodiment of the present disclosure. Figure 25 is a view illustrating the variables used for the BDOF process according to an embodiment of the present disclosure. Figure 26 is a view illustrating a method of generating prediction samples for each subblock in the current CU based on whether to apply BDOF according to an embodiment of the present disclosure. Figure 27 is a view illustrating a method of deriving the gradient, auto-correlation and cross-correlation of the current subblock according to an embodiment of the present disclosure. Figure 28 is a view illustrating a method of deriving motion smoothing (vx, vy), deriving BDOF offsets and generating prediction samples from the current sub-block, according to an embodiment of the present disclosure. Figure 29 is a screen illustrating the process of deriving a prediction sample from the current block by applying PROF. Figure 30 is a view illustrating an example of a PROF process according to the present disclosure. Figure 31 is a view illustrating the case where the reference sample to be taken crosses the boundary of a subimage. Figure 32 is an enlarged view of the capture area of ​​Figure 31. Figure 33 is a view illustrating the reference sampling process according to an embodiment of the present disclosure. Figure 34 is a flowchart illustrating the reference sampling process according to this disclosure. Figure 35 is a view illustrating part of a fractional sample interpolation procedure according to the present disclosure. Figure 36 is a view illustrating part of the sbTMVP derivation method according to the present disclosure. Figure 37 is a view illustrating a method of deriving subimage boundary positions according to the present disclosure. Figure 38 is a view showing a content streaming system, to which embodiments of the present disclosure may be applied. Complete Description of the Invention Hereinafter, embodiments of the present disclosure will be described in detail with reference to the accompanying drawings so that they can be readily implemented by those skilled in the art. However, the present disclosure may be implemented in a variety of different forms, and is not limited to the embodiments described herein. In describing this disclosure, if it is determined that a detailed description of known functions or constructions related to it would unnecessarily ambiguous the scope of this disclosure, such detailed descriptions shall be omitted. In the drawings, portions not related to the description of this disclosure shall be omitted, and the same reference numbers shall be assigned to the same portions. In this disclosure, when a component is connected, coupled or linked to another component, this may include not only a direct connection relationship but also an indirect connection relationship where there is an intermediate component. In addition, when a component includes or has another component, it means that the other component may be further included, rather than excluding the other component unless otherwise stated. In this disclosure, the terms first, second, etc. may be used solely for the purpose of distinguishing one component from another, and do not limit the order or importance of such components unless otherwise indicated. Thus, within the scope of this disclosure, the first component in one embodiment may be referred to as the second component in another embodiment, and similarly, the second component in one embodiment may be referred to as the first component in another embodiment. In this disclosure, components are distinguished from each other for the purpose of clearly describing their respective features, and does not mean that they are necessarily separate. That is, a plurality of components may be integrated and implemented in a single hardware or software unit, or a single component may be distributed and implemented in a plurality of hardware or software units. Therefore, even if not stated otherwise, such embodiments in which components are integrated or components are distributed are also included within the scope of this disclosure. In the present disclosure, the components described in various embodiments are not necessarily essential components, and some components may be optional components. Therefore, an embodiment composed of a subset of the components described in an embodiment is also included within the scope of this disclosure. In addition, embodiments that include components other than the components described in various embodiments are included within the scope of this disclosure. This disclosure relates to image encoding and decoding, and terms used in the disclosure may have the usual meanings commonly used in the technical fields to which this disclosure applies, unless newly defined in this disclosure. In this disclosure, “image” generally refers to a unit representing a single image within a specific time period, and a slice / tile is a coding unit that composes a portion of an image, and an image may be composed of one or more slices / tiles. In addition, a slice / tile may include one or more coding tree units (CTUs). In this disclosure, “pixel” or “pixel” may mean the smallest unit that constitutes a single image (or image). Additionally, “sample” may be used as a term corresponding to pixel. A sample may refer to a pixel or pixel value in general, or may refer to only the pixel / pixel value of the luma component or only the pixel / pixel value of the chroma component. In the present disclosure, “unit” may refer to a basic unit of image processing. A unit may include at least one of a specific region of an image and information associated with the region. Unit may be used interchangeably with terms such as “sample array,” “block,” or “area” in some cases. In general, an M*N block may include samples (or an array of samples) or a set (or array) of transformation coefficients with M columns and N rows. In the present disclosure, “current block” may mean any of “current encoding block”, “current encoding unit”, “encoding target block”, “decoding target block” or “processing target block”. When prediction is performed, “current block” may mean “current prediction block” or “prediction target block”. When transformation (inverse transformation) / quantization (decantization) is performed, “current block” may mean “current transformation block” or “transformation target block”. When filtering is performed, “current block” may mean “filtering target block”. In the disclosure of the present invention, the terms / and , shall be interpreted to indicate “and / or. For example, the expressions A / B and “A, B may mean “A and / or B. Further, A / B / C and “A / B / C may mean “at least one of A, B, and / or C. In the disclosure of the present invention, the term “or” shall be interpreted to indicate “and / or.” For example, the statement “A or B” may include 1) only “A,” 2) only “B,” and / or 3) “A and B.” In other words, in the disclosure of the present invention, the term “or” shall be interpreted to indicate “in addition to or alternatively.” An overview of video coding systems 1 is a display showing a video coding system according to the present disclosure. A video coding system according to an embodiment may include an encoding device (10) and a decoding device (20). The encoding device (10) may transmit encoded video and / or image data or information to the decoding device (20) in the form of a file or stream via a digital storage medium or network. The encoding apparatus (10) according to an embodiment may include a video source generator (11), an encoding unit (12) and a transmitter (13). The decoding apparatus (20) according to an embodiment may include a receiver (21), a decoding unit (22) and a renderer (23). The encoding unit (12) may be called a video / image encoding unit, and the decoding unit (22) may be called a video / image decoding unit. The transmitter (13) may be included in the encoding unit (12). The receiver (21) may be included in the decoding unit (22). The renderer (23) may include a viewer and the viewer may be configured as a separate device or an external component. A video source generator (11) may acquire video / images through the process of capturing, synthesizing or generating video / images. The video source generator (11) may include a video / image capture device and / or a video / image generating device. The video / image capturing device may include, for example, one or more cameras, a video / image archive that includes previously captured video / images, and the like. The video / image generating device may include, for example, a computer, a tablet and a smartphone, and may generate (electronically) video / images. For example, virtual video / images may be generated via a computer or the like. In this case, the video / image capture process may be replaced by a process of generating related data. The encoding unit (12) can encode the input video / image. The encoding unit (12) can perform a series of procedures such as prediction, transformation, and quantization for compression and encoding efficiency. The encoding unit (12) can output the encoded data (encoded video / image information) in the form of a bit stream. The transmitter (13) may transmit encoded video / image information or output data in the form of a bit stream to a receiver (21) of a decoding device (20) via a digital storage medium or network in the form of a file or stream. The digital storage medium may include various storage media such as USB, SD, CD, DVD, Blu-ray, HDD, SSD, and the like. The transmitter (13) may include elements for generating media files via a predetermined file format and may include elements for transmission over a broadcast / communication network. The receiver (21) may extract / receive a bit stream from the storage medium or network and transmit the bit stream to the decoding unit (22). The decoding unit (22) can decode the video / image by performing a series of procedures such as dequantization, inverse transformation, and prediction corresponding to the operations of the encoding unit (12). The renderer (23) can render the decoded video / image. The rendered video / image can be displayed via a viewer. Overview of image encoding equipment Figure 2 is a view schematically showing an image encoding apparatus, to which embodiments of the present disclosure may be applied. As shown in Figure 2, the image encoding apparatus (100) may include an image partitioner (110), a reducer (115), a transformer (120), a quantizer (130), a dequantizer (140), an inverse transformer (150), an adder (155), a filter (160), a memory (170), an intermediate prediction unit (180), an intra-prediction unit (185), and an entropy encoder (190). The intermediate prediction unit (180) and the intra-prediction unit (185) may be referred to collectively as a “prediction unit.” The transformer (120), quantizer (130), dequantizer (140), and inverse transformer (150) may be included in a residual processor. The residual processor may further include a reducer (115). All or at least some of the plurality of components that configure the image encoding apparatus (100) may be configured by a single hardware component (e.g., an encoder or processor) in some embodiments. In addition, the memory (170) may include a decoded picture buffer (DPB) and may be configured by a digital storage medium. An image partitioner (110) may partition an input image (or picture or frame) fed to an image coding device (100) into one or more processing units. For example, the processing units may be called coding units (CUs). The coding units may be obtained by recursively partitioning the coding tree units (CTUs) or the largest coding units (LCUs) according to a quadruple-tree / binary-tree / ternary-tree (QT / BT / TT) structure. For example, one coding unit may be partitioned into a plurality of deeper coding units based on a quadruple-tree structure, a binary tree structure, and / or a ternary structure. For partitioning the coding units, the quadruple-tree structure may be applied first and the binary tree structure and / or the ternary structure may be applied later.The coding procedure according to the present disclosure may be performed based on a final coding unit that is no longer partitioned. The largest coding unit may be used as the final coding unit or a coding unit of a deeper depth obtained by partitioning the largest coding unit may be used as the final coding unit. Here, the coding procedure may include prediction, transformation, and reconstruction procedures, which will be described later. As another example, the processing unit of the coding procedure may be a prediction unit (PU) or a transformation unit (TU). The prediction unit and the transformation unit may be separated or partitioned from the final coding unit. The prediction unit may be a sample prediction unit, and the transformation unit may be a unit for deriving transformation coefficients and / or a unit for deriving residual signals from transformation coefficients. A prediction unit (intermediate prediction unit (180) or intra prediction unit (185)) may perform prediction on a block to be processed (current block) and generate a predicted block that includes prediction samples for the current block. The prediction unit may determine whether the intra prediction or intermediate prediction is applied to the current block or is CU-based. The prediction unit may generate various information related to the prediction of the current block and transmit the generated information to an entropy decoder (190). The information in the prediction may be encoded on an entropy encoder (190) and output in the form of a bit stream. The intra-prediction unit (185) may predict the current block by referring to samples in the current image. The referenced samples may be located in the surrounding area of ​​the current block or may be located separately according to the intra-prediction mode and / or intra-prediction technique. The intra-prediction mode may include a plurality of non-directional modes and a plurality of directional modes. The non-directional mode may include, for example, a DC mode and a planar mode. The directional mode may include, for example, 33 directional prediction modes or 65 directional prediction modes according to the level of directional detail of the prediction. However, this is only an example, more or fewer directional prediction modes may be used depending on the settings of the intra-prediction unit (185). The intra-prediction unit (185) may determine the prediction mode applied to the current block by using the prediction modes applied to adjacent blocks. The intermediate prediction unit (180) may derive a predicted block for the current block based on a reference block (reference sample array) specified by a motion vector in the reference image. In this case, to reduce the amount of motion information transmitted in the intermediate prediction mode, the motion information may be predicted at the block, subblock, or sample unit based on the correlation of the motion information between adjacent blocks and the current block. The motion information may include a motion vector and an index of the reference image. The motion information may further include intermediate prediction direction information (L0 prediction, L1 prediction, biprediction, etc.). In the case of intermediate prediction, the adjacent blocks may include spatially adjacent blocks contained in the current image and temporally adjacent blocks contained in the reference image. The reference image that includes the reference block and the reference image that includes the temporally adjacent blocks may be the same or different.Temporally adjacent blocks can be called collocated reference blocks, collocated CUs (colCU), and the like. A reference image that includes temporally adjacent blocks can be called a collocated image (colPic). For example, the intermediate prediction unit (180) can configure a list of motion information candidates based on adjacent blocks and output information indicating which candidates are used to derive the motion vector and / or index of the reference image of the current block. Intermediate prediction can be performed based on various prediction modes. For example, in the case of skip mode and merge mode, the intermediate prediction unit (180) can use the motion information of adjacent blocks as the motion information of the current block. In the case of skip mode, unlike merge mode, residual signals cannot be transmitted.In the case of motion vector prediction (MVP) mode, the motion vectors of adjacent blocks can be used as motion vector predictors, and the motion vector of the current block can be signaled by encoding the difference of the motion vectors and an indicator for the motion vector predictor. The difference of the motion vectors can mean the difference between the motion vector of the current block and the motion vector predictor. The prediction unit can generate prediction signals based on various prediction methods and prediction techniques described below. For example, the prediction unit can apply not only intra-prediction or inter-prediction but also simultaneously apply both intra-prediction and inter-prediction to predict the current block. A prediction method that simultaneously applies both intra-prediction and inter-prediction to predict the current block can be called combined inter- and intra-prediction (CIIP). In addition, the prediction unit can perform intra-block copy (IBC) to predict the current block. The intra-block copy can be used for encoding image / video content of games or the like, for example, screen content coding (SCC).IBC is a method for predicting the current image using previously reconstructed reference blocks in the current image at locations separated from the current block by a predetermined distance. When IBC is applied, the location of the reference block in the current image can be encoded as a vector (block vector) corresponding to the predetermined distance. The prediction signal generated by the prediction unit can be used to generate a reconstructed signal or to generate a residual signal. The subtractor (115) can generate a residual signal (residual block or residual sample array) by subtracting the prediction signal (predicted block or predicted sample array) output from the prediction unit from the input image signal (original block or original sample array). The resulting residual signal can be transmitted to the transformer (120). The transform (120) may generate transformation coefficients by applying a transformation technique to the residual signal. For example, the transformation technique may include at least one of a discrete cosine transform (DCT), a discrete sine transform (DST), a Karhunen-Loeve transform (KLT), a graph-based transform (GBT), or a conditionally nonlinear transform (CNT). Here, GBT refers to a transform obtained from a graph when the relationship information between pixels is represented by a graph. CNT refers to a transform obtained based on a prediction signal generated using all previously reconstructed pixels. In addition, the transformation process may be applied to a rectangular block of pixels of the same size or may be applied to blocks of varying sizes other than rectangular. The quantizer (130) can quantize the transformation coefficients and transmit them to the entropy decoder (190). The entropy encoder (190) can encode the quantized signal (information about the quantized transformation coefficients) and output a bit stream. The information about the quantized transformation coefficients can be referred to as residual information. The quantizer (130) can rearrange the quantized transformation coefficients in block form into a one-dimensional vector based on the coefficient scan sequence and output the information about the quantized transformation coefficients based on the quantized transformation coefficients in one-dimensional vector form. The entropy encoder (190) may perform various encoding methods such as, for example, exponential Golomb, context-adaptive variable length coding (CAVLC), context-adaptive binary arithmetic coding (CABAC), and the like. The entropy encoder (190) may encode information essential for video / image reconstruction other than quantized transformation coefficients (e.g., syntactic element values, etc.) together or separately. The encoded information (e.g., encoded video / image information) may be transmitted or stored in a network abstraction layer (NAL) unit in the form of a bit stream.Further video / image information may include information about various parameter sets such as adaptation parameter set (APS), picture parameter set (PPS), sequence parameter set (Sequence Parameter Set), or video parameter set (VPS). In addition, video / image information may further include general constraint information. Signaled information, information and / or syntactic elements transmitted are described in the coding procedure described above and are included in the bit stream. The bit stream can be transmitted over a network or can be stored in a digital storage medium. The network can include a broadcast network and / or a communications network, and the digital storage medium can include various storage media such as USB, SD, CD, DVD, Blu-ray, HDD, SSD, and so on. A transmitter (not shown) that transmits the signal output from the entropy encoder (190) and / or a storage unit (not shown) that stores the signal can be included as an internal / external element of the image encoding apparatus (100). Alternatively, the transmitter can be provided as a component of the entropy encoder (190). The quantized transformation coefficients removed from the quantization (130) can be used to generate a residual signal. For example, a residual signal (residual block or residual sample) can be reconstructed by applying dequantization and inverse transformation to the quantized transformation coefficients via dequantization (140) and inverse transformation (150). The adder (155) adds the reconstructed residual signal to the prediction signal output from the intermediate prediction unit (180) or the intra prediction unit (185) to produce a reconstructed signal (reconstructed image, reconstructed block, reconstructed sample array). If there is no residual for the block to be processed, such as the case where the skip mode is applied, the predicted block can be used as the reconstructed block. The adder (155) can be called a reconstructor or reconstructed block generator. The resulting reconstructed signal can be used for the next intra block prediction to be processed on the current image and can be used for the next inter-image prediction through filtering as described below. Meanwhile, as explained below, luma mapping with chroma scaling (LMCS) can be applied to the image encoding process. Filter 160 can improve subjective / objective image quality by applying filtering to the reconstructed signal. For example, filter 160 can generate a modified reconstructed image by applying various filtering methods to the reconstructed image and storing the modified reconstructed image in memory 170, specifically, memory DPB 170. Such various filtering methods can include, for example, block filtering, sample adaptive offset, adaptive loop filter, bilateral filter, and so on. Filter 160 can generate various information related to the filtering and transmit the resulting information to an entropy encoder 190 as described later in the description of each filtering method. The information related to the filtering can be encoded by the entropy encoder 190 and output in the form of a bit stream. The modified reconstructed image transmitted to the memory (170) can be used as a reference image in the intermediate prediction unit (180). When the intermediate prediction is applied through the image encoding device (100), the prediction mismatch between the image encoding device (100) and the image decoding device can be avoided and the coding efficiency can be improved. The DPB memory (170) may store a modified reconstructed image for use as a reference image in the intermediate prediction unit (180). The memory (170) may store block motion information from which motion information in the current image is derived (or encoded) and / or block motion information in the reconstructed image. The stored motion information may be transmitted to the intermediate prediction unit (180) and used as spatial adjacent block motion information or temporal adjacent block motion information. The memory (170) may store reconstructed samples from the reconstructed blocks in the current image and may transfer the reconstructed samples to the intra prediction unit (185). Overview of image decoding equipment Figure 3 is a view schematically showing an image decoding apparatus, to which embodiments of the present disclosure may be applied. As shown in Figure 3, the image decoding apparatus (200) may include an entropy decoder (210), a dequantizer (220), an inverse transformer (230), an adder (235), a filter (240), a memory (250), an intermediate prediction unit (260), and an intra-prediction unit (265). The intermediate prediction unit (260) and the intra-prediction unit (265) may be referred to collectively as a “prediction unit.” The dequantizer (220) and the inverse transformer (230) may be included in a residual processor. All or at least some of the plurality of components that configure the image decoding apparatus (200) may be configured by a hardware component (e.g., a decoder or a processor) according to an embodiment. In addition, the memory (250) may include a decoded image buffer (DPB) or may be configured by a digital storage medium. The image decoding apparatus (200), which has received a bit stream including video / image information, can reconstruct the image by performing a process corresponding to the process performed by the image encoding apparatus (100) of FIG. 2. For example, the image decoding apparatus (200) can perform decoding using a processing unit applied to the image encoding apparatus. Thus, the processing unit of decoding can be a coding unit, for example. The coding unit can be obtained by partitioning the coding tree unit or the largest coding unit. The reconstructed image signal decoded and output through the image decoding apparatus (200) can be reproduced through a reproducing apparatus (not shown). The image decoding apparatus (200) may receive a signal output from the image encoding apparatus of Figure 2 in the form of a bit stream. The received signal may be decoded via an entropy decoder (210). For example, the entropy decoder (210) may parse the bit stream to derive information (e.g., video / image information) necessary for image reconstruction (or image reconstruction). The video / image information may further include information about various parameter sets such as an adaptation parameter set (APS), a picture parameter set (PPS), a sequence parameter set (SPS), or a video parameter set (VPS). In addition, the video / image information may further include general constraint information. The image decoding apparatus may further decode the image based on the information about the parameter set and / or general constraint information.The signaled / received information and / or syntactic elements described in the present disclosure may be decoded via a decoding procedure and obtained from the bit stream. For example, the entropy decoder 210 decodes information about the bit stream based on a coding method such as exponential Golomb coding, CAVLC, or CABAC, and outputs the values ​​of the syntactic elements required for image reconstruction and quantized values ​​of the transformation coefficients for the residuals.More specifically, the CABAC entropy decoding method can receive bins corresponding to each syntactic element in the bit stream, determine the context model using the information of the decoding target syntactic element, decode the information of the adjacent block and the decoding target block or the information of the symbol / bin decoded in the previous stage, and perform arithmetic decoding on the bin by predicting the probability of the bin occurrence with the specified context mode, and generate symbols corresponding to the values ​​of each syntactic element. In this case, the CABAC entropy decoding method can update the context model by using the information of the decoded symbol / bin for the context model of the next symbol / bin after determining the context model.Information related to prediction among the information decoded by the entropy decoder (210) can be provided to the prediction units (intermediate prediction unit (260) and intra prediction unit (265)), and the residual values ​​at which entropy decoding has been performed on the entropy decoder (210), namely, the quantized transformation coefficients and related parameter information, can be input to the dequantizer (220). In addition, information about filtering among the information decoded by the entropy decoder (210) can be provided to the filter (240). Meanwhile, a receiver (not shown) for receiving the signal outputted from the image coding apparatus can be further configured as an internal / external element of the image decoding apparatus (200), or the receiver can be a component of the entropy decoder (210). Meanwhile, the image decoding apparatus according to the present disclosure can be referred to as a video / image / image decoding apparatus. The image decoding apparatus can be classified into an information decoder (video / image / image information decoder) and a sample decoder (video / image / image sample decoder). The information decoder can include an entropy decoder (210). The sample decoder can include at least one of a dequantizer (220), an inverse transformer (230), an adder (235), a filter (240), a memory (250), an intermediate prediction unit (260) or an intra prediction unit (265). The dequantizer (220) can dequantize the quantized transformation coefficients and output the transformation coefficients. The dequantizer (220) can rearrange the quantized transformation coefficients in a two-dimensional block form. In this case, the rearrangement can be performed based on the sequence of coefficient scans performed on the image encoding device. The dequantizer (220) can dequantize the quantized transformation coefficients using quantization parameters (e.g., quantization step size information) and obtain the transformation coefficients. The inverse transform (230) can inversely transform the transform coefficients to obtain a residual signal (residual block, residual sample array). The prediction unit may perform predictions on the current block and generate a predicted block that includes prediction samples for the current block. The prediction unit may determine whether intra-prediction or intermediate-prediction is applied to the current block based on information about the prediction output of the entropy decoder (210) and may determine specific intra- / inter-prediction modes (prediction techniques). As described in the prediction unit of the image coding apparatus (100) that the prediction unit can generate prediction signals based on various prediction methods (techniques) which will be described later. The intra prediction unit (265) can predict the current block by referring to samples in the current image. The description of the intra prediction unit (185) applies equivalently to the intra prediction unit (265). The intermediate prediction unit (260) may derive blocks produced for the current block based on a reference block (reference sample array) specified by a motion vector in the reference image. In this case, to reduce the amount of motion information transmitted in the intermediate prediction mode, motion information may be predicted in a block, subblock, or sample unit based on the correlation of motion information between adjacent blocks and the current block. The motion information may include a motion vector and an index of the reference image. The motion information may further include intermediate prediction direction information (L0 prediction, L1 prediction, Biprediction, etc.). In the case of intermediate prediction, adjacent blocks may include spatially adjacent blocks contained in the current image and temporally adjacent blocks contained in the reference image.For example, the intermediate prediction unit 260 may configure a list of motion information candidates based on adjacent blocks and derive the current block's motion vector and / or reference image index based on the received candidate selection information. The intermediate prediction may be performed based on various prediction modes, and the information about the prediction may include information indicating the intermediate prediction mode for the current block. The adder (235) may generate a reconstructed signal (reconstructed image, reconstructed block, reconstructed sample array) by adding the obtained residual signal to a predicted signal (predicted block, predicted sample array) output from a prediction unit (which includes an intermediate prediction unit (260) and / or an intra prediction unit (265)). The description of the adder (155) may be applied equivalently to the adder (235). Meanwhile, as described below, luma mapping with chroma scaling (LMCS) can be applied in the image decoding process. Filter 240 can improve subjective / objective image quality by applying filtering to the reconstructed signal. For example, filter 240 can produce a modified reconstructed image by applying various filtering methods to the reconstructed image and storing the modified reconstructed image in memory 250, in particular, memory DPB 250. Such various filtering methods can include, for example, block filtering, sample adaptive offset, adaptive loop filter, bilateral filter, and so on. The reconstructed (modified) image stored in the memory DPB (250) may be used as a reference image in the intermediate prediction unit (260). The memory (250) may store block motion information from which the motion information in the current image is derived (or decoded) and / or block motion information in the reconstructed image. The stored motion information may be transmitted to the intermediate prediction unit (260) for use as spatial adjacent block motion information or temporal adjacent block motion information. The memory (250) may store reconstructed samples from the reconstructed blocks in the current image and transfer the reconstructed samples to the intra prediction unit (265). In the present disclosure, the embodiments described in the filter (160), the intermediate prediction unit (180), and the intra prediction unit (185) of the image encoding apparatus (100) may be applied similarly or in conjunction with the filter (240), the intermediate prediction unit (260), and the intra prediction unit (265) of the image decoding apparatus (200). The big picture of the predictions between An image encoding / decoding apparatus may perform intermediate predictions on block units to derive prediction samples. Intermediate predictions may mean predictions derived in a manner dependent on image data elements other than the current image. If an intermediate prediction applies to the current block, the predicted block for the current block may be derived based on a reference block defined by a motion vector in the reference image. In this case, to reduce the amount of motion information transmitted in the intermediate prediction mode, the motion information of the current block can be derived based on the correlation of motion information between adjacent blocks and the current block, and the motion information can be derived in units of blocks, subblocks, or samples. The motion information can include motion vectors and reference image indices. The motion information can further include intermediate prediction type information. Here, intermediate prediction type information can mean intermediate prediction direction information. The intermediate prediction type information can indicate that the current block is predicted using one of L0 prediction, L1 prediction, or biprediction. When applying intermediate prediction to the current block, adjacent blocks of the current block may include spatially adjacent blocks contained in the current image and temporally adjacent blocks contained in the reference image. The reference image that includes the reference blocks for the current block and the reference image that includes the temporally adjacent blocks may be the same or different. The temporally adjacent blocks may be referred to as linked reference blocks or linked CUs (colCU), and the reference image that includes the temporally adjacent blocks may be referred to as linked images (colPic). Meanwhile, a list of motion information candidates can be constructed based on the blocks adjacent to the current block, and, in this case, index or flag information indicating the candidates to be used can be signaled to derive the motion vector of the current block and / or the index of the reference image. The motion information may include L0 motion information and / or L1 motion information according to the intermediate prediction type. The motion vector in the L0 direction may be specified as L0 or MVL0 motion vector, and the motion vector in the L1 direction may be specified as L1 or MVL1 motion vector. The prediction based on the L0 motion vector may be specified as L0 prediction, the prediction based on the L1 motion vector may be specified as L1 prediction, and the prediction based on the L0 motion vector and the L1 motion vector may be specified as bi-prediction. Herein, the L0 motion vector may mean the motion vector corresponding to the reference image list L0 and the L1 motion vector may mean the motion vector corresponding to the reference image list L1. The reference image list (L0) may include the image before the current image in the output sequence as the reference image, and the reference image list (L1) may include the image after the current image in the output sequence. The previous image may be specified as the front (reference) image and the subsequent image may be specified as the rear (reference) image. Meanwhile, the reference image list (L0) may further include the image after the current image in the output sequence as the reference image. In this case, in the reference image list (L0), the previous image may be indexed first and the subsequent image may be indexed later. The reference image list (L1) may further include the image before the current image in the output sequence as the reference image. In this case, in the reference image list (L1), the subsequent image may be indexed first and the previous image may be indexed later.Here, the output order can correspond to the picture order count (POC). Figure 4 is a flowchart illustrating intermediate prediction based on video / image encoding methods. Figure 5 is a view illustrating the configuration of the intermediate predictor (180) according to the present disclosure. The encoding method of Figure 6 can be performed by the image encoding apparatus of Figure 2. Specifically, step S410 can be performed by the intermediate predictor (180), and step S420 can be performed by the residual processor. Specifically, step S420 can be performed by the reducer (115). Step S430 can be performed by the entropy encoder (190). The prediction information of step S630 can be derived by the intermediate predictor (180), and the residual information of step S630 can be derived by the residual processor. The residual information is information about the residual samples. The residual information can include information about the quantized transformation coefficients for the residual samples. As described above, the residual samples can be derived as transformation coefficients through the transformer (120) of the image encoding apparatus, and the transformation coefficients can be derived as quantized transformation coefficients through the quantizer (130).Information about the quantized transformation coefficients can be encoded by the entropy encoder (190) through a residual coding procedure. The image coding apparatus can perform intermediate prediction against the current block (S410). The image coding apparatus can derive intermediate prediction modes and motion information from the current block and generate prediction samples from the current block. Herein, the procedures of determining the intermediate prediction modes, deriving motion information and generating prediction samples can be performed simultaneously or one of them can be performed before the other procedures. For example, as shown in Figure 5, the intermediate prediction unit (180) of the image coding apparatus can include a prediction mode determining unit (181), a motion information deriving unit (182) and a prediction sample deriving unit (183). The prediction mode determining unit (181) can determine the prediction modes of the current block, the motion information deriving unit (182) can derive motion information from the current block, and the prediction sample deriving unit (183) can derive prediction samples from the current block.For example, an intermediate prediction unit (180) of an image coding apparatus can search for blocks similar to the current block within a predetermined area (search area) of a reference image through motion estimation, and derive a reference block whose difference from the current block is equal to or less than a predetermined criterion or a minimum. Based on this, a reference image index indicating the reference image in which the reference block is located can be derived, and a motion vector can be derived based on the position of the difference between the reference block and the current block. The image coding apparatus can determine a mode applied to the current block among various intermediate prediction modes. The image coding apparatus can compare rate-distortion (RD) costs for various prediction modes and determine an optimal intermediate prediction mode of the current block.However, the method of determining the prediction mode between the current blocks by the image coding equipment is not limited to the above examples, and various methods can be used. For example, the current interblock prediction mode may be specified to be at least one of a combined mode, a jump combined mode, a motion vector prediction (MVP) mode, a symmetric motion vector difference (SMVD) mode, an affine mode, a subblock-based combined mode, an adaptive motion vector resolution (AMVR) mode, a history-based motion vector predictor (HMVP) mode, a pairwise average combined mode, a combined mode with motion vector difference (MMVD) mode, a decoder-side motion vector smoothing (DMVR) mode, a combined interblock and intrablock prediction (CIIP) mode or a geometric partitioning (GPM) mode. For example, when a jump mode or a merge mode is applied to the current block, the image encoding device can derive merge candidates from blocks adjacent to the current block and construct a merge candidate list using the derived merge candidates. In addition, the image encoding device can derive a reference block whose difference from the current block is equal to or less than a predetermined criterion or a minimum, among the reference blocks indicated by the merge candidates included in the merge candidate list. In this case, the merge candidates corresponding to the derived reference block can be selected, and merge index information indicating the selected merge candidates can be generated and signaled to the image decoding device. Motion information of the current block can be derived using the motion information of the selected merge candidates. As another example, When MVP mode is applied to the current block, the image encoding tool can derive motion vector predictor (MVP) candidates from the blocks adjacent to the current block and construct a list of MVP candidates using the derived MVP candidates. In addition, the image encoding tool can use the motion vectors of the MVP candidates selected from among the MVP candidates included in the MVP candidate list as the MVP of the current block. In this case, for example, the motion vector indicating the reference block derived by the motion estimation described above can be used as the motion vector of the current block, the MVP candidate with the motion vector having the smallest difference from the motion vector of the current block among the MVP candidates can be the selected MVP candidate. The motion vector difference (MVD) which is the difference obtained by subtracting the MVP from the motion vector of the current block can be derived.In this case, the index information indicating the selected MVP candidates and information about the MVD can be signaled to the image decoding device. In addition, when applying the MVP mode, the value of the reference image index can be constructed as the reference image index information and separately signaled to the image decoding device. The image encoding apparatus may derive residual samples based on the predicted samples (S420). The image encoding apparatus may derive residual samples through a comparison between the original samples of the current block and the predicted samples. For example, the residual samples may be derived by subtracting the predicted samples corresponding to the original samples. An image encoding device can encode image information including prediction information and residual information (S430). The image encoding device can output the encoded image information in the form of a bit stream. The prediction information can include prediction mode information (e.g., a jump marker, a merge marker or a mode index, etc.) and information about motion information as information related to the prediction procedure. Among the prediction mode information, the jump marker indicates whether the jump mode is applied to the current block, and the merge marker indicates whether the merge mode is applied to the current block. Alternatively, the prediction mode information can indicate one of a plurality of prediction modes, such as a mode index. When the jump marker and the merge marker are 0, it can be determined that the mode MVP is applied to the current block. Information about the motion information may include candidate selection information (e.g., a merge index, an mvp flag, or an mvp index) that is information for deriving the motion vector. Among the candidate selection information, the merge index may be signaled when the merge mode is applied to the current block and may be information for selecting one of the merge candidates included in the merge candidate list. Among the candidate selection information, the MVP flag or the MVP index may be signaled when the MVP mode is applied to the current block and may be information for selecting one of the MVP candidates in the MVP candidate list. Specifically, the MVP flag may be signaled using the syntactic elements mvp_10_flag or mvp_11_flag. In addition, information about the motion information may include information about the reference image index information and / or the MVD.In addition, information about the motion information may include information indicating whether to apply L0 prediction, L1 prediction, or Bi-prediction. Residual information is information about the residual samples. Residual information may include information about the quantized transformation coefficients for the residual samples. The output bit stream can be stored in a (digital) storage medium and transmitted to image decoding equipment or can be transmitted to image decoding equipment over a network. As described above, an image encoding device can generate a reconstructed image (an image comprising reconstructed samples and reconstructed blocks) based on reference samples and residual samples. This is for the image encoding device to derive prediction results that are the same as those performed by the image decoding device, thereby improving the encoding efficiency. Thus, the image encoding device can store the reconstructed image (or reconstructed samples and reconstructed blocks) in memory and use the reconstructed image as a reference image for intermediate predictions. As described above, an in-loop filtering procedure can be further applied to the reconstructed image. Figure 6 is a flowchart illustrating intermediate prediction based on video / image decoding methods. Figure 7 is a view illustrating the configuration of the intermediate prediction unit (260) according to the present disclosure. The image decoding device can perform operations corresponding to those performed by the image encoding device. The image decoding device can perform predictions on the current block based on the received prediction information and derive prediction samples. The decoding method of Figure 6 can be performed by the image decoding apparatus of Figure 3. Steps S610 to S630 can be performed by the intermediate prediction unit (260), and the prediction information of step S610 and the residual information of step S640 can be obtained from the bit stream by the entropy decoder (210). The residual processor of the image decoding apparatus can derive residual samples for the current block based on the residual information (S640). Specifically, the dequantizer (220) of the residual processor can perform dequantization based on the quantized transformation coefficients derived based on the residual information to derive the transformation coefficients, and the inverse transformer (230) of the residual processor can perform inverse transformation on the transformation coefficients to derive the residual samples for the current block. Step S650 can be performed by the adder (235) or the reconstructor. Specifically, the image decoding device can determine the prediction mode of the current block based on the received prediction information (S610). The image decoding device can determine which prediction mode is applied to the current block based on the prediction mode information in the prediction information. For example, it can determine whether skip mode is applied to the current block based on the skip marker. Additionally, it can determine whether merge mode or MVP mode is applied to the current block based on the merge marker. Alternatively, one of various intermediate prediction mode candidates can be selected based on the mode index. Intermediate prediction mode candidates can include skip mode, merge mode, and / or MVP mode, or can include various intermediate prediction modes described below. The image decoding apparatus may derive motion information from the current block based on a predetermined intermediate prediction mode (S620). For example, if a jump mode or a merge mode is applied to the current block, the image decoding apparatus may construct a merge candidate list, as described below, and select one of the merge candidates included in the merge candidate list. The selection can be made based on the candidate selection information described above (the combined index). The movement information of the current block can be derived using the movement information of the selected combined candidates. For example, the movement information of the selected combined candidates can be used as the movement information of the current block. As another example, when MVP mode is applied to the current block, the image decoding tool can construct a list of MVP candidates and use the motion vector of the MVP candidate selected from among the MVP candidates included in the MVP candidate list as the MVP of the current block. The selection can be made based on the candidate selection information described above (mvp marker or mvp index). In this case, the MVD of the current block can be derived based on the information about the MVD, and the motion vector of the current block can be derived based on the MVP and the MVD of the current block. In addition, the reference image index of the current block can be derived based on the reference image index information. The image indicated by the reference image index in the reference image list of the current block can be derived as the reference image referred to for intermediate prediction of the current block. The image decoding apparatus can generate prediction samples of the current block based on the motion information of the current block (S630). In this case, a reference image can be derived based on the index of the reference image of the current block, and prediction samples of the current block can be derived using samples of the reference block indicated by the motion vector of the current block in the reference image. In some cases, a prediction sample filtering procedure can be further performed on all or some of the prediction samples of the current block. For example, as shown in Figure 7, an intermediate prediction unit (260) of an image decoding apparatus may include a prediction mode determination unit (261), a motion information derivation unit (262) and a prediction sample derivation unit (263). In the intermediate prediction unit (260) of the image decoding apparatus, the prediction mode determination unit (261) may determine the prediction mode of the current block based on the received prediction mode information, the motion information derivation unit (262) may derive motion information (motion vector and / or reference image index, etc.) of the current block based on the received motion information, and the prediction sample derivation unit (263) may derive prediction samples of the current block. The image decoding apparatus can generate residual samples from the current block based on the received residual information (S640). The image decoding apparatus can generate reconstructed samples from the current block based on the predicted samples and residual samples and generate a reconstructed image based on this (S650). Thereafter, an in-loop filtering procedure can be applied to the reconstructed image as described above. As described above, an intermediate prediction procedure may include the step of determining an intermediate prediction mode, the step of deriving motion information according to the determined prediction mode, and the step of performing a prediction (generating a prediction sample) based on the derived motion information. The intermediate prediction procedure may be performed by an image encoding apparatus and an image decoding apparatus, as described above. Next here, the steps to derive motion information according to the prediction mode will be described in more detail. As described above, intermediate prediction can be performed using motion information from the current block. The image encoding device can derive optimal motion information from the current block through a motion estimation procedure. For example, the image encoding device can search for similar reference blocks with high correlation within a predetermined search range in the reference image using the original blocks in the original image for the current block in fractional pixel units, and derive motion information therefrom. The block similarity can be calculated based on the sum of absolute differences (SAD) between the current block and the reference block. In this case, motion information can be derived based on the reference block with the smallest SAD in the search area. The derived motion information can be signaled to the image decoding device by various methods based on intermediate prediction modes. When the merge mode is applied to the current block, the motion information of the current block is not transmitted directly and the motion information of the current block is derived using the motion information of the adjacent blocks. Accordingly, the motion information of the current prediction block can be indicated by transmitting marker information indicating that the merge mode is used and candidate selection information (e.g., a merge index) indicating which adjacent blocks are used as merge candidates. In the present disclosure, because the current block is a prediction performance unit, the current block can be given the same meaning as the current prediction block, and the adjacent blocks can be used as the same meaning as the adjacent prediction blocks. The image encoding tool can search for the combined candidate blocks used to derive motion information from the current block to perform the combined mode. For example, up to five combined candidate blocks can be used, without any limitation. The maximum number of combined candidate blocks can be transmitted in the chunk header or the tile group header, without any limitation. After finding the combined candidate blocks, the image encoding tool can generate a list of combined candidates and select the combined candidate block with the smallest RD cost as the final combined candidate block. A merge candidate list can use, for example, five merge candidate blocks. For example, four spatial merge candidates and one temporal merge candidate can be used. The big picture of affine mode Next, the affine mode, which is an example of an intermediate prediction mode, will be described in detail. In conventional video encoding / decoding systems, only one motion vector is used to express the motion information of the current block (translational motion model). However, in conventional methods, the optimal motion information is only expressed in block units, but the optimal motion information cannot be expressed in pixel units. To solve this problem, an affine motion mode that defines the motion information of a block in pixel units has been proposed. According to the affine mode, the motion vector for each pixel and / or subblock unit of a block can be determined using two to four motion vectors related to the current block. Compared to existing motion information expressed using translation (or displacement) of pixel values, in affine mode, motion information for each pixel can be expressed using at least one of translation, scaling, rotation or shear. Figure 8 is a view illustrating the motion that can be expressed in affine mode. Among the motions shown in Figure 8, the affine mode in which the motion information for each pixel is expressed using displacement, scaling or rotation can be either a similarity or simplified affine mode. The affine mode in the following description can mean either a similarity or simplified affine mode. Motion information in affine mode can be represented using two or more control point motion vectors (CPMVs). The motion vectors of specific pixel positions of a current block can be derived using CPMVs. In this case, the set of motion vectors for each pixel and / or subblock of a current block can be defined as an affine motion vector field (affine MVF). Figure 9 is a view illustrating the parameter model of the affine mode. When the affine mode is applied to the current block, the affine MVF can be derived using either the 4-parameter model or the 6-parameter model. In this case, the 4-parameter model can mean the type of model where two CPMVs are used and the 6-parameter model can mean the type of model where three CPMVs are used. Figures 9(a) and 9(b) show the CPMVs used in the 4-parameter model and the 6-parameter model, respectively. If the current block position is (x, y) , then the motion vector according to the pixel position can be derived according to Equation 1 or 2 below. For example, the motion vector according to the 4-parameter model can be derived according to Equation 1 and the motion vector according to the 6-parameter model can be derived according to Equation 2. Equation 1 mvx= nwlx-mvox IV \mvy = mviy-mvoyW mvly-mvay IV IV y + W y + mvOy Equation 2 mvr=-------x H--y + mv^ΛW ΗυΛmvlv-mvov mv2v-mv0vmvv=-------x H-------y + mvnvv ywh In Equations 1 and 2, mvO = (mv_0x, mv_0y} can be the CPMV at the top left corner position of the current block, mvl = {mv_lx, mv_ly} can be the CPMV at the top right position of the current block, and mv2 = (mv_2x, mv_2y} can be the CPMV at the bottom left position of the current block. In this case, W and H correspond to the width and height of the current block, respectively, and mv = (mv_x, mv_y} can mean the motion vector of the pixel position {x, y}. In the encoding / decoding process, the affine MVF can be defined in pixel units and / or pre-defined subblocks. When the affine MVF is defined in pixel units, the motion vector can be derived based on each pixel value. Meanwhile, when the affine MVF is defined in subblock units, the corresponding block motion vector can be derived based on the center pixel value of a subblock. The center pixel value can mean a virtual pixel located at the center of the subblock or the bottom right pixel among the four pixels located at the center. In addition, the center pixel value can be a specific pixel in the subblock and can be a pixel representing the subblock. In the disclosure of the present invention, the case where the affine MVF is defined in 4x4 subblock units will be described. However, this is only for convenience of description and the size of the subblock can be varied. That is, if affine prediction is available, the motion models applicable to the current block can include three models, namely, a translational motion model, a 4-parameter affine motion model and a 6-parameter affine motion model. Here, the translational motion model can indicate the model used by the existing block unit motion vector, the 4-parameter affine motion model can indicate the model used by two CPMVs, and the 6-parameter affine motion model can indicate the model used by three CPMVs. The affine mode can be divided into detailed modes according to the encoding / decoding method of motion information. For example, the affine mode can be further divided into affine MVP mode and affine combined mode. When the affine merge mode is applied to the current block, the CPMV can be derived from adjacent blocks of the current block encoded / decoded in affine mode. When at least one of the adjacent blocks of the current block is encoded / decoded in affine mode, the affine merge mode can be applied to the current block. That is, when the affine merge mode is applied to the current block, the CPMV of the current block can be derived using the CPMV of the adjacent block. For example, the CPMV of the adjacent block can be determined to be the CPMV of the current block or the CPMV of the current block can be derived based on the CPMV of the adjacent block. When the CPMV of the current block is derived based on the CPMV of the adjacent block, at least one of the encoding parameters of the current block or the adjacent block can be used. For example, the CPMV of the adjacent block can be modified based on the size of the adjacent block and the size of the current block and used as the CPMV of the current block. Meanwhile, the affine merge in which MV is derived in a subblock unit can be referred to as the subblock merge mode, which can be specified by merge_subblock_flag having the first value (e.g., 1). In this case, the list of affine merge candidates described below can be referred to as the subblock merge candidate list. In this case, the candidate derived as SbTMVP described below can be further included in the subblock merge candidate list. In this case, the candidate derived as sbTMVP can be used as the index candidate #0 of the subblock merge candidate list. In other words, the candidate derived as sbTMVP can be placed before the inherited affine candidate and the constructed affine candidate described below in the subblock merge candidate list. For example, an affine mode flag specifying whether affine mode can be applied to the current block can be specified, which can be signaled at at least one of the current block's higher levels, such as sequence, image, slice, tile, tile group, block, etc. For example, the affine mode flag can be named sps_affine_enabled_flag. When affine merge mode is applied, a list of affine merge candidates can be configured to derive the current block's CPMV. In this case, the list of affine merge candidates can include at least one of legacy affine merge candidates, constructed affine merge candidates, or null merge candidates. Legacy affine merge candidates can mean candidates derived using adjacent block CPMVs if the current block's adjacent blocks are encoded / decoded in affine mode. Constructed affine merge candidates can mean candidates having each CPMV derived based on the motion vector of the adjacent block of each control point (CP). Meanwhile, null merge candidates can mean candidates consisting of CPMVs having size 0. In the following description, CP can mean the specific position of the block used to derive the CPMV. For example, CP can be each vertex position of the block. Figure 10 is a view illustrating the method for generating an affine merge candidate list. Referring to the flowchart of Figure 10, affine merge candidates can be added to the affine merge candidate list in the order of legacy affine merge candidates (S1210), constructed affine merge candidates (S1220), and null merge candidates (S1230). A null merge candidate can be added if the number of candidates covered in the candidate list does not meet the maximum number of candidates even though all legacy affine merge candidates and constructed affine merge candidates are added to the affine merge candidate list. In this case, null merge candidates can be added until the number of candidates of the affine merge candidate list meets the maximum number of candidates. Figure 11 is a view illustrating the control point motion vector (CPMV) derived from adjacent blocks. For example, a maximum of two candidate affine inheritance joins can be derived, each of which can be derived based on at least one of the left adjacent block and the top adjacent block. Figure 12 is a view illustrating the adjacent blocks for deriving inherited affine join candidates. The inherited affine merge candidate derived based on the left adjacent block is derived based on at least one of the adjacent blocks A0 or A1 of Figure 12, and the inherited affine merge candidate derived based on the top adjacent block may be derived based on at least one of the adjacent blocks B0, B1 or B2 of Figure 12. In this case, the scanning order of the adjacent blocks may be A0 to A1 and B0, B1 and B2, but is not limited thereto. For left and top respectively, the inherited affine merge candidate may be derived based on the first available adjacent block in the scanning order. In this case, redundancy check may not be performed between the candidates derived from the left adjacent block and the top adjacent block. For example, as shown in Figure 11, When the left adjacent block (A) is encoded / decoded in affine mode, at least one of the motion vectors v2, v3 and v4 corresponding to the CP of the adjacent block (A) can be derived. When the adjacent block (A) is encoded / decoded via a 4-parameter affine model, the inherited affine combination candidate can be derived using v2 and v3. Conversely, When the adjacent block (A) is encoded / decoded via a 6-parameter affine model, the inherited affine combination candidate can be derived using v2, v3 and v4. Figure 13 is a view illustrating adjacent blocks for derivation of constructed affine combination candidates. A constructed affine candidate can mean a candidate that has a CPMV derived using a combination of common motion information of adjacent blocks. The motion information for each CP can be derived using either spatially adjacent blocks or temporally adjacent blocks of the current block. In the following description, CPMVk can mean the motion vector representing the kth CP. For example, referring to Figure 13, CPMV1 can be determined to be the first available motion vector from motion vectors B2, B3, and A2, and, in this case, the scan order can be B2, B3, and A2. CPMV2 can be determined to be the first available motion vector from motion vectors B1 and B0, and, in this case, the scan order can be B1 and B0. CPMV3 can be determined to be one of motion vectors A1 and A0, and, in this case, the scan order can be A1 and A0.If TMVP can be applied to the current block, CPMV4 can be determined as the motion vector T which is a temporally adjacent block. Once the four motion vectors for each CP are derived, a constructed affine joint candidate can be derived based on these. The constructed affine joint candidate can be configured by including at least two motion vectors selected from among the four motion vectors derived for each CP. For example, the constructed affine joint candidate can consist of at least one of {CPMV1, CPMV2, CPMV3}, {CPMV1, CPMV2, CPMV4}, {CPMV1, CPMV3, CPMV4}, {CPMV2, CPMV3, CPMV4}, {CPMV1, CPMV2} or {CPMV1, CPMV3} in this order. A constructed affine candidate consisting of three motion vectors can be a candidate for a 6-parameter affine model. On the other hand, a constructed affine candidate consisting of two motion vectors can be a candidate for a 4-parameter affine model.To avoid the process of scaling motion vectors, If the CP reference image indices are different from each other, the corresponding CPMV combinations can be removed without being used to derive the constructed affine candidates. When the affine MVP mode is applied to the current block, the encoding / decoding device can derive two or more CPMV predictors and CPMV for the current block and derive the CPMV difference based on the CPMV predictors and CPMV. In this case, the CPMV difference can be signaled from the encoding device to the decoding device. The image decoding device can derive the CPMV predictor for the current block, reconstruct the signaled CPMV difference, and then derive the CPMV of the current block based on the CPMV predictor and CPMV difference. Meanwhile, if the affine merge mode or subblock-based TMVP is not applied to the current block (for example, the value of the affine merge flag or merge_subblock_flag is 0), the affine MVP mode can be applied to the current block. Alternatively, if the value of the inter_affine_flag is 1, the affine MVP mode can be applied to the current block. Meanwhile, the affine MVP mode can be indicated as the affine CP MVP mode. The affine MVP candidate list described below can be referred to as the control point motion vector predictor candidate list. When affine MVP mode is applied to the current block, the affine MVP candidate list can be configured to derive the CPMV for the current block. In this case, the affine MVP candidate list can include at least one of the legacy affine MVP candidates, the constructed affine MVP candidates, the translational motion affine MVP candidates, or the null MVP candidates. For example, the affine MVP candidate list can include a maximum of n (e.g., n=2) candidates. In this case, a legacy affine MVP candidate may mean a candidate derived based on the CPMV of adjacent blocks, if the blocks adjacent to the current block are encoded / decoded in affine mode. A constructed affine MVP candidate may mean a candidate derived by generating a combination of CPMVs based on the CP motion vectors of adjacent blocks. A null MVP candidate may mean a candidate consisting of CPMVs having a value of 0. The derivation method and characteristics of the legacy affine MVP candidate and the constructed affine MVP candidate are the same as those of the above-described legacy affine candidate and the constructed affine candidate and thus their description will be omitted. If the maximum number of candidates of the affine MVP candidate list is 2, then the constructed affine MVP candidates, translational motion affine MVP candidates, and null MVP candidates can be added if the current number of candidates is less than 2. Specifically, the translational motion affine MVP candidates can be derived in the following order. For example, if the number of candidates covered in the affine MVP candidate list is less than 2 and the constructed affine MVP candidate CPMV0 is valid, then CPMV0 can be used as an affine MVP candidate. That is, an affine MVP candidate whose motion vectors CP0, CP1, CP2 are all CPMV0 can be added to the affine MVP candidate list. Furthermore, if the number of affine MVP candidate list is less than 2 and the constructed affine MVP candidate CPMV1 is valid, then CPMV1 can be used as an affine MVP candidate. That is, an affine MVP candidate whose motion vectors CP0, CP1, CP2 are all CPMV1 can be added to the affine MVP candidate list. Furthermore, if the number of affine MVP candidate list is less than 2 and the constructed affine MVP candidate CPMV2 is valid, then CPMV2 can be used as an affine MVP candidate. That is, an affine MVP candidate whose motion vectors CP0, CP1, CP2 are all CPMV2 can be added to the affine MVP candidate list. Despite the conditions described above, if the number of candidates in the affine MVP candidate list is less than 2, then the temporal motion vector predictor (TMVP) in the current block can be added to the affine MVP candidate list. Despite the conditions described above, if the number of candidates in the affine MVP candidate list is less than 2, the zero MVP candidate can be added to the affine MVP candidate list. Figure 14 is a view illustrating the method for generating a list of affine MVP candidates. Referring to the flowchart of Figure 14, candidates can be added to the affine MVP candidate list in the order of legacy affine MVP candidate (S1610), constructed affine MVP candidate (S1620), translational motion affine MVP candidate (S1630), and null MVP candidate (S1640). As described above, steps S1620 to S1640 can be performed depending on whether the number of candidates included in the affine MVP candidate list is less than 2 at each step. The scan order of legacy affine MVP candidates can be the same as the scan order of legacy affine merge candidates. However, in the case of legacy affine MVP candidates, only adjacent blocks that refer to the same reference image as the current block's reference image can be considered. When legacy affine MVP candidates are added to the list of affine MVP candidates, redundancy checking can be omitted. To derive the constructed affine MVP candidates, only the spatially adjacent blocks shown in Figure 13 can be considered. In addition, the scanning order of the constructed affine MVP candidates can be the same as the scanning order of the constructed affine merge candidates. In addition, to derive the constructed affine MVP candidates, the indexes of the adjacent block reference images can be checked, and, in the scanning order, the first adjacent block that is intermediate-coded and refers to the same reference image as the current block reference image can be used. Overview of subblock-based temporal motion vector prediction (SbTMVP) mode Next, the subblock-based TMVP mode, which is an example of an intermediate prediction mode, will be described in detail. According to the subblock-based TMVP mode, the motion vector field (MVF) for the current block can be derived, and the motion vector can be derived in subblock units. Unlike the conventional TMVP mode which is performed in units of coding units, for coding units applying the sub-block-based TMVP mode, motion vectors can be encoded / decoded in units of sub-coding units. In addition, according to the conventional TMVP mode, temporal motion vectors can be derived from collocated blocks in the collocated image, but, in the sub-block-based TMVP mode, the motion vector field can be derived from the reference block in the collocated image specified by the motion vectors derived from adjacent blocks of the current block. Furthermore, the motion vectors derived from adjacent blocks can be called motion shifts or representative motion vectors of the current block. Figure 15 is a view illustrating the adjacent blocks of the subblock-based TMVP mode. When the subblock-based TMVP mode is applied to the current block, the adjacent blocks for determining the motion shift can be determined. For example, scanning for adjacent blocks for determining the motion shift can be performed in the order of blocks A1, B1, B0, and A0 of Figure 15. As another example, the adjacent blocks for determining the motion shift can be limited to specific adjacent blocks of the current block. For example, the adjacent block for determining the motion shift can always be determined to be block A1. If the adjacent blocks have motion vectors that refer to interlocked images, the corresponding motion vectors can be determined to be the motion shifts. The motion vectors determined to be the motion shifts can be called temporal motion vectors. Meanwhile, if the motion vectors described above cannot be derived from the adjacent blocks, the motion shift can be set to be (0, 0). Figure 16 is a view illustrating the method of deriving the motion vector field according to the subblock-based TMVP mode. Next, the reference block in the interlocking image specified by the motion shift can be determined. For example, subblock-based motion information (motion vector or reference image index) can be obtained from the interlocking image by adding the motion shift to the coordinates of the current block. In the example shown in Figure 16, it is assumed that the motion shift is the motion vector of block A1. By applying the motion shift to the current block, the subblocks in the interlocking image (the interlocking subblocks) corresponding to each subblock configuring the current block can be specified. Next, using the motion information of the corresponding subblocks in the interlocking image (the interlocking subblocks), the motion information of each subblock of the current block can be derived. For example, the motion information of the corresponding subblock can be obtained from the center position of the corresponding subblock.In this case, the center position is the position of the bottom right sample placed at the center of the corresponding subblock. If the specific subblock motion information of the linked block corresponding to the current block is not available, then the center subblock motion information of the linked block can be determined to be the corresponding subblock motion information. When the corresponding subblock motion vector is derived, it can be transferred to the reference image index and the current subblock motion vector, similarly to the TMVP process described above. That is, When the subblock-based motion vector is derived, the scaling of the motion vector can be performed by considering the POC of the reference image of the reference block. As described above, subblock-based TMVP candidates for the current block can be derived using the motion vector fields or subblock-derived motion information of the current block. Next, the combined candidate list configured in the subblock unit is defined as the subblock unit combined candidate list. The affine combined candidates described above and the subblock-based TMVP candidates can be combined to construct the subblock unit combined candidate list. Meanwhile, a subblock-based TMVP mode flag specifying whether the subblock-based TMVP mode can be applied to the current block can be defined, which can be signaled at at least one level among the levels higher than the current block such as sequence, image, slice, tile, tile group, block, etc. For example, the subblock-based TMVP mode flag can be named sps_sbtmvp_enabled_flag. If the subblock-based TMVP mode can be applied to the current block, the subblock-based TMVP candidate can be added first to the subblock unit joint candidate list and then the affine joint candidate can be added to the subblock unit joint candidate list. Meanwhile, the maximum number of candidates that can be included in the subblock unit joint candidate list can be signaled. For example, the maximum number of candidates that can be included in the subblock unit joint candidate list is 5. The subblock size used to derive the combined candidate list of subblock units can be signaled or specified to be MxN. For example, MxN is 8x8. Thus, only if the current block size is 8x8 or larger, affine mode or subblock-based TMVP mode can be applied to the current block. Next, embodiments of the prediction implementation method according to the disclosure of the present invention will be described. The following prediction implementation method can be carried out at step S410 of Figure 4 or step S630 of Figure 6. A predicted block for the current block can be generated based on motion information derived according to the prediction mode. The predicted block (prediction block) can include prediction samples (prediction sample array) of the current block. If the motion vector of the current block specifies fractional sample units, an interpolation procedure can be performed and, through this procedure, prediction samples of the current block can be derived based on reference samples in fractional sample units in the reference image. If intermediate affine prediction is applied to the current block, prediction samples can be generated based on the MV of sample units / subblocks.When bi-prediction is applied, the prediction samples derived through the weighted sum or weighted average (by phase) of the prediction samples derived based on L0 prediction (i.e., the prediction using MVL0 and the reference image in the reference image list (L0)) and the prediction samples derived based on L1 prediction (i.e., the prediction using MLV1 and the reference image in the reference image list (L1)) can be used as the prediction samples of the current block. When applying bi-prediction and the reference image used for L0 prediction and the reference image used for L1 prediction are placed in different temporal directions with respect to the current image (i.e., when they correspond to bi-prediction and bi-prediction), it can be called true bi-prediction. In the image decoding apparatus, reconstructed samples and reconstructed images can be generated based on the derived prediction samples and then a filtering procedure in 48 loops can be performed. In addition, in the image encoding apparatus, residual samples can be derived based on the derived prediction samples and image information encoding including prediction information and residual information can be performed. Bi-prediction with CU-level weights, BCW When bi-prediction is applied to the current block as described above, a prediction sample can be derived based on a weighted average. Conventionally, a bi-prediction signal (i.e., a bi-prediction sample) can be derived through a simple average of the L0 prediction signal (L0 prediction sample) and the L1 prediction signal (L1 prediction sample). That is, a bi-prediction sample is derived through an average of the L0 prediction sample based on the L0 and MVLO reference images and the L1 prediction sample based on the L1 and MVL1 reference images. However, according to the present disclosure, when applying bi-prediction, a bi-prediction signal (bi-prediction sample) can be derived through a weighted average of the L0 prediction signal and the L1 prediction signal as follows. Equation 3 ^bi-pred = ((8 - W) * Po+ W * Pi + 4) » 3 In Equation 3 above, Ppi-pred denotes the bi-prediction signal (bi-prediction block) derived by weighted average and Po and Pi denote the L0 prediction sample (L0 prediction block) and Ll prediction sample (Ll prediction block), respectively. In addition, (8-w) and w denote the weights applied to Po and Pi, respectively. In generating bi-prediction signals by weighted average, five weights can be allowed. For example, the weight (w) can be selected from {-2,3,4,5,10}. For each bi-directionally predicted CU, the weight (w) can be determined by one of two methods. As the first of the two methods, if the current CU is not in the combined mode (non-combined CU), the weight index can be signaled together with the motion vector difference. For example, the bit stream can include information about the weight index after the information about the motion vector difference. As the second of the two methods, if the current CU is in the combined mode (combined CU), the weight index can be derived from adjacent blocks based on the combined candidate index (combined index). The bi-prediction signal generation by weighted averaging can be limited to be applied only to CUs that have a size that includes 256 or more samples (luma component samples). That is, bi-prediction by weighted averaging can be performed only on CUs where the product of the current block width and height is 256 or more. In addition, the weight (w) can be used as one of the five weights as described above and any one of a different number of weights can be used. For example, according to the characteristics of the current image, five weights can be used for images with low delay and three weights can be used for images without low delay. In this case, the three weights are {3,4,5}. Image encoding tools can determine the weight index without significantly increasing complexity by applying a fast search algorithm. In this case, the fast search algorithm can be summarized as follows. Furthermore, unequal weights may mean that the weights applied to P0 and P1 are not equal. Equal weights may mean that the weights applied to P0 and P1 may be equal. - In case the AMVR mode whose motion vector resolution is adaptively changed is applied together, if the current image is a low-delay image, only the unequal weights can be checked conditionally for each 1-pel motion vector resolution and 4-pel motion vector resolution respectively. - In case the affine mode is applied together and the affine mode is selected as the optimal mode of the current block, the image encoding tool can perform affine motion estimation (ME) for each unequal weight. - If the two reference images used for bi-prediction are equivalent, only unequal weights can be checked conditionally. - Unequal weights cannot be checked if the specified conditions are met. The specified image can be based on the POC distance between the current image and the reference image, quantization parameters (QP), temporal levels, etc. The BCW weight index can be encoded using one context-encoded bin and one or more subsequent shortcut-encoded bins. The first context-encoded bin determines whether equal weights are used. If unequal weights are used, additional bins can be shortcut-encoded and signaled. Additional bins can be signaled to determine the weights used. Weighted prediction (WP) is a tool for efficiently encoding images including fading. According to the weighted prediction, weighting parameters (weight and offset) can be signaled for each reference image included in each reference image list (L0 and L1). Then, when motion compensation is performed, the weight and offset can be applied to the corresponding reference images. Weighted prediction and BCW can be used for different image types. To avoid the interaction between weighted prediction and BCW, the BCW weight index can be signaled for the CU that uses weighted prediction. In this case, the weight can be summed to 4. That is, equal weights can be applied. In the case of CUs implementing merge mode, the weight index can be inferred from adjacent blocks based on the merge candidate index. This can be applied in both general merge mode and legacy affine merge mode. In the case of the constructed affine merge mode, the affine motion information can be configured based on the motion information of a maximum of three blocks. The BCW weight index for a CU using the constructed affine merge mode can be set at the BCW weight index of the first CP in the combination. That is, BCW can not be applied to a CU encoded in CIIP mode. For example, the BCW weight index of a CU encoded in CIIP mode can be set at a value that specifies an equivalent weight. Bi-directional optical flow (BDOF) According to the present disclosure, BDOF can be used to smooth bi-prediction signals. BDOF is to generate prediction samples by calculating smoothed motion information when bi-prediction is applied to the current block (e.g., CU). Thus, the process of calculating smoothed motion information by applying BDOF can be included in the motion information derivation step described above. For example, BDOF can be applied at the 4x4 sub-block level. That is, BDOF can be performed within the current block in 4x4 sub-block units. BODF can, for example, be applied to CUs that meet at least one or all of the following conditions. - CU is encoded in true bi-prediction mode, that is, one of the two reference images precedes the current image in order of appearance and the other reference image follows the current image in order of appearance - CU is not in affine mode or ATMVP combined mode — CU has more than 64 luma samples - the height and width of the CU is 8 luma samples or more - BCW weight index specifies equal weights, that is, applying equal weights to the L0 prediction samples and the L1 prediction samples weighted prediction (WP, Weighted Prediction) is not applied to the current CU CIIP mode is used for the current CU In addition, BDOF may be applied only to the luma component. However, the present disclosure is not limited thereto and BDOF may be applied to the chroma component or both the luma component and the chroma component. The BDOF mode is based on the concept of optical flow. That is, it assumes that the motion of the object is smooth. When applying BDOF, for each 4x4 sub-block, the motion smoothing (vx, vy) can be calculated. The motion smoothing can be calculated by minimizing the difference between the L0 predicted sample and the L1 predicted sample. The motion smoothing can be used to adjust the value of the bi-prediction samples within the 4x4 sub-block. Hereinafter, the process of performing BDOF will be described in more detail. . di(k),. . First, the horizontal gradient ——(l,j) and vertical gradient(k)—— (l,j) of the two predicted signals can be calculated. In this case, k can be 0 or 1. The gradient can be calculated by directly calculating the difference between two adjacent samples. For example, the gradient can be calculated as follows. Equation 4 di(k) / _____ _______ —(l,j) = ((l(kKi + 1,j) » shiftl) - [( / ](k)(l - 1,j) » shlft1)) (ff) = (G(k)(l,j + 1) » shiftl) - (l(kKl,j - 1) » shiftl)) In Equation 4 above, I(k)(i, j) represents the sample value of coordinate (i, j) of the prediction signal in list k (k = 0, 1). For example, I(0)(i, j) can represent the sample value at position (i, j) in the prediction block L0, and I(1)(i, j) can represent the sample value at position (i, j) in the prediction block L1. In Equation 4 above, the first shift shift1 can be determined based on the bit depth of the luma component. For example, if the bit depth of the luma component is bitDepth, shift1 can be determined to be max(6, bitDepth-6). As described above, after calculating the gradients, the auto-correlations and cross-correlations S1, S2, S3, S5 and S6 between the gradients can be calculated as follows. Equation 5 S1 = ^ij^^bs^O), S3 = Σ^Ο · SignOMiJ)) S2 = Σ ^(l'J) ·Sign(^(i'J)) (i,j)^Ω s5 = ^ij)^Abs (^(lj)) S6= EftjMW · ^y(l,j) of mana x(ij) Ψy(i,D / d!(1)d / (o)\(l,J^ + ^—(l,j^ \ dx dx / (1)(o) (l'j') + -s—(l>j')\ dy dy / » na» aa= ( / ω(ύ;) » nb) - ( / (0)(iJ) »x nb-6 in the window around which Ω is 4x4. In Equation 5 above, the rit pitch can be set to min( 1, bitDepth-11 ) and min( 4, bitDepth-8), respectively. The motion smoothing (vx, vy) can be derived as follows using the auto-correlation and cross-correlation between gradients described above. Equation 6 vx= S±> 0? clip3 (~th'Bl0, th'Bl0, -((S3» Llog2SJ)) : 0 = S5> 0? clip3 (—thBI0, th'BI0, - 2n»-n“ - ^vxS2im) « nS2+ vxS2^ / 2) by) = rnd I iy where S2 m= S2» η$2, S_(2, s) = S_2&(2^(n_(S_2 ) ) — 1), th'BI0= 213 BDand [] is the floor function. In Equation 6 above, ns2 can be 12. Based on the derived motion smoothing and gradient, the following adjustments can be made to each sample in the 4x4 sub-block. Equation 7 dl y) \ ---1 + |--dx dx J * \ dy dy / Finally, the predBDor prediction sample of CU, to which BDOF is applied, can be calculated by setting the biprediction sample of CU as follows. Equation 8 predBD0F(x,y) = (lw(x,y) + Ιω(χ,γ) + b(x,y) + ooffset) » shift In the above equation, na, nt, and ns2 can be 3, 6, and 12, respectively. These values ​​can be selected so that the multiplier does not exceed 15 bits in the BDOF process and the bit-width of the intermediate parameters is kept within 32 bits. To derive the gradient value, a predicted sample I(k)(i, j) in the list k (k=0, 1) that already exists outside the current CU can be generated. Figure 17 is a screen illustrating a CU extended to perform BDOF. As shown in Figure 17, to perform BDOF, rows / columns extending around the boundaries of the CU can be used. To control the computational complexity of generating prediction samples outside the boundaries, prediction samples in the extended region (white area in Figure 17) can be generated using a bilinear filter, and prediction samples in the CU (grey area in Figure 17) can be generated using a normal 8-tap motion-compensated interpolation filter. Sample values ​​at the extended positions can be used only for gradient calculations. If sample values ​​and / or gradient values ​​lying outside the boundaries of the CU are needed to perform the remaining steps of the BDOF process, the nearest adjacent sample values ​​and / or gradient values ​​can be overlaid (repeated) and used. If the width and / or height of a CU is greater than 16 luma samples, the corresponding CU can be split into sub-blocks having a width and / or height of 16 luma samples. The boundaries of the sub-blocks can be treated in the same way as the CU boundaries described above in the BDOF process. The maximum unit size on which the BDOF process is performed can be limited to 16x16. For each subblock, whether BDOF is performed can be determined. That is, the BDOF process for each subblock can be skipped. For example, if the SAD value between the initial L0 prediction sample and the original L1 prediction sample is less than a predetermined threshold, the BDOF process can not be applied to the corresponding subblock. In this case, if the width and height of the corresponding subblock are W and H, the predetermined threshold can be set at (8 * W * ( H >> ). Considering the complexity of additional SAD calculation, the SAD between the initial L0 prediction sample and the original L1 prediction sample calculated in the DMVR process can be reused. If BCW is available for the current block, for example, if the BCW weight index specifies unequal weights, BDOF may not be applied. Similarly, if WP is available for the current block, for example, if luma_weight_lx_flag for at least one of the two reference images is 1, BDOF may not be applied. In this case, luma_weight_lx_flag may be information specifying whether the weighting factor of WP for the luma component with predicted lx (x is 0 or 1) is present in the bit stream or information specifying whether WP is applied to the luma component with predicted lx. If the CU is encoded in symmetric MVD (SMVD) or CIIP mode, BDOF may not be applied. Prediction Refinement with Optical Flow (PROF) Hereinafter, a method of smoothing a predicted block-based affine motion compensation by applying optical flow will be described. Prediction samples generated by performing subblock-based affine motion compensation can be smoothed based on differences derived by optical flow equations. Such smoothing of prediction samples can be called prediction smoothing with optical flow (PROF) in the present disclosure. With PROF, intermediate predictions of pixel-level granularity can be achieved without increasing the bandwidth of memory access. The parameters of the affine motion model can be used to derive the motion vector of each pixel in the CU. However, since pixel-based affine motion compensation prediction results in high complexity and increases the memory access bandwidth, sub-block-based affine motion compensation prediction can be performed. When sub-block-based affine motion compensation prediction is performed, the CU can be divided into 4x4 sub-blocks and the motion vector can be determined for each sub-block. In this case, the motion vector of each sub-block can be derived from the CPMV of the CU. Sub-block-based affine motion compensation is a trade-off between coding efficiency and the complexity and bandwidth of memory access. Since the motion vector is derived in sub-block units, the complexity and bandwidth of memory access are reduced but the prediction accuracy is degraded. Thus, motion compensation of refined granularity can be achieved through smoothing by applying optical flow to the sub-block-based affine motion compensation prediction. As described above, the predicted luma samples can be refined by adding differences derived by the optical flow equations after performing sub-block-based affine motion compensation. More specifically, PROF can be performed in the following four steps. Step 1) The predicted sub-block I(i, j) is generated by performing sub-block-based affine motion compensation. Step 2) The spatial gradients gx(i, j) and gy(i, j) of the predicted subblock are calculated at each sample position. In this case, a 3-tap filter can be used, and the filter coefficients can be [-1, 0, 1]. For example, the spatial gradient can be calculated as follows. Equation 9 gx(i, j) = I(i + 1,j)- / (i -1,j) gy(W = KU +1) - KU - 1) To calculate the gradient, the predicted sub-block can be extended by one pixel on each side. In this case, to reduce bandwidth and memory complexity, pixels of the extended boundary can be copied from the nearest integer pixel in the reference image. This eliminates the need for additional interpolation for the overlay region. Step 3) The luma prediction refinement (AI(i, j)) can be calculated by the optical flow equation. For example, the following equation can be used. Equation 10 M(i,j) = gx(i,j) * Δνχ(ί, j) + gy(i,j) * Δvy(i, j) In the above equation, Av(i, j) represents the difference between the pixel motion vector (pixel MV, v(i, j)) calculated at sample position (i, j) and the sub-block MV of the sub-block, to which sample (i, j) belongs. Figure 18 is a view illustrating the relationship between Av(i, j), v(i, j) and the sub-block motion vectors. In the example shown in Figure 18, for example, the difference between the motion vector v(i, j) at the top-left sample position of the current sub-block and the motion vector vSB of the current sub-block can be represented by a bold dotted arrow, and the vector represented by the bold dotted arrow can correspond to Δv(ί, j). The affine model parameters and the pixel positions of the sub-block centers are not changed. Thus, Δν(ί, j) can be calculated only for the first sub-block and can be reused for other sub-blocks in the same CU. Assuming that the horizontal offset and The vertical offset from the pixel position to the sub-block center is x and y respectively, Δν(χ, y) can be derived as follows. Equation 11 ( hv ^.y-^y { / h In the above, (vox, voy), (vix, viy) and (v2x, v2y) correspond to the top-left CPMV, top-right CPMV and bottom-left CPMV, respectively, and w and h represent the width and height of the CU, respectively. Step 4) Finally, the final prediction block I'(i, j) can be generated based on the calculated luma prediction smoothing ΔΙ(ι, j) and the predicted sub-blocks I(i, j). For example, the final prediction block I' can be generated as follows. Equation 12 I '(i,j) = ((1,) + ^((1,) Big picture sub-picture As described above, quantization and dequantization of the luma and chroma components can be performed based on the quantization parameters. In addition, a single image to be encoded can be divided into units of a number of CTUs, slices, tiles or blocks, and, furthermore, the image can be divided into units of a number of subimages. Within an image, subimages can be encoded or decoded regardless of whether the preceding subimages are encoded or decoded. For example, different quantizations or different resolutions can be applied to a number of subimages. Furthermore, subimages can be processed like individual images. For example, the image to be encoded can be a projected image or an image encapsulated in a 360-degree image / video or an omnidirectional image / video. In this embodiment, a portion of an image may be rendered or displayed based on the display port of a user terminal (e.g., a head-worn display). Thus, to implement low latency, at least one subimage covering the display port among the subimages that construct a single image may be encoded or decoded preferentially or independently of the remaining subimages. The encoded result of a subimage can be referred to as a bitstream, sub-stream, or simply a bitstream. Decoding tools can decode subimages from a bitstream, sub-stream, or bitstream. In this case, high-level syntax (HLS) such as PPS, SPS, VPS, and / or DPS (Decoding Parameter Set) can be used to encode / decode the subimage. In the present disclosure, high-level syntax (HLS) may include at least one of APS syntax, PPS syntax, SPS syntax, VPS syntax, DPS syntax, or chunk header syntax. For example, APS (APS syntax) or PPS (PPS syntax) may include information / parameters that are generally applicable to one or more chunks or images. SPS (SPS syntax) may include information / parameters that are generally applicable to one or more sequences. VPS (VPS syntax) may include information / parameters that are generally applicable to a plurality of layers. DPS (DPS syntax) may include information / parameters that are generally applicable to an entire video. For example, DPS may include information / parameters related to the concatenation of a coded video sequence (CVS). Definition of subimage Subimages can construct rectangular areas of the encoded image. The sizes of subimages can be specified differently within the image. For all images belonging to a sequence, the sizes and positions of specific individual subimages can be specified equally. Individual subimage sequences can be decoded independently. Tiles and slices (and CTBs) can be constrained to not extend along subimage boundaries. For this purpose, the encoding apparatus can perform encoding such that the subimages are decoded independently. For this purpose, semantic constraints in the bit stream may be required. In addition, for each image in a sequence, the arrangement of tiles, slices, and blocks in the subimages can be constructed differently. The purpose of subimage design Subpicture design aims at abstraction or encapsulation of a range less than the picture level or greater than the slice or tile group level. Therefore, the VCL NAL units of a subset of the motion constant tile set (MCTS) can be extracted from one VVC bitstream and relocated to another VVC bitstream without the hassle of modification at the VCL level. Here, MCTS is an encoding technology that enables spatial and temporal independence between tiles. When MCTS is applied, information about tiles not covered by the MCTS to which the current tile belongs cannot be referenced. When the image is split into MCTS and encoded, independent transmission and decoding of the MCTS are possible. The subimage design has the advantage of changing the viewing direction in a mixed-resolution 360° streaming scheme depending on the display port. Subimage use cases The use of subimages is required in 360° schemes that rely on a viewport that provides a true spatial resolution extended at the viewport. For example, a scheme in tiles covering the viewport, derived from a 6K (6144*3072) ERP (Equi Rectangular Projection) or Cube Map Projection (CMP) resolution image that has equivalent 4K (HEVC level 5.1) decoding performance is covered in Sections D.6.3 and D.6.4 of OMAF and used in the VR Industry Forum Guidelines. These resolutions are known to be suitable for head-worn displays using quad-HD (2560x1440) display panels. Encoding: Content can be encoded with two spatial resolutions including a resolution having a cube face size of 1656x1536 and a resolution having a cube face size of 768x768. In both bit streams, a 6x4 tile grid can be used and MCTS can be encoded at each tile position. Streamed MCTS selection: 12 MCTS can be selected from the high-resolution bit stream, and an additional 12 MCTS can be obtained from the low-resolution bit stream. Therefore, a hemisphere (180°*180°) of streamed content can be generated from the high-resolution bit stream. Decoding using MCTS and bitstream concatenation: MCTS from single time events are received, which can be concatenated into a encoded image having a resolution of 1920x4608 subject to HEVC level 5.1. In another option for the concatenated image, four tile columns have a width value of 768, two tile columns have a width value of 384 and three tile rows have a height value of 768, thus constructing an image consisting of 3840x2304 luma samples. Here, the units of width and height can be units of the number of luma samples. Subpicture signaling Signaling of subimages can be done at the SPS level as shown in Figure 19. Figure 19 shows the syntax for signaling subimage syntax elements at the SPS. Hereafter, the syntax elements of Figure 19 will be explained. The pic_width_max_in_luma_samples syntax element can specify the maximum width of each decoded image with respect to SPS in luma sample units. The value of pic_width_max_in_luma_samples is greater than 0, and can have values ​​that are integer multiples of andMinCbSizeY. Here, MinCbSizeY is a variable that specifies the minimum size of the luma component encoding block. The pic_height_max_in_luma_samples syntax element can specify the maximum height of each decoded image with respect to SPS in luma sample units. pic_height_max_in_luma_samples is greater than 0 and can have values ​​that are integer multiples of MinCbSizeY. The subpic_grid_col_width_minus1 syntax element can be used to specify the width of each element of the subpicture identifier grid. For example, subpic_grid_col_width_minus1 can specify the width of each element of the subpicture identifier grid in units of 4 samples, and the value obtained by adding 1 to subpic_grid_col_width_minus1 can specify the width of the individual elements of the subpicture identifier grid in units of 4 samples. The length of the syntax element can be Ceil( Log2( pic_width_max_in_luma_samples / 4) ) bit length. Therefore, the variable NumSubPicGridCols which specifies the number of columns in the subpicture grid can be derived as follows. NumSubPicGridCols = ( pic_width_max_in_luma_samples + subpic_grid_col_width_minus1 * 4 + 3 ) / ( subpic_grid_col_width_minus1 * 4 + 4 ) The subpic_grid_row_height_minus1 syntax element can be used to specify the height of each element of the subpicture identifier grid. For example, subpic_grid_row_height_minus1 can specify the height of each element of the subpicture identifier grid in units of 4 samples. The value obtained by adding 1 to subpic_grid_row_height_minus1 can specify the height of an individual element of the subpicture identifier grid in units of 4 samples. The length of the syntax element can be Ceil( Log2( pic_height_max_in_luma_samples / 4) ) bit length. Therefore, the variable NumSubPicGridRows which specifies the number of rows in the subpicture grid can be derived as follows. NumSubPicGridRows = ( pic_height_max_in_luma_samples + subpic_grid_row_height_minus1 * 4 + 3 ) / ( subpic_grid_row_height_minus1 * 4 + 4 ) The syntax element subpic_grid_idx[ i ][ j ] can specify the subimage index of grid position (i, j). The length of the syntax element can be Ceil( Log2( max_subpics_minus1 + 1 ) ) bits. The variables SubPicTop[ subpic_grid_idx[ i ][ j ] ], SubPicLeft[ subpic_grid_idx[ i ][ j ] ], SubPicWidth[ subpic_grid_idx[ i ][ j ] ], SubPicHeight[ subpic_grid_idx[ i ][ j ] ] and NumSubPics can be derived as in the algorithm of Figure 20. The syntactic element subpic_treated_as_pic_flag[ i ] can specify whether the subimage is treated the same as a normal image in the decoding process other than the in-loop filtering process. For example, the first value (e.g., 0) of subpic_treated_as_pic_flag[ i ] can specify that the i-th subimage of each encoded image in CVS is not treated as an image in the decoding process other than the in-loop filtering process. The second value (e.g., 1) of subpic_treated_as_pic_flag[ i ] can specify that the i-th subimage of each encoded image in CVS is treated as an image in the decoding process other than the in-loop filtering process. If the value of subpic_treated_as_pic_flag[ i ] is not obtained from the bit stream, the value of subpic_treated_as_pic_flag[ i ] can be derived as the first value (e.g., 0). The syntax element loop_filter_across_subpic_enabled_flag[ i ] can specify whether in-loop filtering is performed along the boundary of the i-th subimage belonging to the individual encoding image of the CVS. For example, the first value (e.g., 0) of loop_filter_across_subpic_enabled_flag[ i ] can specify that in-loop filtering is not performed along the boundary of the i-th subimage belonging to the individual encoding image of the CVS. The second value (e.g., 1) of loop_filter_across_subpic_enabled_flag[ i ] can specify that in-loop filtering is allowed along the boundary of the i-th subimage belonging to the individual encoding image of the CVS. If the value of loop_filter_across_subpic_enabled_flag[ i ] is not obtained from the bit stream, the value of loop_filter_across_subpic_enabled_flag[ i ] can be derived as a second value. Meanwhile, for the bit stream conformation, the following constraints can be applied. For any two subimages subpicA and subpicB, if the index of subpicA is less than the index of subpicB, all the NAL units encoded from subpicA can have a lower decoding order than all the NAL units encoded from subpicB. Alternatively, after decoding is performed, the shape of the subimage needs to have a perfect left boundary and a perfect upper boundary that construct the boundary of the image or the boundary of the previously decoded subimage. The big picture of subimage-based encoding and decoding The following disclosure relates to the encoding / decoding of the images and / or subimages described above. The encoding apparatus may encode the current image based on the structure of the subimage. Alternatively, the encoding apparatus may encode at least one subimage that constructs the current image and output a (sub-)bit stream that includes (encoded) information about the at least one (encoded) subimage. The decoding apparatus can decode at least one subimage included in the current image based on a (sub-)bit stream that includes (encoded) information about at least one subimage. Figure 21 is a view illustrating a method of encoding an image using subimages by an encoding apparatus according to an embodiment. The encoding apparatus may separate an input image into a plurality of subimages (S2110). The encoding apparatus may encode at least one subimage using information about the subimages (S2110). For example, each subimage may be independently separated and encoded using information about the subimages. Further, the encoding apparatus may output a bit stream by encoding image information that includes information about the subimages (S2130). Herein, the bit stream for the subimages may be called a sub-stream or bit substream. Information about subimages will be described variously in the present disclosure, and, for example, there may be information about whether in-loop filtering can be performed along the boundaries of the subimages, information about the area of ​​the subimages, information about grid spacing for using the subimages, etc. Figure 22 is a view illustrating a method of decoding an image using a subimage by a decoding apparatus according to an embodiment. The decoding apparatus may obtain information about the subimage from a bit stream (S2210). Further, the decoding apparatus may derive at least one subimage (S2220) and decode at least one subimage (S2230). In this manner, the decoding apparatus may decode at least one subimage and thereby output at least one subimage or a currently decoded image that includes at least one subimage. The bit stream may include substreams or sub-bit streams for the subimages. As described above, information about a subimage can be constructed on HLS from a bit stream. The decoding apparatus can derive at least one subimage based on the information about the subimage. The decoding apparatus can decode the subimage based on a CABAC method, a prediction method, a residual processing method (transformation, quantization), an in-loop filtering method, etc. When decoded subpictures are output, the decoded subpictures can be output together in the form of an OPS (Output Subpicture Set). For example, if the current image corresponds to a 360° or omnidirectional image / video and is rendered partially, only some of the subpictures can be decoded and some or all of the decoded subpictures can be rendered according to the user's display port. If information specifying whether in-loop filtering along subimage boundaries is available specifies availability, the decoding apparatus may perform in-loop filtering (e.g., a block filter) on the subimage boundaries located between two subimages. However, if the subimage boundaries are the same as the image boundaries, the in-loop filtering process for the subimage boundaries may not be applied. The present disclosure relates to subimage-based encoding / decoding. Hereinafter, BDOF and PROF, to which embodiments of the present disclosure may be applied, will be described in more detail. As described above, by applying BDOF to the intermediate prediction process to refine the reference samples in the motion compensation process, it is possible to improve the compression performance of the image. BDOF can be performed in normal mode. That is, BDOF is not performed in the case of affine mode, GPM mode, or CIIP mode. Figure 23 is a screen illustrating the process of deriving a prediction sample from the current block by applying BDOF. The BDOF-based intermediate prediction procedure of Figure 23 can be performed by both image encoding and image decoding equipment. First, in step S2310, motion information of the current block may be derived. The motion information of the current block may be derived by various methods described in the present disclosure. For example, the motion information of the current block may be derived by a regular combined mode, an MMVD mode, or an AMVP mode. The motion information may include bi-prediction motion information (L0 motion information and L1 motion information). For example, the L0 motion information may include MVL0 (L0 motion vector) and refIdxL0 (L0 reference image index), and the L1 motion information MVL1 (L1 motion vector) and refIdxL1 (L1 reference image index). After that, the prediction sample of the current block can be derived based on the information derived from the current block (S2320). Specifically, the L0 prediction sample for the current block can be derived based on the L0 motion information. In addition, the L1 prediction sample for the current block can be derived based on the L1 motion information. Thereafter, a BDOF offset may be derived based on the derived prediction sample (S2330). The BDOF of step S2330 may be performed according to the method described in the present disclosure. For example, the BDOF offset may be derived based on the gradient (by phase) of the prediction sample L0 and the gradient (by phase) of the prediction sample L1. After that, based on the LX (X = 0 or 1) prediction sample and BDOF offset, the refined prediction sample of the current block can be derived (S2340). The refined prediction sample can be used to generate the final prediction block of the current block. The image encoding apparatus can derive residual samples by comparison with the original samples based on the predicted samples of the current block generated according to the method of Figure 23. Information (residual information) about the residual samples can be included and encoded in the image / video information and output in the form of a bit stream as described above. In addition, the image decoding apparatus can generate a reconstructed current block based on the predicted samples of the current block generated according to the method of Figure 23 and the residual samples obtained based on the residual information in the bit stream, as described above. Figure 24 is a view illustrating the input and output of a BDOF process according to an embodiment of the present disclosure. As shown in Figure 24, the input from the process BDOF may include the width nCbW of the current block, height CbH), prediction subblocks predSamplesL0 and predSamplesL1 with a boundary area extended by a predetermined length (e.g., 2), prediction direction indices predFlagL0 and predFlagL1 and reference image indices refIdxL0 and refIdxL1. In addition, the input of the BDOF process may further include the BDOF utilization flag bdofUtilizationFlag. In this case, the BDOF utilization flag may be included in the subblock unit within the current block to specify whether BDOF is applied to the corresponding subblock. In addition, the BDOF process can generate refined prediction blocks pbSamples by applying BDOF based on input information. Figure 25 is a view illustrating the variables used for the BDOF process according to an embodiment of the present disclosure. Figure 25 may be a process following Figure 24. As shown in Figure 25, to perform the BDOF process, the input bit depth bitDepth of the current block can be set to BitDepthY. In this case, BitDepthY can be derived based on the information about the bit depth signaled through the bit stream. In addition, various appropriate shifts can be specified based on the bit depth. For example, the first shift shift1, the second shift shift2, the third shift shift3 and the fourth shift shift4 can be derived based on the bit depth as shown in Figure 24. In addition, the offset offst4 can be specified based on shift4. In addition, the variable mvRefineThres to specify the cutting range of motion smoothing can be specified based on the bit depth. The use of the various variables described in Figure 24 will be explained below. Figure 26 is a view illustrating a method of generating prediction samples for each subblock in the current CU based on whether to apply BDOF according to an embodiment of the present disclosure. Figure 26 may be a process following Figure 25. The process shown in Figure 26 can be performed for each subblock in the current CU and, in this case, the size of the subblock can be 4x4. If the BDOF utilization flag bdofUtilizationFlag for the current subblock is the first value (false, “0), BDOF can not be applied to the current subblock. In this case, the prediction sample of the current subblock is derived by the weighted sum of the L0 prediction sample and the L1 prediction sample, and, in this case, applying the weight to the L0 prediction sample and applying the weight to the L1 prediction sample can be the same. The shift4 and offset used in Equation (1) of Figure 26 can be the values ​​specified in Figure 17. If the BDOF utilization flag bdofUtilizationFlag for the current subblock is the second value (true, “1), BDOF can be applied to the current subblock. In this case, the prediction sample of the current subblock can be generated by the BDOF process according to the present disclosure. Figure 27 is a view illustrating a method of deriving gradients, auto-correlations and cross-correlations from the current subblock according to an embodiment of the present disclosure. Figure 27 may be a process following Figure 26. The process shown in Figure 27 is performed for each subblock in the current CU and, in this case, the subblock size can be 4x4. According to Figure 27, according to Equation (1) and Equation (2), the position (hx, hy) for each sample position (x, y) in the current sub-block can be derived. After that, the horizontal gradient and vertical gradient for each sample position can be derived according to Equation (3) to Equation (6). After that, the variables (the first intermediate parameter diff and the second intermediate parameters tempH and tempV) for deriving auto-correlation and cross-correlation can be derived according to Equation (7) to Equation (9). For example, the first intermediate parameter diff can be derived using the value obtained by applying an appropriate shift to the predicted samples predSamplesL0 and predSamplesL1 of the current block with a second shift shift2. For example, the second intermediate parameter tempH and tempV can be derived by applying an appropriate shift to the sum of the gradients in the direction L0 and the gradient in the L1 direction with the third shift shift3 as in Equation (8) and Equation (9). After that, the auto-correlation and cross-correlation can be derived based on the first intermediate parameter and the second intermediate parameter derived according to Equation (10) to Equation (16). Figure 28 is a view illustrating a method of deriving motion smoothing (vx, vy), deriving BDOF offsets and generating prediction samples from the current sub-block, according to an embodiment of the present disclosure. Figure 28 may be a process following Figure 27. The process shown in Figure 28 is performed for each subblock in the current CU and, in this case, the subblock size can be 4x4. According to Equation 28, the motion smoothing (vx, vy) can be derived according to Equation (1) and Equation (2). The motion smoothing can be truncated within the range specified by mvRefineThres. In addition, based on the motion smoothing and gradient, the BDOF offset bdofOffset can be derived according to Equation (3). The prediction samples pbSamples of the current sub-block can be generated using the BDOF offset derived according to Equation (4). By continuously performing the method described with reference to Figures 24 to 28, the BDOF process according to the first embodiment of the present disclosure can be implemented. In the embodiment according to Figures 24 to 28, the first shift shift1 is set at Max(6, bitDepth-6), and mvRefineThres is set at 1< <Max(, bitDepth-7). Dengan demikian, lebar bit dari predSample dan tiap-tiap parameter BDOF menurut BitDepth dapat diderivasi seperti yang ditunjukkan pada tabel berikut. Table 1 BitDepth predSample Shift1 Gradient Vx, Vy bdofOffset 8 16 6 11 6 17 [-25022, [-779, [-32, [-49856, 24958] 779] 31] 48298] 10 16 6 11 6 17 12 16 6 11 6 17 14 18 8 11 8 19 16 20 10 11 10 21 As described above, by applying BDOF to the intermediate prediction process to refine the reference samples in the motion compensation process, it is possible to improve the compression performance of the image. BDOF can be performed in normal mode. That is, BDOF is not performed in the case of affine mode, GPM mode, or CIIP mode. PROF can be performed on blocks encoded in affine mode, as a method similar to BDOF. As described above, by smoothing the reference samples in each 4x4 sub-block through PROF, it is possible to improve the compression performance of the image. According to the present disclosure, the affine motion information (subblock motion) described above from the current block can be derived, and the affine motion information can be refined through the PROF process described above or the prediction samples derived based on the affine motion information can be refined. Figure 29 is a screen illustrating the process of deriving a prediction sample from the current block by applying PROF. The PROF-based intermediate prediction procedure of Figure 29 can be performed by both image encoding equipment and image decoding equipment. First, in step S2910, motion information of the current block may be derived. The motion information of the current block may be derived by various methods described in the present disclosure. For example, the motion information of the current block may be derived by methods described in the affine mode or the sub-block-based TMVP mode described above. The motion information may include sub-block motion information of the current block. The sub-block motion information may include bi-prediction sub-block motion information (sub-block L0 motion information and sub-block L1 motion information). For example, the sub-block L0 motion information may include sbMVL0 (sub-block L0 motion vector) and refIdxL0 (reference image index L0), and the sub-block L1 motion information may include sbMVL1 (sub-block L1 motion vector) and refIdxL1 (reference image index L1). After that, the prediction samples of the current block can be derived based on the information derived from the current block (S2920). Specifically, the L0 prediction samples for each subblock of the current block can be derived based on the motion information of the L0 subblock. In addition, the L1 prediction samples for each subblock of the current block can be derived based on the motion information of the L1 subblock. Thereafter, a PROF offset can be derived based on the derived prediction sample (S2930). The PROF of step S2930 can be performed according to the method described in the present disclosure. For example, the difference motion vector diffMv and the gradient of LX (X = 0 or 1) of the prediction sample can be calculated and, based on these, a PROF offset di or ΔΙ can be derived according to the method described in the present disclosure. Various examples of the present disclosure relate to the derivation of the difference motion vector, the derivation of the gradient and / or the derivation of the PROF offset. After that, based on the LX (X = 0 or 1) prediction samples and the PROF offset, the refined prediction samples of the current block can be derived (S2940). The refined prediction samples can be used to generate the final prediction block of the current block. For example, the final prediction block of the current block can be generated by weight-summing the refined L0 prediction samples and the refined L1 prediction samples. The image encoding apparatus can derive residual samples by comparison with the original samples based on the predicted samples of the current block generated according to the method of Figure 29. Information (residual information) about the residual samples can be included and encoded in the image / video information and output in the form of a bit stream as described above. In addition, the image decoding apparatus can generate a reconstructed current block based on the predicted samples of the current block generated according to the method of Figure 29 and the residual samples obtained based on the residual information in the bit stream, as described above. Figure 30 is a view illustrating an example of a PROF process according to the present disclosure. According to the example of Figure 30, the PROF process can be performed using the width sbWidth, the height sbHeight of the current sub-block, the prediction subblock predSamples where the boundary area extends by a predetermined length borderExtention and the difference motion vector diffMv as input. In this case, the prediction subblock can be, for example, a prediction subblock generated by performing affine motion compensation. As a result of performing the PROF process, a refined prediction subblock pbSamples can be generated. To perform the PROF process, the first shift of a predetermined shift can be calculated. The first shift can be derived based on the BitDepthy bit depth of the luma component. For example, the first shift can be derived as the maximum value of 6 and (BitDepthy - 6) . After that, the horizontal gradient gradientH,gxand the vertical gradient gradientV,gycan be calculated for each sample position (x,y) of the input prediction subblock. The horizontal gradient and the vertical gradient can be calculated according to Equation (1) and Equation (2) of Figure 30, respectively. After that, based on the horizontal gradient, vertical gradient and the difference motion vector diffMv, the PROF offset di or ΔΙ for each sample position can be calculated. For example, the PROF offset can be calculated according to Equation (3) of Figure 30. In Equation (3), the difference motion vector diffMv used to calculate the PROF offset can mean Δν described with reference to Figure 18. In this case, diffMv can be truncated by dmvLimit as follows, and dmvLimit can be calculated based on BitDepthy as follows. Equation 13 diffMvfx / [y / / i / = Clip3( -dmvLimit, dmvLimit - 7, diffMvf x Hy / / i After that, the refined prediction subblock pbSamples can be derived based on the calculated PROF offset and the prediction subblock predSamples. For example, the refined prediction subblock can be derived according to Equation (4) of Figure 30. According to the example from Figure 30, the first shift (shiftl) can be set at Max(6, bitDepth-6), and dmvLimit can be set at 1< <Max(5, bitDepth-7). Selain itu, diffMv dapat dipotong dalam kisaran dari [-dmvLimit, dmvLimit-1]. Dengan demikian, lebar bit dari predSample dan tiap-tiap parameter PROF menurut BitDepthY dapat diderivasi seperti yang ditunjukkan pada tabel berikut. Table 2 BitDepthY predSample Shift1 Gradient diffMv dI 8 16 [-25022, 24958] 6 11 [-779, 779] 6 [-32, 31] 17 [- 49856, 48298] 10 16 6 11 6 17 12 16 6 11 6 17 14 18 8 11 8 19 16 20 10 11 10 21 The present disclosure may provide various embodiments of the case where subimages are treated as images (e.g., subpic_treated_as_pic_flag == 1) in performing subimage-based encoding / decoding. For example, in a BDOF or PROF process, the reference sampling process may be constrained to not reference reference samples included in a subimage different from the subimage to which the current block belongs. In addition, in a temporal motion vector predictor derivation process, a bilinear interpolation process of luma samples, an 8-tap interpolation filtering process of luma samples and / or a chroma sample interpolation process, a constraint according to which subimages are treated as images may be added. As described above, in BDOF and / or PROF processes, to calculate the gradient, reference samples extended by a predetermined length around the boundary of the current block can be used. However, if the current sub-image, to which the current block belongs, is treated as an image, a range of reference samples for calculating the gradient needs to be limited within the same sub-image as the current sub-image. As described above, the predicted samples can be modified by bdofOffset in the case of BDOF and can be modified by dI in the case of PROF. bdofOffset and dI can be obtained based on the gradient. The gradient can be derived based on the difference between the reference samples in the reference image. To derive the gradient, the reference sampling process of taking reference samples from the reference image and the 8-tap interpolation filtering process can be performed. The output of the reference sampling process can be a luma sample at an integer pixel position. Figure 31 is a view illustrating the case where the reference sample to be taken crosses the boundary of a subimage. In Figure 31, the region (capture area) in the reference image to be captured can be specified according to the motion vector mv of the current block covered by the current subimage [1] in the current image. In this case, the capture area can be along the boundary of the subimage [1] in the reference image. That is, the reference sample to be captured can be covered in a subimage different from the current subimage [1]. Figure 32 is an enlarged view of the capture area of ​​Figure 31. As shown in Figure 32, the reference samples to be captured can be covered in subimages (subimage [0], subimage [2], subimage [3]) that are different from the current subimage [1]. Considering the case shown in Figure 32, when the subimage is treated as an image, the range of the reference samples to be captured needs to be limited within the same subimage as the current subimage. Figure 33 is a view illustrating the reference sampling process according to an embodiment of the present disclosure. In the reference sampling process of Figure 33, information about the motion vector of the current block and the reference image refPicLXL can be input. The information about the motion vector of the current block is to specify the sampling area and can be an integer sample position (xIntL, yIntL) derived from the motion vector of the current block. The output of the reference sampling process in Figure 33 can be a prediction block for the current block. In this case, the prediction block can be a prediction block that will be refined by BDOF or PROF. As shown in Figure 33, whether the current sub-image is treated as an image can be specified. For example, if subpic_treated_as_pic_flag is 1, it can be specified that the current sub-image is treated as an image and the capture region can be limited within the same sub-image as the current sub-image. If the current sub-image is treated as an image, as shown in Equation (1) of Figure 33, the x-coordinate xInt that specifies the position of the sample to be captured can be clipped in the range of [SubPicLeftBoundaryPos, SubPicRightBoundaryPos]. SubPicLeftBoundaryPos can specify the position of the left boundary of the current sub-image. In addition, SubPicRightBoundaryPos can specify the position of the right boundary of the current sub-image. According to Equation (1) above, since the x-coordinate of the reference sample to be taken is in the range from SubPicLeftBoundaryPos to SubPicRightBoundaryPos, reference samples located outside the left or right boundaries of the subimage are not captured. The method for deriving SubPicLeftBoundaryPos and SubPicRightBoundaryPos will be described later. Similarly, if the current sub-image is treated as an image, as shown in Equation (2) of Figure 33, the y-coordinate yInt that specifies the position of the sample to be captured can be clipped in the range of [SubPicTopBoundaryPos, SubPicBotBoundaryPos]. SubPicTopBoundaryPos can specify the position of the top boundary of the current sub-image. Additionally, SubPicBotBoundaryPos can specify the position of the bottom boundary of the current sub-image. According to Equation (2) above, the y-coordinate of the reference sample to be taken is in the range from SubPicTopBoundaryPos to SubPicBotBoundaryPos, reference samples that are already outside the upper or lower boundaries of the subimage are not captured. The method for deriving SubPicTopBoundaryPos and SubPicBotBoundaryPos will be described later. If the current sub-image is not treated as an image, for example, when subpic_treated_as_pic_flag is 0, the coordinates of the reference sample to be taken can be derived according to Equations (3) and (4). According to Equations (3) and (4) above, the coordinates of the reference sample to be taken are not truncated by the boundary position of the sub-image. According to Equations (3) and (4) above, the coordinates of the reference sample to be taken can be truncated in the range of the current image. After that, according to Equation (5), the reference sampling from the reference image can be done based on the coordinates (xInt, yInt) of the sample to be taken. Figure 34 is a flowchart illustrating the reference sampling process according to this disclosure. The reference sampling process of Figures 33 and 34 can be performed by image encoding equipment and image decoding equipment to perform BDOF and / or PROF. Referring to Figure 34, first, whether the current sub-image is treated as an image can be determined (S3410). The determination from step S3410 can be done based on the subpic_treated_as_pic_flag. When the current sub-picture is treated as an image (e.g., subpic_treated_as_pic_flag==1), a capture position can be derived (S3420), and the derived capture position can be cropped (S3430). The derivation and cropping of the capture position can be performed according to Equation (1) and Equation (2) of Figure 33. The cropping of step S3430 can be a process of changing the corresponding capture position to a position on the current sub-picture (e.g., a boundary position of the current sub-picture) when the derived capture position is outside the boundary of the current sub-picture. If the current sub-picture is not treated as an image (e.g., subpic_treated_as_pic_flag==0), the capture position can be derived (S3440). The derivation of the capture position can be done, for example, according to Equations (3) and (4) of Figure 33. After that, reference sampling can be performed based on the sampling position derived in step S3430 or step S3440 (S3450). Reference sampling can be performed, for example, according to Equation (5) of Figure 33. According to the embodiment shown in Figure 34, when a subimage is treated as an image, reference samples outside the boundaries of the current subimage may not be sampled in the reference sampling process. That is, reference samples in the reference image that belong to a subimage different from the current subimage may not be referenced. Further herein, fractional sample interpolation procedures according to other embodiments of the present disclosure will be described. As described above, if the motion vector of the current block specifies a fractional sample unit, an interpolation procedure can be performed, and the predicted sample of the current block can be derived based on the reference sample of the fractional sample unit in the reference image accordingly. Figure 35 is a view illustrating part of a fractional sample interpolation procedure according to the present disclosure. To perform a sample fractional interpolation procedure, the variables fRefWidth and fRefHeight can be derived. As shown in Figure 35, fRefWidth and fRefHeight can be derived differently according to whether the current sub-image is treated as an image (subpic_treated_as_pic_flag). If the current sub-image is treated as an image, for example, if subpic_treated_as_pic_flag is 1, fRefWidth and fRefHeight can be derived as follows. fRefWidth = (SubPicWidth[ SubPicIdx ]* ( subpic_grid_col_width_minus1 + 1 ) *4) fRefHeight = (SubPicHeight[ SubPicIdx ]* ( subpic_grid_row_height_minus1 + 1 ) *4) In the above, SubPicWidth[SubPicIdx] and SubPicHeight[SubPicIdx] can mean the width and height of the current subimage, respectively. In this case, the width and height of the subimage can be expressed by the number of grids that configure the subimage. For example, a subimage width of 4 can mean that the corresponding subimage spans four grids in the horizontal direction. In addition, subpic_grid_col_width_minus1 and subpic_grid_row_height_minus1 can mean the width and height of the grid that configure the subimage, respectively. In this case, the width and height of the grid can be expressed in units of 4 pixels. For example, a grid width of 4 can mean that the grid width is 16 (4x4) pixels. If the current sub-image is not treated as an image, for example, if subpic_treated_as_pic_flag is 0, fRefWidth and fRefHeight can be derived as the width PicOutputWidthL and height PicOutputHeightL of the reference image output image, respectively. The fRefWidth and fRefHeight derived as described above can be used to derive the scaling factors hori_scale_fp and vert_scale_fp according to Equation (1) and Equation (2) of Figure 35. After that, according to Equations (3) to Equation (6), the positions (refxL, refyL) of the reference samples specified by the motion vector can be derived. Based on the derived positions of the reference samples, the integer positions (xIntL, yIntL) and fractional positions (xFracL, yFracL) can be derived. The reference sampling process or the 8-tap interpolation filtering process described above can be performed based on the derived integer positions and / or fractional positions. Further herein, a method of deriving sbTMVP according to another embodiment of the present disclosure will be described. As described with reference to Figure 16, a motion shear can be derived and applied to the current block, thereby specifying the subblocks (subblock col) in the col image corresponding to each subblock that configures the current block. After that, using the motion information of the corresponding subblocks (subblock col) of the col image, the motion information of each subblock of the current block can be derived. Based on the information derived from the subblocks, the sbTMVP can be derived. Figure 36 is a view illustrating part of the sbTMVP derivation method according to the present disclosure. Figure 36 shows part of the process after the shear motion for the current block is derived during the sbTMVP derivation process. Specifically, the process shown in Figure 36 includes the process of deriving the positions of the corresponding col subblocks for each subblock that configures the current block. According to Figure 36, firstly, according to Equation (1) and Equation (2), the position (xSb, ySb) of the bottom-right center sample of the current sub-block can be derived. After that, the position (xColSb, yColSb) of the col sub-block can be derived according to whether the current sub-image is treated as an image (e.g., subpic_treated_as_pic_flag). Specifically, when subpic_treated_as_pic_flag is 1, the y-coordinate yColSb of subblock col can be derived according to Equation (3). In this case, yColSb derived by ySb and the movement shear tempMv can be clipped to the position specified by SubPicTopBoundaryPos and SubPicBotBoundaryPos and the y-coordinate yCtb of the current CTB. By Equation (3), the y-coordinate of subblock col is in the current CTB and the current sub-image. In addition, when subpic_treated_as_pic_flag is 0, according to Equation (4), the y-coordinate yColSb of subblock col can be derived. In this case, yColSb derived by ySb and the movement shear tempMv can be clipped to the position specified by the y-coordinate yCtb of the current CTB and the height of the current image. By equation (4), the y coordinate of subblock col is in the current CTB and the current image. Similarly, when subpic_treated_as_pic_flag is 1, according to Equation (5), the x-coordinate xColSb of the subblock col can be derived. In this case, xColSb derived by xSb and the shift motion tempMv can be clipped to the position specified by SubPicLeftBoundaryPos and SubPicRightBoundaryPos and the x-coordinate xCtb of the current CTB. By equation (5), the x-coordinate of the subblock col is in the current CTB and the current sub-picture. Additionally, if subpic_treated_as_pic_flag is 0, according Equation (6), the x-coordinate yColSb of the subblock col can be derived. In this case, xColSb derived by xSb and the shift motion tempMv can be cut to the position specified by the x-coordinate xCtb of the current CTB and the current image width. By equation (6), the x-coordinate of the subblock col is in the current CTB and the current image. According to the present disclosure, when the current sub-image is treated as an image, the col subblock of the current sub-block for deriving sbTMVP is in the same sub-image as the current sub-image. Further herein, a method of deriving the position of a subimage boundary according to another embodiment of the present disclosure will be described. Figure 37 is a view illustrating a method of deriving subimage boundary positions according to the present disclosure. According to this disclosure, the left boundary position subimage SubPicLeftBoundaryPos, the right boundary position subimage SubPicRightBoundaryPos, subimage top boundary position SubPicTopBoundaryPos and subimage bottom boundary position SubPicBotBoundaryPos may be derived. In this disclosure, SubPicIdx may be an index to identify each subimage in the current image. If the current sub-picture specified by SubPicIdx is treated as an image, SubPicLeftBoundaryPos may be derived based on the position information SubPicLeft of a predetermined unit specifying the left position of the current sub-picture and the width information subpic_grid_col_width_minus1 of the corresponding unit. The predetermined unit may be a grid. However, the present disclosure is not limited thereto and, for example, the predetermined unit may be a CTU. If the predetermined unit is a CTU, the left position of the current sub-picture may be derived as the product of the position information of the CTU unit specifying the left position of the current sub-picture and the CTU size. When the current sub-image is treated as an image, SubPicRightBoundaryPos may be derived based on the position information of a predetermined unit specifying the left position of the current sub-image, the width information of a predetermined unit specifying the width of the current sub-image and the corresponding width information of the unit. For example, the position information of a predetermined unit specifying the right position of the current sub-image may be derived by adding the left position of the current sub-image and the width of the current sub-image. SubPicRightBoundaryPos may be derived by performing a -1 operation on the last calculated value. The predetermined unit may be a grid. However, the present disclosure is not limited thereto and, for example, the predetermined unit may be a CTU.When the specified unit is CTU, the right position of the current sub-picture can be derived by performing a -1 operation on the product of the position information of the CTU unit specifying the right position of the current sub-picture and the CTU size. The position information of the CTU unit specifying the right position of the current sub-picture can be derived as the sum of the position information of the CTU unit specifying the left position of the current sub-picture and the width information of the CTU unit specifying the width of the current sub-picture. Similarly, SubPicTopBoundaryPos may be derived based on the SubPicTop position information of a predetermined unit specifying the top position of the current sub-picture and the height information subpic_grid_low_height_minus1 of the corresponding unit. The predetermined unit may be a grid. However, the present disclosure is not limited thereto and, for example, the predetermined unit may be a CTU. When the predetermined unit is a CTU, the top position of the current sub-picture may be derived as the product of the position information of the CTU unit specifying the top position of the current sub-picture and the size of the CTU. If the current sub-image is treated as an image, SubPicBotBoundaryPos may be derived based on the position information of a predetermined unit specifying the top position of the current sub-image, the height information of a predetermined position specifying the height of the current sub-image and the corresponding height of the unit. For example, the position information of a predetermined unit specifying the bottom position of the current sub-image may be derived by adding the top position of the current sub-image and the height of the current sub-image. SubPicBotBoundaryPos may be derived by performing a -1 operation on the last calculated value. The predetermined unit may be a grid. However, the present disclosure is not limited thereto and, for example, the predetermined unit may be a CTU.When the specified unit is CTU, the bottom position of the current sub-image can be derived by performing a -1 operation on the product of the position information of the CTU unit specifying the bottom position of the current sub-image and the CTU size. The position information of the CTU unit specifying the bottom position of the current sub-image can be derived as the sum of the position information of the CTU unit specifying the top position of the current sub-image and the height information of the CTU unit specifying the height of the current sub-image. The various embodiments described in the present disclosure may be implemented singly or in combination with other embodiments. Alternatively, some embodiments may be added to other embodiments or some embodiments may be replaced by some other embodiments. Although the exemplary method of the present disclosure described above is represented as a series of operations for clarity of description, it is not intended to limit the order in which the steps are performed, and the steps may be performed simultaneously or in a different order as necessary. To implement the method according to the present disclosure, the steps described may further include other steps, may include the remaining steps except for some steps, or may include other additional steps except for some steps. In the present disclosure, an image encoding apparatus or an image decoding apparatus that performs a predetermined operation (step) may perform the operation (step) confirming the execution condition or situation of the corresponding operation (step). For example, if it is described that a predetermined operation is performed When a predetermined condition is satisfied, then the image encoding apparatus or the image decoding apparatus may perform the predetermined operation after determining whether the predetermined condition is satisfied. The various embodiments of the disclosure of the invention are not a list of all possible combinations and are intended to describe representative aspects of the disclosure of the invention, and the subject matter described in the various embodiments may be applied independently or in combination with two or more. Various embodiments of the present disclosure may be implemented in hardware, firmware, software, or a combination thereof. In the case of implementing the present disclosure with hardware, the present disclosure may be implemented with an application-specific integrated circuit (ASIC), a digital signal processor (DSP), a digital signal processing device (DSPD), a programmable logic device (PLD), a field-programmable gate array (FPGA), a general-purpose processor, a controller, a microcontroller, a microprocessor, etc. In addition, image decoding apparatus and image encoding apparatus, to which embodiments of the present disclosure are applied, may include multimedia broadcasting transmitting and receiving apparatus, mobile communication terminals, home cinema video apparatus, digital cinema video apparatus, surveillance cameras, video chat apparatus, real-time communication apparatus such as video communication, mobile streaming apparatus, storage medium, camcorders (cameras and recorders), video on demand (VoD) service provider apparatus, over the top video (OTT) video apparatus, Internet streaming service provider apparatus, three-dimensional (3D) video apparatus, video telephone video apparatus, medical video apparatus, and so on, and may be used to process video signals or data signals. For example, OTT video apparatus may include game consoles, blu-ray players, Internet access TBs, home theater systems, smartphones, tablet PCs, digital video recorders (DVRs), or the like. Figure 38 is a view showing a content streaming system, to which embodiments of the present disclosure may be applied. As shown in Figure 38, a content streaming system, to which embodiments of the present disclosure are applied, may primarily include an encoding server, a streaming server, a web server, and a multimedia input device. The encoding server compresses the content input from multimedia input devices such as smartphones, cameras, camcorders, etc. into digital data to generate a bitstream and transmits the bitstream to the streaming server. As another example, if multimedia input devices such as smartphones, cameras, camcorders, etc., generate a bitstream directly, the encoding server can be omitted. The bit stream may be generated by an image encoding method or image encoding apparatus, to which embodiments of the present disclosure are applied, and the streaming server may temporarily store the bit stream in the process of transmitting or receiving the bit stream. A streaming server transmits multimedia data to a user's device based on the user's request via a web server, and the web server serves as a medium for informing the user about the service. When a user requests a desired service from a web server, the web server delivers it to the streaming server, and the streaming server transmits the multimedia data to the user. In this case, the content streaming system may include a separate control server. In this case, the control server functions to control the commands / responses between devices in the content streaming system. A streaming server can receive content from a media store and / or an encoding server. For example, if content is received from an encoding server, it can be received in real time. In this case, to provide a seamless streaming service, the streaming server can cache the bitstream for a specified period of time. Examples of user devices may include mobile phones, smart phones, laptop computers, digital broadcast terminals, personal digital assistants (PDAs), portable multimedia players (PMPs), navigation devices, slate PCs, tablet PCs, ultrabooks, wearable devices (e.g., smart watches, smart glasses, head-worn displays), digital TVs, desktop computers, digital signs, and the like. Each server in a content streaming system can be operated as a distributed server, where data received from each server can be distributed. The scope of the disclosure includes machine-executable software or instructions (e.g., operating systems, applications, firmware, programs, etc.) to enable operations in accordance with the methods of various embodiments to be executed on an apparatus or computer, a non-transitory computer-readable medium having such software or instructions stored therein and executed on the apparatus or computer. Industrial Applications Embodiments of the present disclosure may be used to encode or decode images.

Claims

Claim 1. An image decoding method performed by an image decoding apparatus, the image decoding method comprising: determining whether bi-directional optical flow (BDOF) or predictive smoothing with optical flow (PROF) is applied to the current block; based on the BDOF or PROF applied to the current block, generating a prediction sample of the current block from a reference image of the current block based on the motion information of the current block; and deriving a smoothed prediction sample for the current block, by applying the BDOF or PROF to the current block based on the generated prediction sample.

2. The image decoding method of claim 1, wherein the generation of prediction samples from the current block is performed based on whether the current sub-image encompassing the current block is treated as an image.

3. The image decoding method of claim 2, wherein whether the current subimage is treated as an image is determined based on marker information signaled through the bit stream.

4. The image decoding method of claim 3, wherein the marker information is signaled via a sequence parameter set (SPS).

5. The image decoding method of claim 2, wherein the generation of prediction samples from the current block is performed based on the positions within the reference image to generate prediction samples, and wherein the positions within the reference image are cropped within a predetermined range.

6. The image decoding method of claim 5, wherein, based on the current subimage being treated as an image, a predetermined range is specified by the boundary positions of the current subimage.

7. The image decoding method of claim 6, wherein the position within the reference image includes an x ​​coordinate and a y coordinate, wherein the x coordinate is truncated within a range from the left boundary position and the right boundary position of the current sub-image, and wherein the y coordinate is truncated within a range from the upper boundary position and the lower boundary position of the current sub-image.

8. The image decoding method of claim 7, wherein the left boundary position of the current sub-block is derived as the product of position information of a predetermined unit specifying the left position of the current sub-image and the width of the predetermined unit, wherein the right boundary position of the current sub-image is derived by performing a -1 operation on the product of position information of a predetermined unit specifying the right position of the current sub-image and the width of the predetermined unit, wherein the upper boundary position of the current sub-image is derived as the product of position information of a predetermined unit specifying the top position of the current sub-image and the height of the predetermined unit,and where the lower boundary position of the current sub-image is derived by performing a -1 operation on the position information product of the specified unit specifying the lower position of the current sub-image and the height of the specified unit., 9. The image decoding method of claim 8, wherein the predetermined unit is a grid or CTU.

10. The image decoding method of claim 5, wherein, based on the current sub-image not being treated as an image, the predetermined range is the range of the current block that includes the current block.

11. An image encoding method performed by an image encoding apparatus, the image encoding method comprising: determining whether bi-directional optical flow (BDOF) or predictive smoothing with optical flow (PROF) is applied to the current block; based on the BDOF or PROF applied to the current block, generating a prediction sample of the current block from a reference image of the current block based on the motion information of the current block; and deriving a smoothed prediction sample for the current block, by applying the BDOF or PROF to the current block based on the generated prediction sample.

12. The image encoding method of claim 11, wherein the generation of prediction samples from the current block is performed based on whether the current sub-image encompassing the current block is treated as an image.

13. A method of transmitting a bit stream generated by the image encoding method of claim 11.