Image Encoding / Decoding Method and Apparatus, and Recording Medium Storing Bitstream
By generating a candidate list for inter prediction using spatial, temporal, and induced candidates, the method addresses the challenge of accurate motion information handling in high-resolution videos, enhancing encoding and decoding efficiency.
Patent Information
- Application Number
- JP2024575696
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2022-07-06
- Filing Date
- 2023-07-03
- Publication Date
- 2025-07-17
AI Technical Summary
Existing video compression technologies face challenges in generating accurate and reliable inter prediction candidates for high-resolution and high-quality videos, particularly in handling motion information of current blocks effectively.
The method generates a candidate list for inter prediction by deriving motion information from spatial and temporal candidates, including induced candidates based on the upper right and lower left blocks of the current block, adaptively determining search directions and ranges, and selectively adding candidates to improve accuracy.
This approach enhances the accuracy and reliability of inter prediction by utilizing motion information within the current picture and adaptively determining temporal neighboring blocks, leading to improved encoding and decoding efficiency.
Smart Images

Figure 2025522759000001_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to a video encoding / decoding method and apparatus, and a recording medium storing a bitstream.
Background Art
[0002] In recent years, the demand for high-resolution and high-quality videos such as HD (High Definition) videos and UHD (Ultra High Definition) videos has been increasing in various application fields, and thus, high-efficiency video compression technologies have been discussed.
[0003] As video compression technologies, there are various technologies such as an inter prediction technology for predicting pixel values included in a current picture from pictures before or after the current picture, an intra prediction technology for predicting pixel values included in the current picture using pixel information within the current picture, and an entropy coding technology for assigning short codes to values with high occurrence frequencies and long codes to values with low occurrence frequencies. Using such video compression technologies, video data can be effectively compressed and transmitted or stored.
Summary of the Invention
Problems to be Solved by the Invention
[0004] This disclosure aims to provide a method and apparatus for generating various candidates for a candidate list for inter prediction.
[0005] This disclosure aims to provide a method and apparatus for deriving a candidate reflecting the motion of the lower right end position of the current block.
[0006] This disclosure aims to provide a method and apparatus for deriving motion information of temporal candidates.
[0007] This disclosure aims to provide a method and apparatus for adding various candidates to a candidate list.
Means for Solving the Problems
[0008] The video decoding method and apparatus according to the present disclosure generate a candidate list for a current block, derive motion information of the current block based on any one of a plurality of candidates belonging to the candidate list, and perform inter prediction on the current block based on the motion information of the current block. Here, the plurality of candidates may include at least one of a spatial candidate, a temporal candidate, or an induced candidate.
[0009] In the video decoding method and apparatus according to the present disclosure, the motion information of the induced candidate may be induced based on at least one of the upper right block of the current block or the lower left block of the current block.
[0010] In the video decoding method and apparatus according to the present disclosure, the upper right block of the current block is a block including an upper right reference position, and the lower left block of the current block may be a block including a lower left reference position. Here, the upper right reference position and the lower left reference position may be determined based on at least one of the width or height of the current block.
[0011] In the video decoding method and apparatus according to the present disclosure, the upper right block of the current block is a block including a first position moved a predetermined first distance in a left or right direction from an upper right reference position, and the lower left block of the current block may be a block including a second position moved a predetermined second distance in a lower or upper direction from a lower left reference position. Here, the first position and the second position may be located on the same straight line.
[0012] In the video decoding method and apparatus according to the present disclosure, the upper right block of the current block is a block including a first position moved a predetermined distance in a left or right direction from an upper right reference position, and the lower left block of the current block may be a block including a second position moved the predetermined distance in a lower or upper direction from a lower left reference position.
[0013] In the video decoding method and apparatus according to the present disclosure, the predetermined distance may be adaptively determined based on the size of the current block.
[0014] In the video decoding method and apparatus according to the present disclosure, the search direction for determining at least one of the upper right block or the lower left block may be determined based on whether the width of the current block is greater than the height of the current block.
[0015] In the video decoding method and apparatus according to the present disclosure, at least one of the upper right block or the lower left block may be determined by searching for an available block having motion information different from at least one of the left block or the upper block of the current block.
[0016] In the video decoding method and apparatus according to the present disclosure, the search range for determining at least one of the upper right block or the lower left block is divided into a plurality of sub - ranges, and the induced candidate may be induced for each of the plurality of sub - ranges.
[0017] In the video decoding method and apparatus according to the present disclosure, the induced candidate may be adaptively added to the candidate list based on whether the temporal candidate is available for the current block.
[0018] In the video decoding method and apparatus according to the present disclosure, the temporal candidate has motion information of a temporal neighboring block corresponding to the current block, and the temporal neighboring block may be specified by the motion information induced based on at least one of the upper right block or the lower left block.
[0019] The video encoding method and apparatus according to the present disclosure can generate a candidate list for a current block and perform inter prediction on the current block based on any one of a plurality of candidates belonging to the candidate list. Here, the plurality of candidates may include at least one of a spatial candidate, a temporal candidate, or an induced candidate.
[0020] There is provided a computer-readable digital storage medium storing encoded video / video information for performing a video decoding method by a decoding apparatus according to the present disclosure.
[0021] There is provided a computer-readable digital storage medium storing video / video information generated by a video encoding method according to the present disclosure.
[0022] There are provided a method and an apparatus for transmitting video / video information generated by a video encoding method according to the present disclosure.
Advantages of the Invention
[0023] According to the present disclosure, by utilizing the motion information within the current picture to induce a candidate in which the motion at the lower right end position of the current block is reflected, the accuracy and reliability of inter prediction can be improved.
[0024] According to the present disclosure, by selectively using at least one of a plurality of induced candidates, the accuracy and reliability of inter prediction can be improved.
[0025] According to the present disclosure, by adaptively determining the position of a temporal neighboring block for a temporal candidate, the reliability for the temporal candidate can be enhanced, and further the accuracy of inter prediction can be improved.
Brief Description of the Drawings
[0026]
Figure 1
[0027]
Figure 2
[0028]
Figure 3
[0029]
Figure 4
[0030]
Figure 5
[0031]
Figure 6
[0032]
Figure 7
[0033]
Figure 8
Embodiments for Carrying Out the Invention
[0034] The present disclosure can be modified in various ways and can have various embodiments. Specific embodiments are illustrated in the drawings and described in detail below. However, this is not intended to limit the present disclosure to specific embodiments, and it should be understood to include all modifications, equivalents, or alternatives within the spirit and technical scope of the present disclosure. In the description of each figure, similar reference numerals are used for similar components.
[0035] Terms such as first, second, etc. may be used to describe various components, but these components should not be limited by these terms. These terms are only used for the purpose of distinguishing one component from another. For example, without departing from the scope of the rights of the present disclosure, the first component can be named the second component, and similarly, the second component can be named the first component. The term "and / or" includes a combination of a plurality of related described items or any one of the plurality of related described items.
[0036] When a component is referred to as being "connected to" or "coupled to" another component, it should be understood that it may be directly connected to or directly coupled to the other component, or there may be still other components in between. On the other hand, when a component is referred to as being "directly connected to" or "directly coupled to" another component, it should be understood that there are no other components in between.
[0037] The terms used in this application are only for the purpose of describing specific embodiments and are not intended to limit the present disclosure. Singular expressions include plural expressions unless otherwise specified in the context. In this application, terms such as "including" or "having" are intended to specify the presence of features, numbers, steps, operations, components, parts, or combinations thereof described in the specification, and should not be construed as precluding the presence or addition of one or more other features, numbers, steps, operations, components, parts, or combinations thereof.
[0038] This disclosure relates to video / video coding. For example, the methods / embodiments disclosed herein may be applied to the methods disclosed in the VVC (versatile video coding) standard. Also, the methods / embodiments disclosed in this specification may be applied to the methods disclosed in the EVC (essential video coding) standard, the AV1 (AOMedia Video 1) standard, the AVS2 (2nd generation of audio video coding standard), or next-generation video / video coding standards (e.g., H.267 or H.268, etc.).
[0039] This specification presents various embodiments related to video / video coding, and unless otherwise specifically mentioned, the above embodiments may be combined with each other.
[0040] In this specification, "video" can mean a collection of a series of images over time. "Picture" generally means a unit representing one image in a specific time period, and "slice" / "tile" is a unit that constitutes a part of a picture in coding. A slice / tile may include one or more CTUs (Coding Tree Units). One picture may be composed of one or more slices / tiles. One tile is a rectangular area composed of multiple CTUs in a specific tile column and a specific tile row of one picture. A tile column is a rectangular area of CTUs having the same height as the height of the picture and a width specified by the syntax requirements of the picture parameter set. A tile row is a rectangular area of CTUs having a height specified by the picture parameter set and the same width as the width of the picture. The CTUs within one tile are continuously arranged by a CTU raster scan, while the tiles within one picture may be continuously arranged by a tile raster scan. One slice may include an integer number of complete tiles or an integer number of consecutive complete CTU rows within a tile of a picture that can be exclusively included in a single NAL unit. On the other hand, one picture may be divided into two or more sub-pictures. A sub-picture may be a rectangular area of one or more slices in a picture.
[0041] A pixel, picture element or pel can mean the smallest unit that constitutes one picture (or image). Also, the term "sample" may be used as a term corresponding to a pixel. A sample can generally indicate a pixel or a pixel value, and may indicate only the pixel / pixel value of the luma component, or may indicate only the pixel / pixel value of the chroma component.
[0042] A "unit" can mean the basic unit of video processing. A unit may include at least one of a specific area of a picture and information related to that area. One unit may include one luma block and two chroma (e.g., cb, cr) blocks. A unit may, in some cases, be used with the same meaning as terms such as "block" or "area". Generally, an MxN block may include a set (or, array) of samples (or, sample array) or transform coefficients consisting of M columns and N rows.
[0043] As used herein, "A or B" can mean "only A", "only B", or "both A and B". In other words, as used herein, "A or B" may be interpreted as "A and / or B". For example, as used herein, "A, B, or C" can mean "only A", "only B", "only C", or "any combination of A, B, and C".
[0044] The slash ( / ) or comma used herein can mean "and / or". For example, "A / B" can mean "A and / or B". Thus, "A / B" can mean "only A", "only B", or "both A and B". For example, "A, B, C" can mean "A, B, or C".
[0045] As used herein, "at least one of A and B" can mean "only A", "only B", or "both A and B". Also, as used herein, the expressions "at least one of A or B" or "at least one of A and / or B" may be interpreted the same as "at least one of A and B".
[0046] Also, in this specification, "at least one of A, B, and C" can mean "only A", "only B", "only C", or "any combination of A, B, and C". Also, "at least one of A, B, or C" or "at least one of A, B, and / or C" can mean "at least one of A, B, and C".
[0047] Also, the parentheses used in this specification can mean "for example". Specifically, when it is shown as "prediction (intra prediction)", "intra prediction" may be proposed as an example of "prediction". In other words, "prediction" in this specification is not limited to "intra prediction", and "intra prediction" may be proposed as an example of "prediction". Also, when it is shown as "prediction (that is, intra prediction)", "intra prediction" may be proposed as an example of "prediction".
[0048] In this specification, the technical features separately described in the same drawing may be embodied separately or simultaneously.
[0049] FIG. 1 is a diagram showing a video / video coding system according to the present disclosure.
[0050] Referring to FIG. 1, the video / video coding system may include a first device (source device) and a second device (receiver device).
[0051] The source device can transmit encoded video / image information or data in the form of a file or a stream to the receiving device through a digital storage medium or a network. The source device may include a video source, an encoding device, and a transmission unit. The receiving device may include a receiving unit, a decoding device, and a renderer. The encoding device may be referred to as a video / image encoding device, and the decoding device may be referred to as a video / image decoding device. A transmitter may be included in the encoding device. A receiver may be included in the decoding device. The renderer may include a display unit, and the display unit may be composed of a separate device or an external component.
[0052] The video source can obtain video / images through processes such as the capture, synthesis, or generation of video / images. The video source may include a video / image capture device and / or a video / image generation device. The video / image capture device may include one or more cameras, a video / image archive containing previously captured video / images, etc. The video / image generation device may include a computer, a tablet, a smartphone, etc., and can (electronically) generate video / images. For example, virtual video / images may be generated through a computer or the like, and in this case, the video / image capture process may be replaced during the process of generating related data.
[0053] The encoding device can encode the input video / image. The encoding device can perform a series of procedures such as prediction, transformation, quantization, etc. for compression and coding efficiency. The encoded data (encoded video / image information) may be output in the form of a bitstream.
[0054] The transmission unit can transmit the encoded video / video information or data output in the form of a bitstream to the receiving unit of the receiving device through a digital storage medium or network in the form of a file or streaming. The digital storage medium may include various storage media such as USB, SD, CD, DVD, Blu-ray, HDD, SSD, etc. The transmission unit may include elements for generating a media file according to a predefined file format and may include elements for transmission through a broadcast / communication network. The receiving unit can receive / extract the bitstream and transmit it to the decoding device.
[0055] The decoding device can decode the video / video by performing a series of procedures such as inverse quantization, inverse transformation, prediction, etc. corresponding to the operation of the encoding device.
[0056] The renderer can render the decoded video / video. The rendered video / video may be displayed from the display unit.
[0057] FIG. 2 is a schematic block diagram of an encoding device to which the application of the embodiment of the present disclosure is applicable and in which encoding of video / video signals is performed.
[0058] Referring to FIG. 2, the encoding apparatus 200 may include an image partitioner (210), a predictor (220), a residual processor (230), an entropy encoder (240), an adder (250), a filter (260), and a memory (270). The predictor 220 may include an inter-prediction unit 221 and an intra-prediction unit 222. The residual processor 230 may include a transformer (232), a quantizer 233, a dequantizer (234), and an inverse transformer (235). The residual processor 230 may further include a subtractor (231). The adder 250 may be referred to as a reconstructor or a reconstructed block generator. The above-described image partitioner 210, predictor 220, residual processor 230, entropy encoder 240, adder 250, and filter 260 may be configured by one or more hardware components (e.g., an encoding apparatus chipset or a processor) according to an embodiment. Also, the memory 270 may include a DPB (decoded picture buffer) and may be configured by a digital storage medium. The hardware component may further include the memory 270 as an internal / external component.
[0059] The video segmentation unit 210 can divide the input video (or picture, frame) input to the encoding device 200 into one or more processing units. As an example, the processing unit may be referred to as a coding unit (CU). In this case, the coding unit may be recursively divided from a coding tree unit (CTU) or a largest coding unit (LCU) by a QTBTTT (Quad-tree binary-tree ternary-tree) structure.
[0060] As an example, one coding unit may be divided into a plurality of coding units with deeper depths based on a quad-tree structure, a binary-tree structure, and / or a ternary structure. In this case, for example, the quad-tree structure may be applied first, and the binary-tree structure and / or the ternary structure may be applied later. Or, the binary-tree structure may be applied earlier than the quad-tree structure. The coding procedure according to this specification may be performed based on the final coding unit that is no longer divided. In this case, based on the coding efficiency according to the video characteristics, etc., the largest coding unit may be immediately used as the final coding unit, or, if necessary, the coding unit may be recursively divided into coding units with lower depths, and the coding unit having an optimal size may be used as the final coding unit. Here, the coding procedure may include procedures such as prediction, transformation, and restoration described later.
[0061] As another example, the processing unit may further include a prediction unit (PU: Prediction Unit) or a transform unit (TU: Transform Unit). In this case, the prediction unit and the transform unit may be respectively divided or partitioned from the final coding unit described above. The prediction unit may be a unit of sample prediction, and the transform unit may be a unit for deriving transform coefficients and / or a unit for deriving a residual signal from the transform coefficients.
[0062] The term "unit" may be used in the same sense as terms such as "block" or "area" in some cases. Generally, an MxN block may represent a set of samples or transform coefficients consisting of M columns and N rows. Samples may generally represent pixels or pixel values, and may represent only the pixels / pixel values of the luma component, or may represent only the pixels / pixel values of the chroma component. A sample may be used in the sense corresponding to one picture (or video), pixel, or pel.
[0063] The encoding device 200 can subtract a prediction signal (prediction block, prediction sample array) output from the inter prediction unit 221 or the intra prediction unit 222 from the input video signal (original block, original sample array) to generate a residual signal (residual block, residual sample array), and the generated residual signal is transmitted to the conversion unit 232. In this case, the unit that subtracts the prediction signal (prediction block, prediction sample array) from the input video signal (original block, original sample array) within the encoding device 200 may be called the subtraction unit 231.
[0064] The prediction unit 220 can perform a prediction on a processing target block (hereinafter referred to as the current block), and generate a predicted block including a prediction sample for the current block. The prediction unit 220 can determine whether intra prediction or inter prediction is applied in units of the current block or CU. As will be described later in the description of each prediction mode, the prediction unit 220 can generate various pieces of information related to prediction, such as prediction mode information, and transmit it to the entropy encoding unit 240. The information related to prediction may be encoded by the entropy encoding unit 240 and output in the form of a bitstream.
[0065] The intra prediction unit 222 can predict the current block by referring to samples within the current picture. The samples to be referred to may be located in the neighborhood of the current block depending on the prediction mode, or may be located at a certain distance from the current block. In intra prediction, the prediction mode may include one or more non-directional modes and a plurality of directional modes. The non-directional mode may include at least one of the DC mode or the Planar mode. The directional mode may include 33 directional modes or 65 directional modes depending on the degree of detail of the prediction direction. However, this is an example, and a greater or lesser number of directional modes may be used depending on the setting. The intra prediction unit 222 can also determine the prediction mode to be applied to the current block using the prediction mode applied to the neighboring blocks.
[0066] The inter prediction unit 221 can derive a prediction block for the current block based on a reference block (reference sample array) specified by a motion vector on a reference picture. At this time, in order to reduce the amount of motion information transmitted in the inter prediction mode, the motion information can be predicted in units of blocks, sub-blocks, or samples based on the correlation of the motion information between the peripheral block and the current block. The motion information may include a motion vector and a reference picture index. The motion information may further include inter prediction direction information (such as L0 prediction, L1 prediction, Bi prediction, etc.). In the case of inter prediction, the peripheral blocks may include spatial neighboring blocks existing in the current picture and temporal neighboring blocks existing in the reference picture. The reference picture including the reference block and the reference picture including the temporal neighboring block may be the same or different. The temporal neighboring block may be called by names such as a collocated reference block, a collocated CU (colCU), etc., and the reference picture including the temporal neighboring block may also be called a collocated picture (colPic). For example, the inter prediction unit 221 can configure a motion information candidate list based on the peripheral blocks and generate information indicating which candidates are used to derive the motion vector and / or the reference picture index of the current block. Inter prediction may be performed based on various prediction modes. For example, in the case of the skip mode and the merge mode, the inter prediction unit 221 can use the motion information of the peripheral blocks as the motion information of the current block. In the case of the skip mode, different from the merge mode, the residual signal does not need to be transmitted.In the case of the motion vector prediction (MVP) mode, the motion vectors of neighboring blocks are used as motion vector predictors, and by signaling the motion vector difference, the motion vector of the current block can be indicated.
[0067] The prediction unit 220 can generate a prediction signal based on various prediction methods described below. For example, the prediction unit can apply not only intra prediction or inter prediction for the prediction of one block, but also apply intra prediction and inter prediction simultaneously. This may be referred to as the combined inter and intra prediction (CIIP) mode. Also, the prediction unit may be based on the intra block copy (IBC) prediction mode or the palette mode for the prediction of a block. The IBC prediction mode or the palette mode may be used for the coding of content videos / movies such as games, like screen content coding (SCC). IBC basically performs prediction within the current picture, but can be performed similarly to inter prediction in terms of deriving a reference block within the current picture. That is, IBC can use at least one of the inter prediction techniques described in this specification. The palette mode may be regarded as an example of intra coding or intra prediction. When the palette mode is applied, the sample values within the picture can be signaled based on the information regarding the palette table and the palette index. The prediction signal generated by the prediction unit 220 may be used to generate a restored signal or may be used to generate a residual signal.
[0068] The conversion unit 232 can generate transform coefficients by applying a conversion method to the residual signal. For example, the conversion method may include at least one of DCT (Discrete Cosine Transform), DST (Discrete Sine Transform), KLT (Karhunen-Loeve Transform), GBT (Graph-Based Transform), or CNT (Conditionally Non-linear Transform). Here, GBT means the transform obtained from a graph when representing the relationship information between pixels as a graph. CNT means the transform obtained based on generating a prediction signal using all previously restored pixels. Also, the conversion process may be applied to pixel blocks having the same size of a square, or may also be applied to blocks of variable size that are not square.
[0069] The quantization unit 233 quantizes the transform coefficients and transmits them to the entropy encoding unit 240, and the entropy encoding unit 240 can encode the quantized signal (information regarding the quantized transform coefficients) and output it as a bitstream. The information regarding the quantized transform coefficients may be called residual information. The quantization unit 233 can rearrange the quantized transform coefficients in block form into a one-dimensional vector based on the coefficient scan order, and can generate the information regarding the quantized transform coefficients based on the quantized transform coefficients in the one-dimensional vector form.
[0070] The entropy encoding unit 240 can perform various encoding methods such as exponential Golomb, CAVLC (context-adaptive variable length coding), CABAC (context-adaptive binary arithmetic coding), etc. The entropy encoding unit 240 can encode, together or separately, not only the quantized transform coefficients but also information necessary for video / image restoration (e.g., values of syntax elements, etc.).
[0071] The encoded information (e.g., encoded video / video information) may be transmitted or stored in units of NAL (network abstraction layer) units in the form of a bitstream. The video / video information may further include information regarding various parameter sets such as an adaptation parameter set (APS), a picture parameter set (PPS), a sequence parameter set (SPS), or a video parameter set (VPS). Also, the video / video information may further include general constraint information. In this specification, the information and / or syntax elements transmitted / signaled from the encoding device to the decoding device may be included in the video / video information. The video / video information is encoded by the encoding procedure described above and may be included in the bitstream. The bitstream may be transmitted through a network or stored in a digital storage medium. Here, the network may include a broadcast network and / or a communication network, etc., and the digital storage medium may include various storage media such as USB, SD, CD, DVD, Blu-ray, HDD, SSD, etc. The signal output from the entropy encoding unit 240 may be configured as an internal / external element of the encoding device 200 by a transmission unit (not shown) for transmission and / or a storage unit (not shown) for storage, or the transmission unit may be included in the entropy encoding unit 240.
[0072] The quantized transform coefficients output from the quantization unit 233 may be used to generate a prediction signal. For example, by applying inverse quantization and inverse transformation to the quantized transform coefficients using the inverse quantization unit 234 and the inverse transformation unit 235, a residual signal (residual block or residual sample) can be restored. The addition unit 250 may generate a reconstructed signal (reconstructed picture, reconstructed block, reconstructed sample array) by adding the restored residual signal to the prediction signal output from the inter prediction unit 221 or the intra prediction unit 222. When there is no residual for the block to be processed, such as when the skip mode is applied, the predicted block may be used as the reconstructed block. The addition unit 250 may be referred to as a restoration unit or a reconstructed block generation unit. The generated reconstructed signal may be used for intra prediction of the next block to be processed within the current picture, and as will be described later, may be used for inter prediction of the next picture after passing through filtering. On the other hand, LMCS (luma mapping with chroma scaling) may be applied during picture encoding and / or the restoration process.
[0073] The filtering unit 260 can apply filtering to the restored signal to improve subjective / objective image quality. For example, the filtering unit 260 can apply various filtering methods to the restored picture to generate a modified restored picture, and the modified restored picture can be stored in the memory 270, specifically in the DPB of the memory 270. The various filtering methods may include deblocking filtering, sample adaptive offset, adaptive loop filter, bilateral filter, etc. The filtering unit 260 can generate various information related to filtering and transmit it to the entropy encoding unit 240. The information related to filtering may be encoded by the entropy encoding unit 240 and output in the form of a bitstream.
[0074] The modified restored picture transmitted to the memory 270 may be used as a reference picture in the inter prediction unit 221. The encoding device can thereby avoid prediction mismatches in the encoding device 200 and the decoding device when inter prediction is applied, and can also improve the encoding efficiency.
[0075] The DPB of the memory 270 can store the modified restored picture for use as a reference picture in the inter prediction unit 221. The memory 270 can store the motion information of the blocks for which the motion information within the current picture has been derived (or encoded), and / or the motion information of the blocks within the already restored picture. The stored motion information can be transmitted to the inter prediction unit 221 for utilization as the motion information of spatially neighboring blocks or temporally neighboring blocks. The memory 270 can store the restored samples of the restored blocks within the current picture and transmit them to the intra prediction unit 222.
[0076] FIG. 3 is a schematic block diagram of a decoding apparatus to which the application of the embodiment of the present disclosure is applicable and in which decoding of a video / video signal is performed.
[0077] Referring to FIG. 3, the decoding apparatus 300 may include an entropy decoder (310), a residual processor (320), a predictor (330), an adder (340), a filtering unit (filter, 350), and a memory (memoery, 360). The predictor 330 may include an inter-prediction unit 331 and an intra-prediction unit 332. The residual processor 320 may include a dequantizer (321) and an inverse transformer (321).
[0078] The above-described entropy decoding unit 310, residual processing unit 320, prediction unit 330, addition unit 340, and filtering unit 350 may be configured by one hardware component (for example, a decoding apparatus chipset or a processor) according to an embodiment. Further, the memory 360 may include a DPB (decoded picture buffer) and may be configured by a digital storage medium. The hardware component may further include the memory 360 as an internal / external component.
[0079] When a bitstream including video / video information is input, the decoding device 300 can restore the video corresponding to the process in which the video / video information was processed by the encoding device in FIG. 2. For example, the decoding device 300 can derive units / blocks based on the block division related information obtained from the bitstream. The decoding device 300 can perform decoding using the processing unit applied in the encoding device. Therefore, the processing unit for decoding may be a coding unit, and the coding unit may be divided from a coding tree unit or a maximum coding unit according to a quadtree structure, a binary tree structure, and / or a ternary tree structure. One or more conversion units may be derived from the coding unit. Then, the restored video signal decoded and output by the decoding device 300 may be played back by a playback device.
[0080] The decoding device 300 can receive the signal output from the encoding device of FIG. 2 in the form of a bitstream, and the received signal may be decoded by the entropy decoding unit 310. For example, the entropy decoding unit 310 can parse the bitstream and derive information (e.g., video / video information) necessary for video restoration (or, picture restoration). The video / video information may further include information regarding various parameter sets such as an Adaptation Parameter Set (APS), a Picture Parameter Set (PPS), a Sequence Parameter Set (SPS), or a Video Parameter Set (VPS). Also, the video / video information may further include general constraint information. The decoding device can decode a picture based on the information regarding the parameter set and / or the general constraint information. The signals / received information and / or syntax elements described later in this specification may be decoded by the decoding procedure and obtained from the bitstream. For example, the entropy decoding unit 310 can decode the information in the bitstream based on a coding method such as exponential Golomb coding, CAVLC, or CABAC, and output the value of the syntax element necessary for video restoration and the quantized value of the transform coefficient regarding the residual. More specifically, the CABAC entropy decoding method receives the bin corresponding to each syntax element in the bitstream, determines a context model using the syntax element information to be decoded, the decoding information of the surrounding and decoding target blocks, or the information of the symbol / bin decoded in the previous stage, predicts the occurrence probability of the bin by the determined context model, and performs arithmetic decoding of the bin to generate a symbol corresponding to the value of each syntax element. At this time, the CABAC entropy decoding method can update the context model using the information of the symbol / bin decoded for the context model of the next symbol / bin after determining the context model.Of the information decoded by the entropy decoding unit 310, the information related to prediction is provided to the prediction unit (inter prediction unit 332 and intra prediction unit 331), and the residual value obtained by performing entropy decoding in the entropy decoding unit 310, that is, the quantized transform coefficient and related parameter information, may be input to the residual processing unit 320. The residual processing unit 320 can derive a residual signal (residual block, residual sample, residual sample array). Also, the information related to filtering among the information decoded by the entropy decoding unit 310 may be provided to the filtering unit 350. On the other hand, a receiving unit (not shown) that receives a signal output from the encoding device may be further configured as an internal / external element of the decoding device 300, or the receiving unit may be a component of the entropy decoding unit 310.
[0081] On the other hand, the decoding device according to the present specification may be referred to as a video / video / picture decoding device, and the decoding device may be classified into an information decoding device (video / video / picture information decoding device) and a sample decoding device (video / video / picture sample decoding device). The information decoding device may include the entropy decoding unit 310, and the sample decoding device may include at least one of the inverse quantization unit 321, inverse transform unit 322, addition unit 340, filtering unit 350, memory 360, inter prediction unit 332, and intra prediction unit 331.
[0082] In the inverse quantization unit 321, the quantized transform coefficients can be inverse quantized to output the transform coefficients. The inverse quantization unit 321 can rearrange the quantized transform coefficients in a two-dimensional block form. In this case, the rearrangement can be performed based on the coefficient scan order performed in the encoding device. The inverse quantization unit 321 can perform inverse quantization on the quantized transform coefficients using quantization parameters (for example, quantization step size information) to obtain the transform coefficients.
[0083] In the inverse transform unit 322, the transform coefficients are inverse transformed to obtain a residual signal (residual block, residual sample array).
[0084] The prediction unit 320 can perform prediction on the current block to generate a predicted block including prediction samples for the current block. The prediction unit 320 can determine whether intra prediction or inter prediction is to be applied to the current block based on the information regarding the prediction output from the entropy decoding unit 310, and can determine a specific intra / inter prediction mode.
[0085] The prediction unit 320 can generate a prediction signal based on various prediction methods described below. For example, the prediction unit 320 can apply intra prediction or inter prediction for the prediction of one block, and can also apply intra prediction and inter prediction simultaneously. This may be referred to as the CIIP (combined inter and intra prediction) mode. Further, the prediction unit may be based on the intra block copy (IBC) prediction mode or the palette mode for the prediction of the block. The IBC prediction mode or the palette mode may be used for content video / moving image coding such as games like SCC (screen content coding). IBC basically performs prediction within the current picture, but can be performed similarly to inter prediction in terms of deriving a reference block within the current picture. That is, IBC can use at least one of the inter prediction methods described in this specification. The palette mode can be regarded as an example of intra coding or intra prediction. When the palette mode is applied, information regarding the palette table and the palette index may be included in and signaled in the video / video information.
[0086] The intra prediction unit 331 can predict the current block by referring to samples within the current picture. The samples to be referred to may be located in the neighborhood of the current block according to the prediction mode, or may be located at a certain distance from the current block. In intra prediction, the prediction mode may include one or more non-directional modes and a plurality of directional modes. The intra prediction unit 331 can determine the prediction mode to be applied to the current block by using the prediction mode applied to the neighboring blocks.
[0087] The inter prediction unit 332 can derive a prediction block for the current block based on a reference block (reference sample array) specified by a motion vector on a reference picture. At this time, in order to reduce the amount of motion information transmitted in the inter prediction mode, the motion information can be predicted in units of blocks, sub-blocks, or samples based on the correlation of the motion information between the peripheral block and the current block. The motion information may include a motion vector and a reference picture index. The motion information may further include inter prediction direction information (such as L0 prediction, L1 prediction, Bi prediction, etc.). In the case of inter prediction, the peripheral blocks may include spatial neighboring blocks existing within the current picture and temporal neighboring blocks existing in the reference picture. For example, the inter prediction unit 332 can construct a motion information candidate list based on the peripheral blocks and derive the motion vector and / or reference picture index of the current block based on the received candidate selection information. Inter prediction may be performed based on various prediction modes, and the information related to the prediction may include information indicating the inter prediction mode for the current block.
[0088] The addition unit 340 can generate a restored signal (restored picture, restored block, restored sample array) by adding the obtained residual signal to the prediction signal (prediction block, prediction sample array) output from the prediction unit (including the inter prediction unit 332 and / or the intra prediction unit 331). When there is no residual for the processing target block as in the case where the skip mode is applied, the prediction block may be used as the restored block.
[0089] The addition unit 340 may be referred to as a restoration unit or a restoration block generation unit. The generated restoration signal may be used for intra prediction of the next processing target block within the current picture, and as will be described later, may be output after passing through filtering, or may be used for inter prediction of the next picture. On the other hand, LMCS (luma mapping with chroma scaling) may be applied during the picture decoding process.
[0090] The filtering unit 350 can apply filtering to the restoration signal to improve subjective / objective image quality. For example, the filtering unit 350 can apply various filtering methods to the restored picture to generate a modified restored picture, and can transmit the modified restored picture to the memory 360, specifically, to the DPB of the memory 360. The various filtering methods may include deblocking filtering, sample adaptive offset, adaptive loop filter, bilateral filter, and the like.
[0091] The (modified) restored picture stored in the DPB of the memory 360 may be used as a reference picture in the inter prediction unit 332. The memory 360 can store the motion information of the blocks for which the motion information within the current picture has been derived (or decoded), and / or the motion information of the blocks within the already restored picture. The stored motion information can be transmitted to the inter prediction unit 260 for utilization as the motion information of spatially adjacent blocks or temporally adjacent blocks. The memory 360 can store the restored samples of the restored blocks within the current picture and transmit them to the intra prediction unit 331.
[0092] In this specification, the examples described in the filtering unit 260, inter prediction unit 221, and intra prediction unit 222 of the encoding device 200 may be applied identically or correspondingly to the filtering unit 350, inter prediction unit 332, and intra prediction unit 331 of the decoding device 300, respectively.
[0093] FIG. 4 is a diagram showing an inter prediction method performed by a decoding device according to an embodiment of the present disclosure.
[0094] Referring to FIG. 4, a candidate list for predicting the motion information of the current block can be generated (S400).
[0095] The candidate list may be for the merge mode or the AMVP (advanced motion vector prediction) mode. However, it is not limited thereto, and the candidate list may be for the affine merge mode or the affine inter mode.
[0096] The motion information may include at least one of a motion vector, a reference picture index, inter prediction direction information, or weight value information for bi-directional weighted prediction.
[0097] The candidate list may include at least one of a spatial candidate, a temporal candidate, an induced candidate, or a history-based candidate.
[0098] The spatial candidate may be induced based on the motion information of a peripheral block (hereinafter referred to as a spatial peripheral block) that is spatially adjacent to the current block. The spatial peripheral block may include at least one of an upper peripheral block, a left peripheral block, a lower left peripheral block, an upper right peripheral block, or an upper left peripheral block of the current block.
[0099] Alternatively, the spatial candidate may be derived based on the motion information of non-adjacent blocks. Here, the non-adjacent blocks are blocks decoded before the current block, which belong to the same picture as the current block but do not mean blocks adjacent to the current block. The non-adjacent blocks may be blocks at positions separated from the current block by K sample units in at least one of the vertical, horizontal, or diagonal directions. Here, K may be an integer of 4, 8, 16, 32 or more.
[0100] The temporal candidate may be derived based on the motion information of surrounding blocks temporally adjacent to the current block (hereinafter referred to as temporal surrounding blocks). The temporal surrounding blocks may belong to a reference picture (collocated picture) decoded before the current picture and be blocks at the same position as the current block. The blocks at the same position may be blocks including the position of the upper left corner sample of the current block, the position of the central sample, or the position of the sample adjacent to the lower right corner of the current block. Alternatively, the temporal surrounding blocks may mean blocks at positions shifted by a predetermined displacement vector from the blocks at the same position.
[0101] The displacement vector may be the motion vector of a specific block preset identically in the encoding device and the decoding device. As an example, the preset specific block may be any one of the left surrounding block, upper surrounding block, upper left surrounding block, upper right surrounding block, or lower left surrounding block of the current block. However, when the motion vector of the specific block is not available, the displacement vector may be derived based on the motion vector corresponding to the lower right position of the current block (or the motion vector of the derived candidate). The method of deriving the motion vector corresponding to the lower right position of the current block will be described later, and detailed description is omitted here.
[0102] Alternatively, the transition vector may be determined as any one of a plurality of motion vector candidates belonging to a motion vector candidate list. The plurality of motion vector candidates may include at least one of the motion vectors of the above-described spatial neighboring blocks or the motion vectors corresponding to the lower right end positions of the current blocks described below (or the induced candidate motion vectors).
[0103] The above-described temporal candidates are available only when the temporal_mvp_enabled_flag signaled in a high-level syntax such as a sequence parameter set (SPS), a picture parameter set (PPS), a picture header (PH), a slice header (SH), etc. is TRUE, and are unavailable when the motion information of the block at the same position is unavailable. Also, since the reliability of the motion information in the current picture is higher than that in the reference picture, a method for inducing a candidate (hereinafter referred to as an induced candidate) in which the motion of the lower right end position of the current block is reflected from the motion information of a specific block of the current block will be described.
[0104] The derived candidate may be derived based on the motion information of at least one of the upper right block or the lower left block of the current block.
[0105] Determination Method 1 of the Upper-Right / Lower-Left Block
[0106] The upper right block of the current block may be a coding block including samples at the upper right reference position or a sub-block having a predetermined size. Here, the sub-block may be any one of the sub-blocks within the coding block divided into a predetermined size, and the predetermined size may be 2x2, 2x4, 4x2, 4x4, or more. The sub-blocks in the embodiments described below may also be understood in the same sense. Similarly, the lower left block of the current block may be a coding block including samples at the lower left reference position or a sub-block having a predetermined size.
[0107] Assuming that the position of the sample at the upper left end of the current block is (0, 0), the upper right reference position and the lower left reference position may be defined as (N*W, -1) and (-1, N*H), respectively. Here, W and H can respectively represent the width and height of the current block, and N may be at least one integer belonging to the range of 1 to 32. However, it is not limited thereto, and N may be at least one integer belonging to the range of -32 to 0.
[0108] For the sake of convenience of explanation, assume that N is 2. In this case, based on the upper right block including the upper right reference position of (2W, -1) and the lower left block including the lower left reference position of (-1, 2H), the induced candidate motion information may be induced. As an example, the induced candidate motion information may be induced as shown in Equation 1 below.
[0109] [Equation 1]
[0110] MV RB (x)=(W*MV RT (x)+W*MV LB (x)) / 2W
[0111] MV RB (y)=(H*MV RT (y)+H*MV LB (y)) / 2H
[0112] In Equation 1, MV RT (x) means the motion vector of the x component of the upper right block including the upper right reference position of (2W, -1), and MV RT (y) can mean the motion vector of the y component of the upper right block including the upper right reference position of (2W, -1). MV LB (x) means the motion vector of the x component of the lower left block including the lower left reference position of (-1, 2H), and MV LB (y) can mean the motion vector of the y component of the lower left block including the lower left reference position of (-1, 2H). MV RB (x) and MV RB(y) is the motion vector corresponding to the lower right corner position of the current block, and this may be used as the induced candidate motion vector. As shown in Equation 1, in order to calculate the induced candidate motion vector, the horizontal distance (W) between the current block and the upper right block, and the vertical distance (H) between the current block and the lower left block may be considered.
[0113] The upper right / lower left reference positions may be identically applied not only when the current block is square but also when the current block is non-square. That is, even when the width of the current block is greater than the height or the width of the current block is less than the height, the upper right reference position and the lower left reference position may be defined as (N*W, -1) and (-1, N*H), respectively.
[0114] Alternatively, depending on whether the current block is square or not, the upper right / lower left reference positions may be defined to be different from each other. As an example, when the current block is square, as described above, the upper right reference position and the lower left reference position may be defined as (N*W, -1) and (-1, N*H), respectively. On the other hand, when the current block is non-square, the upper right reference position and the lower left reference position may be defined as (M*(W + H), -1) and (-1, M*(W + H)), respectively. Here, W and H may respectively represent the width and height of the current block, and M may be at least one integer belonging to the range of 1 to 16. However, it is not limited thereto, and M may be at least one integer belonging to the range of -16 to 0.
[0115] For the sake of convenience of explanation, assume that M is 1. In this case, the induced candidate motion information may be induced based on the upper right block including the upper right reference position of ((W + H), -1) and the lower left block including the lower left reference position of (-1, (W + H)). As an example, the induced candidate motion information may be induced as shown in the following Equation 2.
[0116] [Equation 2]
[0117] MV RB (x)=(W*MVRT (x) + H * MV LB (x)) / (W + H)
[0118] MV RB (y) = (W * MV RT (y) + H * MV LB (y)) / (W + H)
[0119] In Equation 2, MV RT (x) means the motion vector of the x - component of the upper - right block including the upper - right reference position of ((W + H), - 1), and MV RT (y) can mean the motion vector of the y - component of the upper - right block including the upper - right reference position of ((W + H), - 1). MV LB (x) means the motion vector of the x - component of the lower - left block including the lower - left reference position of (-1, (W + H)), and MV LB (y) can mean the motion vector of the y - component of the lower - left block including the lower - left reference position of (-1, (W + H)). MV RB (x) and MV RB (y) are the motion vectors corresponding to the lower - right position of the current block, and these may be used as the induced candidate motion vectors. As in Equation 2, in order to calculate the induced candidate motion vectors, the horizontal distance (H) between the current block and the upper - right block, and the vertical distance (W) between the current block and the lower - left block may be considered.
[0120] The above - mentioned ranges of N and M may be variably determined by the area that the current block can reference. As an example, the area that the current block can reference may be limited within one coding tree unit (CTU) to which the current block belongs. In this case, when the size of the CTU is 128x128 and the size of the current block is 4x4, N may be at least one integer belonging to the range of 1 to 31.
[0121] Determination Method 2 of the Upper-Right / Lower-Left Block
[0122] The upper right block of the current block may be a coding block or sub-block that includes a position (hereinafter referred to as the RT position) that has moved a predetermined distance (α) in the left or right direction from the upper right reference position described above. The α may be an integer such as 4, 8, 16, 32, or more. The α may be a value preset identically in the encoding device and the decoding device. Alternatively, the α may be variably determined based on at least one of the size of the current block (for example, width, height, ratio of width to height, product of width and height), whether the current block is square, whether the current block touches the boundary of the CTU, or the position of the current block in the CTU. As an example, when the size of the current block is smaller than 32x32, α may be set to 4, and when the size of the current block is larger than or equal to 32x32, α may be set to 8. On the other hand, the lower left block of the current block may be determined depending on the position of the upper right block of the current block. The lower left block of the current block may be a coding block or sub-block that includes a position (hereinafter referred to as the LB position) that has moved a predetermined distance in the lower or upper direction from the lower left reference position described above corresponding to the RT position. At this time, the LB position corresponding to the RT position may be determined to be located on a straight line passing through the RT position and the lower right position of the current block.
[0123] As an example, when moving a predetermined distance (α) in the left direction from the upper right reference position of (2W, -1), the equation of the straight line passing through the aforementioned RT position, the lower right position of the current block, and the LB position may be expressed as Equation 3 below.
[0124] [Equation 3]
[0125] JPEG2025522759000002.jpg1269
[0126] The RT position is variably determined by the value of α, may be defined as (2W - α, -1), and the LB position may be calculated using the equation of a straight line that changes according to the value of α. In the unlikely event that the size of the current block is 8x8 and the value of α is 4, the RT position and the LB position may be determined as (12, -1) and (-1, 28) respectively.
[0127] Alternatively, for the convenience of calculation, the upper right reference position and the lower left reference position may be simplified to (2W, 0) and (0, 2H) respectively. In this case, the equation of the straight line in Equation 3 may be changed as in Equation 4.
[0128] [Equation 4]
[0129] JPEG2025522759000003.jpg1266
[0130] When using Equation 4, under the same conditions, the initial RT position may be determined as (12, 0), and the initial LB position may be determined as (0, 24). Based on the initial RT / LB positions, the final RT position and the LB position may be determined as (12, -1) and (-1, 24) respectively.
[0131] Alternatively, the lower-left block of the current block may be a coding block or a sub-block including a position (hereinafter referred to as the LB position) that has moved a predetermined distance (β) in the upward or downward direction from the aforementioned lower-left reference position. The β may be an integer such as 4, 8, 16, 32, or more. The β may be a value preset identically in the encoding device and the decoding device. Alternatively, the β may be variably determined based on at least one of the size of the current block (e.g., width, height, ratio of width to height, product of width and height), whether the current block is square, whether the current block touches the boundary of the CTU, or the position of the current block in the CTU. The β may be a value different from the aforementioned α. On the other hand, the upper-right block of the current block may be determined depending on the position of the lower-left block of the current block. The upper-right block of the current block may be a coding block or a sub-block including a position (hereinafter referred to as the RT position) that has moved a predetermined distance in the rightward or leftward direction from the aforementioned upper-right reference position corresponding to the LB position. At this time, the RT position corresponding to the LB position may be determined to be located on a straight line passing through the LB position and the lower-right position of the current block.
[0132] Determination Method 3 of the Upper-Right / Lower-Left Block
[0133] The upper right block of the current block may be determined by searching for available blocks within a predetermined search range. The search range may be from the upper right position of the current block (e.g., (W, -1)) to the upper right reference position. Specifically, while moving in units of a predetermined distance (α), it is possible to check whether the coding block or sub-block including the moving position is available. Here, the predetermined distance (α) is as described in "Method 2 for Determining the Upper Right / Lower Left Blocks". The first available block discovered during the search process may be set as the upper right block of the current block. At this time, the search may be performed in the rightward direction from the upper right position of the current block, or may be performed in the leftward direction from the upper right reference position of the current block. On the other hand, the lower left block of the current block may be determined depending on the position of the previously determined upper right block, as described in "Method 2 for Determining the Upper Right / Lower Left Blocks".
[0134] Alternatively, the lower left block of the current block may be determined by searching for available blocks within a predetermined search range. The search range may be from the lower left position of the current block (e.g., (-1, H)) to the lower left reference position. Specifically, while moving in units of a predetermined distance (β), it is possible to check whether the coding block or sub-block including the moving position is available. Here, the predetermined distance (β) is as described in "Method 2 for Determining the Upper Right / Lower Left Blocks". The first available block discovered during the search process may be set as the lower left block of the current block. At this time, the search may be performed in the downward direction from the lower left position of the current block, or may be performed in the upward direction from the lower left reference position of the current block. On the other hand, the upper right block of the current block may be determined depending on the position of the previously determined lower left block, as described in "Method 2 for Determining the Upper Right / Lower Left Blocks".
[0135] Determination Method 4 of the Upper-Right / Lower-Left Block
[0136] The upper-right block and the lower-left block of the current block may be coding blocks or sub-blocks respectively including positions that are moved by a predetermined distance (α) from the above-described upper-right reference position and lower-left reference position. Specifically, the upper-right block of the current block may be a coding block or sub-block including a position moved by a predetermined distance (α) to the left or right from the upper-right reference position, and the lower-left block of the current block may be a coding block or sub-block including a position moved by a predetermined distance (α) downward or upward from the lower-left reference position. Here, the predetermined distance (α) is as described in the "Method 2 for Determining the Upper-Right / Lower-Left Block".
[0137] Determination Method 5 of the Upper-Right / Lower-Left Block
[0138] The upper-right block of the current block may be determined by searching for available blocks in the left or right direction from a predetermined search start position. The search start position may be the upper-right reference position and / or the upper-right position of the current block. Specifically, while moving in units of a predetermined distance (α) from the search start position, it is possible to check whether a coding block or sub-block including the moving position is available. Here, the predetermined distance (α) is as described in the "Method 2 for Determining the Upper-Right / Lower-Left Block". The first available block discovered in the search process may be set as the upper-right block of the current block.
[0139] On the other hand, the lower-left block of the current block may be determined by searching for available blocks in the downward or upward direction from a predetermined search start position. The search start position may be the lower-left reference position and / or the lower-left position of the current block. Specifically, while moving in units of the same distance (α) from the search start position, it is possible to check whether a coding block or sub-block including the moving position is available. The first available block discovered in the search process may be set as the lower-left block of the current block.
[0140] The search direction may be preset to be the same for both the encoding device and the decoding device. As an example, the search direction for the upper right block may be the leftward direction from the upper right reference position, and the search direction for the lower left block may be the downward direction from the lower left reference position or the lower left position. Alternatively, the search direction for the upper right block may be the rightward direction from the upper right reference position or the upper right position, and the search direction for the lower left block may be the upward direction from the lower left reference position.
[0141] Alternatively, the search direction may vary according to the size and / or shape of the current block. As an example, when the width of the current block is smaller than the height, the search direction for the upper right block may be the rightward direction from the upper right reference position, and the search direction for the lower left block may be the upward direction with respect to the lower left reference position. Conversely, when the width of the current block is larger than the height, the search direction for the upper right block may be the leftward direction from the upper right reference position, and the search direction for the lower left block may be the downward direction with respect to the lower left reference position.
[0142] When determining the upper right / lower left block of the current block, different methods may be applied. As an example, the upper right block may be determined based on any one of methods 1 to 5 for determining the upper right / lower left block, and the lower left block may be determined based on any other one of methods 1 to 5 for determining the upper right / lower left block.
[0143] The upper right / lower left block may be determined by at least one of the methods for determining the upper right / lower left block described above. Based on at least one of the position or motion information of the determined upper right / lower left block, the motion information corresponding to the lower right position of the current block may be induced, which may be used as the induced candidate motion information.
[0144] As an example, the motion vector corresponding to the lower right position of the current block may be induced as shown in Equation 5 below.
[0145] [Equation 5]
[0146] MV RB (x)=(dx0*MV RT (x)+dx1*MV LB (x)) / (dx0+dx1)
[0147] MV RB (y)=(dy1*MV RT (y)+dy0*MV LB (y)) / (dy1+dy0)
[0148] In Equation 5, MV RT (x) means the motion vector of the x component of the upper - right - hand block, and MV RT (y) can mean the motion vector of the y component of the upper - right - hand block. MV LB (x) means the motion vector of the x component of the lower - left - hand block, and MV LB (y) can mean the motion vector of the y component of the lower - left - hand block. MV RB (x) means the motion vector of the x component corresponding to the lower - right - hand position of the current block, and MV RB (y) can mean the motion vector of the y component corresponding to the lower - right - hand position of the current block. On the other hand, in Equation 5, dx0 and dy0 can respectively mean the width and height of the current block. Also, dx1 means the horizontal distance between the current block and the upper - right - hand block, and dy1 can mean the vertical distance between the current block and the lower - left - hand block. As an example, dx1 may be determined based on the difference between the position of the upper - right - hand block (i.e., the RT position, the moving position in the left / right direction, or the upper - right - hand reference position) determined by at least one of the above - mentioned determination methods 2 - 5 for the upper - right - hand block and the width of the current block. dy1 may be determined based on the difference between the position of the lower - left - hand block (i.e., the LB position, the moving position in the up / down direction, or the lower - left - hand reference position) determined by at least one of the above - mentioned determination methods 2 - 5 for the lower - left - hand block and the height of the current block.
[0149] Any one of the determination methods 1, 2, or 4 for the upper right / lower left block may be applied when the other one is not available. As an example, when there is no upper right / lower left block determined by the determination method 1 for the upper right / lower left block or it has no motion information, the upper right / lower left block may be determined by the determination method 2 or 4 for the upper right / lower left block. Or, when there is no upper right / lower left block determined by the determination method 2 for the upper right / lower left block or it has no motion information, the upper right / lower left block may be determined by the determination method 1 or 4 for the upper right / lower left block. Or, when there is no upper right / lower left block determined by the determination method 4 for the upper right / lower left block or it has no motion information, the upper right / lower left block may be determined by the determination method 1 or 2 for the upper right / lower left block.
[0150] A plurality of induced candidates may be induced for the current block. For this purpose, a plurality of upper right blocks and a plurality of lower left blocks may be determined for the current block respectively. As an example, assume that the first to Nth upper right blocks and the first to Nth lower left blocks are determined. In this case, the ith induced candidate may be induced based on the motion information of the ith upper right block and the motion information of the corresponding ith lower left block (1 ≤ i ≤ N). Hereinafter, the method for determining a plurality of upper right blocks and lower left blocks will be described in detail.
[0151] Determination Method 1 of Multiple Upper-Right / Lower-Left Blocks
[0152] By searching for available blocks within a predetermined search range, the upper-right block of the current block can be determined, and a plurality of lower-left blocks can be determined based on the position of the determined upper-right block. The search range may be from the upper-right position of the current block (e.g., (W, -1)) to the upper-right reference position. Specifically, while moving in units of a predetermined distance (α), it is possible to check whether the coding block or sub-block including the moving position is available. Here, the predetermined distance (α) is as described in "Method for Determining Upper-Right / Lower-Left Blocks 2". The search direction may be the right direction from the upper-right position of the current block, or the left direction from the upper-right reference position of the current block.
[0153] Based on the available blocks discovered in the search process, the upper-right block of the current block may be determined. As an example, all the available blocks discovered in the search process may be set as the upper-right block of the current block. Alternatively, the search may be performed until a predetermined number (N) of available blocks are discovered, and the N available blocks discovered in the search process may be set as the upper-right block of the current block. Here, N is a value preset identically in the encoding device and the decoding device, and may be an integer such as 2, 3, 4, or more. Alternatively, N may be variably determined based on the information signaled in the bitstream or the attributes of the current block (e.g., size, shape, inter prediction mode, etc.).
[0154] On the other hand, the lower-left blocks of the current block may be determined corresponding to the positions of the previously determined upper-right blocks respectively. The positions of the lower-left blocks corresponding to the positions of the upper-right blocks may be determined as described in "Method for Determining Upper-Right / Lower-Left Blocks 2", and the repeated description is omitted here.
[0155] Alternatively, by searching for available blocks within a predetermined search range, the lower-left block of the current block can be determined, and a plurality of upper-right blocks can be determined based on the position of the determined lower-left block. The search range may be from the lower-left position of the current block (e.g., (-1, H)) to the lower-left reference position. Specifically, while moving in units of a predetermined distance (β), it is possible to check whether the coding block or sub-block including the moving position is available. Here, the predetermined distance (β) is as described in "Method 2 for Determining Upper-Right / Lower-Left Blocks". The search direction may be the downward direction from the lower-left position of the current block, or the upward direction from the lower-left reference position of the current block.
[0156] Based on the available blocks discovered during the search process, the lower-left block of the current block may be determined. As an example, all the available blocks discovered during the search process may be set as the lower-left block of the current block. Alternatively, the search may be performed until a predetermined number (M) of available blocks are discovered, and the M available blocks discovered during the search process may be set as the upper-right blocks of the current block. Here, M is a value preset identically in the encoding device and the decoding device, and may be an integer such as 2, 3, 4, or more. Alternatively, M may be variably determined based on the information signaled in the bitstream or the attributes of the current block (e.g., size, shape, inter prediction mode, etc.).
[0157] On the other hand, the upper-right blocks of the current block may be determined corresponding to the positions of the previously determined lower-left blocks respectively. The positions of the upper-right blocks corresponding to the positions of the lower-left blocks may be determined as described in "Method 2 for Determining Upper-Right / Lower-Left Blocks", and the overlapping explanations are omitted here.
[0158] Determination Method 2 of Multiple Upper-Right / Lower-Left Blocks
[0159] The upper right block of the current block may be determined by searching for available blocks in the left or right direction from a predetermined search start position. The search start position may be the upper right reference position and / or the upper right position of the current block. Specifically, while moving in units of a predetermined distance (α) from the search start position, it is possible to check whether a coding block or sub-block including the moving position is available. Here, the predetermined distance (α) is as described in "Method 2 for Determining the Upper Right / Lower Left Block".
[0160] The upper right block of the current block may be determined based on the available blocks discovered during the search process. As an example, all the available blocks discovered during the search process may be set as the upper right block of the current block. Or, the search may be performed until a predetermined number (N) of available blocks are discovered, and the N available blocks discovered during the search process may be set as the upper right block of the current block. Here, N is a value preset identically in the encoding device and the decoding device, and may be an integer of 2, 3, 4, or more. Or, N may be variably determined based on information signaled in the bitstream or attributes of the current block (e.g., size, shape, inter prediction mode, etc.).
[0161] On the other hand, the lower left block of the current block may be determined by searching for available blocks in the lower or upper direction from a predetermined search start position. The search start position may be the lower left reference position and / or the lower left position of the current block. Specifically, while moving in units of the same distance (α) from the search start position, it is possible to check whether a coding block or sub-block including the moving position is available.
[0162] Similarly, the lower left block of the current block may be determined based on the available blocks discovered in the search process. As an example, all the available blocks discovered in the search process may be set as the lower left block of the current block. Alternatively, the search may be performed until a predetermined number (N) of available blocks are discovered, and the N available blocks discovered in the search process may be set as the lower left block of the current block.
[0163] The search direction for the upper right / lower left block is as described in "Determination Method 5 of the Upper Right / Lower Left Block", and duplicate explanations are omitted here.
[0164] Determination Method 3 of Multiple Upper-Right / Lower-Left Blocks
[0165] The search range for the upper right block of the current block may be divided into a plurality of sub-ranges, and the search for available blocks may be performed for each sub-range. Here, the search range may be from the upper right position of the current block (for example, (W, -1)) to the upper right reference position. Alternatively, the search range may include one or more sub-ranges in the left direction and one or more sub-ranges in the right direction from the upper right reference position. The range allowed for the search may be restricted within the current coding tree unit (CTU) to which the current block belongs, two or more coding tree units including the current coding tree unit, or a coding tree unit row (CTU row) including the current coding tree unit. The sub-range is defined by an L sample interval, and L may be 4, 8, 16, 32, 64, 128 or more. L may be a value preset identically in the encoding device and the decoding device, or may be variably determined based on the attributes (for example, size, shape, etc.) of the current block.
[0166] Specifically, while moving in units of a predetermined distance (α) within the current sub-range, it is possible to check whether a coding block or sub-block including the moving position is available. Here, the predetermined distance (α) is as described in the "Method 2 for Determining the Upper-Right / Left-Lower Blocks". The available block first discovered in the search process may be set as the upper-right block of the current block. Thereafter, the search within the current sub-range is completed, and the search within the next sub-range may be performed in the same manner. By the above-described process, at least one upper-right block may be derived for each sub-range.
[0167] On the other hand, the search range for the lower-left block of the current block may be divided into a plurality of sub-ranges, and the search for available blocks may be performed for each sub-range. Here, the search range may be from the lower-left position of the current block (e.g., (-1, H)) to the lower-left reference position. Alternatively, the search range may include one or more sub-ranges in the upper direction from the lower-left reference position and one or more sub-ranges in the lower direction from the lower-left reference position. The range within which the search is allowed may be limited within the current coding tree unit (CTU) to which the current block belongs. The sub-range may be defined as L sample intervals, as described above.
[0168] Specifically, while moving in units of a predetermined distance (α or β) within the current sub-range, it is possible to check whether a coding block or sub-block including the moving position is available. Here, the predetermined distance (α or β) is as described in the "Method 2 for Determining the Upper-Right / Left-Lower Blocks". The available block first discovered in the search process may be set as the lower-left block of the current block. Thereafter, the search within the current sub-range is completed, and the search within the next sub-range may be performed in the same manner. By the above-described process, at least one lower-left block may be derived for each sub-range.
[0169] Determination Method 4 of Multiple Upper-Right / Lower-Left Blocks
[0170] A plurality of upper-right blocks may be determined based on at least two combinations among the above-described "Methods 1 to 5 for Determining Upper-Right / Lower-Left Blocks". Similarly, a plurality of lower-left blocks may be determined based on at least two combinations among the above-described "Methods 1 to 5 for Determining Upper-Right / Lower-Left Blocks".
[0171] When determining a plurality of upper-right / lower-left blocks, different methods may be applied to each other. As an example, a plurality of upper-right blocks may be determined based on any one of Methods 1 to 4 for determining a plurality of upper-right / lower-left blocks, and a plurality of lower-left blocks may be determined based on any other one of Methods 1 to 4 for determining a plurality of upper-right / lower-left blocks.
[0172] The i-th upper-right block / lower-left block may be determined by at least one of the above-described methods for determining a plurality of upper-right / lower-left blocks. Based on at least one of the positions or motion information of the determined i-th upper-right block / lower-left block, motion information corresponding to the lower-right position of the current block may be induced, which may be used as the i-th induced candidate motion information.
[0173] As an example, the motion vector corresponding to the lower-right position of the current block may be induced as described in Equation 5 above, and detailed description is omitted here.
[0174] The plurality of induced candidates may be added to a candidate list for predicting the motion information of the current block. Alternatively, at least one of the plurality of induced candidates may be selectively added to the candidate list.
[0175] As an example, a template matching-based cost may be calculated for the plurality of induced candidates, and the top T induced candidates in ascending order of the calculated cost may be added to the candidate list. The template matching-based cost may be defined as the sample difference between the template region of the current block and the template region of the reference block, where the reference block may be identified by the motion information of the induced candidate. T may be an integer of 1, 2, 3, or more.
[0176] Alternatively, index information for identifying at least one induced candidate to be added to the candidate list among the plurality of induced candidates may be signaled. Only at least one induced candidate identified by the index information may be added to the candidate list.
[0177] In the above-described embodiment, the upper right / lower left block of the current block may be determined by searching for available blocks. However, even when an available block is found by the search, whether the available block is determined as the upper right / lower left block of the current block may be based on whether the motion information of the available block is the same as the motion information of the spatial neighboring blocks of the current block.
[0178] As an example, when the available block and the spatial neighboring blocks have the same motion vector, the available block may be restricted so as not to be determined as the upper right block (or lower left block) of the current block. Alternatively, when the available block and the spatial neighboring blocks have the same motion vector but different reference picture indices, the available block may be determined as the upper right block (or lower left block) of the current block. Alternatively, when the available block and the spatial neighboring blocks have the same motion vector and reference picture index, the available block may be restricted so as not to be determined as the upper right block (or lower left block) of the current block.
[0179] When the available block is restricted so as not to be determined as the upper-right block (or the lower-left block) of the current block, an available block having motion information different from that of the spatial surrounding block may be re-searched, and the available block found by the re-search may be determined as the upper-right block (or the lower-left block) of the current block.
[0180] When the available block found in the search (or re-search) process for the upper-right block and the available block found in the search (or re-search) process for the lower-left block both have the same motion vector and / or reference picture index as the spatial surrounding block, the available block may be restricted so as not to be determined as the upper-right block and the lower-left block of the current block, respectively.
[0181] In the search (or re-search) process for the upper-right block, an available block having motion information different from that of the spatial surrounding block may be found, but in the search (or re-search) process for the lower-left block, no available block having motion information different from that of the available block or the spatial surrounding block may be found. In such a case, a coding block or a sub-block including a position moved by a predetermined distance (d) from the lower-left position of the current block may be determined as the lower-left block of the current block. Here, the predetermined distance (d) may be the distance from the upper-right position of the current block to a predetermined upper-right block. Alternatively, a coding block or a sub-block including a first default position preset identically in the encoding device and the decoding device may be determined as the lower-left block of the current block. Here, the first default position may be the left position of the current block (for example, (-1, H-1)).
[0182] Conversely, in the search (or re-search) process for the lower left block, an available block having motion information different from that of the spatial neighboring blocks may be found, but in the search (or re-search) process for the upper right block, an available block having motion information different from that of the available blocks or the spatial neighboring blocks may not be found. In such a case, a coding block or sub-block including a position moved by a predetermined distance (d) from the upper right position of the current block may be determined as the upper right block of the current block. Here, the predetermined distance (d) may be the distance from the lower left position of the current block to a predetermined lower left block. Alternatively, a coding block or sub-block including a second default position preset identically in the encoding device and the decoding device may be determined as the upper right block of the current block. Here, the second default position may be the upper position of the current block (e.g., (W-1, -1)).
[0183] Alternatively, in the search (or re-search) process for the upper right block and the lower left block, an available block having motion information different from that of the available blocks or the spatial neighboring blocks may not be found. In such a case, a coding block or sub-block including the first default position may be determined as the lower left block of the current block, and a coding block or sub-block including the second default position may be determined as the upper right block of the current block.
[0184] The spatial neighboring blocks to be compared with the available block may include at least one of the upper neighboring block or the left neighboring block of the current block. As an example, in the search process for the upper right block, the available block may be compared with the upper neighboring block and may not be compared with the left neighboring block. Also, in the search process for the lower left block, the available block may be compared with the left neighboring block and may not be compared with the upper right neighboring block. Alternatively, regardless of whether it is the search process for the upper right block or the search process for the lower left block, the available block may be compared with the left and upper neighboring blocks.
[0185] The induced candidate according to the present disclosure may not be added to the candidate list if the temporal candidate is available, and may be added to the candidate list if the temporal candidate is not available. Alternatively, the induced candidate may be added to the candidate list as a new candidate replacing the temporal candidate. Alternatively, the induced candidate may be added to the candidate list independently of the temporal candidate regardless of whether the temporal candidate is available. That is, both the temporal candidate and the induced candidate may be added to the candidate list of the current block. Alternatively, the induced candidate may be adaptively added to the candidate list based on a flag (e.g., temporal_mvp_enabled_flag) indicating the availability of the temporal candidate. The flag may be signaled at a high level such as a sequence parameter set, a picture parameter set, a picture header, a slice header, etc. As an example, when the flag indicates that the temporal candidate is available for the video sequence, the current picture, or the current slice, the induced candidate is not added to the candidate list, and when the flag indicates that the temporal candidate is not available for the video sequence, the current picture, or the current slice, the induced candidate may be added to the candidate list.
[0186] A history-based candidate can mean a candidate having motion information of a block decoded before the current block (hereinafter referred to as a previous block). The motion information of the previous block may be sequentially stored in a buffer having a predetermined size according to the decoding order. The previous block may be a spatially adjacent block adjacent to the current block, or may be a block not adjacent to the current block.
[0187] The induced candidates according to the present disclosure may be added before or after the spatial candidates in the candidate list. The induced candidates may be added between the spatial candidates and the temporal candidates in the candidate list. The induced candidates may be added between the temporal candidates and the history-based candidates in the candidate list. The induced candidates may be added after the history-based candidates.
[0188] Referring to FIG. 4, motion information of the current block may be induced based on the candidate list and the candidate index (S410).
[0189] The candidate index may mean information encoded to induce motion information of the current block. The candidate index may identify one or more candidates among a plurality of candidates belonging to the candidate list.
[0190] The motion vector of the motion information may mean a motion vector in block units. Alternatively, the motion vector of the motion information may mean a motion vector induced in sub-block units of the current block. For this purpose, the current block may be divided into a plurality of NxM sub-blocks. Here, the NxM sub-blocks may be in the form of a rectangle (N>M or N<M) or a square (N = M). The N and M values may be 2, 4, 8, 16, 32 or more.
[0191] Referring to FIG. 4, inter prediction can be performed on the current block based on the induced motion information (S420).
[0192] Specifically, a reference block can be identified using the motion information of the current block. The reference block may be identified for each sub-block of the current block. The reference blocks of each sub-block may belong to one reference picture. That is, the sub-blocks belonging to the current block can share one reference picture. Alternatively, a reference picture index may be independently set for each sub-block of the current block. The identified reference block may be set as the predicted block of the current block.
[0193] FIG. 5 is a diagram showing a schematic configuration of an inter prediction unit 332 that performs an inter prediction method according to the present disclosure.
[0194] The inter prediction method performed by the decoding device has been described with reference to FIG. 4, which may be identically performed by the inter prediction unit 332 of the decoding device, and the detailed description thereof will be omitted below.
[0195] Referring to FIG. 5, the inter prediction unit 332 may include a candidate list generation unit 500, a motion information derivation unit 510, and a prediction sample generation unit 520.
[0196] The candidate list generation unit 500 can generate a candidate list for predicting / deriving the motion information of the current block. The candidate list may be for the merge mode or the AMVP mode, or may be for the affine merge mode or the affine inter mode. The motion information may include at least one of a motion vector, a reference picture index, inter prediction direction information, or weight value information for bi-directional weighted prediction.
[0197] The candidate list generation unit 500 can derive a plurality of candidates including at least one of spatial candidates, temporal candidates, derived candidates, or history-based candidates, and add these to the candidate list. Also, the candidate list generation unit 500 can derive one or more derived candidates and add these to the candidate list. The method of deriving a plurality of candidates and the method of adding a plurality of candidates to the candidate list are as described in FIG. 4, and duplicate descriptions are omitted here.
[0198] The motion information derivation unit 510 can derive the motion information of the current block based on the candidate list and the candidate index. Here, the candidate index is information encoded to derive the motion information of the current block, and can identify one or more candidates among the plurality of candidates belonging to the candidate list.
[0199] When the motion information guiding unit 510 guides the motion information of the current block, it can also guide the motion vector in units of blocks, or can also guide the motion vector in units of sub-blocks of the current block. To guide the motion vector in units of sub-blocks, the motion information guiding unit 510 may divide the current block in units of NxM sub-blocks. Here, the NxM sub-blocks may be in the form of a rectangle (N>M or N<M) or a square (N = M). The values of N and M may be 2, 4, 8, 16, 32 or more.
[0200] The prediction sample generation unit 520 can perform inter prediction based on the motion information obtained by the motion information guiding unit 510 and generate a prediction sample of the current block.
[0201] Specifically, the prediction sample generation unit 520 can identify a reference block using the motion information of the current block, and can generate a prediction block of the current block based on the identified reference block.
[0202] Alternatively, when the motion information guiding unit 510 guides the motion vector in units of sub-blocks, the prediction sample generation unit 520 can identify the reference block for each sub-block of the current block, and generate a prediction block of the current block based on the identified reference block. At this time, the identified reference block may belong to one reference picture. Alternatively, the reference picture index may be independently set for each sub-block of the current block. In this case, any one of the identified reference blocks may belong to a reference picture different from the other one.
[0203] FIG. 6 is an embodiment according to the present disclosure, and is a diagram showing an inter prediction method performed by an encoding device.
[0204] Referring to FIG. 6, a candidate list for determining the motion information of the current block can be generated (S600).
[0205] The candidate list may be for merge mode or AMVP mode, or may be for affine merge mode or affine inter mode. The motion information may include at least one of a motion vector, a reference picture index, inter prediction direction information, or weighted value information for bi-directional weighted prediction.
[0206] A plurality of candidates including at least one of a spatial candidate, a temporal candidate, an induced candidate, or a history-based candidate may be induced, and the plurality of induced candidates may be added to the candidate list. Also, one or more induced candidates may be induced, and these may be added to the candidate list. The method of inducing a plurality of candidates and the method of adding a plurality of candidates to the candidate list are as described in FIG. 4, and duplicate descriptions are omitted here.
[0207] Referring to FIG. 6, inter prediction can be performed on the current block based on the candidate list (S610).
[0208] The predicted sample of the current block may be generated based on at least one optimal candidate among the plurality of candidates belonging to the candidate list.
[0209] Specifically, the reference block may be identified using the motion information of the optimal candidate among the plurality of candidates. Based on the identified reference block, a predicted block of the current block may be generated. Or, the reference block may be identified separately for each sub-block of the current block, and based on the identified reference blocks, a predicted block of the current block may be generated. In this case, the reference blocks corresponding to the sub-blocks of the current block may belong to one reference picture. Or, a reference picture index may be independently set for each sub-block of the current block, and in this case, any one of the identified reference blocks may belong to a different reference picture from another one.
[0210] Also, among a plurality of candidates belonging to the candidate list, at least one optimal candidate for generating a predicted sample of the current block can be determined. A candidate index identifying the determined optimal candidate can be encoded. The encoded candidate index may be inserted into a bit stream.
[0211] FIG. 7 is a diagram showing a schematic configuration of an inter prediction unit 221 that performs an inter prediction method according to the present disclosure.
[0212] The inter prediction method performed by the encoding apparatus with reference to FIG. 6 has been described, and this may be identically performed by the inter prediction unit 221 of the encoding apparatus, and detailed description thereof will be omitted below.
[0213] Referring to FIG. 7, the inter prediction unit 221 may include a candidate list generation unit 700 and a predicted sample generation unit 710.
[0214] The candidate list generation unit 700 can generate a candidate list for determining motion information of the current block. The candidate list may be for a merge mode or an AMVP mode, or may be for an affine merge mode or an affine inter mode. The motion information may include at least one of a motion vector, a reference picture index, inter prediction direction information, or weight value information for bidirectional weighted prediction.
[0215] The candidate list generation unit 700 may derive a plurality of candidates including at least one of a spatial candidate, a temporal candidate, an induced candidate, or a history-based candidate, and add this to the candidate list. Also, the candidate list generation unit 700 may derive one or more induced candidates and add this to the candidate list. The method of deriving a plurality of candidates and the method of adding a plurality of candidates to the candidate list are as described in FIG. 4, and overlapping descriptions are omitted here.
[0216] The prediction sample generation unit 710 can perform inter prediction on the current block based on the candidate list. The prediction sample of the current block may be generated based on at least one optimal candidate among a plurality of candidates belonging to the candidate list.
[0217] Specifically, the prediction sample generation unit 710 can identify a reference block using the motion information of the optimal candidate among a plurality of candidates, and can generate a predicted block of the current block based on the identified reference block. Alternatively, the reference block may be identified separately for each sub-block of the current block, and a predicted block of the current block may be generated based on the identified reference block. In this case, the reference blocks corresponding to the sub-blocks of the current block may belong to one reference picture. Alternatively, a reference picture index may be independently set for each sub-block of the current block. In this case, any one of the identified reference blocks may belong to a different reference picture from another one.
[0218] Also, the prediction sample generation unit 710 may further include a motion information determination unit (not shown). The motion information determination unit can determine at least one optimal candidate for generating the prediction sample of the current block among a plurality of candidates belonging to the candidate list. At this time, the entropy encoding unit 240 can encode a candidate index specifying the determined optimal candidate and insert it into the bit stream.
[0219] In the above-described embodiments, although the method is described based on a sequence diagram by a series of steps or blocks, the embodiments are not limited to the order of the steps, and a certain step may occur in a different order or simultaneously with a step different from that described above. Also, those skilled in the art will understand that the steps shown in the sequence diagram are not exclusive, and other steps may be included, or one or more steps of the sequence diagram may be deleted without affecting the scope of the embodiments of this document.
[0220] The method according to the embodiment of the above-described document may be embodied in the form of software, and the encoding device and / or decoding device according to the document may be included in, for example, a device that performs video processing such as a TV, a computer, a smartphone, a set-top box, a display device, etc.
[0221] When the embodiment in this document is embodied as software, the above-described method may be embodied as a module (process, function, etc.) that executes the above-described functions. The module may be stored in a memory and executed by a processor. The memory may be provided inside or outside the processor and may be connected to the processor by various well-known means. The processor may include an ASIC (application-specific integrated circuit), other chip sets, logic circuits, and / or data processing devices. The memory may include a ROM (read-only memory), a RAM (random access memory), a flash memory, a memory card, a storage medium, and / or other storage devices. That is, the embodiments described in this document may be embodied and executed on a processor, a microprocessor, a controller, or a chip. For example, the functional units shown in each figure may be embodied and executed on a computer, a processor, a microprocessor, a controller, or a chip. In this case, information for the embodiment (for example, information on instructions) or an algorithm may be stored in a digital storage medium.
[0222] In addition, the decoding device and encoding device to which the embodiments of this specification are applied may be included in a multimedia broadcast transmission / reception device, a mobile communication terminal, a home cinema video device, a digital cinema video device, a surveillance camera, a video conversation device, a real-time communication device such as video communication, a mobile streaming device, a storage medium, a camcorder, an on-demand video (VoD) service providing device, an OTT video (Over the top video) device, an Internet streaming service providing device, a three-dimensional (3D) video device, a VR (virtual reality) device, an AR (argumente reality) device, a video phone video device, a transportation means terminal (for example, a vehicle terminal including an autonomous driving vehicle, an airplane terminal, a ship terminal, etc.) and a medical video device, etc., and may be used to process a video signal or a data signal. For example, the OTT video (Over the top video) device may include a game console, a Blu-ray player, an Internet-connected TV, a home theater system, a smartphone, a tablet PC, a DVR (Digital Video Recorder), etc.
[0223] In addition, the processing method to which the embodiments of this specification are applied may be produced in the form of a program executed by a computer and may be stored in a computer-readable recording medium. Similarly, the multimedia data having the data structure according to the embodiments of this specification may also be stored in a computer-readable recording medium. The computer-readable recording medium includes all kinds of storage devices and distributed storage devices in which computer-readable data is stored. The computer-readable recording medium may include, for example, a Blu-ray Disc (BD), a Universal Serial Bus (USB), a ROM, a PROM, an EPROM, an EEPROM, a RAM, a CD-ROM, a magnetic tape, a floppy disk, and an optical data storage device. In addition, the computer-readable recording medium includes a medium embodied in the form of a carrier wave (for example, transmission via the Internet). Also, the bitstream generated by the encoding method may be stored in a computer-readable recording medium or transmitted through a wired or wireless communication network.
[0224] Also, the embodiments of this specification may be embodied as a computer program product by program code, and the program code may be executed by a computer according to the embodiments of this specification. The program code may be stored on a computer-readable carrier.
[0225] FIG. 8 shows an example of a content streaming system to which the embodiments of the present disclosure are applicable.
[0226] Referring to FIG. 8, the content streaming system to which the embodiments of this specification are applied may generally include an encoding server, a streaming server, a web server, a media storage, a user device, and a multimedia input device.
[0227] The encoding server has the role of compressing the content input from a multimedia input device such as a smartphone, a camera, or a camcorder into digital data to generate a bitstream and transmitting this to the streaming server. As another example, when a multimedia input device such as a smartphone, a camera, or a camcorder directly generates a bitstream, the encoding server may be omitted.
[0228] The bitstream may be generated by an encoding method or a bitstream generation method to which the embodiments of this specification are applied, and the streaming server can temporarily store the bitstream in the process of transmitting or receiving the bitstream.
[0229] The streaming server transmits multimedia data to a user device based on a user request via a web server, and the web server serves as a medium to inform the user of what services are available. If the user requests a desired service from the web server, the web server transmits it to the streaming server, and the streaming server transmits multimedia data to the user. At this time, the content streaming system may include a separate control server. In this case, the control server is responsible for controlling commands / responses between each device in the content streaming system.
[0230] The streaming server can receive content from a media storage and / or encoding server. For example, when receiving content from the encoding server, the content can be received in real time. In this case, in order to provide a smooth streaming service, the streaming server can store the bitstream for a certain period of time.
[0231] Examples of the user device include mobile phones, smartphones, laptop computers, digital broadcast terminals, PDAs (personal digital assistants), PMPs (portable multimedia players), navigation devices, slate PCs, tablet PCs, ultrabooks, wearable devices (such as smartwatches, smart glasses, HMDs (head mounted displays)), digital TVs, desktop computers, digital signage, and the like.
[0232] Each server in the content streaming system may be operated as a distributed server, and in this case, the data received by each server may be processed distributively.
[0233] The claims described in this specification may be combined in various ways. For example, the technical features of the method claims in this specification may be combined and embodied as an apparatus, and the technical features of the apparatus claims in this specification may be combined and embodied as a method. Also, the technical features of the method claims in this specification and the technical features of the apparatus claims may be combined and embodied as an apparatus, and the technical features of the method claims in this specification and the technical features of the apparatus claims may be combined and embodied as a method.
Claims
1. generating a candidate list for the current block; deriving motion information of the current block based on any one of a plurality of candidates belonging to the candidate list; performing inter prediction on the current block based on the motion information of the current block, comprising: the plurality of candidates includes at least one of a spatial candidate, a temporal candidate, or an induced candidate; a video decoding method, wherein the motion information of the induced candidate is derived based on at least one of the upper right block of the current block or the lower left block of the current block.
2. the upper right block of the current block is a block including an upper right reference position; the lower left block of the current block is a block including a lower left reference position; the video decoding method according to claim 1, wherein the upper right reference position and the lower left reference position are determined based on at least one of the width or height of the current block.
3. the upper right block of the current block is a block including a first position moved a predetermined first distance in a left or right direction from the upper right reference position; the lower left block of the current block is a block including a second position moved a predetermined second distance in a lower or upper direction from the lower left reference position; the video decoding method according to claim 1, wherein the first position and the second position are located on the same straight line.
4. the upper right block of the current block is a block including a first position moved a predetermined distance in a left or right direction from the upper right reference position; the lower left block of the current block is a block including a second position moved the predetermined distance in a lower or upper direction from the lower left reference position, the video decoding method according to claim 1.
5. the video decoding method according to claim 4, wherein the predetermined distance is adaptively determined based on the size of the current block.
6. the video decoding method according to claim 4, wherein a search direction for determining at least one of the upper right block or the lower left block is determined based on whether the width of the current block is greater than the height of the current block.
7. The video decoding method according to claim 1, wherein at least one of the upper right block or the lower left block is determined by searching for an available block having motion information different from that of at least one of the left block or the upper block of the current block.
8. The search range for determining at least one of the upper right block or the lower left block is divided into a plurality of sub-ranges, The induced candidate is induced for each of the plurality of sub-ranges, and the video decoding method according to claim 1.
9. The video decoding method according to claim 1, wherein the induced candidate is adaptively added to the candidate list based on whether the temporal candidate is available for the current block.
10. The temporal candidate has motion information of a temporal neighboring block corresponding to the current block, The video decoding method according to claim 1, wherein the temporal neighboring block is specified by motion information induced based on at least one of the upper right block or the lower left block.
11. Generating a candidate list for the current block; Performing an inter prediction on the current block based on any one of a plurality of candidates belonging to the candidate list, including: The plurality of candidates includes at least one of a spatial candidate, a temporal candidate, or an induced candidate, A video encoding method, wherein the motion information of the induced candidate is induced based on at least one of the upper right block of the current block or the lower left block of the current block.
12. A computer-readable storage medium for storing a bitstream generated by the video encoding method according to claim 11.
13. Generating a candidate list for the current block, wherein the plurality of candidates belonging to the candidate list includes at least one of a spatial candidate, a temporal candidate, or an induced candidate, and the motion information of the induced candidate is induced based on at least one of the upper right block of the current block or the lower left block of the current block; Generating a predicted block of the current block based on any one of the plurality of candidates; Encoding the current block based on the predicted block to generate a bitstream; A step of transmitting data including the bitstream, and a data transmission method for video information including the same.