Video decoding device, video encoding device, and recording medium
Patent Information
- Application Number
- JP2026102243
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2012-01-31
- Filing Date
- 2026-06-19
- Publication Date
- 2026-09-08
AI Technical Summary
【0027】 本発明によると、クリップされた動きベクトルを利用して映像を符号化/復号化することができる。
Smart Images

Figure 2026143826000001_ABST
Abstract
Description
[Technical Field]
[0001] The present invention relates to video processing, and more particularly, to an encoding / decoding method and apparatus using motion vectors. [Background Art]
[0002] Recently, as broadcasting systems supporting High Definition (HD) resolution have expanded globally not only in the Republic of Korea, many users have become accustomed to high-resolution, high-quality video, and accordingly many organizations have accelerated the development of next-generation video equipment. In addition, as interest in Ultra High Definition (UHD), which supports resolution four or more times higher than that of HDTV along with HDTV, has increased, compression techniques for video with even higher resolution and higher image quality are required.
[0003] For video compression, inter prediction techniques that predict pixel values included in a current picture from previous and / or subsequent pictures, intra prediction techniques that predict pixel values using pixel information within a picture, and / or entropy encoding techniques that assign a short code to a symbol with high occurrence frequency and a long code to a symbol with low occurrence frequency may be used. [Summary of the Invention] [Problem to be Solved by the Invention]
[0004] The present invention provides a method and apparatus for encoding / decoding video using clipped motion vectors.
[0005] The present invention provides a method of clipping a motion vector of a reference picture.
[0006] The present invention provides a method of transmitting information on a motion vector. [Means for Solving the Problem]
[0007] [1] According to one embodiment of the present invention, a video encoding method is provided. The video encoding method includes the steps of: generating a clipped motion vector by clipping the motion vector of a reference picture within a predetermined dynamic range; storing the clipped motion vector in a buffer; and encoding the motion vector of a block to be encoded using the motion vector stored in the buffer.
[0008] In [2][1], the dynamic range is defined by the level of the video codec.
[0009] In [3][1], the dynamic range is determined by a predetermined bit depth, which is defined by the level of the video codec.
[0010] In [4][1], the X and Y components of the motion vector of the reference picture are clipped at different dynamic ranges.
[0011] [5] According to one embodiment of the present invention, a video decoding method is provided. The video decoding method includes the steps of: generating a clipped motion vector by clipping the motion vector of a reference picture within a predetermined dynamic range; storing the clipped motion vector in a buffer; deriving the motion vector of a block to be decoded using the motion vector stored in the buffer; and performing interpredictive decoding using the motion vector of the block to be decoded.
[0012] In [6][5], the dynamic range is defined by the level of the video codec.
[0013] In [7][5], the dynamic range is determined by a predetermined bit depth, which is defined by the level of the video codec.
[0014] In [8][5], the dynamic range is determined by a predetermined bit depth, which is obtained via a sequence parameter set transmitted from the video encoding device.
[0015] In [9][8], the sequence parameter set includes a flag indicating whether the motion vector of the reference picture was clipped and a parameter for obtaining the bit depth.
[0016] In
[10] [9], the video decoding method further includes a step of compressing the motion vector of a reference picture, and the sequence parameter set further includes a flag indicating whether the motion vector of the reference picture has been compressed and a parameter for obtaining the compression ratio of the motion vector of the reference picture.
[0017] In
[11] [5], the video decoding method further includes the step of limiting the representation resolution of the motion vector of the reference picture.
[0018] In
[12] [5], the clipped motion vectors are stored according to their priority.
[0019] In
[13] [5], the clipped motion vector is the motion vector of the block encoded in interprediction mode.
[0020] In
[14] [5], the video decoding method further includes the step of performing scaling with respect to the motion vector of the reference picture.
[0021] In
[15] [5], the X and Y components of the motion vector of the reference picture are clipped at different dynamic ranges.
[0022] In
[16]
[15] , the dynamic range of the X component and the dynamic range of the Y component are defined by the level of the video codec.
[0023]
[17] According to one embodiment of the present invention, there is provided a video decoding device including a reference picture buffer and a motion compensation unit. The reference picture buffer stores a reference picture. The motion compensation unit generates a prediction block using the reference picture and a motion vector of the reference picture. Here, the motion vector of the reference picture is clipped within a predetermined dynamic range.
[0024]
[18] In
[17] , the dynamic range is defined by the level of a video codec.
[0025]
[19] In
[17] , the dynamic range is determined by a predetermined bit depth, and the bit depth is defined by the level of a video codec.
[0026]
[20] In
[17] , the dynamic range is determined by a predetermined bit depth, and the bit depth is obtained via a sequence parameter set transmitted from a video encoding device. Effects of the Invention
[0027] According to the present invention, a video can be encoded / decoded using clipped motion vectors.
[0028] According to the present invention, the size of memory space required for storing motion vectors can be reduced.
[0029] According to the present invention, the memory access bandwidth required when reading data from a memory can be reduced. Brief Description of the Drawings
[0030] [Figure 1] It is a block diagram showing an example of the structure of a video encoding device. [Figure 2] It is a block diagram showing an example of the structure of a video decoding device. [Figure 3] An example of an encoding / decoding target picture and a reference picture is shown. [Figure 4] This is an example of limiting the dynamic range of a motion vector. [Figure 5] This is a sequence diagram showing how to store the motion vector of a reference picture. [Figure 6] This is a sequence diagram showing how to store the motion vector of a reference picture. [Figure 7] This is a sequence diagram showing how to store the motion vector of a reference picture. [Figure 8] This is a sequence diagram showing how to store the motion vector of a reference picture. [Figure 9] This is an example of quantizing motion vectors. [Figure 10] This example demonstrates how to read motion information from a reference picture. [Figure 11] This example demonstrates how to read motion information from a reference picture. [Figure 12] This example demonstrates how to read motion information from a reference picture. [Figure 13] This example demonstrates how to read motion information from a reference picture. [Figure 14] This is a sequence diagram showing a video encoding method according to one embodiment of the present invention. [Figure 15] This is a sequence diagram showing a video decoding method according to one embodiment of the present invention. [Modes for carrying out the invention]
[0031] Hereinafter, embodiments of the present invention will be specifically described with reference to the drawings. However, in describing embodiments of the present invention, if it is determined that a specific description of a known configuration or function would obscure the gist of the present invention, such detailed description will be omitted.
[0032] When one component is described as being "connected" or "linked" to another component, it means that it may be directly connected to or linked to the other component, but another component may also be present in between. Furthermore, when a specific component is described as "including" in this invention, it does not exclude other components, but rather means that additional components may be included within the scope of the embodiments or technical ideas of this invention.
[0033] Terms such as “first,” “second,” etc., can be used to describe various components, but the components are not limited by these terms. That is, the terms are used solely for the purpose of distinguishing one component from another. For example, as long as it does not fall outside the scope of the rights of the present invention, the first component may be named the second component, and similarly, the second component may be named the first component.
[0034] Furthermore, the components shown in the embodiments of the present invention are illustrated independently to demonstrate that they perform distinct and characteristic functions, and this does not mean that each component cannot be embodied in a single piece of hardware or software. That is, each component is distinguished for the sake of explanation, and multiple components can be integrated to operate as a single component, or a single component can be divided into multiple components to operate, and this is included within the scope of the present invention as long as it does not deviate from the essence of the present invention.
[0035] Furthermore, some components are not essential for performing the essential functions of the present invention, but are optional components for improving performance. The present invention can also be embodied in a structure that includes only the essential components, excluding the optional components, and such a structure is also included within the scope of the rights of the present invention.
[0036] Figure 1 is a block diagram showing an example of the structure of a video encoding device.
[0037] Referring to Figure 1, the video encoding device 100 includes a motion prediction unit 111, a motion compensation unit 112, an intra prediction unit 120, a switch 115, a subtractor 125, a conversion unit 130, a quantization unit 140, an entropy encoding unit 150, an inverse quantization unit 160, an inverse conversion unit 170, an adder 175, a filter unit 180, and a reference picture buffer 190.
[0038] The video encoding device 100 encodes the input video into intra prediction mode or inter prediction mode and outputs a bit stream. Intra prediction means prediction within a single frame, and inter prediction means prediction between frames. The video encoding device 100 transitions between intra prediction mode and inter prediction mode via the switching of switch 115. After generating prediction blocks for the input blocks of the input video, the video encoding device 100 encodes the residual between the input blocks and the prediction blocks.
[0039] In intra-prediction mode, the intra-prediction unit 120 generates predicted blocks by performing spatial prediction using the pixel values of already encoded blocks surrounding the current block.
[0040] In interprediction mode, the motion prediction unit 111 searches for the reference block that best matches the input block within the reference picture stored in the reference picture buffer 190 during the motion prediction process and obtains a motion vector. The motion compensation unit 112 generates a predicted block by performing motion compensation using the motion vector. Here, the motion vector is a two-dimensional vector used for interprediction and indicates the offset between the block currently being encoded / decoded and the reference block.
[0041] The subtractor 125 generates a residual block based on the residual between the input block and the predicted block, and the transformer 130 transforms the residual block and outputs a transform coefficient. The quantization unit 140 quantizes the transform coefficient and outputs a quantized coefficient.
[0042] The entropy coding unit 150 outputs a bitstream by performing entropy coding based on the information obtained during the coding / quantization process. Entropy coding reduces the size of the bit sequence for a symbol to be coded by representing frequently occurring symbols with a small number of bits. Therefore, an improvement in video compression performance can be expected through entropy coding. The entropy coding unit 150 can use coding methods such as exponential golomb, CAVLC (Context-Adaptive Variable Length Coding), and CABAC (Context-Adaptive Binary Arithmetic Coding) for entropy coding.
[0043] The encoded picture, to be used as a reference picture for performing interpredictive coding, needs to be decoded and stored again. Therefore, the inverse quantization unit 160 inversely quantizes the quantized coefficients, and the inverse transform unit 170 inversely transforms the inversely quantized coefficients to output the restored residual block. The adder 175 adds the restored residual block to the prediction block to generate a restored block.
[0044] The filter section 180, also known as an adaptive in-loop filter, applies at least one of the following to the restored block: deblocking filtering, SAO (Sample Adaptive Offset) compensation, or ALF (Adaptive Loop Filtering). Deblocking filtering removes block distortion that occurs at the boundaries between blocks, while SAO compensation adds an appropriate offset to the pixel value to compensate for coding errors. ALF performs filtering based on a comparison between the restored image and the original image.
[0045] Meanwhile, the reference picture buffer 190 stores the restored block that has passed through the filter unit 180.
[0046] Figure 2 is a block diagram showing an example of the structure of a video decoding device.
[0047] Referring to Figure 2, the video decoding device 200 includes an entropy decoding unit 210, an inverse quantization unit 220, an inverse transformation unit 230, an intra prediction unit 240, a motion compensation unit 250, an adder 255, a filter unit 260, and a reference picture buffer 270.
[0048] The video decoder 200 outputs a reconstructed video by decoding the bitstream into intra-prediction mode or inter-prediction mode. The video decoder 200 transitions between intra-prediction mode and inter-prediction mode via a switch. The video decoder 200 obtains residual blocks from the bitstream to generate prediction blocks, and then adds the residual blocks and prediction blocks to generate a reconstructed block.
[0049] The entropy decoding unit 210 performs entropy decoding based on a probability distribution. The entropy decoding process is the reverse process of the entropy coding process described above. That is, the entropy decoding unit 210 generates symbols containing quantized coefficients from a bitstream in which frequently occurring symbols are represented by a small number of bits.
[0050] The inverse quantization unit 220 inversely quantizes the quantized coefficients, and the inverse transformation unit 230 inversely transforms the inversely quantized coefficients to generate residual blocks.
[0051] In intra-prediction mode, the intra-prediction unit 240 generates predicted blocks by performing spatial prediction using the pixel values of already decoded blocks surrounding the current block.
[0052] In interpretation mode, the motion compensation unit 250 generates prediction blocks by performing motion compensation using motion vectors and reference pictures stored in the reference picture buffer 270.
[0053] The adder 255 adds the prediction block to the residual block, and the filter unit 260 applies at least one of the following to the block that has passed through the adder: deblocking filtering, SAO compensation, and ALF, to output the reconstructed image.
[0054] Meanwhile, the restored image can be stored in the reference picture buffer 270 and used for motion compensation.
[0055] Hereafter, "block" refers to a unit of encoding / decoding. During the encoding / decoding process, the video is divided into predetermined sizes and encoded / decoded. Therefore, blocks may also be called coding units (CU), prediction units (PU), or transform units (TU), and a single block can be divided into smaller subblocks.
[0056] Here, a prediction unit refers to the basic unit of prediction and / or motion compensation execution. A prediction unit can be divided into multiple partitions, each of which is called a prediction unit partition. When a prediction unit is divided into multiple partitions, each prediction unit partition can become the basic unit of prediction and / or motion compensation execution. In the embodiments of the present invention below, a prediction unit may also refer to a prediction unit partition.
[0057] On the other hand, HEVC (High Efficiency Video Coding) uses a motion vector prediction method based on improved motion vector prediction (AMVP: Advanced Motion Vector Prediction).
[0058] In motion vector prediction methods based on improved motion vector prediction, it is possible to utilize not only the motion vector (MV) of the reconstructed block located around the block to be encoded / decoded, but also the motion vector of a block located at the same or corresponding position as the block to be encoded / decoded within the reference picture. In this case, a block located at the same or spatially corresponding position as the block to be encoded / decoded within the reference picture is called a collocated block, and the motion vector of a collocated block is called a collocated motion vector or temporal motion vector. However, a collocated block is not necessarily a block located at the exact same position as the block to be encoded / decoded in the reference picture; it may also be a block located at a similar position, i.e., a corresponding position, to the block to be encoded / decoded.
[0059] In the motion information merge method, motion information is inferred not only from the reconstructed blocks located nearby but also from equivalent position blocks, and used as the motion information for the block to be encoded / decoded. At that time, the motion information includes at least one of the following: the reference picture index required during inter-prediction, the motion vector, inter-prediction mode information indicating whether it is uni-direction or bi-direction, the reference picture list, and prediction mode information indicating whether it is encoded in intra-prediction mode or inter-prediction mode.
[0060] The predicted motion vector for the block to be encoded / decoded may be not only the motion vector of spatially adjacent blocks to the block to be encoded / decoded, but also the motion vector of equivalent blocks that are temporally adjacent to the block to be encoded / decoded.
[0061] Figure 3 shows an example of a picture to be encoded / decoded and a reference picture.
[0062] In Figure 3, block X represents the block to be encoded / decoded within the picture 310, blocks A, B, C, D, and E represent the restored blocks located around the block to be encoded / decoded. Block T in the reference picture 320 represents the equivalent position block located at the position corresponding to the block to be encoded / decoded.
[0063] For the blocks to be encoded / decoded, the motion vectors to be used as predicted motion vectors can be determined via the motion vector predictor index.
[0064] [Table 1]
[0065] As shown in Table 1, motion vector predictor indices (mvp_idx_l0, mvp_idx_l1) for each reference picture list are sent to the decoder, and the decoder uses the motion vector that is the same as the motion vector predicted by the encoder as the predicted motion vector.
[0066] When encoding / decoding a block using the motion vectors of spatially adjacent surrounding blocks, the motion vectors can be stored using relatively small memory. However, when using temporal motion vectors, all motion vectors of the reference picture must be stored in memory, requiring a relatively large amount of memory, and increasing the memory access bandwidth required when reading data from memory. Therefore, in application environments where memory space is insufficient or power consumption must be minimized, such as in portable receiver terminals, it is necessary to store temporal motion vectors more efficiently.
[0067] On the other hand, a conventional technique for storing motion vectors in memory involves reducing the spatial resolution of the motion vectors. In this method, the motion vectors are compressed by an arbitrary ratio before being stored in memory. For example, motion vectors that are normally stored in 4x4 block units are stored in blocks larger than 4x4, reducing the number of motion vectors stored. In this case, information regarding the compression ratio is transmitted to adjust the block size of the stored motion vectors. This information is transmitted via a Sequence Parameter Set (SPS), as shown in Table 2.
[0068] [Table 2]
[0069] Referring to Table 2, if motion_vector_buffer_comp_flag is 1, the motion vector buffer compression process is performed.
[0070] `motion_vector_buffer_comp_ratio_log2` indicates the compression ratio of the motion vector buffer compression process. If `motion_vector_buffer_comp_ratio_log2` does not exist, it is inferred to be 0, and the motion vector buffer compression ratio is expressed as shown in Equation 1.
[0071]
number
[0072] For example, if all 4x4 blocks of a 1920x1080 (1080p) picture have different motion vectors, and two reference picture lists are used, with two reference pictures in each list, then a total of 3.21 Mbytes of memory space is required to store the temporal motion vectors, as shown below.
[0073] 1. Bit depth of 26bis per motion vector
[0074] (1) Dynamic range of the X component of the motion vector: -252 to +7676 (bit depth: 13 bits)
[0075] (2) Dynamic range of the Y component of the motion vector: -252 to +4316 (bit depth: 13 bits)
[0076] (3) (The dynamic range of each component of the motion vector was calculated relative to the first prediction unit in the picture.)
[0077] If all 2.4x4 block units have different motion vectors: 480x270 = 129600 blocks
[0078] 3. Use two motion vectors for each block.
[0079] 4. Number of reference picture lists: 2
[0080] 5. Use two reference pictures per reference picture list.
[0081] ⇒ 26 bits × 129,600 blocks × 2 motion vectors × 2 reference picture lists × 2 reference pictures = 2,695,6800 bits = 3.21 Mbytes
[0082] The aforementioned method for reducing the spatial resolution of motion vectors can reduce the required memory space and memory access bandwidth by utilizing the spatial correlation of motion vectors. However, this method for reducing the spatial resolution of motion vectors did not limit the dynamic range of the motion vectors.
[0083] If the memory space size is reduced to 1 / 4, the required memory space in the example above will decrease to approximately 0.8 Mbytes. Furthermore, if the dynamic range of the motion vector is limited to only 6 bits per component of the motion vector, the required memory space can be further reduced to 0.37 Mbytes.
[0084] Therefore, in this invention, the dynamic range of the motion vector is limited in order to reduce the size of the memory space required to store the motion vector and the memory access bandwidth required when reading data from memory. The motion vector of the reference picture with a limited dynamic range can be used as a temporal motion vector in the block to be encoded / decoded.
[0085] Below, dynamic range refers to the interval between the minimum and maximum values that the negative and positive components of a motion vector can have, relative to 0. Bit depth indicates the size of the space required to store the motion vector, and may also refer to bit width. Unless otherwise specified, motion vector refers to the motion vector of the reference picture, i.e., the temporal motion vector.
[0086] If each component of a motion vector falls outside its dynamic range, it is represented by the minimum or maximum value of that dynamic range. For example, if the X component of a motion vector is 312 and the maximum value of the dynamic range for each component of the motion vector is 256, then the X component of the motion vector is limited to 256.
[0087] Similarly, if the bit depth of each component of the motion vector is 16 bits and the motion vector is (-36, 24), then limiting the bit depth of each component of the motion vector to 6 bits results in each component of the motion vector having a dynamic range of -32 to +31, and the motion vector is represented as (-32, 24), which is within the dynamic range.
[0088] Furthermore, if the bit depth of each component of the motion vector is 16 bits and the motion vector is (-49, 142), then limiting the bit depth of each component of the motion vector to 9 bits will result in each component of the motion vector having a dynamic range of -256 to +255, and the motion vector will remain unchanged and be represented as (-49, 142).
[0089] Figure 4 shows an example of limiting the dynamic range of a motion vector.
[0090] Referring to Figure 4, if the dynamic range of a motion vector with a dynamic range of -4096 to +4095 is limited to -128 to +127, the bit depth can be reduced from 13 bits to 8 bits.
[0091] Each component of the time motion vector is clipped as shown in Equations 2 and 3 so that it can be stored in a bit depth of N bits(s), where N is a positive integer.
[0092]
number
[0093]
number
[0094] Here, MV_X is the X component of the motion vector, MV_Y is the Y component of the motion vector, min(a,b) is an operation that outputs the smaller value of a and b, and max(a,b) is an operation that outputs the larger value of a and b. clippedMV_X and clippedMV_Y are the X and Y components of the clipped temporal motion vector, respectively, and are stored in memory and used as the temporal motion vector of the block to be encoded / decoded.
[0095] For example, as shown in Table 3, if the memory space is 48 bytes in size and a bit depth of 16 bits is used for each component of the motion vector, a total of 12 motion vectors can be stored.
[0096] [Table 3]
[0097] However, by using only 8 bits of bit depth for each component of the motion vector, a total of 24 motion vectors can be stored, as shown in Table 4.
[0098] [Table 4]
[0099] Therefore, according to the present invention, when the image restored by the encoding and / or decoding device is stored in a decoded picture buffer (DPB) after passing through an in-loop filtering process such as a deblocking filter and an adaptive loop filter, the dynamic range of the motion vector is limited to store the motion vector of the reference picture. Here, the restored picture buffer may also mean the reference picture buffer shown in Figure 1 or Figure 2.
[0100] I. Motion Vector Clipping Process
[0101] The process of clipping each component of the motion vector is performed when the slice type is not an I-picture. The motion vector clipping process is performed in tree block or Largest Coding Unit (LCU) units after the filtering process is complete.
[0102] The inputs to the motion vector clipping process are (xP, yP), which is the pixel position of the top-left corner of the prediction unit in the current picture, and the motion vector matrices MvL0 and MvL1, and the outputs are the clipped motion vector matrices CMvL0 and CMvL1.
[0103] The operations of equations 4 through 7 are performed on matrices MvL0, MvL1, CMvL0, and CMvL1.
[0104]
number
[0105]
number
[0106]
number
[0107]
number
[0108] Here, TMVBitWidth represents the bit depth of the motion vector, and Clip3(a,b,c) is a function that clips c so that it lies within the range between a and b.
[0109] II. Motion Vector Storage Process
[0110] Figures 5 through 8 are sequential diagrams showing how to store the motion vector of a reference picture.
[0111] Referring to Figure 5, the motion vector of the reference picture can be stored by using both the video buffer that stores the restored image and the motion vector buffer that stores the motion vector. At that time, the restored image undergoes an in-loop filtering process (S510), and the motion vector is stored (S540) with its dynamic range limited (S520).
[0112] Furthermore, referring to Figure 6, both the video buffer and the motion vector buffer are used, and the motion vector is stored (S640) after undergoing a dynamic range limiting process (S620) and a spatial resolution reduction process (S630).
[0113] Also, referring to Figure 7, the restored video is stored in the video buffer (S740) after going through an in-loop filtering process (S710), and the motion vectors are stored in the motion vector buffer (S750) with their dynamic range limited (S720).
[0114] Also, referring to Figure 8, the restored video is stored in the video buffer (S840) after an in-loop filtering process (S810), and the motion vector is stored (S850) after a dynamic range limiting process (S820) and a spatial resolution reduction process (S830).
[0115] On the other hand, in the embodiments shown in Figures 6 and 8, the dynamic range limiting process (S620, S820) and the spatial resolution reduction process (S630, S830) are not limited to a specific order and can be changed.
[0116] Furthermore, to further reduce memory access bandwidth, the dynamic ranges for each component of the motion vector can be restricted to be different from each other. For example, only one of the dynamic ranges, either the X component or the Y component, can be restricted, or the dynamic range of the Y component can be restricted more than the dynamic range of the X component.
[0117] The limited dynamic range of the motion vector is transmitted via a sequence parameter set, picture parameter set (PPS), or slice header, and the decoder similarly limits the dynamic range of the temporal motion vector within the sequence, picture, or slice. At the same time, the bit depth, which is the size of the memory space required to store the motion vector represented within the dynamic range, can also be transmitted. Alternatively, instead of storing the motion vector using a fixed bit depth, the temporal motion vector can be efficiently stored to match the motion characteristics of the image by utilizing the dynamic range transmitted via the sequence parameters, picture parameter set, or slice header.
[0118] On the other hand, motion vectors can be quantized and stored. When motion vectors are quantized and stored, their precision decreases. Quantization methods include uniform quantization, where the step size is uniform, and non-uniform quantization, where the step size is not uniform. The step size for quantization is set to a fixed value predetermined by the encoding and decoding devices, or it is transmitted from the encoding device to the decoding device via a sequence parameter set, picture parameter set, or slice header. The decoding device can use the quantized motion vector as is, or dequantize it. Figure 9 shows an example of quantizing a motion vector. Referring to Figure 9, if the motion vector has component values between 32 and 48, the motion vector is quantized to 40.
[0119] Furthermore, motion vectors can be stored with a limited representation resolution. Representation resolution refers to integer pixel units (1 pixel unit), fractional pixel units (1 / 2 pixel unit, 1 / 4 pixel unit, etc.). For example, the resolution of a motion vector processed in 1 / 4 pixel units can be stored in integer pixels. The representation resolution of the motion vector is set to a fixed value predetermined by the encoding and decoding devices, or it is transmitted from the encoding device to the decoding device via a sequence parameter set, picture parameter set, or slice header, etc.
[0120] Furthermore, for only a portion of the temporal motion vectors stored in memory, the dynamic range limiting process, spatial resolution reduction process, and quantization process can be performed.
[0121] When storing motion vectors with a limited dynamic range, additional information about the motion vector's dynamic range can be stored in memory. For example, to limit the motion vector's dynamic range to -128 to +127, a flag of 1 can be stored; to limit it to -32 to +31, a flag of 0 can be stored. In this case, the flag information can be stored together with the motion vector, or in a different memory location than the one where the motion vector is stored. If the flag information and the motion vector are stored in different memories, the flag information can be arbitrarily approached when a particular motion vector is known to be within its stored dynamic range. Furthermore, information about the dynamic range in which some motion vectors are stored can be transmitted via a sequence parameter set, picture parameter set, or slice header, allowing the decoder to operate similarly to the encoder.
[0122] When storing motion vectors with reduced spatial resolution, information regarding the block size of the motion vectors can be added and stored in memory. For example, a flag of 1 can be added when the block size of the motion vectors is 4x4, and a flag of 0 can be added when it is 16x16. In this case, the flag information can be stored together with the motion vectors, or in a different memory location than the one where the motion vectors are stored. If the flag information and motion vectors are stored in different memories, the flag information can be arbitrarily approached when knowing the block size in which a particular motion vector is stored. Furthermore, information regarding the block size in which some motion vectors are stored can be transmitted via a sequence parameter set, picture parameter set, or slice header, allowing the decoder to operate similarly to the encoder.
[0123] When quantizing and storing motion vectors, information regarding the precision of the motion vectors can be added and stored in memory. For example, if the quantization step size is set to 4, a flag of 1 can be added, and if the quantization step size is set to 1, a flag of 0 can be added. In this case, the flag information can be stored together with the motion vector, or in a different memory than the one where the motion vector is stored. If the flag information and the motion vector are stored in different memories, when a particular motion vector is quantized and stored, it can be made to arbitrarily approach the flag information. Furthermore, information regarding the quantized and stored step size of some motion vectors can be transmitted via a sequence parameter set, picture parameter set, or slice header, so that the decoder operates similarly to the encoder.
[0124] Furthermore, when storing motion information in memory, the spatial resolution of the motion vector can be reduced during storage. In this case, the motion information includes at least one of the following: the reference picture index required during interprediction, the motion vector, interprediction mode information indicating whether it is uni-direction or bi-direction, the reference picture list, and prediction mode information indicating whether it is encoded in intraprediction mode or interprediction mode.
[0125] For example, the motion information of the prediction unit with the largest partition size among multiple motion information units in a specific region can be stored in memory as representative motion information. In this case, the specific region may include the region within the block to be encoded / decoded and the regions of surrounding blocks. Also, if the entire picture or slice is divided into a fixed size, the specific region may include the region containing the block where the motion information is stored.
[0126] For example, after removing motion information encoded using methods such as motion information merging or encoding information skipping from multiple motion information contained in a specific region, representative motion information can be stored in memory.
[0127] For example, the most frequently occurring motion information among multiple motion pieces contained in a specific region can be stored in memory as representative motion information. At that time, the number of occurrences of the motion information can be calculated for each block size.
[0128] For example, motion information for a specific location can be stored among multiple motion information items contained within a specific region. In this case, the specific location is a location contained within the specific region, and may be a fixed location within that region. Furthermore, the specific location can be selected from among multiple locations. When multiple locations are used, a priority can be determined for each location, and the motion information can be stored in memory according to that priority.
[0129] For example, when storing multiple motion information items contained within a specific region in memory, motion information for blocks encoded in intra-prediction mode, blocks encoded in PCM (Pulse Coded Modulation) mode, slices, or outside picture boundaries is not stored in memory because no motion information exists in those areas.
[0130] If, in the aforementioned examples, motion information for a specific location is not available, motion information from an equivalent block, motion information from a previously encoded block, or motion information from a surrounding block can be used as the motion information for that location. In this case, the specific location is the location of one sample or block within a block surrounding the block being encoded / decoded. For example, if motion information for a specific location is not available, the median or average of the motion information from the inter-predicted encoded blocks around that location can be stored in memory. For example, if motion information for a specific location is not available, the average of the motion information from the surrounding blocks can be stored in memory. When calculating the median and average, if the motion information of the surrounding blocks differs from one or more of the reference picture index, reference picture list, and inter-prediction mode information, the motion vector can be adjusted in size by the reference picture index, reference picture list, inter-prediction mode information, and picture display order count.
[0131] III. Derivation Process of Motion Vectors
[0132] When motion information is stored in memory using the motion information method described above, and the motion information of a reference picture is used with the motion vector prediction method, the improved motion vector prediction method, or the motion information merging method, the stored motion information can be read.
[0133] For example, it is possible to read the movement information of the corresponding position within a reference picture to the block to be encoded / decoded. In this case, the position corresponding to the block to be encoded / decoded within the reference picture is either a fixed position within a specific region or a position relative to the block to be encoded / decoded.
[0134] Figures 10 to 13 show examples of reading motion information from a reference picture.
[0135] In Figures 10 to 13, block X represents the blocks to be encoded / decoded within the encoded / decoded pictures 1010, 1110, 1210, and 1310, while blocks A, B, C, D, and E represent the restored blocks located around the encoded / decoded blocks. Block T in reference pictures 1020, 1120, 1220, and 1320 represents the equivalent location block corresponding to the encoded / decoded block. Block Y in reference picture 1320 in Figure 13 represents the block corresponding to a location outside the encoded / decoded block.
[0136] Referring to Figure 10, motion information corresponding to the upper-left pixel position within the block X to be encoded / decoded can be read.
[0137] Referring to Figure 11, motion information corresponding to the central pixel position within the block X to be encoded / decoded can be read.
[0138] Referring to Figure 12, motion information corresponding to the position of the bottom-right pixel in the block X to be encoded / decoded within the reference picture can be read.
[0139] Referring to Figure 13, motion information corresponding to the external position of the block X to be encoded / decoded can be read within the reference picture.
[0140] Using motion information stored in memory, i.e., motion information of a reference picture, encoding / decoding methods such as motion vector prediction, improved motion vector prediction, motion information merging, and motion information merging skip can be performed.
[0141] A motion vector can be stored in memory using at least one of the following methods: a method for limiting the dynamic range of the motion vector, a method for reducing the spatial resolution of the motion vector, a method for quantizing the motion vector, or a method for reducing the representation resolution of the motion vector. The stored motion vector can then be used for predicting the motion vector of the block to be encoded / decoded and for merging motion information.
[0142] The process of reading the motion vector of a reference picture from memory is called the derivation process of the temporal motion vector. In the derivation process of the temporal motion vector, TMVbitWidth indicates the bit width of the temporal motion vector stored in memory.
[0143] The inputs to the temporal motion vector derivation process are (xP, yP), which is the pixel position of the top-left corner of the prediction unit in the current picture; nPSW and nPSH, which are the horizontal and vertical dimensions of the brightness prediction unit; and refIdxLX, which is the reference picture index of the current prediction unit partition. The outputs are the motion vector prediction value mxLXCl and the availability flag availableFlagLXCol.
[0144] RefPicOrderCnt(pic,refidx,LX) is a function that outputs the PicOrderCnt of the referenced picture RefPicListX[refidx] of pic. Here, X can be 0 or 1. The PicOrderCnt of a referenced picture persists until the picture is treated as "non-existing". Clip3(a,b,c) is a function that clips c so that it lies within the range between a and b.
[0145] A colPic containing an equilocated partition becomes RefPicList1[0] if its slice type is B-slice and collocated_from_l0_flag is 0. Otherwise, i.e., if its slice type is P-slice and collocated_from_l0_flag is 1, it becomes RefPicList0[0].
[0146] The positions of colPu and colPu (xPCol, yPCol) are induced in the following order.
[0147] 1. The position of the luminance component at the lower right corner of the current prediction unit (xPRb, yPRb) is defined as shown in equations 8 and 9.
[0148]
number
[0149]
number
[0150] 2. colPu is encoded in intra-predictive mode, and colPu does not exist,
[0151] (1) The position of the luminance component at the center of the current prediction unit (xPCtr, yPCtr) is defined as shown in equations 10 and 11.
[0152]
number
[0153]
number
[0154] (2) colPu is set to a prediction unit in colPic that includes the position ((xPCtr>>4)<<4,(yPCtr>>4)<<4).
[0155] 3. (xPCol, yPCol) is the value obtained from the position of the luminance component at the top left corner of colPic to the position of the luminance component at the top left corner of colPu.
[0156] mvLXCol and availableFlagLXCol are derived as follows:
[0157] 1. If colPu is encoded in intra-predictive mode and colPu does not exist, each component of mvLXCol becomes 0, and availableFlagLXCol also becomes 0.
[0158] 2. Otherwise, i.e., colPu is not encoded in intra-predictive mode and colPu exists, mvLXCol and refIdxCol are induced as follows:
[0159] (1) If PredFlagL0[xPCol][yPCol] is 0, the motion vector mvCol is determined to be MvL1[xPCol][yPCol], and the reference picture index refIdxCol is determined to be RefIdxL1[xPCol][yPCol].
[0160] (2) Otherwise, i.e., if PredFlagL0[xPCol][yPCol] is 1, the following process is performed.
[0161] 1) If PredFlagL1[xPCol][yPCol] is 0, the motion vector mvCol is determined to be MvL0[xPCol][yPCol], and the reference picture index refIdxCol is determined to be RefIdxL0[xPCol][yPCol].
[0162] 2) Otherwise, i.e., if PredFlagL1[xPCol][yPCol] is 1, the following process is executed.
[0163] aX becomes either 0 or 1, and the following assignment process is executed.
[0164] i.RefIdxColLX is assigned to RefIdxLX[xPCol][yPCol].
[0165] ii. If PicOrderCnt(colPic) is less than PicOrderCnt(currPic) and RefPicOrderCnt(colPic,RefIdxColLX,LX) is greater than PicOrderCnt(currPic), or if PicOrderCnt(colPic) is greater than PicOrderCnt(currPic) and RefPicOrderCnt(colPic,RefIdxColLX,LX) is less than PicOrderCnt(currPic), then MvXCross is assigned to 1.
[0166] iii. Otherwise, i.e., if PicOrderCnt(colPic) is less than PicOrderCnt(currPic) and RefPicOrderCnt(colPic,RefIdxColLX,LX) is less than or equal to PicOrderCnt(currPic), or if PicOrderCnt(colPic) is greater than PicOrderCnt(currPic) and RefPicOrderCnt(colPic,RefIdxColLX,LX) is greater than or equal to PicOrderCnt(currPic), then MvXCross is assigned to 1.
[0167] b. If any of the following conditions are met, the motion vector mvCol, reference picture index refIdxCol, and ListCol are determined to be MvL1[xPCol][yPCol], RefIdxColL1, and L1, respectively.
[0168] i.Mv0Cross is 0, Mv1Cross is 1.
[0169] ii. Mv0Cross and Mv1Cross are the same, and the reference picture list is L1.
[0170] c. Otherwise, the motion vector mvCol, reference picture index refIdxCol, and ListCol are determined to be MvL0[xPCol][yPCol], RefIdxColL0, and L0, respectively.
[0171] 3) availableFlagLXCol becomes 1, and the calculation in formula 12 or formulas 13 through 18 is executed.
[0172] a. If PicOrderCnt(colPic)-RefPicOrderCnt(colPic,refIdxCol,ListCol) is PicOrderCnt(currPic)-RefPicOrderCnt(currPic,refIdxLX,LX),
[0173]
number
[0174] b. Otherwise,
[0175]
number
[0176]
number
[0177]
number
[0178]
number
[0179] Here, td and tb are as shown in equations 17 and 18.
[0180]
number
[0181]
number
[0182] In other words, referring to equations 13 through 16, mvLXCol is derived from the scaled version of the motion vector mvCol.
[0183] On the other hand, even if a motion vector is clipped within the dynamic range, it can be moved out of the dynamic range again if the clipped motion vector is scaled. Therefore, after deriving the scaled motion vector, the dynamic range of the motion vector can be limited. In that case, equations 15 and 16 can be replaced by equations 19 and 20, respectively.
[0184]
number
[0185]
number
[0186] IV. Method for transmitting information to clip temporal motion vectors in a decoding device.
[0187] The following describes how to transmit the information necessary to clip the temporal motion vector in the decoding device in the same way as the encoding device.
[0188] The TMVBitWidth obtained during the aforementioned temporal motion vector derivation process can be transmitted from the encoding device to the decoding device via a sequence parameter set, picture parameter set, or slice header, etc.
[0189] [Table 5]
[0190] In Table 5, bit_width_temporal_motion_vector_minus8 represents the bit width of the temporal motion vector component. If bit_width_temporal_motion_vector_minus8 does not exist, it is inferred to 0, and the bit width of the temporal motion vector component is expressed as shown in Equation 21.
[0191]
number
[0192] 1. Information transmission method 1 - When the motion vector is compressed and the bit depth of the motion vector is limited.
[0193] [Table 6]
[0194] Referring to Table 6, if motion_vector_buffer_comp_flag is 1, the motion vector buffer compression process is performed.
[0195] `motion_vector_buffer_comp_ratio_log2` indicates the compression ratio of the motion vector buffer compression process. If `motion_vector_buffer_comp_ratio_log2` does not exist, it is inferred to 0, and the motion vector buffer compression ratio is expressed as shown in Equation 22.
[0196]
number
[0197] Furthermore, referring to Table 6, if bit_depth_temporal_motion_vector_constraint_flag is 1, the temporal motion vector bit depth limiting process is executed.
[0198] bit_depth_temporal_motion_vector_minus8 indicates the bit depth of the temporal motion vector. If bit_depth_temporal_motion_vector_minus8 does not exist, it is by analogy to 0, and the bit depth of the temporal motion vector is expressed as shown in Equation 23.
[0199]
number
[0200] 2. Information transmission method 2 - When limiting the bit depth of the motion vector
[0201] [Table 7]
[0202] Referring to Table 7, if bit_depth_temporal_motion_vector_constraint_flag is 1, the temporal motion vector bit depth limiting process is performed.
[0203] bit_depth_temporal_motion_vector_minus8 indicates the bit depth of the temporal motion vector. If bit_depth_temporal_motion_vector_minus8 does not exist, it is by analogy to 0, and the bit depth of the temporal motion vector is expressed as shown in Equation 24.
[0204]
number
[0205] 3. Information transmission method 3 - Limiting the bit depth of the motion vector
[0206] [Table 8]
[0207] Table 8's bit_depth_temporal_motion_vector_minus8 indicates the bit depth of the temporal motion vector. If bit_depth_temporal_motion_vector_minus8 does not exist, it is inferred to 0, and the bit depth of the temporal motion vector is expressed as shown in Equation 25.
[0208]
number
[0209] 4. Information transmission method 4 - Limiting the bit depth for each of the X and Y components of the motion vector.
[0210] [Table 9]
[0211] Referring to Table 9, if bit_depth_temporal_motion_vector_constraint_flag is 1, the temporal motion vector bit depth limiting process is performed.
[0212] `bit_depth_temporal_motion_vector_x_minus8` represents the bit depth of the X component of the temporal motion vector. If `bit_depth_temporal_motion_vector_x_minus8` does not exist, it is by analogy to 0, and the bit depth of the temporal motion vector is expressed as shown in Equation 26.
[0213]
number
[0214] bit_depth_temporal_motion_vector_y_minus8 represents the bit depth of the Y component of the temporal motion vector. If bit_depth_temporal_motion_vector_x_minus8 does not exist, it is by analogy to 0, and the bit depth of the temporal motion vector is expressed as shown in Equation 27.
[0215]
number
[0216] 5. Information transmission method 5 - When compressing motion vectors and limiting the bit depth of motion vectors
[0217] [Table 10]
[0218] Referring to Table 10, if motion_vector_buffer_comp_flag is 1, the motion vector buffer compression process is performed.
[0219] `motion_vector_buffer_comp_ratio_log2` indicates the compression ratio of the motion vector buffer compression process. If `motion_vector_buffer_comp_ratio_log2` does not exist, it is inferred to 0, and the motion vector buffer compression ratio is expressed as shown in Equation 28.
[0220]
number
[0221] V. Defining Dynamic Range via Video Codec Levels
[0222] The dynamic range of a temporal motion vector is not transmitted via the sequence parameter set, picture parameter set, or slice header, but can be defined via the video codec's level. Encoders and decoders can use the level information to determine the limited dynamic range of the motion vector.
[0223] Furthermore, at the level level, the dynamic range and / or bit depth of the X and Y components of the motion vector can be defined to be different from each other, and the minimum and maximum values of each component can also be defined.
[0224] Tables 11 and 12 show an example of defining TMVBitWidth as a level in the temporal motion vector derivation process described above.
[0225] [Table 11]
[0226] Referring to Table 11, TMVBitWidth is set to MaxTMVBitWidth as defined by the level. At that time, MaxTMVBitWidth indicates the maximum bit width of the motion vector when the temporal motion vector is stored in memory.
[0227] On the other hand, TMVBitWidth can be defined at a level, and the difference (delta value) from the defined value can also be sent as a sequence parameter set, picture parameter set, or slice header. That is, TMVBitWidth can be set to a value obtained by adding the difference sent as a sequence parameter set, picture parameter set, or slice header to MaxTMVBitWidth defined at the level. In that case, TMVBitWidth indicates the bit width of the motion vector when the temporal motion vector is stored in memory.
[0228] [Table 12]
[0229] [Table 13]
[0230] delta_bit_width_temporal_motion_vector_minus8 in Table 13 indicates the difference in the bit width of temporal motion vector components. If delta_bit_width_temporal_motion_vector_minus8 does not exist, it is inferred to be 0, and the bit width of the temporal motion vector component is expressed as shown in Equation 29.
[0231] [Equation]
[0232] Further, as shown in Table 14, the dynamic range of each component of the temporal motion vector can also be defined for each level.
[0233] [Table 14]
[0234] Further, as shown in Tables 15 to 17, the bit width of each component of the temporal motion vector can also be defined for each level.
[0235] [Table 15]
[0236] [Table 16]
[0237] [Table 17]
[0238] Furthermore, as shown in Table 18, the bit width of the Y component of the time motion vector can also be defined for each level.
[0239] [Table 18]
[0240] Furthermore, the dynamic range of the temporal motion vector can be defined to a fixed value predetermined by the encoding and decoding devices without transmitting information regarding the limitations of the motion vector, or it can be stored in the form of a fixed bit depth.
[0241] When TMVBitWidth is fixed to the same value in both the encoding and decoding devices, TMVBitWidth is a positive integer such as 4, 6, 8, 10, 12, 14, or 16. In this case, TMVBitWidth indicates the bit width of the motion vector when the temporal motion vector is stored in memory.
[0242] Figure 14 is a sequence diagram showing a video encoding method according to one embodiment of the present invention. Referring to Figure 14, the video encoding method includes a clipping step (S1410), a storage step (S1420), and an encoding step (S1430).
[0243] A video encoding apparatus and / or decoding apparatus clips a motion vector of a reference picture within a predetermined dynamic range (S1410). As described above, through the "I. Motion vector clipping process", a motion vector falling outside the dynamic range is represented by the minimum value or the maximum value of the corresponding dynamic range. Therefore, as described above, through "IV. Information transmission method for clipping temporal motion vector in decoding apparatus" and "V. Definition of dynamic range through video codec level", by limiting the bit depth through a video codec level and / or a sequence parameter set or the like, or limiting the dynamic range through the video codec level, the motion vector of the reference picture can be clipped within the predetermined dynamic range.
[0244] The video encoding apparatus and / or decoding apparatus stores the clipped motion vector of the reference picture in a buffer through "II. Motion vector storing process" as described above (S1420). The motion vector can be stored in the buffer together with the restored video or separately therefrom.
[0245] The video encoding apparatus encodes the motion vector of a current block to be encoded using the stored motion vector of the reference picture (S1430). As described above, through "III. Motion vector deriving process", the improved motion vector prediction method used in HEVC uses not only motion vectors of restored blocks located around the block to be encoded / decoded, but also a motion vector of a block existing at the same position as or a position corresponding to the block to be encoded / decoded in the reference picture. Therefore, the motion vector of the block to be encoded may be not only motion vectors of peripheral blocks adjacent to the block to be encoded, but also the motion vector of the reference picture, that is, a temporal motion vector.
[0246] On the other hand, since the dynamic ranges of the X component and Y component of the motion vector of the reference picture can be defined differently from each other, each component of the motion vector of the reference picture can be clipped within its respective dynamic range.
[0247] In addition to limiting the dynamic range of the reference picture's motion vector, methods for compressing the reference picture's motion vector can also be used. When limiting the dynamic range of the reference picture's motion vector or compressing it, a flag indicating this and corresponding parameters can be defined in the video codec's level and / or sequence parameter set, etc.
[0248] Furthermore, by using motion information stored in memory, i.e., motion information of the reference picture, it is possible to perform encoding methods such as motion vector prediction, improved motion vector prediction, motion information merging, and motion information merging skip.
[0249] Figure 15 is a sequence diagram showing a video decoding method according to one embodiment of the present invention. Referring to Figure 15, the video decoding method includes a clipping step (S1510), a storage step (S1520), a derivation step (S1530), and a decoding step (S1540).
[0250] The clipping step (S1510) and storage step (S1520) in Figure 15 are the same as the clipping step (S1410) and storage step (S1420) in Figure 14, which utilize the previously described "I. Motion Vector Clipping Process" and "II. Motion Vector Storage Process". Furthermore, the derivation step (S1530) in Figure 15 utilizes the previously described "III. Motion Vector Derivation Process" and is symmetrical to the encoding step (S1430) in Figure 14. Therefore, a detailed explanation is omitted.
[0251] The video decoding device performs interpredictive decoding using the motion vector of the block to be decoded (S1540). The video decoding device stores the motion vector in memory using at least one of the following methods: a method for limiting the dynamic range of the motion vector, a method for reducing the spatial resolution of the motion vector, a method for quantizing the motion vector, or a method for reducing the representation resolution of the motion vector. The stored motion vector can then be used for motion vector prediction and motion information merging of the block to be decoded.
[0252] Furthermore, it is possible to use motion information stored in memory, i.e., motion information of the reference picture, to perform decoding methods such as motion vector prediction, improved motion vector prediction, motion information merging, and motion information merging skip.
[0253] Although the embodiments described above are illustrated with sequence diagrams representing a series of steps or blocks, the present invention is not limited to the order of steps described above, and some steps may occur with different steps, in different orders, or simultaneously. Furthermore, a person with ordinary skill in the art to which the present invention belongs will understand that the steps shown in the sequence diagrams are not exclusive, and other steps may be included, or some steps may be omitted.
[0254] Furthermore, the embodiments described above include examples of various embodiments. While it is not possible to describe all possible combinations in order to illustrate the diverse embodiments, a person with ordinary skill in the art to which the present invention pertains will recognize that other combinations are possible. Therefore, the present invention includes all substitutions, modifications, and changes that fall within the scope of the claims.
Claims
1. A reference picture buffer for storing a reference picture and the motion vector of the reference picture, wherein the motion vector of the reference picture is stored in the reference picture buffer, and each component of the stored motion vector is within a predetermined dynamic range. An interpretation unit for generating a predicted block for the current block in the current picture using the motion vector of the current block, Includes, The motion vector of the current block is predicted using a candidate for a temporal motion vector, or the motion vectors of surrounding blocks spatially adjacent to the current block. The candidate for the temporal motion vector is derived by scaling the stored motion vector of the reference picture according to the temporal distance between the pictures, and clipping the scaled motion vector to the predetermined dynamic range. The aforementioned temporal motion vector candidate corresponds to the central position or the lower right position of the current block, in the video decoding device.
2. The video decoding apparatus according to claim 1, wherein the candidate for the temporal motion vector is obtained from the lower right position of the current block, and if no candidate is available at the lower right position, it is obtained from the central position.
3. The video decoding apparatus according to claim 1, wherein the temporal distance is the difference in output order between pictures, and the scaling uses the ratio of the temporal distance to a second temporal distance associated with the reference picture.
4. A reference picture buffer for storing a reference picture and the motion vector of the reference picture, wherein the motion vector of the reference picture is stored in the reference picture buffer, and each component of the stored motion vector is within a predetermined dynamic range. An interpretation unit for generating a predicted block for the current block in the current picture using the motion vector of the current block, Includes, The motion vector of the current block is encoded using a candidate for a temporal motion vector, or the motion vector of a surrounding block spatially adjacent to the current block. The candidate for the temporal motion vector is derived by scaling the stored motion vector of the reference picture according to the temporal distance between the pictures, and clipping the scaled motion vector to the predetermined dynamic range. The aforementioned temporal motion vector candidate corresponds to the central position or the lower right position of the current block, in the video encoding device.
5. A method for transmitting a video bitstream, The steps of generating the aforementioned bitstream, The steps include transmitting the bitstream, Equipped with, The aforementioned bitstream is A step of storing a reference picture and the motion vector of the reference picture in a reference picture buffer, wherein each component of the stored motion vector is within a predetermined dynamic range. The steps of deriving candidate time motion vectors corresponding to the center or lower-right position of the current block in the current picture by scaling the stored motion vector of the reference picture according to the time distance between the pictures, and clipping the scaled motion vector to the predetermined dynamic range, The steps include determining the motion vector of the current block using the aforementioned temporal motion vector candidate or the motion vectors of surrounding blocks spatially adjacent to the current block, and A step of generating a predicted block for the current block using the motion vector of the current block, A method generated by