Image and video coding and decoding

By deriving values from frames with shared time IDs or quantization parameters, the method enhances video coding efficiency, addressing the limitations of existing standards for high dynamic range and ultra-high definition videos.

JP2026513441APending Publication Date: 2026-04-27CANON KK
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
CANON KK
Filing Date
2024-04-08
Publication Date
2026-04-27

AI Technical Summary

Technical Problem

Existing video coding standards like HEVC face challenges in achieving significant improvements in compression efficiency, particularly for high dynamic range and ultra-high definition videos, necessitating further enhancements to achieve better coding efficiency.

Method used

The method involves deriving values from a first frame in a bitstream to determine variables or context increments related to a second frame, where the frames have the same time ID or quantization parameter, and setting the size of the first area based on time or quantization differences, to enhance encoding and decoding processes.

Benefits of technology

This approach improves compression efficiency by leveraging frame relationships, allowing for more effective encoding and decoding of video data, particularly in high dynamic range and ultra-high definition scenarios.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2026513441000001_ABST
    Figure 2026513441000001_ABST
Patent Text Reader

Abstract

A method for encoding / decoding video data into / from a bitstream, wherein the bitstream comprises video data corresponding to a plurality of frames arranged in decoding order, the method comprising: deriving a value from a first area in a first frame among the plurality of frames; and determining from the value a variable, syntax element, or context increment of a syntax element relating to a second area in a second frame, wherein the first frame precedes the second frame in the decoding order. Apparatus for performing the method is also disclosed.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to encoding and decoding video data from a bitstream, and more particularly to encoding and decoding image and video segmentation data. An apparatus for decoding video data from a bitstream and encoding video data into a bitstream, as well as a computer program configured to perform encoding or decoding of video data at runtime, are also provided.

Background Art

[0002] A collaborative team formed by MPEG and VCEG of ITU-T Study Group 16, the Joint Video Experts Team (JVET), announced a new video coding standard called VVC (Versatile Video Coding). The goal of VVC is to provide a significant improvement in compression performance that exceeds the existing HEVC standard (i.e., typically twice that of the previous one). The main target applications and services include, but are not limited to, 360-degree and high dynamic range (HDR) videos. It has shown particular effectiveness for ultra-high definition (UHD) video test materials. Therefore, an improvement in compression efficiency far exceeding the target 50% of the final standard can be expected.

[0003] Since the completion of the standardization of VVC v1, JVET has started the exploration phase by establishing exploration software (ECM). This collects additional tools and improvements to existing tools in addition to the VVC standard in order to target better coding efficiency.

Summary of the Invention

[0004] According to a first aspect of the present invention, a method for decoding video data from a bitstream is provided, the bitstream including video data corresponding to a plurality of frames arranged in decoding order, the method comprising The process involves deriving a value from the first area in the first frame among the aforementioned multiple frames, From the aforementioned values, determine the variables, syntax elements, or context increments of syntax elements related to the second area in the second frame. It has, The first frame precedes the second frame in the decoding order.

[0005] A second aspect of the present invention provides a method for encoding video data into a bitstream, wherein the bitstream includes video data corresponding to a plurality of frames arranged in decoding order, and the method is To derive a value for the first area in the first frame among the aforementioned multiple frames, From the aforementioned values, determine the variables, syntax elements, or context increments of syntax elements related to the second area in the second frame. It has, The first frame precedes the second frame in the decoding order.

[0006] Optionally, each of the plurality of frames has an associated time ID, and the first frame and the second frame have the same time ID. Optionally, the first frame corresponds to the frame closest to the second frame having the same time ID in the decoding order.

[0007] Optionally, each of the plurality of frames has an associated quantization parameter QP, and the first frame and the second frame have the same QP.

[0008] Optionally, the aforementioned first frame is a reference frame.

[0009] Optionally, the first area is larger than or equal to the size of the second area.

[0010] Optionally, the second area corresponds to a coding tree unit (CTU).

[0011] Optionally, the size of the first area is set based on the time distance between the first frame and the second frame. Optionally, the time difference is calculated based on the difference between the picture order count (POC) of the first frame and the POC of the second frame.

[0012] Optionally, each of the plurality of frames has an associated quantization parameter QP, and the size of the first area is set based on the difference between the QP of the first frame and the QP of the second frame.

[0013] Optionally, each of the plurality of frames has an associated time ID, and the size of the first area is set based on the difference between the time ID of the first frame and the time ID of the second frame.

[0014] Optionally, the size of the first area is set based on a value transmitted in one of the sequence parameter set, picture parameter set, picture header, and slice header included in the bitstream.

[0015] Optionally, the size of the first area is set based on the size of the second area.

[0016] Optionally, the step of deriving a value from a first area in the first frame of the plurality of frames includes deriving a value from a block having at least one sample within the first area.

[0017] Optionally, the step of deriving a value from a first area in the first frame of the plurality of frames includes deriving a value from all blocks having at least one sample within the first area.

[0018] Optionally, the step of deriving a value from a first area in the first frame of the plurality of frames includes weighting the block or the value derived from each block based on the number of samples of each block within the first area.

[0019] Optionally, the step of deriving a value from a first area in the first frame of the plurality of frames includes deriving a value only from blocks that are entirely contained within the first area.

[0020] Optionally, the step of deriving a value from a first area in the first frame of the plurality of frames includes deriving values ​​only from blocks located on an NxN grid within the first area, where N is an integer. Optionally, N=16.

[0021] Optionally, the center of the grid is located in the same position as the center of the first area of ​​the first frame within the first frame.

[0022] Optionally, the step of deriving a value from a first area in the first frame of the plurality of frames includes deriving values ​​only from blocks located on points of the pattern within the first area. Optionally, the points of the pattern are unevenly spaced.

[0023] Optionally, the step of determining from the aforementioned value a variable, syntax element, or contextual increment of a syntax element relating to a second area in the second frame includes determining a variable, syntax element, or contextual increment of a syntax element relating to a block in the second area.

[0024] Optionally, the step of determining, from the value, a variable, a syntax element, or a context increment of a syntax element, related to a second area within a second frame, includes determining a variable, a syntax element, or a context increment of a syntax element, related to a block located on an MxM grid within the second area, where M is an integer. Optionally, the MxM grid is shifted horizontally by M / 2 and vertically by M / 2 with respect to the upper left position of the second area.

[0025] Optionally, the center of the first area is located at the same position within the first frame as the center of the second area within the second frame.

[0026] Optionally, the center of the first area is located at the position within the first frame corresponding to the center of the second area within the second frame, shifted by an amount corresponding to a motion vector derived from an area adjacent to the second area.

[0027] Optionally, the method further includes deriving a value from the upper left position of the first area when the center of the first area is outside the first frame.

[0028] Optionally, the step of deriving a value from a first area in a first frame of the plurality of frames includes accessing each block within the first area only once.

[0029] Optionally, the step of deriving a value from a first area in a first frame of the plurality of frames includes deriving values from a first area in a first frame of the plurality of frames and a third area in a third frame of the plurality of frames.

[0030] Optionally, the step of determining from the aforementioned value a variable, syntax element, or context increment of a syntax element relating to a second area in a second frame includes determining a predictor to be added to the residual derived from the bitstream.

[0031] Optionally, the value derived from the first area includes maxMttDepth, and the step of determining from the value a variable, syntax element, or contextual increment of a syntax element relating to the second area in the second frame includes determining the maxMttDepth of the block in the second area.

[0032] Optionally, the value derived from the first area is used to limit the context increment of a syntax element or variable related to the second area.

[0033] Optionally, the value derived from the first area is the minimum value of the quadtree depth value minQTDepth from the block in the first area, and the step of determining from the value a variable, syntax element, or context increment of a syntax element relating to the second area in the second frame includes comparing the derived minQTDepth with the quadtree depth of the block in the second area to determine whether only the quadtree split is permitted.

[0034] Optionally, the value derived from the first area is the maximum value of the multitree depth value maxMttDepth from the block in the first area, and the step of determining from the value a variable, syntax element, or contextual increment of a syntax element relating to the second area in the second frame includes comparing the derived maxMttDepth with the maxMttDepth of the block in the second area, and modifying maxMttDepth based on the comparison.

[0035] Optionally, the value derived from the first area is the average of the quadtree depth values ​​from the blocks in the first area, and the step of determining from the value a variable, syntax element, or context increment of a syntax element relating to the second area in the second frame includes comparing the derived quadtree depth value with the quadtree depth of the blocks in the second area. Optionally, the method further comprises the step of determining an acceptable partition based on the comparison. Optionally, the method further comprises the step of varying maxMttDepth based on the comparison.

[0036] According to a third aspect of the present invention, an apparatus for decoding video data from a bitstream is provided, the apparatus being configured to perform the method of the first aspect.

[0037] According to a fourth aspect of the present invention, an apparatus for encoding video data into a bitstream is provided, the apparatus being configured to perform the method of the second aspect.

[0038] According to a fifth aspect of the present invention, a computer program is provided which is configured to perform the method of the first or second aspect at runtime. [Brief explanation of the drawing]

[0039] For example, please refer to the attached drawing: [Figure 1] Figure 1 is a diagram illustrating the coding structure used in HEVC. [Figure 2] Figure 2 is a schematic block diagram showing a data communication system in which one or more embodiments of the present invention may be implemented. [Figure 3] Figure 3 is a block diagram showing the components of a processing apparatus in which one or more embodiments of the present invention may be implemented. [Figure 4] Figure 4 is a schematic diagram showing the functional elements of an encoder according to an embodiment of the present invention. [Figure 5] Figure 5 is a schematic diagram showing the functional elements of a decoder according to an embodiment of the present invention. [Figure 6] Figure 6 shows the blocks placed relative to the current block, which includes the placed blocks. [Figure 7] Figures 7(a) and (b) show the affine (subblock) mode. [Figure 8] Figure 8 shows candidate subblock time merges. [Figure 9] Figure 9 shows the time-random-access GOP structure for 33 frames with associated time IDs and POCs. [Figure 10] Figure 10 shows the six possible splitting modes of VVC. [Figure 11] Figure 11 shows MaxBTSize and MaxMttDepth. [Figure 12] Figure 12 shows an example of the MinQTSize variable. [Figure 13] Figure 13 shows the possible partitioning constraints. [Figure 14] Figure 14 shows an incomplete CTU at the frame boundary. [Figure 15] Figure 15 shows the encoding achieved by setting MaxMttDepth based on the time ID. [Figure 16] Figure 16 shows an embodiment of the present invention. [Figure 17] Figure 17 shows an example of one embodiment of the present invention. [Figure 18] Figure 18 shows an embodiment of the present invention. [Figure 19] Figure 19 shows an embodiment of the present invention. [Figure 20] Figure 20 shows a system according to an embodiment of the present invention, comprising an encoder or decoder and a communication network. [Figure 21] Figure 21 is a schematic block diagram of a computing device for implementing one or more embodiments of the present invention. [Figure 22]Figure 22 shows a network camera system. [Figure 23] Figure 23 shows a smartphone. [Modes for carrying out the invention]

[0040] Figure 1 illustrates the coding structure used in HEVC (High Efficiency Video Coding) and VVC (Versatile Video Coding) standards. Video sequence 1 consists of a series of digital images i. Each such digital image is represented by one or more matrices. Matrix coefficients represent pixels.

[0041] Image 2 of the sequence can be divided into slice 3. In some examples, a slice may constitute an entire image. These slices are divided into non-overlapping coding tree units (CTUs). A coding tree unit (CTU) is a fundamental processing unit in the HEVC (High Efficiency Video Coding) and VVC (Versatile Video Coding) video standards and conceptually and structurally corresponds to the macroblock units used in some earlier video standards. A CTU is sometimes called an LCU (Largest Coding Unit). A CTU has a luminous component portion and a chroma component portion, and each component portion is called a CTB (Coding Tree Block). These different color components are not shown in Figure 1.

[0042] A CTU is generally 64x64 pixels in the case of HEVC, but in the case of VVC, this size can be 128x128 pixels. Each CTU can be iteratively divided into smaller variable-size coding units (CUs) 5 using quadtree (QT) decomposition.

[0043] A coding unit is a fundamental coding element and consists of two types of subunits called PUs (Prediction Units) and TUs (Transform Units). The maximum size of a PU or TU is equal to the size of the CU. Prediction units correspond to partitions of the CU for predicting pixel values. As shown in 6, various different partitions of the CU into PUs can include partitions into four square PUs and two different partitions into two rectangular PUs. Transform units are fundamental units that undergo spatial transformations using discrete cosine transforms (DCTs). A CU can be partitioned into TUs based on its quadtree representation 7.

[0044] Each slice is embedded in a single Network Abstraction Layer (NAL) unit. Furthermore, the coding parameters of a video sequence are stored in a dedicated NAL unit called a parameter set. HEVC and H.264 / AVC use two types of parameter set NAL units: firstly, the SPS (Sequence Parameter Set) NAL unit, which collects all parameters that do not change throughout the entire video sequence. Typically, it handles the coding profile, video frame size, and other parameters. Secondly, the PPS (Picture Parameter Set) NAL unit contains parameters that may change from one image (or frame) in the sequence to another. HEVC also includes a VPS (Video Parameter Set) NAL unit, which contains parameters describing the overall structure of the bitstream. VPS is a type of parameter set defined in HEVC and applies to all layers of the bitstream. A layer can contain multiple time sublayers, while all version 1 bitstreams are limited to a single layer. HEVC has specific layered extensions for scalability and multi-view, which enable multiple layers with backward-compatible version 1 base layers.

[0045] Another method of dividing images is introduced in VVCs, which include subpictures, which are independently coded groups of one or more slices.

[0046] Figure 2 shows a data communication system in which one or more embodiments of the present invention may be implemented. The data communication system has a transmitting device, in this case a server 201, which is capable of transmitting data packets of a data stream to a receiving device, in this case a client terminal 202, via a data communication network 200. The data communication network 200 may be a WAN (Wide Area network) or a LAN (Local Area Network). Such a network may be, for example, a wireless network (Wifi / 802.11a or b or g), an Ethernet® network, an Internet network, or a mixed network consisting of several different networks. In a particular embodiment of the present invention, the data communication system may be a digital television broadcasting system in which the server 201 transmits the same data content to multiple clients.

[0047] The data stream 204 provided by the server 201 may consist of multimedia data representing video and audio data. In some embodiments of the present invention, the audio and video data streams may be captured by the server 201 using a microphone and a camera, respectively. In some embodiments, the data streams may be stored in the server 201, received by the server 201 from another data provider, or generated in the server 201. The server 201 particularly includes encoders for encoding the video and audio streams to provide compressed bitstreams for transmission in a more compact representation of the data presented as input to the encoder.

[0048] To obtain a better ratio between the quality of the transmitted data and the amount of data transmitted, video data compression may follow, for example, the HEVC format, H.264 / AVC format, VVC format, or the format of the data generated by ECM.

[0049] Client 202 receives the transmitted bitstream, decodes the reconstructed bitstream, and plays the video image on the display device and the audio data through the loudspeaker.

[0050] While the example in Figure 2 considers a streaming scenario, it will be understood that in some embodiments of the present invention, data communication between the encoder and decoder may be performed using a storage medium such as an optical disc.

[0051] In one or more embodiments of the present invention, a video image is transmitted along with data representing a compensation offset, which is applied to reconstructed pixels of the image to provide filtered pixels in the final image.

[0052] Figure 3 schematically shows a processing unit 300 configured to carry out at least one embodiment of the present invention. The processing unit 300 may be a device such as a microcomputer, a workstation, or a light portable device. The device 300 has a communication bus 313 to which the following is connected: - A central processing unit such as a microprocessor labeled CPU 311; - A read-only memory 306, indicated by ROM, for storing a computer program for carrying out the present invention; - Random access memory 312, designated RAM, which stores registers adapted to record variables and parameters necessary to carry out the executable code for the method of an embodiment of the present invention, and the method for encoding a sequence of digital images and / or decoding a bitstream according to an embodiment of the present invention; - A communication interface 302 connected to a communication network 303 through which the digital data to be processed is transmitted and received. Optionally, the device 300 may also include the following components: - A computer program for carrying out a method of one or more embodiments of the present invention, and data storage means such as a hard disk for storing data used or generated during the implementation of one or more embodiments of the present invention; - A disk drive 305 for disk 306, the disk drive being adapted to read data from disk 306 or write data to the disk; - A screen 309 that serves as a graphical interface with the user and / or displays data, via a keyboard 310 or any other pointing means.

[0053] The device 300 can be connected to various peripheral devices, such as a digital camera 320 or a microphone 308, each of which is connected to an input / output card (not shown) to supply multimedia data to the device 300.

[0054] The communication bus provides communication and interoperability between various elements included in or connected to the device 300. The representation of the bus is not limited, and in particular, the central processing unit can operate to communicate instructions directly to any element of the device 300 or through another element of the device 300.

[0055] The disk 306 can be replaced with any information medium, such as a compact disc (CD-ROM), rewritable or non-rewritable, ZIP disc, or memory card, and can generally be replaced with information storage means that can be read by a microcomputer or microprocessor, and which may be incorporated into or not incorporated into the device, or which may be removable and adapted to store one or more programs, and which, by execution, enables the implementation of the method for decoding a bitstream and / or encoding a series of digital images according to the present invention.

[0056] The executable code may be stored in read-only memory 306, hard disk 304, or a removable digital medium such as disk 306 as described above. In a modified example, the executable code of a program may be received by the communication network 303 via interface 302 to be stored in one of the storage means of device 300, such as hard disk 304, before being executed.

[0057] The central processing unit 311 is adapted to control and direct the execution of a set of programs or a portion of the program instructions or software code according to the present invention, based on instructions stored in one of the above-described storage means. When power is turned on, the set of programs or programs stored in non-volatile memory, such as the hard disk 304 or read-only memory 306, is transferred to random access memory 312, which includes the executable code of the program or set of programs, as well as registers for storing variables and parameters necessary to carry out the present invention.

[0058] In this embodiment, the device is a programmable device that uses software to carry out the present invention. However, the present invention may also be implemented in hardware (for example, in the form of an application-specific integrated circuit or ASIC).

[0059] Figure 4 shows a block diagram of an encoder according to at least one embodiment of the present invention. The encoder is represented by connected modules, each module adapted to perform at least one corresponding step of a method for performing at least one embodiment of encoding an image of a sequence of images according to one or more embodiments of the present invention, for example, in the form of program instructions executed by the CPU 311 of the device 300.

[0060] The original sequence of digital images i0 to in401 is received as input by encoder 400. Each digital image is represented by a set of samples, sometimes also called pixels (hereinafter referred to as pixels).

[0061] The bitstream 410 is output by the encoder 400 after the encoding process is performed. The bitstream 410 has multiple encoding units or slices, each slice having a slice header for transmitting encoded values ​​of encoding parameters used to encode the slice, a slice body, and encoded video data.

[0062] Input digital images i0~i n 401 is divided into blocks of pixels by module 402. The blocks correspond to parts of the image and can be of variable size (for example, 4x4, 8x8, 16x16, 32x32, 64x64, 128x128 pixels, and several rectangular block sizes may also be considered). A coding mode is selected for each input block. Two families of coding modes are provided: coding modes based on spatial prediction coding (intra prediction) and coding modes based on temporal prediction (intercoding, merge, SKIP). Possible coding modes are tested.

[0063] Module 403 performs an intra-prediction process in which a given block to be coded is predicted by predictors calculated from neighboring pixels of the block to be coded. If intra-coding is selected, the selected intra-predictors and the indication of the difference between a given block and its predictors are coded to provide a residual.

[0064] Time prediction is performed by motion estimation module 404 and motion compensation module 405. First, a reference image is selected from the set of reference images 416, and the portion of the reference image, also called the reference area or image portion, that is closest to the given block to be encoded (closest in terms of pixel value similarity), is selected by motion estimation module 404. Next, motion compensation module 405 uses the selected area to predict the block to be encoded. The difference between the selected reference area and the given block, also called the residual block, is calculated by motion compensation module 405. The selected reference area is represented using a motion vector.

[0065] Therefore, in both cases (spatial and temporal prediction), the residual is calculated by subtracting the predictor from the original block.

[0066] In the INTRA prediction performed by module 403, the prediction direction is encoded. In the inter prediction performed by modules 404, 405, 416, 418, and 417, at least one motion vector or data is encoded for time prediction to identify such motion vectors.

[0067] If interpretation is selected, information related to the motion vector and residual block is encoded. To further reduce the bitrate, assuming uniform motion, the motion vector is encoded by the difference with respect to the motion vector predictor. The motion vector predictor from the set of motion information predictor candidates is obtained from the motion vector field 418 by the motion vector prediction coding module 417.

[0068] The encoder 400 further includes a selection module 406 for selecting a coding mode by applying coding cost criteria such as rate distortion criteria. To further reduce redundancy, a transformation (such as DCT) is applied to the residual block by a transformation module 407, and the acquired transformed data is then quantized by a quantization module 408 and entropy coded by an entropy coding module 409. Finally, the coded residual block of the currently coded block is inserted into the bitstream 410.

[0069] The encoder 400 also decodes the encoded image to generate a reference image (e.g., a reference image in reference image / picture 416) for motion estimation of subsequent images. This allows the encoder and the decoder receiving the bitstream to have the same reference frame (a reconstructed image or a portion of the image is used). The inverse quantization ("inverse quantization") module 411 performs inverse quantization ("inverse quantization") of the quantized data, followed by the inverse transform module 412. The intra-prediction module 413 uses the prediction information to determine which predictor should be used for a given block, and the motion compensation module 414 actually adds the residuals obtained by module 412 to the reference area obtained from the set of reference images 416.

[0070] Next, post-filtering is applied by module 415 to filter the reconstructed frame (image or portion of an image) of pixels. In embodiments of the present invention, an SAO loop filter is used, and a compensation offset is added to the pixel values ​​of the reconstructed pixels in the reconstructed image. It is understood that post-filtering is not necessarily required. In addition to or instead of SAO loop filtering, any other type of post-filtering may be performed.

[0071] Figure 5 shows a block diagram of a decoder 60 that may be used to receive data from an encoder, according to one embodiment of the present invention. The decoder is represented by connected modules, each module adapted to perform the corresponding steps of the method performed by the decoder 60, for example, in the formation of program instructions executed by the CPU 311 of the device 300.

[0072] Decoder 60 receives a bitstream 61 having encoded units (e.g., data corresponding to blocks or coding units), each consisting of a header containing information about encoding parameters and a body containing encoded video data. As described with reference to Figure 4, the encoded video data is entropy encoded, and the indices of the motion vector predictors are encoded for a given block with a predetermined number of bits. The received encoded video data is entropy decoded by module 62. The residual data is then inversely quantized by module 63, and then the inverse transform is applied by module 64 to obtain the pixel values.

[0073] Mode data indicating the coding mode is also entropy-decoded, and based on that mode, INTRA-type decoding or INTER-type decoding is performed on the coded blocks (units / sets / groups) of the image data.

[0074] In INTRA mode, the INTRA predictor is determined by the INTRA prediction module 65 based on the INTRA prediction mode specified in the bitstream.

[0075] When the mode is INTER, motion prediction information is extracted from the bitstream to find (identify) the reference area used by the encoder. The motion prediction information includes the reference frame index and the motion vector residual. The motion vector predictor is added to the motion vector residual by the motion vector decoding module 70 to obtain the motion vector. The various motion prediction tools used in VVC are described in more detail below with reference to Figures 6 to 10.

[0076] The motion vector decoding module 70 applies motion vector decoding to each current block encoded by motion prediction. Once the index of the motion vector predictor for the current block is obtained, the actual value of the motion vector associated with the current block is decoded and can be used by module 66 to apply motion compensation. The reference image portion indicated by the decoded motion vector is extracted from the reference image 68 and motion compensation 66 is applied. The motion vector field data 71 is updated with the decoded motion vector for use in predicting subsequent decoded motion vectors. Note that in VVCs like HEVC, motion vectors are stored at a 16x16 level rather than 4x4 for time predictors. Decimation means that it is applied only to time predictors and not to spatial predictors. In practice, the goal is to reduce the buffer required to store time motion vectors after coding each frame. This negatively impacts coding efficiency, and generally, this decimation is removed from search software.

[0077] Finally, the decoded block is obtained. If appropriate, post-filtering is applied by the post-filtering module 67. The decoded video signal 69 is finally obtained and provided by the decoder 60.

[0078] VVC Merge Mode VVC includes several additional intermodes compared to HEVC. In particular, a new merge mode is added to the standard merge mode in HEVC.

[0079] Affine mode (subblock mode) In HEVC, only translational motion models are applied for motion compensation prediction (MCP). In the real world, there are many types of motion, such as zooming in / out, rotation, perspective motion, and other irregular motions.

[0080] In JEM, a simplified affine transform motion compensation prediction is applied, and the general principle of affine modes is described below based on an extract from document JVET-G1001, presented at the JVET conference held in Turin from July 13-21, 2017. This entire document is incorporated herein by reference to the extent that it describes other algorithms used in JEM.

[0081] As shown in Figure 7(a), the affine motion field of the block is described by two control point motion vectors.

[0082] Affine mode is a motion compensation mode similar to intermode (AMVP, "classic" merge, or "classic" merge skip). Its principle is to generate one motion information per pixel according to two or three adjacent motion information. In JEM, affine mode derives one motion information for every 4x4 block, as shown in Figure 7(a) (each square is a 4x4 block, and the entire block in Figure 7(a) is a 16x16 block, which is divided into 16 such blocks of 4x4 size squares, and each 4x4 square block has a motion vector associated with it). Affine mode is available in AMVP mode and merge mode (i.e., conventional merge mode, also called "non-affine merge mode," and conventional merge skip mode, also called "non-affine merge skip mode") by enabling flagged affine mode.

[0083] In the VVC specification, affine mode is also known as subblock mode, and these terms are used interchangeably in this specification.

[0084] The VVC subblock merge mode includes a subblock-based time merge candidate, which inherits the motion vector field of the block in the previous frame pointed to by the spatial motion vector candidate A1, as shown in Figure 8. In this figure, the current predictor is not in the block where it was placed, but in the block shifted by the motion vector value of A1.

[0085] If adjacent blocks are coded in inter-affine mode for subblock merging, then inherited affine motion candidates follow this subblock candidate, and then some as constructed affine candidates are derived before some zero Mv candidates.

[0086] Context index increment Context-based adaptive binary arithmetic coding (CABAC) uses context to isolate the probabilities of one or more bins. A context index increment ctxInc is computed to obtain the corresponding context and its relative state for each bin.

[0087] For example, the following expression gives an example of context index increment: ctxInc=(condL && availableL)||(condA && availableA) Here, condL is the value of the associated left syntax element, condA is the value of the associated upper syntax element, and availableL and availableA are the availability values ​​of the left and upper blocks, respectively.

[0088] Random access configuration Figure 9 shows the time random access GOP structure for 33 consecutive frames 0-32. The length of the vertical line representing each frame corresponds to its time ID (for example, the longest length corresponds to time ID 0, and the shortest length corresponds to time ID 5). A frame with time ID 0 is the highest in the time hierarchy because it can be decoded independently of all other frames lower in the time hierarchy, i.e., frames with numerically higher time ID values. Similarly, a frame with time ID 1 is second in the time hierarchy, and they can be decoded independently of all other frames lower in the time hierarchy, i.e., frames with higher time IDs for other time IDs, etc. In other words, a frame with a particular time ID can be decoded independently of frames with higher time ID values, but may depend on frames with lower time IDs. This is known as time scalability.

[0089] This parameter is similar to the hierarchical depth, but the hierarchical depth does not imply decoding independence for all other frames with higher depths.

[0090] VVC split VVC partitioning has a specific block partitioning. For a single tree node, six possible partitions are possible, as shown in Figure 10: - Quad division QT,801 divides a block into four equal-sized square blocks. -Binary partition BT with its two possible subdivisions 802, 803: - Vertical binary split, 802, SPLIT_BT_VER - Horizontal binary split, 803, SPLIT_BT_HOR - The block is divided into three blocks with a larger bandwidth in the center, and the two possible subdivisions TT have 804 and 805: - Vertical ternary split, 804, SPLIT_TT_VER - Horizontal ternary split, 805, SPLIT_TT_HOR - Terminate tree nodes without splitting, No Split, 806.

[0091] In this explanation, a block can be a CTU and / or CU in the coding tree, or more generally, any unit.

[0092] VVC partitioning control variables For the current block, not all possible partitions are always allowed. Which partitions are available depends on several conditions. These conditions depend on several defined partition control variables. The first set of variables defines the maximum and minimum block / node sizes: • CTU size: Corresponds to the root node size of a quadtree (e.g., 256x256, 128x128, 64x64, 32x32, 16x16 sample); • maxBtSize: This is the maximum allowable size of a bilingual root node, i.e., the maximum size of a leaf quadtree node that can be partitioned by a binary partition. If both the height and width of the current block are less than or equal to maxBtSize, the current block can be partitioned by a BT partition. Figure 11 illustrates the concept of maxBtSize when maxBtSize is the size of quadtree leaf node 902 of CTU901.

[0093] • minBtSize: This is the minimum allowable size of a bilingual leaf node, i.e., the minimum width or height of a binary leaf node. Therefore, the current block can be partitioned by a horizontal BT partition if its height is greater than minBtSize. Also, the current block can be partitioned by a vertical BT partition if its width is greater than minBtSize.

[0094] • maxTtSize: The maximum allowable trinity root node size, i.e., the maximum size of a leaf quadtree node that can be partitioned by a ternary partition. If both the height and width of the current block are less than or equal to maxTtSize, the current block can be partitioned by a TT partition. • minTtSize: Represents the minimum allowable trinity (TT) leaf node size, i.e., the minimum width or height of a binary leaf node. However, in contrast to BT partitioning, a minimum TT partition size is considered. Therefore, the current block can be partitioned by a horizontal TT partition if its height is greater than twice minTtSize. Also, the current block can be partitioned by a vertical TT partition if its width is strictly greater than twice minTtSize.

[0095] • minQtSize: This is the minimum allowed quadtree (QT) leaf node size, and therefore, if the current block width is not greater than minQtSize, the QT partitioning mode is not allowed for the current block. Figure 12 shows an example of minQtSize. Considering CTU128, in the example shown, minQtsize is equal to 16.

[0096] Since maxQtSize is not defined, it corresponds to the CTU size.

[0097] The minimum allowable block size for width and height is 4.

[0098] A set of depths is also defined.

[0099] • Depth: This is the depth within the tree. In the VVC specification, a leaf is the end node of a tree, which is the root node of a tree with a depth of 0. This means that this value is incremented (by 1) with each split.

[0100] • mttDepth: This is the depth of the multitree. The multitree includes BT partitioning and TT partitioning. • maxMttDepth is the maximum allowable multitree depth, as defined in the VVC specification. Therefore, mttDepth must be greater than or equal to maxMttDepth. Figure 11 illustrates the concept of maxMttDepth.

[0101] In VVC, these variables are defined independently for Luma and Chroma.

[0102] VTM and ECM software have several other variables that correspond to depth.

[0103] The variable `currBtDepth` is the current number of BT partitions used to reach the current tree node (or current block). The variable `currMttDepth` is the current number of BT and TT partitions used to reach the current tree node (or current block). The variable `maxBtDepth` corresponds to the variable `maxMttDepth` in the VVC specification. `currQtDepth` is the current number of QT partitions used to reach the current tree node (or current block). MaxBtDepth: The maximum allowable div tree depth, i.e., the minimum level at which a binary partition can occur, where the quadtree leaf node is the root (e.g., 3).

[0104] VVC partition control syntax elements To set the values ​​of these different variables, several high-level syntax elements are sent in SPS, as shown in the following table of SPS syntax elements.

[0105] [Table 1]

[0106] When sps_partition_constraints_override_enabled_flag is enabled in SPS, several picture header syntax elements are sent to update partition variables, as shown in the following table of PH syntax elements.

[0107] [Table 2]

[0108] VVC coding split mode In VVC, the coding split mode is sent within the coding_tree, with conditionally parsed flags, split_cu_flag, split_qt_flag, mtt_split_cu_vertical_flag, and mtt_split_cu_binary flags defining the CU split as shown in the following syntax table.

[0109] [Table 3]

[0110] VVC partitioning limits VVC partitioning has several limitations. These limitations are primarily to avoid the same partitioning after several consecutive partitions. Figure 13 illustrates some of these constraints. The idea is to avoid the same partitioning as BT and TT. As shown in Figure 13(a), two consecutive vertical BT partitions are permitted, but as shown in Figure 13(b), a vertical TT partition following a vertical BT partition in the center block is not permitted.

[0111] Similarly, as shown in Figure 13(c), two consecutive horizontal BT divisions are permitted, but as shown in Figure 13(d), a horizontal TT following a horizontal BT division in the center block is not permitted.

[0112] In VVC, there are additional constraints on the interblock size, including the maximum TT and BT block sizes and the minimum chroma block size. These constraints are removed in ECM software.

[0113] Chroma splitting In VVC, chroma partitioning can be inferred based on Luma partitioning, but this can be disabled. For example, in Dual Tree mode, the chroma partitioning tree does not depend on the Luma tree. However, some limitations exist.

[0114] The tree can partially rely on Luma partitioning for CCLM mode, or it can be independent otherwise.

[0115] Picture Boundary Frame resolution is not necessarily equal to an integer multiple of the CTU size. As a result, as shown in Figure 14, incomplete CTUs may exist at frame boundaries, with CTUs 1201-1206 being incomplete due to the lower boundary 1207 and right boundary 1208 of the frame. In VVC, in contrast to previous standards, splitting signaling is permitted at picture boundaries. Splitting at boundaries is applied until the coding tree node represents a CU that is fully located within the picture. However, some splits are inferred (not transmitted). Therefore, different variables such as maxMttDepth, minQtDepth, and minQtSize may be incremented or decremented and may differ from the variables used for splits that are not within boundaries.

[0116] QT BT TT coding selection VTM and ECM software employ several encoder-side optimizations for QT BT TT encoding selection.

[0117] One such optimization involves determining whether the QT partition is tested before the BT partition.

[0118] The conditions are that at least one CU to the left or above the current coding tree node has a QT depth greater than the current coding tree node's QT depth, and the CU width represented by the current coding tree node is greater than minQtSize*2.

[0119] If this condition is true, QT is tested before BT, and the split is treated in the following order: - No splitting -QT -BT horizontal -BT Vertical -TT horizontal -TT vertical Otherwise, the order is as follows: - No splitting -BT horizontal -BT Vertical -TT horizontal -TT vertical -QT According to some optimizations, this order is important because some partitions are not tested depending on the results of the first test mode. Therefore, if QT is tested last, there are many opportunities for it to not be evaluated.

[0120] maxMttDepth The maximum MTT depth significantly impacts the complexity of the encoder. The common test conditions for the ECM have been updated to reduce encoding by using different maxMttDepth settings, as shown in Figure 15. This setting results in a lower maxMttDepth for some time IDs in high-resolution or low-QP settings.

[0121] Adaptable maxBtSize VTM and ECM have a frame-level coding select that sets maxBtSize according to the average block size of previous coded frames with the same depth (=> same time ID within a CTC RA case). The average block size is compared to a threshold as the following pseudocode:

[0122] [Table 4]

[0123] If AMAXBT_TH32 is equal to 15, then AMAXBT_TH64 is equal to 30, and AMAXBT_TH128 is equal to 60. This method decreases the maximum BT size if the average block size is small, and increases it if it is large.

[0124] CABAC time forecast: In ECM, time-based CABAC prediction is used (JVET-Y181). In this method, the previous slice is used for CABAC initialization of the current frame. The stochastic state of each context model is first acquired and stored after coding the CTU up to a specified position. The stored stochastic state is then used as the initial stochastic state for the corresponding context model in the next B slice or P slice coded with the same quantization parameter (QP) or the same corresponding time ID.

[0125] Problems that the invention aims to solve In conventional video coding, temporal redundancy is utilized between samples by intermodes and between motion information by different time candidates or predictors. In ECM, temporal redundancy between CABAC probabilities is also used to improve coding efficiency. Furthermore, other data parameters, variables, and syntax elements should have temporal correlations. However, conventional methods of utilizing these redundancies seem to be either unhelpful or impossible. For example, solutions that use a large number of predictors for samples, such as in different intermodes, are not suitable for data with small coding possibilities. Signaling of several candidates for motion information is useful because motion is information from real words, and it is highly specific. However, this is not useful for predicting parameters, variables, or syntax elements. Similarly, predicting CABAC states is not suitable because it resembles more frame-based QP prediction and does not apply to different content within a video sequence.

[0126] This invention proposes a method for efficiently determining values ​​used to predict data using temporal correlations between data.

[0127] Embodiment Main Embodiments The use of time areas to derive at least one value for predicting, inferring, or determining the contextual increment of an Emb.Main variable, syntax element, or syntax element.

[0128] In one embodiment, a time area is used to derive at least one value to predict, infer, limit, or determine a variable, or to code a syntax element, or to calculate a contextual increment of a syntax element. The value represents a similar variable or syntax element, or another variable or another related syntax element. The time area is derived from a time frame already encoded on the encoder side or already decoded on the decoder side. As an example, Figure 16 illustrates this embodiment. In the figure, the time area (1601) includes several blocks and several parts of blocks within the boundaries of the time area. These blocks are considered for deriving at least one value.

[0129] The advantage of this embodiment is improved coding efficiency through better derivation of predictors or constraints or context increments. Compared to methods used in the prior art, this method is adapted by having a reduced number of variables or syntax elements.

[0130] Emb.Tempo_from1 Frames with the same time ID In one embodiment, the time area is brought about by frames having the same time ID.

[0131] In the random access example, if the current frame has a time ID equal to 4, then in the configuration shown in Figure 9, another encoded / decoded frame with the same time ID 4 is used to determine the value associated with the time area.

[0132] Frames with the same time ID often have the same coding parameters, and in particular, they have the same or similar QP and the same spatial distance to their reference frame. Therefore, these are very interesting for predicting QT depth, as this data correlates with QP and spatial distance between frames.

[0133] Emb.Tempo_from1.1: The nearest frame with the same time ID. In one embodiment, the time area is brought about by the nearest frame having the same time ID.

[0134] As shown in Figure 9, an example of a random access configuration, the nearest frame with the same time ID is (generally) more correlated than other frames. Therefore, the results are better.

[0135] Emb.Tempo_from2 A reference frame or frame with the same QP In one embodiment, the time area is provided by a reference frame or frame having the same QP. Ideally, it is a reference frame having the same QP.

[0136] As mentioned above, QP has a significant impact on block partitioning, and since many syntax elements and variables have similar values, frames with the same QP will have better values ​​determined from time.

[0137] The Emb.Tempo_from3 reference frame is the same one used for time motion vector prediction. In one embodiment, the time area is provided by a reference frame used for time motion vector prediction. This may be the first reference in reference list 0 or the first reference frame in list 1, according to a flag sent in the picture header or slice header.

[0138] Surprisingly, this embodiment provides the best coding efficiency even when this reference frame has a lower QP. However, it is closer to the current frame compared to all frames with the same time ID.

[0139] Emb.Tempo_from4: The nearest reference frame In one embodiment, the time area is provided by the nearest reference frame.

[0140] As described in the previous embodiment, the distance to the current frame appears to be more interesting in terms of the compromise between encoder time reduction and coding efficiency, even if frames with the same QP have statistically more correlation between their QP depths.

[0141] Emb.Tempo_from5 Multiple reference frames In one embodiment, two time frames are considered, and therefore two time areas are used to determine one value. Three or more reference frames can also be considered.

[0142] The advantage of this is better coding efficiency, but it increases the amount of memory access.

[0143] Sending Emb.Decim HLS headers of size N or M In one embodiment, the grid value N for the decimation of a block in the time area and / or the grid value M for the decimation of possible positions in the block are transmitted within the header. In addition, or alternatively, other parameters may also be transmitted as an irregular grid. Ideally, the values ​​are transmitted in a sequence parameter set SPS, as they have an impact on the memory buffer. However, if the values ​​between N and M maintain the same required memory size, these values ​​may be transmitted alternatively in a PPS, picture header, or slice header.

[0144] size Emb.Larger size is greater than or equal to the current block size. In one embodiment, the size of the time area is, if possible, larger than the current block. The objective of this embodiment is to determine a more useful value than the value that can be obtained with the placed block.

[0145] One advantage is that the value determined is based on more values ​​than a single placed block, making it a better value for variables for predicting, inferring, determining, or deriving contextual increments. Of course, this is adapted to some data, as opposed to motion information. This advantage leads to improved coding efficiency.

[0146] The second advantage is that, compared to solutions where blocks are shifted according to motion information (as time subblocks), the motion information does not need to be determined. Therefore, parsing does not rely on motion information that cannot be obtained without complete decoding.

[0147] Another advantage is that, compared to being placed from only one block, several values ​​can be considered, and other values ​​can be derived as minimum, maximum, average, etc., which gives more information for limiting, predicting, or deriving the contextual increment.

[0148] Emb.Larger2 The size of the time area is greater than or equal to the maximum possible block size. In one embodiment, the time area is greater than or equal to the maximum possible block size. For example, the time area is greater than or equal to the CTU size.

[0149] This offers an interesting compromise between coding efficiency and the complexity required to determine the value.

[0150] The Emb.AdaptSize size is adapted according to the variable / parameter. In one embodiment, the size of the time area is adapted according to at least one variable or at least one parameter.

[0151] The advantage of this embodiment is increased coding efficiency, especially when the time area is used to compensate for motion between two frames.

[0152] Emb.AdaptSize1 Time distance In one embodiment, the size of the time area is determined based on the time distance between the current frame and the frame containing the time area. In this embodiment, as the time distance increases, the time area increases. For example, the absolute difference between the picture order count (POC) of the current frame "currPOC" and the frame containing the time area "tempoPOC" is calculated taking into account the time distance between the frames. For example, the size of the time area (widthTempo, heightTempo) may be determined according to the following pseudocode: widthTempo=widthTempoFix+8*abs(currPOC-tempoPOC) heightTempo=heightTempoFix+8*abs(currPOC-tempoPOC) Here, abs() is a function given an absolute value, and widthTempoFix and heightTempoFix are predetermined. For example, they are set to be equal to the CTU size. The number "8" in this formula is just an example; other values ​​can be considered.

[0153] Furthermore, the sequence frame rate may be taken into consideration when applying weights to the absolute difference between points of content (POCs).

[0154] The advantages of this approach include the ability of the time area to compensate for movement between both frames and to maintain temporal correlations between predicted variables or syntax elements.

[0155] Emb.AdaptSize2 QP difference In one embodiment, the size of the time area is determined based on the quantization parameter (QP) of the current frame and the QP of the frame containing the time area. Furthermore, the size is determined based on the difference in QP between these two frames. For example, if the QP of the current frame is lower than the QP of the time frame, the size of the time area increases. Conversely, if the QP of the current frame is higher than the QP of the time frame, it decreases. Moreover, it can be proportional.

[0156] The advantage is improved coding efficiency. In fact, the block size within a frame is related to QP. Indeed, for the same coded frame, a higher QP corresponds to a larger block size. Therefore, a larger time area, where QP is higher for the time frame than for the current frame, increases the chances of finding the correct value.

[0157] Emb.AdaptSize3 time ID, depth In one embodiment, the size of the time area is determined based on the time ID of the current frame and the time ID of the time frame containing the time area. Furthermore, this size can be proportional to the difference between the two time IDs. For example, the size of the time area increases when the time ID of the time frame is smaller than the time ID of the current frame.

[0158] Alternatively, hierarchy depth can be considered instead of time ID.

[0159] The advantage is improved coding efficiency. In fact, frames with small time IDs are often coded over larger time distances between frames. In this case, it is better to increase the time area.

[0160] Value sent in the Emb.AdaptSize4 header In one embodiment, the size of the time area is determined based on a value transmitted in the header. This value may be transmitted alternatively or additionally in the SPS, PPS, picture header, or slice header. For example, the value may be a prediction transmitted in the SPS and estimated in the PPS. An override flag may indicate whether this value is updated for the picture header or not compared to the value in the PPS or SPS.

[0161] The advantage of this is that the implementation of the encoder is not constrained and can be adapted to choose between coding efficiency and complexity.

[0162] Emb.AdaptSize5 Current block size In one embodiment, the size of the time area is determined based on the current block size. Therefore, for larger blocks, the time area is larger, and for smaller blocks, it is smaller. For example, when considering the minimum time area size corresponding to the minimum possible block size, the size of the current block is added to this minimum size in order to obtain the final time area size corresponding to the current block.

[0163] This is adapted to multiple block sizes.

[0164] Blocks within the time area Emb.All_Blocks: All blocks in the time area In one embodiment, the blocks considered for determining the value are all the blocks in the time area. In the example in Figure 16, the time area (1601) does not have its boundaries aligned by the divisions. This corresponds to all blocks (1602-1620) that have at least one sample within the time area (1601). Therefore, in this figure, 19 blocks are considered.

[0165] This is the simplest way to consider blocks within a time area.

[0166] Emb.KeepProportion preserves the proportion of blocks within the time area. In one embodiment, the blocks considered for determining the value are all blocks in the time area, but the values ​​extracted from the blocks are weighted to consider only the portion of the block that is within the time area. In the example in Figure 16, for block 1617, the weights corresponding to portion 1630 are determined, for example, to calculate the average of several values.

[0167] Compared to the previous embodiment, this embodiment is more complex because it requires some additional calculations, but it increases coding efficiency because it is more locally adapted to the current block.

[0168] Emb.FullyContain: Blocks within a time area In one embodiment, instead of the previous block, the block considered for determining the value is the block entirely inside the time area. In the example in Figure 16, if the time area (1601) does not have a boundary aligned by the division, this corresponds to all blocks 1601-1604. Thus, four blocks are considered.

[0169] The advantage of this embodiment is that fewer values ​​need to be used in the calculation, thus reducing the complexity of determining the values. However, this is not the worst case if the time area coincides with the division segment.

[0170] Position of the time area compared to the current block Emb.Center: Center of the block position in the current frame. In one embodiment, the center position of the current block is the center of the time area within the time frame.

[0171] The center is, on average, the best representation of a block. Therefore, coding efficiency is better.

[0172] Alternatively, if the center of the block is outside the frame, the top-left position may be considered.

[0173] Embed.shiftedBasedMV: The time area can be shifted according to motion information. Even if the time area could compensate for the use of motion information for parsing, in one embodiment, the time area is shifted according to a motion vector, for example, an adjacent motion vector. This is particularly efficient when the motion information is large or when the time distance between the current frame and the time frame of the area in which it is located is large.

[0174] Decimation of time area One position on Emb.Decimation N In one embodiment, only blocks present in the grid are considered for determining a value from the time area. In this embodiment, only blocks whose height and width are multiples of N are considered for determining a value.

[0175] Figure 17 shows an example of this embodiment. In this figure, only blocks (1702-1710) within a grid NxN, represented by dots, are used to determine values ​​from time area 1701. Note that in this figure, the time area is aligned by divisions for the sake of simplicity.

[0176] The main advantage is the reduction in the buffers required to store the values ​​used to determine the value from the time area. In fact, all the values ​​needed to determine this value would need to be held in memory for each frame that could be used as a frame in the time area. In the case of a hardware implementation, the worst-case scenario is considered when designing the buffer. The worst-case scenario is the minimum block size in that case. As the minimum size for 8x4 or 4x8, in the worst case, it can be assumed that the values ​​are stored every 4x4 blocks. Therefore, for a 1080p frame, (1920 / 4)*(1080 / 4) = 129600 related values ​​would need to be stored for the time frame. For example, if we assume that the value N is equal to 16, then only (1920 / 16)*(1080 / 16) = 8100 pieces of related information would need to be stored for the time frame. Thus, it reduces the information that needs to be stored by 16.

[0177] Another advantage is the reduction in complexity when the divisions contain small blocks. In fact, it reduces the worst-case complexity because we only need to consider the largest block on the grid.

[0178] Furthermore, it gives greater importance to blocks with larger sizes, which is a better representation of what happens in the time area, in contrast to smaller blocks corresponding to some areas with less frequency. As a result, the decimation is better. Surprisingly, this decimation also leads to an improvement in coding efficiency, especially when the values ​​obtained from the time area are used to restrict or infer the variables of QT, BT, and TT. A decimation where N equals 16 gives the best coding efficiency. If the time area is a CTU of 256x256 luma samples, in the worst case only 256 blocks need to be considered. Conversely, without this decimation, in the worst case 2048 blocks need to be considered, since the minimum block is 4x8 or 8x4.

[0179] Emb.SumPropor is proportional to the block size (especially for data representing segments). In one embodiment, to determine a value from the time area, only blocks present within the grid are considered, and the size of the blocks is considered, for example, when calculating the average. Thus, the value is obtained by considering the proportionality of the blocks considered. This also applies to blocks within the boundaries of the time area, as described in the previous embodiment.

[0180] The advantage of this approach is that it improves coding efficiency compared to methods that do not apply this proportionality.

[0181] Emb.SumCenter grid position N is centered. In one embodiment, the grid under consideration is centered relative to the time area. Considering that the top-left corner of the time area has position (0,0), the top-left position of the grid is (N / 2,N / 2). Figure 18(a) shows a grid that is not centered in the time area, and Figure 18(b) shows a grid that is centered in the time area.

[0182] The advantage is that the closer the value being held is (on average) to the center of the current block, the better the coding efficiency. Emb.SumNotReg Irregular Patterns In one embodiment, the positions are in an irregular pattern. For example, more positions are considered to be in the center of the time area, and the corners of the time area are also considered. Figure 19 shows this irregular pattern.

[0183] One of the advantages of this approach is improved coding efficiency.

[0184] Thinning out the current block positions The possible positions in the Emb.currentPosDecim block are thinned out. In one embodiment, the possible positions of the current block are thinned out. In this embodiment, all possible positions of the block within the current frame are impossible. Only the positions on the grid for each M sample in height and width are considered.

[0185] For example, if we consider the center of the current block as the position in the time area, this position is the center of the time area only if both PosCenter.x and PosCenter.y are multiples of M. Otherwise, the position used is one of the multiples of M around the initial PosCenter. For example, the position of the center of the time area, PosTempo, can be calculated as follows: PosTempo.x = M * (PosCenter.x / M) PosTempo.y = M * (PosCenter.y / M) In these formulas, the partition is an integer partition. Alternatively, it can be obtained by the following shift operation: PosTempo.x=(PosCenter.x>>S)< PosTempo.y=(PosCenter.y>>S)< Here, >> is the left shift operator, << is the right shift operator, and S = Log2(M).

[0186] The main advantage of this embodiment is the reduction in the memory buffer required to store the values ​​determined for the time area. This is of particular interest to encoder implementations. In fact, in an encoder, the values ​​determined from the time area may be determined several times for several block sizes having the same center, for example. If all values ​​need to be determined, this is expensive in terms of memory. Similarly, consider the decimation of the blocks considered in the time area. This decimation of the possible center positions of the time area significantly reduces the buffer (a reduction similar to N=M).

[0187] Furthermore, since fewer values ​​need to be determined on the encoder side, the encoding time is reduced. ​​

[0188] Surprisingly, this decimation leads to improved coding efficiency. This is especially true when values ​​obtained from the time area are used to restrict or infer variables in the QT, BT, and TT sections. A decimation where M is equal to 16 gives the best coding efficiency.

[0189] The Emb.DecimCenter grid position is centered. In one embodiment, the grid of possible block positions has a (M / 2, M / 2) shift for better position decimation, rather than starting from the top-left position (0,0) of the frame. This embodiment is similar to a grid centered on the center of the time area.

[0190] The advantage is improved coding efficiency.

[0191] Emb.DecimBoth Decimation of both time area and block position In one embodiment, both possible positions and decimation of time area blocks are used together.

[0192] This also leads to improved coding efficiency and reduced complexity.

[0193] Embed.Multiplecenter for larger block: For larger blocks, this buffer considers multiple locations. Using buffers significantly reduces memory usage, but this can be constrained if the time area size depends on the current block size. In that case, the value cannot be stored only once on the encoder side.

[0194] In one embodiment, when the size of the current block is used to determine the size of the time area, and when a decimation of possible positions for the current is used, all possible available positions contained within the block are taken into consideration, and all relevant time area values ​​are taken into consideration to determine the value from the time area for the current block.

[0195] This embodiment offers the same advantages as using an adapted time area size according to the current block size. Furthermore, it requires less memory for buffering than using a fixed size for the time area.

[0196] implementation In many implementations, accessing an area in the previous frame requires considering the location of that area, and traversing the area's location is necessary to retrieve information about the blocks or CUs present within that area. This is due to the structure of the buffer containing the information. Therefore, if it is necessary to access a time area and determine its value, the block structure cannot be known without traversing each location. Thus, the basic solution lies in traversing each location and calculating the value associated with each location.

[0197] Emb.Imple Current Implementation In one embodiment, when a value is determined from the time area, the size of the blocks within the time area is taken into consideration to avoid multiple accesses to the same block. Thus, each block has only one access. For example, a Boolean table representing all possible positions within the time area is initialized to the value false. The possible positions in the time area are the NxN grid positions when time area positions are thinned out. Impossible positions in this table are set to equal to 1. Impossible positions are positions outside the time frame. Each position has a corresponding position in this table, which is set to equal to false. When a position is checked in the time area, a value is extracted, and the block is considered extracted. The associated position in the Boolean table is set to equal to true, as are all associated positions corresponding to the current block associated with this position. The next position to be checked is a block that has not yet been checked, and the associated Boolean in the table is equal to false.

[0198] This implementation will speed up the determination of values ​​in the time area.

[0199] Emb.Pred prediction In one embodiment, the time area is used to derive predictors for syntax elements or variables. For example, instead of directly coding the syntax element, the residual is extracted from the bitstream and the predictor is added to this residual. In an alternative example, the first bit is extracted from the bitstream to determine whether the current syntax element is set to equal to the corresponding value obtained from the time area. For example, if this flag is equal to 1, the syntax element is equal to the corresponding value from the time area. Otherwise, this flag is equal to 0, and the other bits are decoded to determine the value of this syntax element.

[0200] The use of time areas is very interesting from the perspective of coding efficiency for obtaining the corresponding values.

[0201] Emb.Infer Guess In one embodiment, instead of the previous method, the value of the syntax element is inferred according to the corresponding value from the time area.

[0202] In one embodiment, the values ​​of variables in the current frame's block are inferred from corresponding values ​​in the time area. For example, the current block's maxMttDepth is determined based on the maxMttDepth determined from the time area "maxMttDepthTempo". Then, according to several rules and the current block's initial maxMttDepth compared to "maxMttDepthTempo", the value of the current block's maxMttDepth is increased, decreased, or left unchanged.

[0203] The advantage of this example is improved coding efficiency, which leads to reduced coding time, by efficiently limiting the maximum multitree depth of the current block.

[0204] Emb.Limit Limit In one embodiment, a value determined from the time area is used to limit a syntax element or variable. For example, a maximum value is determined, and the syntax element value of the current block is limited to this maximum value. Thus, its coding is adapted to this limited number of values ​​in order to reduce the number of bits that need to be transmitted. In the same manner, a minimum value may be considered, or both the maximum and minimum values ​​may be considered.

[0205] Emb.Limit1.QTBT In another example, variables are restricted. For instance, a minimum QT depth from the time area is determined. This value is then used to determine the minimum QT depth relative to the current block QT depth. As a result, fewer bits need to be sent for QT partitioning, and fewer coding possibilities need to be tested, thus reducing encoding time.

[0206] Derived from Emb.Ctx context index increment ctxInc In one embodiment, a value determined from the time area is used to determine the context index increment of a syntax element. For example, the value obtained from the time area "condTempo", if available (availableTempo=1), is added to other context increments obtained from the upper spatial position (condA) and the left spatial position (condL), as shown in the following equation: ctxInc=(condL && availableL)||(condA && availableA)||(condTempo && availableTempo) The advantage of this is improved coding efficiency, as syntax elements generally have spatial and temporal correlations.

[0207] The value Emb.Min Minimum value In one embodiment, the value to be determined from the time area is the minimum value from the blocks to be considered within the time area. Emb.minQTDepth minQTDepth specific cases In certain embodiments, a value determined from the time area is the minimum QT depth value from the blocks considered within the time area. For example, this minQTDepth is then compared to the current QT depth of the current block to determine whether only QT partitioning is permitted.

[0208] This solution reduces encoding execution time and improves coding efficiency by limiting the number of possible divisions to be tested and the associated bits that are not signaled.

[0209] Emb. Max Maximum In one embodiment, the value determined from the time area is the maximum value from the blocks considered within the time area.

[0210] Emb.MaxMttDepth MaxMttDepth specific cases In certain embodiments, a value determined from the time area is the maximum multitree depth value from the blocks considered within the time area. For example, this maxMttDepthTempo is then compared to the current maxMttDepth depth of the current block to determine whether maxMttDepth should be increased or decreased or remain the same, according to other conditions.

[0211] This improves coding efficiency.

[0212] Emb.Median In one embodiment, the value determined from the time area is the median from the blocks considered within the time area.

[0213] Emb.Average In one embodiment, the value determined from the time area is the average value determined from the blocks considered within the time area.

[0214] Maintain the proportion of block size in the time area. As described in other embodiments, the proportion of block sizes should be maintained. In this case, the average value should consider whether each block contains the minimum block unit (minimum block size (4x4)), or alternatively, the number of samples containing the minimum block unit. In this case, the associated value BLVal_i for block i yields the average value AverageVal as follows:

[0215] [Table 5]

[0216] Here, Nb_samples is the total number of samples in all blocks within the time area, nb_blocks is the number of blocks, and height_i and width_i are the height and width of block number i. Alternatively, height_i and width_i can be the height and width, respectively, with respect to the time position, according to the decimation. For example, if the time decimation considered is set to equal 16 (one position for every 16 samples vertically and horizontally), the height and width are divided by 16, or right-shifted by log2(16)=4.

[0217] Furthermore, rounding can be applied to the following formula: AverageVal=(AverageVal+(0.5)) / Nb_samples This rounding improves coding efficiency.

[0218] instead

[0219] [Table 6]

[0220] Keep only the blocks completely within the time area. For blocks that cross the boundary of a time area, only the portion inside the time area is considered, and the height_i and width_i of the relevant block correspond to the height and width inside the time area.

[0221] This can be implemented as the following algorithm:

[0222] [Table 7]

[0223] Here, height_i and width_i are the height and width, respectively, with respect to the time position corresponding to the decimation within the time area.

[0224] Therefore, when a part of the block is outside the temporal area, the number of positions outside the temporal area is subtracted from the height_i and width_i of block i according to the decimation, respectively.

[0225] The blocksize is equal to the number of samples of the current block i. Therefore, it is the multiplication of height and width.

[0226] In hardware implementation, integer division needs to be used. Therefore, the formula needs to be adapted for integer implementation. Thus, in one embodiment, the formula is then as follows:

[0227] [Table 8]

[0228] For the case of the rounded value, roundVal is set equal as follows: roundVal=(tempoAreaWidth>>2)*(tempoAreaHeight>>2)*(tempoAreaWidth / TempoRes)*(tempoAreaHeight / TempoRes) Here, TempoRes corresponds to the decimation of the temporal area. Thus, in some of the above examples, it is 16.

[0229] tempoAreaWidth and tempoAreaHeight are the height and width of the temporal area. Thus, in the above example, they can be equal to the CTU size.

[0230] According to the above example, when the temporal area is set equal to the CTU size and the decimation is equal to 16, roundVal is set equal as follows: roundVal=(CTUSize>>2)*(CTUSize>>2)*(CTUSize / TempoRes)*(CTUSize / TempoRes) This rounding result yields the same result as floating-point division.

[0231] Furthermore, by considering the log2 value, all partitions can be replaced with right shifts.

[0232] Emb.QTDepthTempo QTDepthTempo specific cases In one embodiment, the value determined from the time area is the average value of the QT depth values ​​from the blocks considered within the time area. For example, this QTDepthTempo is then compared to the QT depth of the current block, allowing QT, no splitting, and TT only, or additionally, or alternatively, compared to the QT depth of the current block, and if they are equal, maxMttDepth is increased according to other conditions.

[0233] This results in improved coding efficiency and reduced coding time.

[0234] Emb.Variance (Variance Value) In one embodiment, the value determined from the time area is the variance of the values ​​from the blocks considered within the time area. The variance is the average distance to the mean. This requires first determining the mean and then the variance. Thus, the blocks within the time area are considered twice.

[0235] others Emb.OTHER1 All of these embodiments can be combined.

[0236] Unless otherwise specified, all embodiments described can be combined. In fact, many combinations are synergistic and can result in greater efficiency improvements than the sum of their individual components.

[0237] Emb.Contrib Contribution In particular, the minimum QT depth "minQTDepthTempo", the maximum multitree depth "maxMttDepthTempo", and the average depth value "QTDepthTempo" are determined from the time area. minQTDepthTempo is then compared to the current QT depth of the current block to determine whether only QT splitting is allowed. maxMttDepthTempo is compared to the current maxMttDepth depth of the current block. If it is lower than maxMttDepth, maxMttDepth is decreased. If it is higher and QTDepthTempo is equal to the QT depth of the current block, maxMttDepth is increased. Furthermore, QTDepthTempo is compared to the QT depth of the current block to allow only QT, no splitting, and TT. The time area is based on the reference frame used for time-motion vector prediction and corresponds to the center of the current block. The time area is equal to the CTU size. Only blocks present in a 16x16 grid are considered to determine the three values, and the size of the blocks is taken into account when calculating the average QTDepthTempo. Therefore, the values ​​are obtained by considering the proportionality between blocks. This also applies to blocks within the boundaries of the time area. Finally, the positions of the current blocks are thinned out in a 16x16 grid.

[0238] Implementation of the invention Figure 20 shows systems 191, 195 according to embodiments of the present invention, each having at least one of an encoder 150 or decoder 100 and a communication network 199. According to one embodiment, system 195 is for processing and providing content (e.g., video and audio content for display / output, or streaming video / audio content) to a user who accesses the decoder 100, for example, via the user interface of a user terminal having the decoder 100 or a user terminal capable of communicating with the decoder 100. Such a user terminal may be a computer, mobile phone, tablet, or any other type of device capable of providing / displaying (provided / streamed) content to the user. System 195 acquires / receives a bitstream 101 (e.g., in the form of a continuous stream or signal while previous video / audio is being displayed / output) via the communication network 199. According to one embodiment, system 191 is for processing content and storing the processed content, for example, processed video and audio content for later display / output / streaming. System 191 acquires / receives content having the original sequence 151 of images that have been received and processed by encoder 150 (including filtering by a deblocking filter according to the present invention), and encoder 150 generates a bitstream 101 that is communicated to decoder 100 via communication network 191. The bitstream 101 is then communicated to decoder 100 in several ways, for example, it may be pre-generated by encoder 150 and stored as data in a storage device (e.g., on a server or cloud storage) within communication network 199 until a user requests the content (i.e., bitstream data) from the storage device, at which point the data is communicated / streamed from the storage device to decoder 100.System 191 may also have a content provider for receiving and processing user requests for content so that the requested content can be delivered / streamed from the storage device to the user, and the requested content can be delivered / streamed from the storage device to the user, and the requested content can be delivered / streamed from the storage device to the user terminal, by (for example, by communicating data for a user interface displayed on the user terminal). Alternatively, encoder 150 generates a bitstream 101 and communicates / streams it directly to decoder 100 when the user requests content. Decoder 100 then receives the bitstream 101 (or signal) and performs filtering using a deblocking filter according to the present invention to obtain / generate a video signal 109 and / or an audio signal, which are then used by the user terminal to provide the requested content to the user.

[0239] Any step of the method / process according to the present invention or any function described herein may be implemented in hardware, software, firmware, or any combination thereof. If implemented in software, the step / function may be stored or transmitted as one or more instructions or code or programs, or as computer-readable media, and may be executed by one or more hardware-based processing units, such as a PC ("Personal Computer"), a DSP ("Digital Signal Processor"), a circuit, a processor and memory, a general-purpose microprocessor or central processing unit, a microcontroller, an ASIC ("Application-Specific Integrated Circuit"), a field-programmable logic array (FPGA), or other equivalent integrated or discrete logic circuits. Accordingly, the term "processor" as used herein may refer to any of the aforementioned structures or any other structure suitable for implementing the techniques described herein.

[0240] Embodiments of the present invention can also be implemented by a wide variety of devices or apparatus, including wireless handsets, integrated circuits (ICs), or sets of JCs (e.g., chipsets). While various components, modules, or units are described herein to illustrate functional aspects of devices / apparatus configured to perform their embodiments, implementation by different hardware units is not necessarily required. Rather, the various modules / units may be combined in a codec hardware unit, or provided by a collection of interoperable hardware units including one or more processors in conjunction with appropriate software / firmware.

[0241] Embodiments of the present invention may be realized by a computer in a system or device that includes one or more processing units or circuits for reading and executing computer-executable instructions (e.g., one or more programs) recorded on a storage medium, executing one or more modules / units / functions of the embodiments described above, and / or for executing one or more functions of the embodiments described above, and by controlling, for example, one or more processing units or circuits for executing one or more functions of the embodiments described above. The computer may include separate processing units or a network of separate computers for reading and executing computer-executable instructions. Computer-executable instructions may be provided to the computer from a computer-readable medium, such as a communication medium via a network or tangible storage medium. The communication medium may be a signal / bitstream / carrier wave. Tangible storage media are “non-temporary computer-readable storage media” which may include one or more of the following: hard disks, random access memory (RAM), read-only memory (ROM), storage devices for distributed computing systems, optical discs (such as Compact Discs (CDs), Digital Multipurpose Discs (DVDs), or Blu-ray Discs (BDs) (trademarks)), flash memory devices, memory cards, etc. At least some of the steps / functions may also be implemented in hardware by devices or dedicated components such as FPGAs (“Field-Programmable Gate Arrays”) or ASICs (“Application-Specific Integrated Circuits”).

[0242] Figure 21 is a schematic block diagram of a computing device 3600 for implementing one or more embodiments of the present invention. The computing device 3600 may be a device such as a microcomputer, workstation, or light portable device. The computing device 3600 has a communication bus connected to: - a central processing unit (CPU) 3601 such as a microprocessor; - random access memory (RAM) 3602 for storing executable code of the methods of embodiments of the present invention, and registers adapted to record variables and parameters necessary to implement the methods for encoding or decoding at least a portion of an image according to embodiments of the present invention, the memory capacity of which may be expanded, for example, by optional RAM connected to an expansion port; - read-only memory (ROM) 3603 for storing computer programs for implementing embodiments of the present invention; - a network interface (NET) 3604, which is typically connected to a communication network through which digital data to be processed is transmitted or received. The network interface (NET) 3604 may be a single network interface or may consist of a set of different network interfaces (e.g., wired and wireless interfaces, or different types of wired or wireless interfaces). Data packets are written to the network interface for transmission or read from the network interface for reception under the control of a software application running on the CPU 3601; a user interface (UI) 3605 may be used to receive input from the user or to display information to the user; a hard disk (HD) 3606 may be provided as mass storage; and an input / output module (IO) 3607 may be used to receive / transmit data to and from external devices such as a video source or display. Executable code may be stored in ROM 3603, HD 3606, or on a removable digital medium such as a disk.In a modified version, the executable code of a program may be received via a communication network through NET3604 to be stored in one of the storage means of a communication device 3600, such as HD3606, before execution. The CPU3601 is adapted to control and direct the execution of instructions or parts of instructions of a set of programs or software code of a program according to embodiments of the present invention, and these instructions are stored in one of the aforementioned storage means. After power-up, and after those instructions have been loaded, for example, from the program ROM3603 or HD3606, the CPU3601 can execute instructions from the main RAM memory 3602 relating to a software application. When such a software application is executed by the CPU3601, it causes the steps of the method according to the present invention to be performed.

[0243] Furthermore, according to another embodiment of the present invention, it is understood that the decoder according to the above embodiment is provided in a user terminal such as a computer, a mobile phone, a table, or any other type of device (e.g., a display device) that can provide / display content to a user. According to yet another embodiment, the encoder according to the above embodiment is provided in an image capture device which also has a camera, video camera, or network camera (e.g., closed-circuit television or video surveillance camera) that captures and provides content for the encoder to encode. Two such examples are provided below with reference to Figures 37 and 38.

[0244] Figure 22 shows a network camera system 3700, which includes a network camera 3702 and a client device 202.

[0245] The network camera 3702 includes an imaging unit 3706, an encoding unit 3708, a communication unit 3710, and a control unit 3712.

[0246] The network camera 3702 and the client device 202 are interconnected via the network 200 so that they can communicate with each other.

[0247] The imaging unit 3706 includes a lens and an image sensor (e.g., a CCD (charge coupled device) or a CMOS (complementary metal oxide semiconductor)), captures an image of a subject, and generates image data based on the image. This image can be a still image or a video image.

[0248] The encoding unit 3708 encodes the image data using the above-described encoding method. Or a combination of the above-described encoding methods.

[0249] The communication unit 3710 of the network camera 3702 transmits the encoded image data encoded by the encoding unit 3708 to the client device 202.

[0250] Also, the communication unit 3710 receives a command from the client device 202. The command includes a command for setting parameters for encoding by the encoding unit 3708.

[0251] The control unit 3712 controls other units within the network camera 3702 according to the command received by the communication unit 3712.

[0252] The client device 202 includes a communication unit 3714, a decoding unit 3716, and a control unit 3718.

[0253] The communication unit 3714 of the client device 202 transmits a command to the network camera 3702.

[0254] Also, the communication unit 3714 of the client device 202 receives the encoded image data from the network camera 3712.

[0255] The decoding unit 3716 decodes the encoded image data using the above-described decoding method, or a combination of the above-described decoding methods.

[0256] The control unit 3718 of the client device 202 controls other units within the client device 202 according to user operations and commands received by the communication unit 3714.

[0257] The control unit 3718 of the client device 202 controls the display device 2120 to display the image decoded by the decoding unit 3716.

[0258] Furthermore, the control unit 3718 of the client device 202 controls the display device 2120 to display a GUI (Graphical User Interface) and specifies the parameter values ​​of the network camera 3702, including the encoding parameters of the encoding unit 3708.

[0259] Furthermore, the control unit 3718 of the client device 202 controls other units within the client device 202 according to user input to the GUI displayed by the display device 2120.

[0260] The control unit 3718 of the client device 202 controls the communication unit 3714 of the client device 202 to send a command to the network camera 3702 specifying the parameter values ​​of the network camera 3702, in accordance with user input to the GUI displayed by the display device 2120.

[0261] Figure 23 shows the smartphone 3800.

[0262] The smartphone 3800 includes a communication unit 3802, a decoding unit 3804, a control unit 3806, and a display unit 3808.

[0263] The communication unit 3802 receives encoded image data via the network 200.

[0264] The decoding unit 3804 decodes the encoded image data received by the communication unit 3802.

[0265] The decoding / encoding unit 3804 decodes and encodes the encoded image data using the decoding method described above.

[0266] The control unit 3806 controls other units within the smartphone 3800 according to user operations or commands received by the communication unit 3806.

[0267] For example, the control unit 3806 controls the display unit 3808 to display the image decoded by the decoding unit 3804. The smartphone 3800 may also have a sensor 3812 and an image recording device 3810. In this way, the smartphone 3800 can record images and encode them (using the method described above).

[0268] The smartphone 3800 can then decode the encoded images (using the method described above) and display them via the display unit 3808, or transmit the encoded images to another device via the communication unit 3802 and the network 200.

[0269] Substitute and change While the present invention has been described with reference to embodiments, it should be understood that the present invention is not limited to the disclosed embodiments. Those skilled in the art will understand that various changes and modifications can be made without departing from the scope of the invention, as defined in the appended claims. All features disclosed in this specification (including any appended claims, abstract, and drawings) and / or all steps of any method or process so thus disclosed can be combined in any combination, except for any combination in which at least some of such features and / or steps are mutually exclusive. Each feature disclosed in this specification (including any appended claims, abstract, and drawings) may be replaced by an alternative feature serving the same, equivalent, or similar purpose, unless otherwise expressly stated. Therefore, unless specifically stated, each disclosed feature is merely an example of a general set of equivalent or similar features.

[0270] Furthermore, any result of the above comparison, judgment, evaluation, selection, execution, performing, or consideration, for example, a selection made during the encoding or filtering process, may be indicated in data within the bitstream, for example, in a flag or data indicating the result, or may be decidable / inferred from there, and as a result, the indicated or determined / inferred result may be used in the process instead of actually performing the comparison, judgment, evaluation, selection, execution, performing, or consideration, for example, during the decoding process.

[0271] In the claims, the word “comprising” does not exclude other elements or steps, and the indefinite article “a” or “an” does not exclude plurals. The mere fact that different features are described in different dependent claims does not imply that combinations of these features cannot be used to one's advantage.

[0272] The reference numerals appearing in the claims are for illustrative purposes only and do not have any limiting effect on the scope of the claims.

Claims

1. A method for decoding video data from a bitstream, wherein the bitstream includes video data corresponding to a plurality of frames arranged in decoding order, and the method is The process involves deriving a value from the first area in the first frame among the aforementioned multiple frames, From the aforementioned values, determine the variables, syntax elements, or contextual increments of syntax elements related to the second area in the second frame. It has, A method wherein the first frame precedes the second frame in the decoding sequence.

2. The method according to claim 1, wherein each of the plurality of frames has an associated time ID, and the first frame and the second frame have the same time ID.

3. The method according to claim 2, wherein the first frame corresponds to the frame closest to the second frame having the same time ID in the decoding order.

4. The method according to any one of claims 1 to 3, wherein each of the plurality of frames has a related quantization parameter QP, and the first frame and the second frame have the same QP.

5. The method according to any one of claims 1 to 4, wherein the first frame is a reference frame.

6. The method according to any one of claims 1 to 5, wherein the first area is equal to or greater than the size of the second area.

7. The method according to claim 6, wherein the second area corresponds to a coding tree unit CTU.

8. The method according to any one of claims 1 to 7, wherein the size of the first area is set based on the temporal distance between the first frame and the second frame.

9. The method according to claim 8, wherein the time difference is calculated based on the difference between the picture order count POC of the first frame and the POC of the second frame.

10. The method according to any one of claims 1 to 9, wherein each of the plurality of frames has an associated quantization parameter QP, and the size of the first area is set based on the difference between the QP of the first frame and the QP of the second frame.

11. The method according to any one of claims 1 to 10, wherein each of the plurality of frames has an associated time ID, and the size of the first area is set based on the difference between the time ID of the first frame and the time ID of the second frame.

12. The method according to any one of claims 1 to 7, wherein the size of the first area is set based on a value transmitted in one of the sequence parameter set, picture parameter set, picture header, and slice header included in the bitstream.

13. The method according to any one of claims 1 to 12, wherein the size of the first area is set based on the size of the second area.

14. The method according to any one of claims 1 to 13, wherein the step of deriving a value from a first area in a first frame of the plurality of frames includes deriving a value from a block having at least one sample in the first area.

15. The method according to any one of claims 1 to 13, wherein the step of deriving a value from a first area in a first frame of the plurality of frames includes deriving values ​​from all blocks having at least one sample in the first area.

16. The method according to claim 14 or 15, wherein the step of deriving a value from a first area in a first frame of the plurality of frames includes weighting the block or the value derived from each block based on the number of samples of each block having in the first area.

17. The method according to any one of claims 14 to 16, wherein the value is derived from the block or each block using integer arithmetic.

18. The method according to any one of claims 1 to 13, wherein the step of deriving a value from a first area in the first frame of the plurality of frames comprises deriving a value only from blocks that are entirely contained within the first area.

19. The method according to any one of claims 1 to 13, wherein the step of deriving a value from a first area in the first frame of the plurality of frames comprises deriving a value only from blocks located on an NxN grid within the first area, where N is an integer.

20. The method according to claim 19, wherein N = 16.

21. The method according to claim 19 or 20, wherein the center of the grid is located in the same position as the center of the first area of ​​the first frame within the first frame.

22. The method according to any one of claims 1 to 13, wherein the step of deriving a value from a first area in the first frame of the plurality of frames includes deriving a value only from blocks located on points of a pattern in the first area.

23. The method according to claim 22, wherein the points of the pattern are unevenly spaced.

24. The method according to any one of claims 1 to 23, wherein the step of determining a variable, syntax element, or contextual increment of a syntax element relating to a second area in a second frame from the aforementioned value comprises determining a variable, syntax element, or contextual increment of a syntax element relating to a block in the second area.

25. The method according to any one of claims 1 to 23, wherein the step of determining a variable, syntax element, or context increment of a syntax element relating to a second area in a second frame from the aforementioned value comprises determining a variable, syntax element, or context increment of a syntax element relating to a block located on an MxM grid in the second area, where M is an integer.

26. The method according to claim 25, wherein the MxM grid is shifted by M / 2 horizontally and M / 2 vertically with respect to the upper left position of the second area.

27. The method according to any one of claims 1 to 26, wherein the center of the first area is located in the same position within the first frame as the center of the second area within the second frame.

28. The method according to any one of claims 1 to 26, wherein the center of the first area is located at the position in the first frame corresponding to the center of the second area in the second frame, which is shifted by an amount corresponding to a motion vector derived from an area adjacent to the second area.

29. The method according to any one of claims 1 to 26, further comprising deriving a value from the upper left position of the first area if the center of the first area is outside the first frame.

30. The method according to any one of claims 1 to 29, wherein the step of deriving a value from a first area in the first frame of the plurality of frames includes accessing each block in the first area only once.

31. The step of deriving a value from the first area in the first frame among the plurality of frames is: The method according to any one of claims 1 to 30, comprising deriving a value from a first area in a first frame among the plurality of frames and a third area in a third frame among the plurality of frames.

32. The method according to any one of claims 1 to 31, wherein the step of determining a variable, syntax element, or context increment of a syntax element relating to a second area in a second frame from the aforementioned value includes determining a predictor to be added to the residual derived from the bitstream.

33. The method according to any one of claims 1 to 31, wherein the value derived from the first area includes maxMtDepth, and the step of determining from the value a variable, syntax element, or contextual increment of a syntax element relating to the second area in the second frame includes determining maxMtDepth of the block of the second area.

34. The method according to any one of claims 1 to 31, wherein the value derived from the first area is used to limit the context increment of a syntax element or variable relating to the second area.

35. The method according to any one of claims 1 to 31, wherein the value derived from the first area is the minimum value of the quadtree depth value minQTDepth from the block in the first area, and the step of determining from the value a variable, syntax element, or context increment of a syntax element relating to the second area in the second frame, comprises comparing the derived minQTDepth with the quadtree depth of the block in the second area to determine whether only the quadtree split is permitted.

36. The method according to any one of claims 1 to 31, wherein the value derived from the first area is the maximum value of the multitree depth value maxMtDepth from the block in the first area, and the step of determining from the value a variable, syntax element, or contextual increment of a syntax element relating to the second area in the second frame includes comparing the derived maxMtDepth with the maxMtDepth of the block in the second area, and modifying maxMtDepth based on the comparison.

37. The method according to any one of claims 1 to 31, wherein the value derived from the first area is the average of the quadtree depth values ​​from the block in the first area, and the step of determining from the value a variable, syntax element, or context increment of a syntax element relating to the second area in the second frame includes comparing the derived quadtree depth value with the quadtree depth of the block in the second area.

38. The method according to claim 37, further comprising the step of determining an acceptable division based on the comparison.

39. The method according to claim 37 or claim 38, further comprising the step of changing maxMtDepth based on the comparison.

40. A device for decoding image data from a bitstream, the device is configured to perform the method described in any one of claims 1 to 39.

41. A method for encoding video data into a bitstream, wherein the bitstream includes video data corresponding to a plurality of frames arranged in decoding order, and the method is To derive a value for the first area in the first frame among the aforementioned plurality of frames, From the aforementioned values, determine the variables, syntax elements, or contextual increments of syntax elements related to the second area in the second frame. It has, A method wherein the first frame precedes the second frame in the decoding sequence.

42. The method according to claim 41, wherein each of the plurality of frames has an associated time ID, and the first frame and the second frame have the same time ID.

43. The method according to claim 42, wherein the first frame corresponds to the frame closest to the second frame having the same time ID in the decoding order.

44. The method according to any one of claims 41 to 43, wherein each of the plurality of frames has a related quantization parameter QP, and the first frame and the second frame have the same QP.

45. The method according to any one of claims 41 to 44, wherein the first frame is a reference frame.

46. The method according to any one of claims 41 to 45, wherein the first area is equal to or greater than the size of the second area.

47. The method according to claim 46, wherein the second area corresponds to a coding tree unit CTU.

48. The method according to any one of claims 41 to 47, wherein the size of the first area is set based on the temporal distance between the first frame and the second frame.

49. The method according to claim 48, wherein the time difference is calculated based on the difference between the picture order count POC of the first frame and the POC of the second frame.

50. The method according to any one of claims 41 to 49, wherein each of the plurality of frames has an associated quantization parameter QP, and the size of the first area is set based on the difference between the QP of the first frame and the QP of the second frame.

51. The method according to any one of claims 41 to 50, wherein each of the plurality of frames has an associated time ID, and the size of the first area is set based on the difference between the time ID of the first frame and the time ID of the second frame.

52. The method according to any one of claims 41 to 47, wherein the size of the first area is set based on a value transmitted in one of the sequence parameter set, picture parameter set, picture header, and slice header included in the bitstream.

53. The method according to any one of claims 41 to 52, wherein the size of the first area is set based on the size of the second area.

54. The method according to any one of claims 41 to 53, wherein the step of deriving a value from a first area in a first frame of the plurality of frames includes deriving a value from a block having at least one sample in the first area.

55. The method according to any one of claims 40 to 52, wherein the step of deriving a value from a first area in a first frame of the plurality of frames comprises deriving a value from all blocks having at least one sample in the first area.

56. The method according to claim 54 or 55, wherein the step of deriving a value from a first area in a first frame of the plurality of frames includes weighting the block or the value derived from each block based on the number of samples of each block having in the first area.

57. The method according to any one of claims 54 to 56, wherein the value is derived from the block or each block using integer arithmetic.

58. The method according to any one of claims 41 to 53, wherein the step of deriving a value from a first area in the first frame of the plurality of frames comprises deriving a value only from blocks that are entirely contained within the first area.

59. The method according to any one of claims 41 to 53, wherein the step of deriving a value from a first area in the first frame of the plurality of frames comprises deriving a value only from blocks located on an NxN grid within the first area, where N is an integer.

60. The method according to claim 59, wherein N = 16.

61. The method according to claim 59 or 60, wherein the center of the grid is located in the same position as the center of the first area of ​​the first frame within the first frame.

62. The method according to any one of claims 41 to 53, wherein the step of deriving a value from a first area in the first frame of the plurality of frames includes deriving a value only from blocks located on points of a pattern in the first area.

63. The method according to claim 62, wherein the points of the pattern are unevenly spaced.

64. The method according to any one of claims 41 to 63, wherein the step of determining a variable, syntax element, or contextual increment of a syntax element relating to a second area in a second frame from the aforementioned value comprises determining a variable, syntax element, or contextual increment of a syntax element relating to a block in the second area.

65. The method according to any one of claims 41 to 63, wherein the step of determining a variable, syntax element, or context increment of a syntax element relating to a second area in a second frame from the aforementioned value comprises determining a variable, syntax element, or context increment of a syntax element relating to a block located on an MxM grid in the second area, where M is an integer.

66. The method according to claim 65, wherein the MxM grid is shifted by M / 2 horizontally and M / 2 vertically with respect to the upper left position of the second area.

67. The method according to any one of claims 41 to 66, wherein the center of the first area is located in the same position within the first frame as the center of the second area within the second frame.

68. The method according to any one of claims 41 to 66, wherein the center of the first area is located at the position in the first frame corresponding to the center of the second area in the second frame, which is shifted by an amount corresponding to a motion vector derived from an area adjacent to the second area.

69. The method according to any one of claims 41 to 66, further comprising deriving a value from the upper left position of the first area if the center of the first area is outside the first frame.

70. The method according to any one of claims 41 to 69, wherein the step of deriving a value from a first area in the first frame of the plurality of frames includes accessing each block in the first area only once.

71. The step of deriving a value from the first area in the first frame among the plurality of frames is: The method according to any one of claims 41 to 70, comprising deriving a value from a first area in a first frame among the plurality of frames and a third area in a third frame among the plurality of frames.

72. The method according to any one of claims 41 to 71, wherein the step of determining from the aforementioned value a variable, syntax element, or context increment of a syntax element relating to a second area in a second frame includes determining a predictor to be added to the residual derived from the bitstream.

73. The method according to any one of claims 41 to 71, wherein the value derived from the first area includes maxMtDepth, and the step of determining from the value a variable, syntax element, or contextual increment of a syntax element relating to the second area in the second frame includes determining maxMtDepth of the block of the second area.

74. The method according to any one of claims 41 to 71, wherein the value derived from the first area is used to limit the context increment of a syntax element or variable relating to the second area.

75. The method according to any one of claims 41 to 71, wherein the value derived from the first area is the minimum value of the quadtree depth value minQTDepth from the block in the first area, and the step of determining from the value a variable, syntax element, or context increment of a syntax element relating to the second area in the second frame, comprises comparing the derived minQTDepth with the quadtree depth of the block in the second area to determine whether only the quadtree split is permitted.

76. The method according to any one of claims 41 to 71, wherein the value derived from the first area is the maximum value of the multitree depth value maxMtDepth from the block in the first area, and the step of determining from the value a variable, syntax element, or contextual increment of a syntax element relating to the second area in the second frame includes comparing the derived maxMtDepth with the maxMtDepth of the block in the second area, and modifying maxMtDepth based on the comparison.

77. The method according to any one of claims 41 to 71, wherein the value derived from the first area is the average of the quadtree depth values ​​from the block in the first area, and the step of determining from the value a variable, syntax element, or context increment of a syntax element relating to the second area in the second frame comprises comparing the derived quadtree depth value with the quadtree depth of the block in the second area.

78. The method according to claim 77, further comprising the step of determining an acceptable division based on the comparison.

79. The method according to claim 77 or claim 78, further comprising the step of changing maxMtDepth based on the comparison.

80. An apparatus for encoding image data into a bitstream, wherein the apparatus is configured to perform the method described in any one of claims 41 to 79.

81. A computer program configured to perform the method described in any one of claims 1 to 39 or 41 to 79 during execution.