Encoding and decoding of images and videos

By optimizing the encoding and decoding process using multi-frame correlation information in the VVC standard, the problem of improving the compression performance of HEVC in high-efficiency video coding is solved, achieving a significant improvement in compression efficiency, and is suitable for 360-degree and HDR videos.

CN121128167APending Publication Date: 2025-12-12CANON KK
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202480024271.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2023-06-08
Filing Date
2024-04-08
Publication Date
2025-12-12

AI Technical Summary

Technical Problem

Existing video coding standards such as HEVC have room for improvement in compression performance, especially when processing 360-degree and high dynamic range (HDR) videos, where it is difficult to achieve significant improvements in compression efficiency.

Method used

By introducing new tools and improving existing ones in the Multifunctional Video Coding (VVC) standard, the encoding and decoding processes are optimized by deriving values ​​from the correlation information between multiple frames, such as temporal ID, quantization parameter QP, and temporal distance, to determine the contextual increments or variables of syntactic elements.

Benefits of technology

It improves the compression efficiency of video encoding, achieving a compression performance improvement of over 50%, and is suitable for ultra-high-definition (UHD) video testing materials.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121128167A_ABST
    Figure CN121128167A_ABST
Patent Text Reader

Abstract

A method for encoding image data in / decoding video data from a bitstream, the bitstream comprising video data corresponding to a plurality of frames arranged in a decoding order, the method comprising: deriving a value from a first region in a first frame of the plurality of frames; and determining a context increment or a syntax element or a variable of a syntax element related to a second region in a second frame from the value, where the first frame precedes the second frame in decoding order. Apparatus for performing these methods are also disclosed.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to encoding and decoding data from a bitstream, and more particularly to encoding and decoding image and video partition data. Means for decoding video data from a bitstream and encoding video data into a bitstream are also provided, as well as a computer program configured to encode or decode video data at execution. Background Technology

[0002] The Joint Video Experts Group (JVET) (a collaborative team comprised of MPEG and ITU-T Study Group 16's VCEG) has released a new video coding standard called Multifunctional Video Coding (VVC). VVC aims to provide a significant improvement in compression performance (i.e., typically twice as much) over the existing HEVC standard. Key target applications and services include, but are not limited to, 360-degree and High Dynamic Range (HDR) video. Specific effects have been demonstrated on Ultra High Definition (UHD) video test material. Therefore, for the final standard, we can expect an improvement in compression efficiency well beyond the target of 50%.

[0003] Since the end of the VVC v1 standard, JVET has initiated the exploration phase by establishing Explorer Software (ECM). JVET has collected additional tools and improved existing tools based on the VVC standard to achieve better coding efficiency. Summary of the Invention

[0004] According to a first aspect of the invention, a method is provided for decoding video data from a bitstream, the bitstream comprising video data corresponding to a plurality of frames arranged in a decoding order, the method comprising: deriving a value from a first region in a first frame of the plurality of frames; and determining, based on the value, a context increment or syntactic element or variable relating to a second region in a second frame, wherein the first frame precedes the second frame in the decoding order.

[0005] According to a second aspect of the invention, a method is provided for encoding video data into a bitstream, the bitstream comprising video data corresponding to a plurality of frames arranged in a decoding order, the method comprising: deriving a value for a first region in a first frame of the plurality of frames; and determining, based on the value, a context increment or syntactic element or variable relating to a second region in a second frame, wherein the first frame precedes the second frame in the decoding order.

[0006] Optionally, each of the plurality of frames has an associated temporal ID, and the first frame and the second frame have the same temporal ID. Optionally, the first frame corresponds to the frame with the same temporal ID that is closest to the second frame in the decoding order.

[0007] Optionally, each of the plurality of frames has an associated quantization parameter, namely QP, and wherein the first frame and the second frame have the same QP.

[0008] Optionally, the first frame is a reference frame.

[0009] Optionally, the first region is equal to or larger than the second region in size.

[0010] Optionally, the second region corresponds to a coding tree unit, i.e., a CTU.

[0011] Optionally, the size of the first region is set based on the temporal distance between the first frame and the second frame. Optionally, the temporal difference is calculated based on the difference between the image sequence count (POC) of the first frame and the POC of the second frame.

[0012] Optionally, each of the plurality of frames has an associated quantization parameter, namely QP, and wherein the size of the first region is set based on the difference between the QP of the first frame and the QP of the second frame.

[0013] Optionally, each of the plurality of frames has an associated temporal ID, and the size of the first region is set based on the difference between the temporal ID of the first frame and the temporal ID of the second frame.

[0014] Optionally, the size of the first region is set based on values ​​transmitted in one of the following: a sequence parameter set, an image parameter set, an image header, and a strip header contained in the bit stream.

[0015] Optionally, the size of the first region can be set based on the size of the second region.

[0016] Optionally, the step of deriving a value from a first region in a first frame of the plurality of frames includes: deriving a value from a block having at least one sample in the first region.

[0017] Optionally, the step of deriving values ​​from a first region in a first frame of the plurality of frames includes: deriving values ​​from all blocks that have at least one sample in the first region.

[0018] Optionally, the step of deriving values ​​from a first region in a first frame of the plurality of frames includes: weighting the values ​​derived from the block or blocks based on the number of samples each block has in the first region.

[0019] Optionally, the step of deriving a value from a first region in a first frame of the plurality of frames includes: deriving a value only from a block that is completely contained within the first region.

[0020] Optionally, the step of deriving values ​​from a first region in a first frame of the plurality of frames includes: deriving values ​​only from blocks located on an N×N grid within the first region, where N is an integer. Optionally, N = 16.

[0021] Optionally, the center of the grid is located at the same position within the first frame as the center of the first region within the first frame.

[0022] Optionally, the step of deriving values ​​from a first region in a first frame of the plurality of frames includes: deriving values ​​only from blocks of points of a pattern located within the first region. Optionally, the points of the pattern are spaced out in a non-uniform manner.

[0023] Optionally, the step of determining the context increment or syntactic element or variable of a syntactic element associated with a second region in the second frame based on the value includes: determining the context increment or syntactic element or variable of a syntactic element associated with a block in the second region.

[0024] Optionally, the step of determining the context increment, syntactic element, or variable of a syntactic element associated with a second region in the second frame based on the value includes: determining the context increment, syntactic element, or variable of a syntactic element associated with a block located on an M×M grid within the second region, where M is an integer. Optionally, the M×M grid is shifted by M / 2 in the horizontal direction and by M / 2 in the vertical direction relative to the upper-left position of the second region.

[0025] Optionally, the center of the first region is located at the same position within the first frame as the center of the second region within the second frame.

[0026] Optionally, the center of the first region is located within the first frame at a position corresponding to the position within the second frame after the center of the second region has been shifted by an amount corresponding to the motion vector derived from the region adjacent to the second region.

[0027] Optionally, the method further includes: deriving a value from the upper left position of the first region when the center of the first region is outside the first frame.

[0028] Optionally, the step of deriving a value from a first region in a first frame of the plurality of frames includes: accessing each block within the first region only once.

[0029] Optionally, the step of deriving a value from a first region in a first frame of the plurality of frames includes:

[0030] Values ​​are derived from the first region in the first frame of the plurality of frames and the third region in the third frame of the plurality of frames.

[0031] Optionally, the step of determining the context increment or syntactic element or variable of the syntactic element associated with the second region in the second frame based on the value includes: determining a predictor to be added to the residual derived from the bitstream.

[0032] Optionally, the value derived from the first region includes maxMttDepth, and the step of determining a context increment or syntactic element or variable of a syntactic element associated with the second region in the second frame based on the value includes: determining maxMttDepth for blocks of the second region.

[0033] Optionally, values ​​derived from the first region can be used to limit the contextual increment of variables or syntactic elements associated with the second region.

[0034] Optionally, the value derived from the first region is the minimum value minQTDepth of the quadtree depth of the block from the first region, and the step of determining the context increment or syntactic element or variable of the syntactic element associated with the second region in the second frame based on the value includes: comparing the derived minQTDepth with the quadtree depth of the block in the second region to determine whether only quadtree splitting is allowed.

[0035] Optionally, the value derived from the first region is the maximum value of the multi-tree depth value of the block from the first region, maxMttDepth, and the step of determining the context increment or syntactic element or variable of the syntactic element associated with the second region in the second frame based on the value includes: comparing the derived maxMttDepth with the maxMttDepth of the block in the second region and changing maxMttDepth based on the comparison.

[0036] Optionally, the value derived from the first region is the average of the quadtree depth values ​​of blocks from the first region, and the step of determining the context increment or syntactic element or variable of a syntactic element related to the second region in the second frame based on the value includes: comparing the derived quadtree depth value with the quadtree depth of a block in the second region. Optionally, the method further includes the step of: determining an allowable split based on the comparison. Optionally, the method further includes the step of: changing maxMttDepth based on the comparison.

[0037] According to a third aspect of the invention, an apparatus is provided for decoding video data from a bitstream, wherein the apparatus is configured to perform the method of the first aspect.

[0038] According to a fourth aspect of the invention, there is provided an apparatus for encoding video data into a bitstream, wherein the apparatus is configured to perform the method of the second aspect.

[0039] According to a fifth aspect of the invention, a computer program is provided, which is configured to perform the method of the first aspect or the second aspect when executed. Attached Figure Description

[0040] Now, let's refer to the attached diagram as an example, in which:

[0041] Figure 1 This is a diagram used to illustrate the coding structure used in HEVC;

[0042] Figure 2 This is a block diagram schematically illustrating a data communication system that can implement one or more embodiments of the present invention;

[0043] Figure 3 This is a block diagram illustrating components of a processing apparatus that can implement one or more embodiments of the present invention;

[0044] Figure 4 This is a schematic diagram illustrating the functional elements of an encoder according to an embodiment of the present invention;

[0045] Figure 5 This is a schematic diagram illustrating the functional elements of a decoder according to an embodiment of the present invention;

[0046] Figure 6 Shows the block positioned relative to the current block, which includes adjacent blocks;

[0047] Figure 7 (a) and Figure 7 (b) illustrates the radial (sub-block) pattern;

[0048] Figure 8 Examples of candidates for temporal merging of sub-blocks;

[0049] Figure 9 This example illustrates a temporal random access GOP structure with associated temporal IDs and POCs for 33 frames.

[0050] Figure 10 Examples of six possible splitting patterns for VVC;

[0051] Figure 11 Examples of MaxBTSize and MaxMttDepth;

[0052] Figure 12 Example of the MinQTSize variable;

[0053] Figure 13 Examples of possible partition constraints;

[0054] Figure 14 An example of an incomplete CTU within the frame border;

[0055] Figure 15 Example of setting the encoding of MaxMttDepth based on time domain ID;

[0056] Figure 16 An embodiment of the present invention is illustrated;

[0057] Figure 17 An example illustrating one embodiment of the present invention;

[0058] Figure 18 An embodiment of the present invention is illustrated;

[0059] Figure 19 An embodiment of the present invention is illustrated;

[0060] Figure 20 This is a diagram illustrating a system including an encoder or decoder and a communication network according to an embodiment of the present invention;

[0061] Figure 21 This is a schematic block diagram of a computing device for implementing one or more embodiments of the present invention;

[0062] Figure 22 This is a diagram illustrating a webcam system;

[0063] Figure 23 This is a diagram illustrating a smartphone. Detailed Implementation

[0064] Figure 1 This relates to the coding structures used in the High Efficiency Video Coding (HEVC) and Variety Video Coding (VVC) standards. Video sequence 1 consists of a series of digital images i. Each such digital image is represented by one or more matrices. The matrix coefficients represent pixels.

[0065] The sequence of images 2 can be segmented into strips 3. In some cases, a strip can constitute the entire image. These strips are segmented into non-overlapping coding tree units (CTUs). A coding tree unit (CTU) is the basic processing unit in the High Efficiency Video Coding (HEVC) and Multi-Functional Video Coding (VVC) video standards, and conceptually corresponds structurally to the macroblock unit used in several previous video standards. A CTU is sometimes also called a maximum coding unit (LCU). A CTU has luma and chroma component parts, each component part being called a coding tree block (CTB). These different color components are not... Figure 1 As shown in the image.

[0066] For HEVC, the CTU is typically 64 pixels × 64 pixels, while for VVC, the size can be 128 pixels × 128 pixels. Quadtree (QT) decomposition can be used to iteratively divide each CTU into smaller variable-size coding units (CUs).

[0067] The coding unit is the basic coding element and consists of two seed units called the prediction unit (PU) and the transform unit (TU). The maximum size of the PU or TU is equal to the size of the CU. The prediction unit corresponds to the partition of the CU used for predicting pixel values. It is possible to partition the CU into various different partitions of the PU, as shown in Figure 6, including partitions divided into 4 square PUs and two different partitions divided into 2 rectangular PUs. The transform unit is the basic unit for spatial transformation using the discrete cosine transform (DCT). The CU can be based on a quadtree representation divided into 7 partitions of the TU.

[0068] Each stripe is embedded in a Network Abstraction Layer (NAL) unit. Additionally, the encoding parameters of the video sequence are stored in dedicated NAL units called parameter sets. In HEVC and H.264 / AVC, two types of parameter set NAL units are used: First, the Sequence Parameter Set (SPS) NAL unit, which collects all parameters that remain unchanged throughout the entire video sequence. Typically, it handles the encoding profile, video frame size, and other parameters. Second, the Picture Parameter Set (PPS) NAL unit, which includes parameters that can be changed from one picture (or frame) in the sequence to another picture (or frame). HEVC also includes the Video Parameter Set (VPS) NAL unit, which contains parameters describing the overall structure of the bitstream. VPS is a type of parameter set defined in HEVC and applied to all layers of the bitstream. A layer can contain multiple temporal sublayers, and all version 1 bitstreams are confined to a single layer. HEVC has certain layer extensions for scalability and multi-view functionality, and these extensions will allow multiple layers with a backward-compatible version 1 base layer.

[0069] In VVC that includes sub-images, other ways of splitting images have been introduced, where a sub-image is an independent coding group of one or more stripes.

[0070] Figure 2An example is provided of a data communication system that can implement one or more embodiments of the present invention. The data communication system includes a transmission device (in this case, server 201) operable to transmit data packets of a data stream to a receiving device (in this case, client terminal 202) via a data communication network 200. The data communication network 200 can be a wide area network (WAN) or a local area network (LAN). Such a network can be, for example, a wireless network (Wifi / 802.11a or b or g), an Ethernet network, an Internet network, or a hybrid network consisting of several different networks. In a particular embodiment of the invention, the data communication system can be a digital television broadcasting system, wherein server 201 sends the same data content to multiple clients.

[0071] The data stream 204 provided by server 201 may consist of multimedia data representing video and audio data. In some embodiments of the invention, the audio and video data streams may be captured by server 201 using a microphone and a camera, respectively. In some embodiments, the data stream may be stored on server 201 or received by server 201 from other data providers, or generated at server 201. Server 201 is provided with encoders for encoding the video and audio streams, particularly for providing compressed bitstreams for transmission, which are more compact representations of the data presented as input to the encoder.

[0072] To achieve a better ratio of data quality to data volume, video data can be compressed, for example, according to HEVC, H.264 / AVC, VVC, or the format of data generated by ECM.

[0073] Client 202 receives the transmitted bit stream and decodes the reconstructed bit stream to reproduce video images on a display device and reproduce audio data using a speaker.

[0074] Despite Figure 2 The examples consider streaming scenarios, but it will be appreciated that in some embodiments of the invention, media storage devices such as optical discs can be used for data communication between the encoder and decoder.

[0075] In one or more embodiments of the invention, video images are transmitted together with data representing compensation offsets of reconstructed pixels to be applied to the image, so as to provide filtered pixels in the final image.

[0076] Figure 3 A processing device 300 configured to implement at least one embodiment of the present invention is illustrated schematically. The processing device 300 may be a device such as a microcomputer, workstation, or lightweight portable device. The device 300 includes a communication bus 313 connected to:

[0077] - This refers to the central processing unit 311 of the CPU, such as a microprocessor;

[0078] - Read-only memory 306, denoted as ROM, is used to store computer programs that implement the present invention;

[0079] - A random access memory 312, represented as RAM, for storing executable code of the methods according to embodiments of the present invention, and registers suitable for recording variables and parameters required to implement the methods for encoding digital image sequences and / or decoding bit streams according to embodiments of the present invention; and

[0080] - A communication interface 302 connected to the communication network 303, through which digital data to be processed is transmitted or received.

[0081] Optionally, the device 300 may also include the following components:

[0082] - A data storage component 304, such as a hard disk, is used to store a computer program for implementing one or more embodiments of the present invention, as well as data used or generated during the implementation of one or more embodiments of the present invention;

[0083] - A disk drive 305 for disk 306, the disk drive being adapted to read data from disk 306 or write data to said disk;

[0084] - Screen 309 is used to display data and / or serve as a graphical interface for user interaction via keyboard 310 or any other indicating device.

[0085] Device 300 can be connected to various peripheral devices such as digital camera 320 or microphone 308, each of which is connected to an input / output card (not shown) to provide multimedia data to device 300.

[0086] The communication bus provides communication and interoperability between various elements included in or connected to device 300. The representation of the bus is not limiting, and in particular, the central processing unit is operable to communicate instructions directly or by means of other elements of device 300 to any element of device 300.

[0087] Disk 306 may be replaced by any information medium such as a rewritable or non-rewritable compact disc (CD-ROM), ZIP disc, or memory card, and generally by an information storage component that can be read by a microcomputer or microprocessor. Disk 306 may be integrated into or not integrated into the device, may be portable, and is adapted to store one or more programs that execute to enable the implementation of the method for encoding digital image sequences and / or the method for decoding bit streams according to the present invention.

[0088] Executable code can be stored in read-only memory 306, hard disk 304, or removable digital media (such as, for example, disk 306 as described above). According to a variation, the executable code of the program can be received via interface 302 through communication network 303 to be stored in one of the storage components of device 300 (such as hard disk 304) before execution.

[0089] The central processing unit 311 is adapted to control and direct the execution of instructions or software code portions of one or more programs according to the invention, stored in one of the aforementioned storage components. Upon power-up, one or more programs stored in non-volatile memory (e.g., on hard disk 304 or in read-only memory 306) are transferred to random access memory 312 (which then contains executable code for one or more programs) and registers for storing variables and parameters necessary for implementing the invention.

[0090] In this embodiment, the device is a programmable device that implements the invention using software. However, alternatively, the invention can be implemented in hardware (e.g., in the form of an application-specific integrated circuit or ASIC).

[0091] Figure 4 A block diagram illustrating an encoder according to at least one embodiment of the present invention is shown. The encoder is represented by connected modules, each module being adapted to implement, for example, in the form of programming instructions executed by the CPU 311 of the device 300, at least one corresponding step of a method for encoding images in an image sequence according to one or more embodiments of the present invention.

[0092] The encoder 400 receives the raw sequence 401 of digital images i0 to in as input. Each digital image is represented by a set of samples (sometimes also called pixels) (hereinafter referred to as pixels).

[0093] After performing the encoding process, the encoder 400 outputs a bit stream 410. The bit stream 410 includes multiple encoding units or stripes, each strip including a strip header for transmitting the encoded values ​​of the encoding parameters used for strip encoding, and a strip body including encoded video data.

[0094] Module 402 will input digital images i0 to i nThe image is divided into 401 pixel blocks. These blocks correspond to portions of the image and can have variable sizes (e.g., 4×4, 8×8, 16×16, 32×32, 64×64, 128×128 pixels, and several rectangular block sizes can also be considered). An encoding mode is selected for each input block. Two families of encoding modes are provided: a spatial prediction-based encoding mode (intra-frame prediction) and a temporal prediction-based encoding mode (inter-frame coding, merging, skipping). Possible encoding modes were tested.

[0095] Module 403 implements intra-frame prediction processing, wherein the block to be encoded is predicted by a predictor calculated based on the neighboring pixels of the given block. If intra-frame coding is selected, the selected intra-frame predictor and an indication of the difference between the given block and its predictor are encoded to provide a residual.

[0096] Temporal prediction is implemented by motion estimation module 404 and motion compensation module 405. First, a reference image from reference image set 416 is selected, and motion estimation module 404 selects a portion of the reference image (also referred to as a reference region or image portion) that is the closest region relative to the given block to be encoded (most similar in pixel values). Motion compensation module 405 then uses the selected region to predict the block to be encoded. The difference between the selected reference region and the given block (also referred to as the residual block) is calculated by motion compensation module 405. The selected reference region is indicated using motion vectors.

[0097] Therefore, in both cases (spatial and temporal predictions), the residuals are calculated by subtracting the predictor from the original block.

[0098] In intra-frame prediction implemented by module 403, the prediction direction is encoded. In inter-frame prediction implemented by modules 404, 405, 416, 418, and 417, at least one motion vector or data used to identify such a motion vector is encoded for temporal prediction.

[0099] If inter-frame prediction is selected, information related to motion vectors and residual blocks is encoded. To further reduce the bit rate, motion is assumed to be homogeneous, and motion vectors are encoded by the difference relative to the motion vector predictor. The motion vector predictor is obtained from the set of motion information predictor candidates by the motion vector prediction and encoding module 417 from the motion vector field 418.

[0100] The encoder 400 also includes a selection module 406, which selects the encoding mode by applying encoding cost criteria (such as rate-distortion criteria). To further reduce redundancy, a transform (such as DCT) is applied to the residual block by the transform module 407. The resulting transform data is then quantized by the quantization module 408 and entropy encoded by the entropy encoding module 409. Finally, the encoded residual block of the current block being encoded is inserted into the bit stream 410.

[0101] Encoder 400 also decodes the encoded image to produce a reference image for motion estimation of subsequent images (e.g., a reference image in reference image / picture 416). This allows the encoder and decoder receiving the bitstream to have the same reference frame (using the reconstructed image or a portion of the image). Inverse quantization (dequantization) module 411 performs inverse quantization (dequantization) of the quantized data, followed by an inverse transform by inverse transform module 412. Intra-frame prediction module 413 uses the prediction information to determine which predictor to use for a given block, and motion compensation module 414 actually adds the residual obtained by module 412 to the reference region obtained from reference image set 416.

[0102] Then, module 415 applies post-filtering to filter the reconstructed pixel frames (image or image portion). In embodiments of the invention, a SAO loop filter is used, wherein a compensation offset is added to the pixel values ​​of the reconstructed pixels in the reconstructed image. It should be understood that post-filtering is not always necessary. Furthermore, any other type of post-filtering may be performed instead of SAO loop filtering or in addition to SAO loop filtering.

[0103] Figure 5 A block diagram illustrating a decoder 60 according to an embodiment of the present invention is shown. The decoder 60 can be used to receive data from an encoder. The decoder is represented by connected modules, each module being adapted to implement corresponding steps of the method implemented by the decoder 60, for example, in the form of programming instructions to be executed by the CPU 311 of the device 300.

[0104] Decoder 60 receives bitstream 61 including encoding units (e.g., data corresponding to blocks or decoding units), each encoding unit consisting of a header containing information related to encoded parameters and a body containing encoded video data. (See also: Regarding...) Figure 4 As described, for a given block, entropy encoding is performed on the encoded video data at a predetermined number of bits, and the index of the motion vector predictor is also encoded. The received encoded video data is entropy decoded by module 62. The residual data is then dequantized by module 63, and subsequently, an inverse transform is applied by module 64 to obtain the pixel values.

[0105] The pattern data used to indicate the encoding mode is also entropy decoded, and based on this pattern, intra-frame type decoding or inter-frame type decoding is performed on the encoded blocks (units / sets / groups) of image data.

[0106] In intra-frame mode, intra-frame prediction module 65 determines the intra-frame predictor based on the intra-frame prediction mode specified in the bit stream.

[0107] If the mode is inter-frame, motion prediction information is extracted from the bitstream to locate (identify) the reference region used by the encoder. The motion prediction information includes the reference frame index and motion vector residuals. The motion vector predictor is added to the motion vector residuals by the motion vector decoding module 70 to obtain the motion vectors. See below for further details. Figures 6 to 10 This section discusses in more detail the various motion predictor tools used in VVC.

[0108] Motion vector decoding module 70 applies motion vector decoding to each current block encoded via motion prediction. Once the index of the motion vector predictor for the current block is obtained, the actual values ​​of the motion vectors associated with the current block can be decoded, and these actual values ​​are used to apply motion compensation via module 66. A portion of the reference image indicated by the decoded motion vectors is extracted from the reference image 68 to apply motion compensation 66. The motion vector field data 71 is updated using the decoded motion vectors for subsequent motion vector prediction. Note that in VVC, as in HEVC, motion vectors are stored at a 16×16 level instead of a 4×4 level for the temporal predictor. This means that decimation is applied only to the temporal predictor and not to the spatial predictor. In practice, the aim is to reduce the buffer required to store temporal motion vectors after the encoding of each frame. This negatively impacts coding efficiency, and this decimation is typically removed from exploration software.

[0109] Finally, the decoded block is obtained. Post-filtering is applied by post-filtering module 67 where appropriate. Decoder 60 ultimately obtains and provides the decoded video signal 69.

[0110] VVC merging mode

[0111] Compared to HEVC, several inter-frame modes have been added to VVC. In particular, a new merging mode has been added to HEVC's regular merging mode.

[0112] Affine mode (sub-block mode)

[0113] In HEVC, only the translational motion model is applied to motion compensation prediction (MCP). However, in the real world, there are various types of motion, such as zooming in / out, rotation, perspective motion, and other irregular motions.

[0114] In JEM, a simplified affine transformation motion compensation prediction is applied, and the general principles of affine modes are described below based on an excerpt from document JVET-G1001, submitted to the JVET conference in Turin from July 13 to 21, 2017. The entire document is incorporated herein by reference because it describes other algorithms used in JEM.

[0115] like Figure 7 As shown in (a), the affine motion field of the block is described by two control point motion vectors.

[0116] Affine mode is a motion compensation mode similar to inter-frame modes (AMVP, "classic" merge, or "classic" merge skip). Its principle is to generate motion information per pixel based on 2 or 3 adjacent motion information. In JEM, as... Figure 7 As depicted in (a), the affine pattern derives motion information for each 4×4 block (each square is a 4×4 block, and...). Figure 7 The entire block in (a) is a 16×16 block of 16 such squares of size 4×4 (each 4×4 square block has its associated motion vector)). By enabling affine mode using a flag, affine mode can be used in AMVP mode and merge mode (i.e., classic merge mode (also known as "non-affine merge mode") and classic merge skip mode (also known as "non-affine merge skip mode")).

[0117] In the VVC specification, the affine pattern is also called the subblock pattern; these terms are used interchangeably in this specification.

[0118] like Figure 8 The depicted VVC sub-block merging pattern includes temporal merging candidates based on sub-blocks, which inherit the motion vector field of blocks in previous frames pointed to by spatial motion vector candidate A1. In this figure, the current predictor is not a side-by-side block, but a block with the motion vector value of A1 shifted.

[0119] If adjacent blocks have been encoded using inter-frame affine patterns with sub-block merging, then the sub-block candidate is followed by inherited affine motion candidates, which are then derived as part of the constructed affine candidates before some zero Mv candidate.

[0120] Context index increment

[0121] Context-based adaptive binary arithmetic coding (CABAC) uses context to separate the probabilities of one or more intervals. To obtain the corresponding context of an interval and its relative state, the context index increment ctxInc is computed.

[0122] For example, the following formula gives an example of context index increment.

[0123] ctxInc=(condL&&availableL)||(condA&&availableA)

[0124] Where: condL is the value of the associated left-hand syntactic element, and condA is the value of the associated top-hand syntactic element. availableL and availableA are the availability values ​​of the left-hand and top-hand blocks, respectively.

[0125] Random access configuration

[0126] Figure 9 The diagram illustrates the temporal random access GOP structure for 33 consecutive frames 0 to 32. The length of the vertical line representing each frame corresponds to its temporal ID. (For example, the longest length corresponds to temporal ID 0, and the shortest length corresponds to temporal ID 5). Frames with temporal ID 0 are at the highest level in the temporal hierarchy because they can be decoded independently of all other frames in the lower temporal hierarchy (i.e., all other frames with numerically higher temporal ID values). Similarly, frames with temporal ID 1 are second in the temporal hierarchy, and they can be decoded independently of all other frames in the lower temporal hierarchy (i.e., all other frames with higher temporal IDs), and so on for the other temporal IDs. In other words, a frame with a particular temporal ID can be decoded independently of frames with higher temporal ID values, but can depend on frames with lower temporal IDs. This is known as temporal scalability.

[0127] This parameter is similar to layer depth, but layer depth does not imply independence in decoding all other frames with higher depths.

[0128] VVC partition

[0129] VVC partitions have specific block partitions. For a tree node, such as Figure 10 The six possible splits described are possible:

[0130] - Quad-splitter QT 801, which divides a block into four equal-sized square blocks;

[0131] -Binary splitting of BT, with its two possible subdivisions 802 and 803:

[0132] - Vertical binary split 802, Split_BT_VER

[0133] - Horizontal binary split 803, Split_BT_HOR

[0134] - Tri-branch split TT, with its two possible subdivisions 804 and 805, where the block is split into 3 blocks with a larger band in the middle:

[0135] - Vertical ternary split 804, Split_TT_VER

[0136] - Horizontal ternary split 805, SPLIT_TT_HOR

[0137] - No split 806, its terminating tree node, therefore there is no split.

[0138] In this specification, a block can be a CTU and / or a CU or more generally any unit in the coding tree.

[0139] VVC split control variables

[0140] For the current block, all possible splits are always permitted. Which splits are available depends on several conditions. These conditions depend on several defined split control variables. The first set of variables defines the maximum and minimum block / node sizes:

[0141] -CTU size: It corresponds to the size of the root node of the quadtree (e.g., 256×256, 128×128, 64×64, 32×32, 16×16 luminance samples);

[0142] -maxBtSize: This is the maximum allowed size of the root node of the binary tree, that is, the maximum size of the leaf quadtree node that can be partitioned by binary splitting. If the height and width of the current block are both less than or equal to maxBtSize, then the current block can be split due to BT splitting. Figure 11 Example of the concept of maxBtSize, where maxBtSize is the size of the quadleaf leaf node 902 of CTU 901.

[0143] `-maxBtSize` is the minimum allowed size of a binary leaf node; that is, the minimum width or height of a binary leaf node. Therefore, if the height of the current block is greater than `minBtSize`, the current block can be split due to horizontal BT splitting. And if the width of the current block is greater than `minBtSize`, the current block can be split due to vertical BT splitting.

[0144] -maxTtSize: This is the maximum allowed size of the root node of the ternary tree, that is, the maximum size of the leaf quadtree node that can be partitioned by ternary splitting. If the height and width of the current block are both less than or equal to maxTtSize, then the current block can be split due to TT splitting.

[0145] -minTtSize: Represents the minimum allowed size of a leaf node in a ternary tree (TT); that is, the minimum width or height of a leaf node in a binary tree. However, instead of allowing BT splits, the minimum TT partition size is considered. Therefore, if the height of the current block is greater than twice minTtSize, the current block can be split due to horizontal TT splits. And if the width of the current block is strictly greater than twice minTtSize, the current block can be split due to vertical TT splits.

[0146] -minQtSize: is the minimum allowed size of a leaf node in a quadtree (QT); therefore, for the current block, QT splitting mode is not allowed if the width of the current block is not greater than minQtSize. Figure 12 An example illustrating minQtSize. By considering CTU128, in the illustrated example, minQtsize equals 16.

[0147] There is no definition for maxQtSize, therefore it corresponds to the CTU size.

[0148] The minimum allowed block size for both width and height is 4.

[0149] It also defines depth sets.

[0150] - Depth: This is the depth within the tree. In the VVC specification, a leaf is the terminating node of the tree, which is the root node of a tree at depth 0. This means that for each split, this value is incremented (by 1).

[0151] -mttDepth: This is the depth of the multi-tree. Multi-trees include BT splitting and TT splitting.

[0152] `maxMttDepth` is defined in the VVC specification as the maximum allowed multi-tree depth. Therefore, `mttDepth` is greater than or equal to `maxMttDepth`. Figure 11 Example of the concept of maxMttDepth.

[0153] In VVC, these variables are defined independently for luminance and chrominance.

[0154] In VTM and ECM software, there are several other variables corresponding to depth.

[0155] The variable `currBtDepth` is the current number of BT splits used to reach the current tree node (or current block). The variable `currMttDepth` is the current number of both BT and TT splits used to reach the current tree node (or current block). The variable `maxBtDepth` corresponds to the variable `maxMttDepth` in the VVC specification. `currQtDepth` is the current number of QT splits used to reach the current tree node (or current block). `maxBtDepth`: is the maximum allowed binary tree depth, i.e., the lowest level at which binary splits can occur, where the leaf node of a quadtree is the root (e.g., 3).

[0156] VVC splits control syntax elements

[0157] To set the values ​​of these different variables, as depicted in the table below for SPS syntax elements, some high-level syntax elements are transmitted in SPS.

[0158]

[0159]

[0160] When the sps_partition_constraints_override_enabled_flag is enabled in SPS, as depicted in the following table of PH syntax elements, some image header syntax elements are transferred to update the partition variables.

[0161]

[0162]

[0163]

[0164] VVC encoding splitting mode

[0165] In VVC, as depicted in the following syntax table, the encoding split pattern is transmitted in the coding_tree, where the conditional resolution flags split_cu_flag, split_qt_flag, mtt_split_cu_vertical_flag, and mtt_split_cu_binary_flag define the split of the CU.

[0166]

[0167]

[0168]

[0169] VVC splitting limitations

[0170] VVC partitioning has several limitations. These limitations are primarily designed to prevent the creation of identical partitions after several consecutive splits. Figure 13 Here are some examples of these constraints. The idea is to avoid using the same partitions for BT and TT. For example... Figure 13 As depicted in (a), two consecutive vertical BT splits are allowed, but as Figure 13 As depicted in (b), vertical BT splitting after vertical TT in the center block is not allowed.

[0171] In the same way, such as Figure 13 As described in (c), two consecutive horizontal BT splits are allowed, but as Figure 13 As depicted in (d), horizontal BT splitting after horizontal TT in the center block is not allowed.

[0172] In VVC, there are additional constraints regarding the minimum chroma block size for inter-frame block size and the maximum block sizes for TT and BT. These constraints have been eliminated for ECM software.

[0173] Color partitioning

[0174] In VVC, chroma partitioning can be inferred from luma partitioning, but this can be disabled. For example, in dual-tree mode, the chroma partitioning tree is independent of the luma tree. However, some limitations exist.

[0175] The tree can also depend in part on the luminance partition of the CCLM mode, otherwise it is independent.

[0176] Image boundary

[0177] Frame resolution is not always an integer multiple of CTU size. Therefore, as Figure 14 The depicted image shows that incomplete CTUs may exist within the frame borders, with CTUs 1201 to 1206 being incomplete due to the bottom boundary 1207 and right boundary 1208 of the frame. In VVC, signaling for splitting at image boundaries is allowed, compared to previous standards. Split processing at boundaries is applied until the coded tree node representation is entirely within the image's CU. However, some splits are inferred (not transmitted). Therefore, different variables such as maxMttDepth, minQtDepth, and minQtsize increase or decrease and may differ from the variables used for splits not at boundaries.

[0178] QT BT TT encoding selection

[0179] In the VTM and ECM software, several encoder-side optimizations are used for QT BT TT encoding selection.

[0180] One such optimization involves determining whether QT splitting is tested before BT splitting.

[0181] The conditions are: at least one CU to the left or above the current coding tree node has a QT depth greater than the QT depth of the current coding tree node; and whether the width of the CU represented by the current coding tree node is greater than minQtSize*2.

[0182] If this condition is true, QT will be tested before BT, and the split will be considered in the following order.

[0183] - No splitting

[0184] -QT

[0185] -BT level

[0186] -BT Vertical

[0187] -TT level

[0188] -TT Vertical

[0189] Otherwise, the order would be as follows.

[0190] - No splitting

[0191] -BT level

[0192] -BT Vertical

[0193] -TT level

[0194] -TT Vertical

[0195] -QT

[0196] This order is important because, based on some optimizations, depending on the results of the first tested pattern, several splits will not be tested. Therefore, when Qt is tested last, there are many cases that will not be evaluated.

[0197] maxMttDepth

[0198] The maximum MTT depth has a significant impact on encoder complexity. For example... Figure 15 The common test conditions for ECM described have been updated to reduce encoding by setting different maxMttDepth values. In this setting, maxMttDepth is lower for a given time-domain ID at large resolutions or small QP settings.

[0199] Adaptive maxBtSize

[0200] In VTM and ECM, there is a frame-level coding selection that sets maxBtSize based on the average block size of previously encoded frames with the same depth (in the case of CTC RA => the same temporal ID). The average block size is compared to a threshold, as shown in the following pseudocode.

[0201] if (dBlkSize < AMAXBT_TH32)

[0202] {

[0203] newMaxBtSize = 32;

[0204] }

[0205] else if (dBlkSize) <AMAXBT_TH64)

[0206] {

[0207] newMaxBtSize = 64;

[0208] }

[0209] else if(dBlkSize<AMAXBT_TH128)

[0210] {

[0211] newMaxBtSize = 128;

[0212] }

[0213] else

[0214] f

[0215] newMaxBtSize = 256;

[0216] }

[0217] Where AMAXBT_TH32 equals 15, AMAXBT_TH64 equals 30, and AMAXBT_TH128 equals 60. This method decreases the maximum BT size when the average block size is small and increases the maximum BT size when it is large.

[0218] Time-domain prediction of CABAC

[0219] In ECM, temporal CABAC prediction (JVET-Y0181) is used. In this method, previous slices are used for CABAC initialization of the current frame. After encoding the CTU up to a specified position, the probability states of each context model are first obtained and stored. Then, the stored probability states are used as the initial probability states of the corresponding context models in the next B-slice or P-slice encoded with the same quantization parameter (QP) or the same corresponding temporal ID.

[0220] The problem solved by the present invention

[0221] In traditional video coding, temporal redundancy is utilized between samples due to inter-frame patterns, and between motion information due to different temporal candidates or predictors. In ECM, temporal redundancy between CABAC probabilities is also used to improve coding efficiency. However, other data parameters, variables, and syntactic elements should also have temporal relevance. Traditional methods utilizing this redundancy seem useless or impossible. For example, solutions that use a large number of predictors for samples, such as different inter-frame patterns, are unsuitable for data with a small number of coding possibilities. Several candidate signalings for motion information are useful because motion is information from the real world and it is very specific. But this is useless for predicting parameters, variables, or syntactic elements. Similarly, predicting CABAC states is unsuitable because it looks more like frame-based QP prediction and is not appropriate for different content within a video sequence.

[0222] This invention proposes a method for efficiently determining the values ​​used to predict these data by utilizing the temporal correlation between them.

[0223] Example

[0224] Main Implementation Examples

[0225] Emb.Main uses temporal regions to derive at least one value to predict, infer, or determine the contextual increment of a syntactic element, syntactic element, or variable.

[0226] In embodiments, a temporal region is used to derive at least one value to predict, infer, constrain, or determine a variable or to encode or compute the contextual increment of a syntactic element. This value represents a similar variable or syntactic element, or another related variable or syntactic element. The temporal region is derived from a temporal frame that has already been encoded on the encoder side or decoded on the decoder side. In one example, Figure 16 This embodiment is illustrated. In this figure, the time domain region (1601) comprises several blocks and portions of blocks within the boundaries of the time domain region. These blocks are considered to derive at least one value.

[0227] The advantage of this embodiment is improved coding efficiency due to a better derivation of the predictor, constraint, or context increment. Compared to methods used in the prior art, this method is more suitable for syntactic elements or variables with a reduced number of values.

[0228] Frames with the same temporal ID as Emb.Tempo_from1

[0229] In this embodiment, the time domain region comes from frames with the same time domain ID.

[0230] For example Figure 9 The example of the random access configuration shown uses another encoded / decoded frame with the same time domain ID 4 to determine the value associated with the time domain region if the current frame has a time domain ID equal to 4.

[0231] Frames with the same temporal ID typically have the same coding parameters, especially since they share the same or similar QP and spatial distance with their reference frames. Therefore, these frames are of great interest for predicting QT depth, as this data is related to both the QP and the spatial distance between frames.

[0232] The closest frame with the same temporal ID as Emb.Tempo_from1.1

[0233] In this embodiment, the time domain region comes from the nearest frame with the same time domain ID.

[0234] For an example of a random access configuration, such as Figure 9 As shown, the closest frame with the same temporal ID is (typically) more relevant than other frames. Therefore, the results are better.

[0235] Emb.Tempo_from2 has a frame or reference frame with the same QP.

[0236] In this embodiment, the time-domain region comes from a frame or a reference frame that has the same QP. Ideally, the reference frame has the same QP.

[0237] As mentioned above, QP has a significant impact on block partitioning, and many syntactic elements and variables have similar values. Therefore, for frames with the same QP, values ​​determined from the temporal domain are better.

[0238] The Emb.Tempo_from3 reference frame is the same as the reference frame used for temporal motion vector prediction.

[0239] In this embodiment, the temporal region is derived from a reference frame used for temporal motion vector prediction. This reference frame can be a first reference frame from reference list 0 or a first reference frame from list 1, depending on the flags transmitted in the image header or strip header.

[0240] Surprisingly, this embodiment achieves optimal coding efficiency even with a lower QP for the reference frame. However, the reference frame is closer to the current frame than all frames with the same temporal ID.

[0241] Emb.Tempo_from4 is the closest reference frame.

[0242] In this embodiment, the time domain region is derived from the closest reference frame.

[0243] As illustrated with respect to the previous embodiments, even though frames with the same QP have statistically greater correlation between their QP depths, the distance to the current frame seems to be of greater interest in the trade-off between encoder time reduction and coding efficiency.

[0244] Emb.Tempo_from5 has more than one reference frame.

[0245] In this embodiment, two time-domain frames are considered, and therefore two time-domain regions are used to determine a value. More than two reference frames may also be considered.

[0246] The advantage is better coding efficiency, but it increases memory access volume.

[0247] Emb.Decim HLS transmission of size N or M in the header

[0248] In an embodiment, the value N of the grid used for extracting blocks of the time-domain region and / or the value M of the grid used for extracting possible positions of the blocks are transmitted in the header. Additionally or alternatively, other parameters such as irregular grids may also be transmitted. Ideally, this value is transmitted in the sequence parameter set (SPS) because it affects the memory buffer. However, if the values ​​between N and M maintain the same required memory size, these values ​​can alternatively be transmitted in the PPS, image header, or strip header.

[0249] size

[0250] The size of Emb.Larger is greater than or equal to the size of the current block.

[0251] In this embodiment, the size of the time-domain region is larger than the current block, if possible. The purpose of this embodiment is to determine a value that is more useful than the value that can be obtained using parallel blocks.

[0252] One advantage is that it provides better values ​​for variables that need to be predicted, inferred, determined, or derived for contextual increments, because the determined values ​​are based on more values ​​than a single side-by-side block, which is, of course, suitable for some data rather than motion information. This advantage results in improved coding efficiency.

[0253] The second advantage is that, compared to solutions that shift blocks based on motion information (as temporal sub-blocks), it is not necessary to determine the motion information. Therefore, the resolution does not depend on motion information that is unavailable without complete decoding.

[0254] Another advantage is that, compared to a single block from only one block, several values ​​can be considered and other values ​​can be derived as minimum, maximum average, etc., which provides more information to constrain, predict, or derive contextual increments.

[0255] The size of the Emb.Larger2 time-domain region is greater than or equal to the maximum possible block size.

[0256] In this embodiment, the time-domain region is greater than or equal to the maximum possible block size. For example, the time-domain region is equal to or greater than the CTU size.

[0257] This presents an interesting trade-off between coding efficiency and the complexity of deterministic values.

[0258] Emb.AdaptSize adapts the size based on variables / parameters.

[0259] In the embodiments, the size of the time domain region is adapted according to at least one variable or at least one parameter.

[0260] The advantage of this embodiment is improved coding efficiency, especially when used in the temporal domain for motion compensation between two frames.

[0261] Emb.AdaptSize1 temporal distance

[0262] In this embodiment, the size of the temporal region is determined based on the temporal distance between the current frame and the frame containing the temporal region. In this embodiment, the temporal region increases as the temporal distance increases. For example, the absolute difference between the picture sequence count (POC) "currPOC" of the current frame and the picture sequence count "tempoPOC" of the frame containing the temporal region is calculated to account for the temporal distance between frames. For example, the size of the temporal region (widthTempo, heightTempo) can be determined according to the following pseudocode.

[0263] widthTempo=widthTempoFix+8*abs(currPOC-tempoPOC)

[0264] heightTempo=heightTempoFix+8*abs(currPOC-tempoPOC)

[0265] Where: abs() is a function that returns the absolute value. widthTempoFix and heightTempoFix are predetermined. For example, they are set to be equal to the CTU size. The number "8" in the formula is an example, but other values ​​can be considered.

[0266] Additionally, the frame rate of the sequence can be considered to apply weights to the absolute difference between POCs.

[0267] The advantage is that the temporal domain can compensate for motion between two frames and maintain the temporal correlation between variables or syntactic elements to be predicted, etc.

[0268] Emb.AdaptSize2 QP difference

[0269] In this embodiment, the size of the temporal region is determined based on the quantization parameter (QP) of the current frame and the QP of the frame containing the temporal region. Furthermore, the size is determined based on the QP difference between the two frames. For example, the size of the temporal region increases when the QP of the current frame is lower than the QP of the temporal frame, and conversely, decreases when the QP of the current frame is higher than the QP of the temporal frame. Additionally, the size of the temporal region can be proportional.

[0270] The advantage is improved coding efficiency. In fact, the block size within a frame is related to the QP (Quickness of Quantity). Specifically, for the same encoded frame, a higher QP results in a larger block size. Therefore, for a temporal frame, a higher QP leads to a larger temporal region, increasing the likelihood of finding the correct value.

[0271] Emb.AdaptSize3 temporal ID, depth

[0272] In this embodiment, the size of the time-domain region is determined based on the time-domain ID of the current frame and the time-domain ID of the time-domain frame containing the time-domain region. Furthermore, this size can be proportional to the difference between the two time-domain IDs. For example, the size of the time-domain region increases when the time-domain ID of the time-domain frame is less than the time-domain ID of the current frame.

[0273] Alternatively, instead of time-domain IDs, hierarchy depth can be considered.

[0274] The advantage is improved coding efficiency. In practice, frames with hourly domain IDs are typically encoded using a large temporal distance between frames. For this situation, it's best to increase the temporal domain area.

[0275] The value transmitted in the header of Emb.AdaptSize4

[0276] In this embodiment, the size of the temporal region is determined based on a value transmitted in the header. This value may be transmitted alternatively or additionally in the SPS, PPS, image header, or strip header. For example, the value may be transmitted in the SPS and predicted or inferred in the PPS. An overlay flag indicates whether the value has been updated for the image header compared to the value in the PPS or SPS.

[0277] The advantage is that the encoder implementation is unrestricted and can be adapted to choose between coding efficiency and complexity.

[0278] Emb.AdaptSize5 is the size of the current block.

[0279] In this embodiment, the size of the time-domain region is determined based on the current block size. Therefore, a larger time-domain region corresponds to a larger block size, and a smaller time-domain region corresponds to a smaller block size. For example, if a minimum time-domain region size corresponding to the smallest possible block size is considered, the current block size is added to that minimum size to obtain the final time-domain region size corresponding to the current block.

[0280] This is more suitable for multiple block sizes.

[0281] Blocks within the time domain

[0282] Emb.All_Blocks: All blocks in the time-domain region.

[0283] In this embodiment, the block considered to determine the value is all blocks in the time-domain region. Figure 16 In the example, the border of the time-domain region (1601) is not aligned with the split partition. This corresponds to all blocks (1602 to 1620) that have at least one sample within the time-domain region (1601). Therefore, 19 blocks are considered in this figure.

[0284] This is the simplest way to consider blocks within a time-domain region.

[0285] Emb.KeepProportion maintains the proportion of blocks within a time-domain region.

[0286] In this embodiment, the blocks considered to determine the values ​​are all blocks in the time-domain region, but the values ​​extracted from the blocks are weighted to consider only the portions of the blocks within the time-domain region. Figure 16 For example, for block 1617, the weights corresponding to part 1630 are determined to calculate, for example, the average of several values.

[0287] This embodiment is more complex than the previous one because it requires some additional computation, but it improves coding efficiency because it is more locally adapted to the current block.

[0288] Emb.FullyContain is a block entirely within the time domain.

[0289] In this embodiment, as an alternative to the previous embodiment, the block considered for determining the value is a block that is entirely within the time domain. Figure 16 In the example, this corresponds to all blocks 1601 to 1604 where the border of the time domain region (1601) is not aligned with the split partition. Therefore, four blocks are considered.

[0290] The advantage of this embodiment is that it reduces the complexity of determining the values ​​because fewer values ​​are required for calculation. However, this excludes the worst-case scenario of time-domain region matching and partitioning.

[0291] Position of the time domain region compared to the current block

[0292] Emb.Center is the center of the block location in the current frame.

[0293] In this embodiment, the center of the current block is the center of the temporal region in the temporal frame.

[0294] This center is usually the optimal representation of the block. Therefore, the encoding efficiency is better.

[0295] Alternatively, when the center of the block is outside the frame, the top-left position can be considered.

[0296] Emb.shiftedBasedMV can shift time-domain regions based on motion information.

[0297] Even though the temporal region can compensate for the use of motion information in the parsing process, in this embodiment, the temporal region is also shifted based on motion vectors (e.g., adjacent motion vectors). This is particularly efficient if there is a large amount of motion information or a large temporal distance between the current frame and temporal frames in adjacent regions.

[0298] Extraction of time domain regions

[0299] Emb.Decimation takes one position for every N.

[0300] In this embodiment, only blocks existing in the grid are considered to determine values ​​from the time domain region. Specifically, only blocks whose height and width are multiples of N are considered to determine the values.

[0301] Figure 17 An example of this embodiment is shown. In this figure, to determine the value from time domain region 1701, only blocks (1702 to 1710) in an N×N grid represented by points are used. Note that in this figure, the time domain region is aligned with the split partition for simplicity of description.

[0302] The main advantage is the reduction in the buffer required to store the values ​​used to determine values ​​from the time-domain region. In practice, for each frame that can be used as a time-domain region, all the values ​​needed to determine that value need to be held in memory. For hardware implementation, the buffer is designed considering the worst-case scenario. In this case, the worst-case scenario is the minimum block size. As the minimum size in the case of 8×4 or 4×8, the worst-case scenario for storing values ​​for each 4×4 block can be considered. Therefore, for a 1080p frame, (1920 / 4)*(1080 / 4) = 129600 related values ​​need to be stored for the time-domain frame. If, for example, we consider the value N to be equal to 16, then for the time-domain frame, only (1920 / 16)*(1080 / 16) = 8100 related values ​​need to be stored. Therefore, this reduces the amount of information that needs to be stored to 1 / 16 of the original.

[0303] Another advantage is the reduced complexity when the partition contains small blocks. In fact, it reduces the worst-case complexity because it only needs to consider blocks on the grid at most.

[0304] Furthermore, instead of small blocks corresponding to less frequent regions, greater emphasis is placed on larger blocks that serve as a better representation of what is happening in the temporal region. As a result, decimation is better. And, surprisingly, this decimation yields an improvement in coding efficiency. In particular, when using values ​​obtained from the temporal region to constrain or infer variables for QT, BT, and TT, decimation with N equal to 16 gives the best coding efficiency. If the temporal region is a CTU of 256×256 luminance samples, then in the worst case, only 256 blocks need to be considered. Conversely, without this decimation, since the minimum block size is 4×8 or 8×4, 2048 blocks need to be considered in the worst case.

[0305] Emb.SumPropor is proportional to the block size (especially for data representing partitions).

[0306] In one embodiment, for example, only blocks existing in the grid are considered to determine the value from the time-domain region, and the size of the blocks is taken into account to calculate the average value. Therefore, the value is obtained by taking into account the scale of the blocks considered. This also applies to blocks within the boundaries of the time-domain region as described in previous embodiments.

[0307] Compared to methods that do not apply this proportionality, the advantage is improved coding efficiency.

[0308] The position N of the Emb.SumCenter grid is centered.

[0309] In this embodiment, the grid under consideration is centered relative to the time domain region. If the top-left corner of the time domain region under consideration has a position (0,0), then the top-left position of the grid is (N / 2,N / 2). Figure 18Example (a) illustrates a grid that is not centered in the time domain, and Figure 18 Example (b) illustrates a grid centered in the time domain.

[0310] The advantage is improved encoding efficiency, because the preserved values ​​will (usually) be closer to the center of the current block.

[0311] Emb.SumNotReg irregular pattern

[0312] In this embodiment, the positions are presented in an irregular pattern. For example, more positions are considered at the center of the time domain region, and the corners of the time domain region are also considered. Figure 19 This is an example of an irregular pattern.

[0313] The advantage is that it can sometimes improve coding efficiency.

[0314] Extraction of the current block position

[0315] Emb.currentPosDecim extracts the possible positions of the block.

[0316] In this embodiment, possible locations of the current block are extracted. In this embodiment, all possible locations of the block in the current frame are impossible. Only the location on a grid of M samples per M samples in height and width is considered.

[0317] For example, if the center of the current block is considered the location of the time-domain region, then that location is the center of the time-domain region only if both PosCenter.x and PosCenter.y are multiples of M. Otherwise, the location used is one of the multiples of M around the initial PosCenter. For example, the location of the center of the time-domain region PosTempo is obtained as follows.

[0318] PosTempo.x=M*(PosCenter.x / M)

[0319] PosTempo.y=M*(PosCenter.y / M)

[0320] In these expressions, the division is integer division. Alternatively, it can be obtained using the following bitwise shift operations.

[0321] PosTempo.x=(PosCenter.x>>S)<

[0322] PosTempo.y=(PosCenter.y>>S)<

[0323] Where: >> is the left shift operator, and << is the right shift operator, and S = Log2(M).

[0324] ​​The main advantage of this embodiment is the reduction in the memory buffer required to store the values ​​determined for the time-domain region. This is particularly interesting for encoder implementations. In practice, at the encoder, the values ​​determined from the time-domain region can be determined several times for several block sizes, for example, with the same center. This is memory-intensive if all values ​​need to be determined. The same applies to the extraction of blocks considered within the time-domain region. By extracting the possible center locations of the time-domain region, the buffer is significantly reduced (a similar reduction exists for N=M).

[0325] Additionally, this reduces encoding time because fewer values ​​need to be determined at the encoder.

[0326] Furthermore, surprisingly, this extraction yields improved coding efficiency. This is especially true when using values ​​obtained from the time-domain region to constrain or infer variables for QT, BT, and TT partitions. Extraction with M equal to 16 yields the best coding efficiency.

[0327] The Emb.DecimCenter grid is centered.

[0328] In this embodiment, the grid for a possible block location may not start from the top-left position (0,0) of the frame, but may have a shift of (M / 2,M / 2) to achieve better location extraction. This embodiment is similar to a grid centered on the center of the temporal region.

[0329] The advantage is improved coding efficiency.

[0330] Extraction of both temporal regions and block locations in Emb.DecimBoth

[0331] In this embodiment, both the block of the time domain region and the extraction of possible locations are used together.

[0332] This also demonstrates improved coding efficiency and reduced complexity.

[0333] For larger blocks, Emb.Multiplecenter considers multiple locations within the buffer.

[0334] The use of buffers significantly reduces memory usage, but can be constrained if the time-domain region size depends on the current block size. In such cases, the value cannot be stored only once on the encoder side.

[0335] In the embodiments, when the size of the current block is used to determine the size of the time domain region, and when the possible locations of the current block are extracted, all possible available locations contained in the block are considered, and the values ​​of all relevant time domain regions are considered to determine the values ​​from the time domain regions for the current block.

[0336] This embodiment has the same advantages as using a time-domain region size that adapts to the current block size. Furthermore, compared to using a fixed size for the time-domain region, this embodiment does not require more memory for the buffer.

[0337] accomplish

[0338] In many implementations, to access a region in a previous frame, the region's location is considered, and it's necessary to traverse that region to obtain information about the blocks or CUs present within it. This is due to the structure of the buffer containing this information. Therefore, when accessing a temporal region to determine a value, the block structure is unknown without reaching each location. Thus, a basic solution involves traversing each location and calculating the value associated with that location.

[0339] Emb.Imple current implementation

[0340] In this embodiment, when determining values ​​from a time-domain region, the size of blocks within the time-domain region is considered to avoid multiple accesses to the same block. Therefore, each block is accessed only once. For example, a Boolean table representing all possible locations within the time-domain region is initialized to the value false. When extracting locations within the time-domain region, the possible locations are N×N grid locations. Impossible locations in this table are set to 1. Impossible locations are locations outside the time-domain frame. Each location has its corresponding position in the table, which is also set to false. When checking a location within the time-domain region, a value is extracted and a block is considered extracted. The relevant positions within the Boolean table are set to true, and also equal to all relevant positions corresponding to the current block associated with that location. The next location checked is then a block that has not yet been checked, and the relevant Boolean in the table is set to false.

[0341] Because of this implementation, the values ​​of the time domain region can be determined more quickly.

[0342] Emb.Pred Pred Prediction

[0343] In this embodiment, a temporal region is used to derive a predictor for the syntactic element or variable. For example, instead of directly encoding the syntactic element, a residual is extracted from the bitstream, and the predictor is added to that residual. In an alternative example, the first bit is extracted from the bitstream to determine whether the current syntactic element is set to equal to the corresponding value obtained from the temporal region. For example, if the flag is 1, the syntactic element is equal to the corresponding value from the temporal region. Otherwise, the flag is 0, and the remaining bits are decoded to determine the value of the syntactic element.

[0344] The use of time-domain regions is of great interest in terms of coding efficiency for obtaining corresponding values.

[0345] Emb.Infer inference

[0346] In this embodiment, as an alternative to the previous embodiment, the value of the syntactic element is inferred based on the corresponding value from the time domain region.

[0347] In this embodiment, the values ​​of variables for the current frame's block are inferred from the corresponding values ​​from the temporal region. For example, the current block's maxMttDepth is determined based on the maxMttDepth "maxMttDepthTempo" determined from the temporal region. Furthermore, the value of the current block's maxMttDepth is increased, decreased, or left unchanged according to certain rules and the current block's initial maxMttDepth compared to "maxMttDepthTempo".

[0348] The advantage of this example is that it improves coding efficiency by efficiently limiting the maximum multi-tree depth of the current block, thereby reducing coding time.

[0349] Emb.Limit Limit

[0350] In an embodiment, values ​​determined from a time-domain region are used to constrain syntactic elements or variables. For example, a maximum value is determined and the syntactic element values ​​of the current block are constrained to that maximum value. Therefore, the encoding is adapted to this constrained number of values ​​to reduce the number of bits that need to be transmitted. Similarly, a minimum value, or both a maximum and a minimum value, can be considered.

[0351] Emb.Limit1.QTBT

[0352] In another example, the variables are constrained. For instance, the minimum QT depth from the time-domain region is determined. And this value is used to determine the minimum QT depth of the current block QT depth. Therefore, fewer bits need to be transmitted for QT splitting, and the encoding time is reduced because fewer encoding possibilities need to be tested.

[0353] Emb.Ctx exports context index increment ctxInc.

[0354] In an embodiment, the context index increment of a syntactic element is determined using a value determined from the temporal region. For example, the value “condTempo” obtained from the temporal region (if available (availableTempo=1)) is added to other context increments obtained from the spatial positions above (condA) and to the left (condL), as in the following formula.

[0355] ctxInc=(condL&&availableL)||(condA&&availableA)||(condTempo&&availableTempo)

[0356] The advantage is improved encoding efficiency, because syntactic elements usually have spatial and temporal relevance.

[0357] The value is:

[0358] Emb.Min minimum value

[0359] In an embodiment, the value to be determined from the time domain region is the minimum value from the block to be considered within the time domain region.

[0360] Emb.minQTDepth Special case of minQTDepth

[0361] In a particular embodiment, the value to be determined from the time-domain region is the minimum QT depth value of the block to be considered within the time-domain region. For example, this minQTDepth is then compared with the current QT depth of the current block to determine whether only QT splitting is allowed.

[0362] This reduces code runtime and improves coding efficiency by limiting the number of possible splits to be tested and the number of relevant bits that are not signaled due to the solution.

[0363] Emb.Max maximum value

[0364] In an embodiment, the value to be determined from the time domain region is the maximum value from the block to be considered within the time domain region.

[0365] A special case of Emb.MaxMttDepth

[0366] In a particular embodiment, the value to be determined from the time domain region is the maximum value of the multi-tree depth values ​​of the blocks to be considered within the time domain region. For example, this maxMttDepthTempo is then compared with the current maxMttDepth depth of the current block to determine, based on other conditions, whether maxMttDepth needs to be increased, decreased, or kept the same.

[0367] This improves coding efficiency.

[0368] Emb.Median median

[0369] In an embodiment, the value to be determined from the time domain region is the median value of the block to be considered within the time domain region.

[0370] Emb.Average Average

[0371] In an embodiment, the value to be determined from the time domain region is the average value determined from the blocks to be considered within the time domain region.

[0372] Maintain the block size ratio within the time domain region.

[0373] As previously mentioned with respect to other embodiments, the block size proportion should be maintained. In this case, the average value should take into account the number of samples contained in each block or, alternatively, the smallest block unit (minimum block size (4×4)). In this case, for block i and the associated value BLVal_i, the average value AverageVal is obtained as follows.

[0374] AverageVal=0

[0375] Nb_samples = 0

[0376] For (i = 0 to nb_blocks)

[0377]

[0378] Where: Nb_samples is the number of samples in all blocks of the time-domain region, nb_blocks is the number of blocks, and height_i and width_i are the height and width of block number i. In an alternative example, height_i and width_i can be the height and width expressed in terms of the extracted time-domain position, respectively. For example, if the considered time-domain decimation is set to equal 16 (taking a position for 16 samples vertically and horizontally), then the height and width are divided by 16. Or, they are shifted right by log2(16) = 4.

[0379] Alternatively, a rounding process can be added to the expression as follows.

[0380] AverageVal=(AverageVal+(0.5)) / Nb_samples

[0381] This rounding improves coding efficiency.

[0382] In the alternative example,

[0383] AverageVal=0

[0384] Nb_samples = 0

[0385] For (i = 0 to nb_blocks)

[0386]

[0387] Only blocks that are completely within the time domain region are preserved.

[0388] When a block crosses the boundary of a time domain region, only the portion inside the time domain region is considered, and the height_i and width_i of the relevant block correspond to the height and width inside the time domain region.

[0389] This can be implemented using the following algorithm.

[0390] AverageVal=0

[0391] Nb_samples = 0

[0392] For (i = 0 to nb_blocks)

[0393]

[0394] Where: height_i and width_i are the height and width represented by the extracted time-domain position within the time-domain region, respectively.

[0395] Therefore, if a portion of a block is outside the time domain region, the number of positions outside the time domain region is subtracted from the height_i and width_i of block i, respectively, according to the extraction.

[0396] blocksize equals the number of samples in the current block i. Therefore, blocksize is the product of height and width.

[0397] For hardware implementation, integer division is required. Therefore, the expression needs to be suitable for integer implementation. Thus, in one embodiment, the expression becomes as follows.

[0398] AverageVal=0

[0399] Nb_samples = 0

[0400] For (i = 0 to nb_blocks)

[0401]

[0402] Wherein: the rounding value roundVal is set to be equal to the following.

[0403] roundVal=(tempoAreaWidth>>2)*(tempoAreaHeight>>2)*(tempoAreaWidth / TempoRes)*(tempoAreaHeight / TempoRes)

[0404] Here, TempoRes corresponds to the extraction of the time-domain region. Therefore, in the previously mentioned examples, TempoRes is 16. And tempoAreaWidth and tempoAreaHeight are the height and width of the time-domain region. Therefore, in the previously described examples, they can be equal to the CTU size.

[0405] Based on the previously mentioned example (where the time domain region is set to equal the CTU size and the decimation is equal to 16), roundVal is set to equal to the following.

[0406] roundVal=(CTUSize>>2)*(CTUSize>>2)*(CTUSize / TempoRes)*(CTUSize / TempoRes)

[0407] Using this rounded value, the same result as floating-point division was obtained.

[0408] Furthermore, by considering the log2 value, all divisions can be replaced by right shift.

[0409] Emb.QTDepthTempo Special case of QtDepthTempo

[0410] In an embodiment, the value to be determined from the time-domain region is the average of the QT depth values ​​of the blocks to be considered within the time-domain region. For example, this QTDepthTempo is then compared with the QT depth of the current block to allow only QT, no splits, and TT, or additionally or alternatively, the QTDepthTempo is compared with the QT depth of the current block, and maxMttDepth is increased if the QTDepthTempo is equal to the QT depth of the current block and according to other conditions.

[0411] This results in improved coding efficiency and reduced coding time.

[0412] Emb.Variance

[0413] In this embodiment, the value to be determined from the time-domain region is the variance of the values ​​from the blocks to be considered within the time-domain region. The variance is the average of the distances to the mean. This requires determining the mean first, and then the variance. Therefore, the blocks within the time-domain region are considered twice.

[0414] other

[0415] Emb.OTHER1 can combine all these embodiments.

[0416] Unless otherwise explicitly stated, all described embodiments may be combined. In fact, many combinations are synergistic and can produce greater efficiency gains than the sum of their parts.

[0417] Emb.Contrib for contributions

[0418] Specifically, the minimum QT depth “minQTDepthTempo”, the maximum multi-tree depth “maxMttDepthTempo”, and the average depth value “QTDepthTempo” are determined from the temporal region. Then, minQTDepthTempo is compared with the current QT depth of the current block to determine whether only QT splitting is allowed. maxMttDepthTempo is compared with the current maxMttDepth depth of the current block. If maxMttDepthTempo is lower than maxMttDepth, maxMttDepth is decreased. If maxMttDepthTempo is higher, and if QTDepthTempo is equal to the QT depth of the current block, maxMttDepth is increased. Furthermore, QTDepthTempo is compared with the QT depth of the current block to allow only QT, no splitting, and TT. The temporal region is based on the reference frame used for temporal motion vector prediction, and the temporal region corresponds to the center of the current block. The temporal region is equal to the CTU size. Three values ​​are determined by considering only the blocks present in a 16×16 grid, and the average value QTDepthTempo is calculated taking into account the size of the blocks. Therefore, this value is obtained by considering the proportionality between blocks. This also applies to blocks at the boundaries of the time-domain region. Finally, the position of the current block is extracted using a 16×16 grid.

[0419] Realization of the invention

[0420] Figure 20Systems 191 and 195 according to embodiments of the present invention are illustrated, comprising at least one of an encoder 150 or a decoder 100 and a communication network 199. According to an embodiment, system 195 is used to process and provide content (e.g., video and audio content for display / output or streaming video / audio content) to a user, who accesses decoder 100, for example, through a user interface of a user terminal including decoder 100 or a user terminal capable of communicating with decoder 100. Such a user terminal may be a computer, mobile phone, tablet computer, or any other type of device capable of providing / displaying (provided / streamed) content to the user. System 195 receives / receives bitstream 101 (in the form of a continuous stream or signal (e.g., when displaying / outputting earlier video / audio)) via communication network 199. According to an embodiment, system 191 is used to process and store processed content, such as video and audio content processed for display / output / streaming at a later time. System 191 acquires / receives content comprising an original image sequence 151, which is received and processed by encoder 150 (including filtering using a deblocking filter according to the invention), and encoder 150 generates a bitstream 101 to be transmitted to decoder 100 via communication network 191. The bitstream 101 is then transmitted to decoder 100 in various ways, for example, it may be pre-generated by encoder 150 and stored as data in a storage device within communication network 199 (e.g., on a server or cloud storage device) until a user requests content (i.e., bitstream data) from the storage device, at which point the data is transferred / streamed from the storage device to decoder 100. System 191 may also include a content providing device for providing / streaming (e.g., by transmitting data of a user interface to be displayed on a user terminal) content information (e.g., a title of the content and other metadata / storage location data for identifying, selecting, and requesting the content) of the content stored in the storage device to the user, and for receiving and processing user requests for content such that the requested content can be transferred / streamed from the storage device to the user terminal. Alternatively, encoder 150 generates bitstream 101 and transmits / streams it directly to decoder 100 when the user requests content. Decoder 100 then receives bitstream 101 (or signal) and filters it using the deblocking filter according to the invention to obtain / generate video signal 109 and / or audio signal, which the user terminal then uses to provide the requested content to the user.

[0421] Any step of the method / process according to the invention or the function described herein can be implemented in hardware, software, firmware, or any combination thereof. If implemented in software, the step / function can be stored on or transmitted via one or more hardware-based processing units as one or more instructions, code, programs, or computer-readable media, and executed by one or more hardware-based processing units, such as programmable computing machines, which can be PCs (“personal computers”), DSPs (“digital signal processors”), circuits, circuit systems, processors and memories, general-purpose microprocessors or central processing units, microcontrollers, ASICs (“application-specific integrated circuits”), field-programmable logic arrays (FPGAs), or other equivalent integrated or discrete logic circuit systems. Therefore, the term “processor” as used herein can refer to any of the foregoing structures or any other structures suitable for implementing the techniques described herein.

[0422] Embodiments of the present invention can also be implemented by various devices or apparatuses, including wireless mobile phones, integrated circuits (ICs), or IC sets (e.g., chipsets). Various components, modules, or units are described herein to illustrate functional aspects of the apparatus / apparatus configured to perform these embodiments, but they do not necessarily need to be implemented by different hardware units. Rather, various modules / units may be combined in a codec hardware unit or provided by a collection of interoperable hardware units, which include one or more processors incorporating suitable software / firmware.

[0423] Embodiments of the present invention can be implemented by a computer of a system or device that reads and executes computer-executable instructions (e.g., one or more programs) recorded on a storage medium to perform one or more modules / units / functions in the above embodiments and / or includes one or more processing units or circuits for performing one or more functions in the above embodiments. Furthermore, the invention can be implemented by a method performed by the computer of the system or device, such as reading and executing computer-executable instructions from a storage medium to perform one or more functions in the above embodiments and / or controlling one or more processing units or circuits to perform one or more functions in the above embodiments. The computer may include a network of separate computers or separate processing units to read and execute the computer-executable instructions. The computer-executable instructions may be provided to the computer, for example, via a network or tangible storage medium from a computer-readable medium such as a communication medium. The communication medium may be a signal / bit stream / carrier. Tangible storage media are “non-transitory computer-readable storage media”, which may include, for example, hard disks, random access memory (RAM), read-only memory (ROM), storage devices for distributed computing systems, optical discs (e.g., compact discs (CDs), digital versatile optical discs (DVDs), or Blu-ray discs (BDs). TM One or more of the following: flash memory devices, memory cards, etc. At least some steps / functions may also be implemented in hardware by a machine or a dedicated component (such as an FPGA (“Field Programmable Gate Array”) or an ASIC (“Application-Specific Integrated Circuit”)).

[0424] Figure 21This is a schematic block diagram of a computing device 3600 for implementing one or more embodiments of the present invention. The computing device 3600 may be a device such as a microcomputer, workstation, or lightweight portable device. The computing device 3600 includes a communication bus connected to: - a central processing unit (CPU) 3601, such as a microprocessor; - a random access memory (RAM) 3602 for storing executable code of methods according to embodiments of the present invention and registers adapted to record variables and parameters required for implementing methods for encoding or decoding at least a portion of an image according to embodiments of the present invention, the storage capacity of which may be expanded, for example, by an optional RAM connected to an expansion port; - a read-only memory (ROM) 3603 for storing a computer program for implementing embodiments of the present invention; - a network interface (NET) 3604, typically connected to a communication network through which digital data to be processed is transmitted or received. The network interface (NET) 3604 may be a single network interface or a set of different network interfaces (e.g., wired and wireless interfaces, or different types of wired or wireless interfaces). Under the control of a software application running on the CPU 3601, data packets are written to or read from the network interface for transmission or reception; a user interface (UI) 3605 can be used to receive input from a user or display information to a user; a hard disk (HD) 3606 can be configured as a mass storage device; and an input / output module (IO) 3607 can be used to receive / send data from / to external devices (such as video sources or displays). Executable code can be stored in ROM 3603, on HD 3606, or on a removable digital medium such as a disk. According to a variant, the executable code of the program can be received via NET 3604 through a communication network to be stored in one of the storage components (such as HD 3606) of the computing device 3600 before execution. The CPU 3601 is adapted to control and direct the execution of instructions or portions of software code of one or more programs according to embodiments of the present invention, the instructions being stored in one of the aforementioned storage components. For example, after power-on, the CPU 3601 can execute software application-related instructions from the main RAM memory 3602 after the instructions have been loaded from the program ROM 3603 or HD 3606. This software application, when executed by the CPU 3601, causes the steps of the method according to the invention to be performed.

[0425] It should also be understood that, according to other embodiments of the invention, a decoder according to the above embodiments is provided in a user terminal such as a computer, mobile phone (cellular phone), tablet, or any other type of device capable of providing / displaying content to a user (e.g., a display device). According to yet another embodiment, an encoder according to the above embodiments is provided in an image capture device, which further includes a camera, video camera, or webcam (e.g., a closed-circuit television or video surveillance camera) for capturing and providing content for encoding by the encoder. See below. Figure 22 and Figure 23 Here are two such examples.

[0426] Figure 22 This is a diagram illustrating a network camera system 3700 including a network camera 3702 and a client device 202.

[0427] The network camera 3702 includes a camera unit 3706, an encoding unit 3708, a communication unit 3710, and a control unit 3712.

[0428] The network camera 3702 and the client device 202 are interconnected via network 200 so that they can communicate with each other.

[0429] The camera unit 3706 includes a lens and an image sensor (e.g., a charge-coupled device (CCD) or complementary metal-oxide-semiconductor (CMOS)), and captures an image of the object and generates image data based on the image. The image can be a still image or a video image.

[0430] The encoding unit 3708 encodes the image data by using the encoding method described above or a combination of the encoding methods described above.

[0431] The communication unit 3710 of the network camera 3702 transmits the encoded image data encoded by the encoding unit 3708 to the client device 202.

[0432] In addition, the communication unit 3710 receives commands from the client device 202. These commands include those for setting parameters for encoding by the encoding unit 3708.

[0433] The control unit 3712 controls other units in the network camera 3702 according to the commands received by the communication unit 3710.

[0434] The client device 202 includes a communication unit 3714, a decoding unit 3716, and a control unit 3718.

[0435] The communication unit 3714 of the client device 202 transmits commands to the network camera 3702.

[0436] In addition, the communication unit 3714 of the client device 202 receives encoded image data from the webcam 3702.

[0437] The decoding unit 3716 decodes the encoded image data by using the decoding method described above or a combination of the decoding methods described above.

[0438] The control unit 3718 of the client device 202 controls other units in the client device 202 based on user operations or commands received by the communication unit 3714.

[0439] The control unit 3718 of the client device 202 controls the display device 2120 to display the image decoded by the decoding unit 3716.

[0440] The control unit 3718 of the client device 202 also controls the display device 2120 to display the values ​​of the parameters for specifying the network camera 3702 (including the parameters for encoding the encoding unit 3708) in a GUI (graphical user interface).

[0441] The control unit 3718 of the client device 202 also controls other units in the client device 202 based on user operation input to the GUI displayed on the display device 2120.

[0442] The control unit 3718 of the client device 202 controls the communication unit 3714 of the client device 202 based on user operation input to the GUI displayed on the display device 2120, so as to transmit commands for specifying the values ​​of parameters of the network camera 3702 to the network camera 3702.

[0443] Figure 23 This is a diagram illustrating a smartphone 3800.

[0444] The smartphone 3800 includes a communication unit 3802, a decoding / encoding unit 3804, a control unit 3806, and a display unit 3808.

[0445] The communication unit 3802 receives encoded image data via the network 200.

[0446] The decoding / encoding unit 3804 decodes the encoded image data received by the communication unit 3802.

[0447] The decoding / encoding unit 3804 decodes / encodes the encoded image data using the decoding method described above.

[0448] The control unit 3806 controls other units in the smartphone 3800 based on user operations or commands received from the communication unit 3802.

[0449] For example, control unit 3806 controls display unit 3808 to display images decoded by decoding / encoding unit 3804. Smartphone 3800 may also include sensor 3812 and image recording device 3810. In this way, smartphone 3800 can record images and encode them (using the methods described above).

[0450] The smartphone 3800 can then (using the methods described above) decode the encoded image and display it via the display unit 3808, or transmit it to another device via the communication unit 3802 and the network 200.

[0451] Replacement and modification

[0452] While the invention has been described with reference to embodiments, it should be understood that the invention is not limited to the disclosed embodiments. Those skilled in the art will understand that various changes and modifications can be made without departing from the scope of the invention as defined in the appended claims. All features disclosed in this specification (including any appended claims, abstract, and drawings) and / or all steps of any method or process so disclosed can be combined in any combination other than mutually exclusive combinations of at least some of such features and / or steps. Unless otherwise expressly stated, the features disclosed in this specification (including any appended claims, abstract, and drawings) can be replaced by alternative features serving the same, equivalent, or similar purpose. Thus, unless otherwise expressly stated, the disclosed features are merely one example of equivalent or similar features in a general series.

[0453] It should also be understood that any result of the above comparisons, determinations, evaluations, selections, executions, processes, or considerations (e.g., selections made during encoding or filtering processes) may be indicated in or determined / inferred from data in the bitstream (e.g., flags or data indicating the result), such that the indicated or determined / inferred result may be used for processing rather than actually being compared, determined, evaluated, selected, executed, processed, or considered, for example, during decoding processes.

[0454] In the claims, the word "comprising" does not exclude other elements or steps, and the indefinite articles "a" or "an" do not exclude plural. The fact that different features are recited only in mutually different dependent claims does not indicate that a combination of these features cannot be used advantageously.

[0455] The reference numerals appearing in the claims are for illustrative purposes only and should not be construed as limiting the scope of the claims.

Claims

1. A method for decoding video data from a bitstream, the bitstream comprising video data corresponding to a plurality of frames arranged in decoding order, the method comprising: Derive a value from the first region in the first frame of the plurality of frames; as well as The context increment, syntactic element, or variable associated with the second region in the second frame is determined based on the value. The first frame precedes the second frame according to the decoding order.

2. The method according to claim 1, wherein, Each of the plurality of frames has an associated temporal ID, and wherein the first frame and the second frame have the same temporal ID.

3. The method according to claim 2, wherein, The first frame corresponds to the frame with the same time domain ID that is closest to the second frame in the decoding order.

4. The method according to any one of the preceding claims, wherein, Each of the plurality of frames has an associated quantization parameter, namely QP, and wherein the first frame and the second frame have the same QP.

5. The method according to any one of the preceding claims, wherein, The first frame is a reference frame.

6. The method according to any one of the preceding claims, wherein, The first region is equal to or larger than the second region in size.

7. The method according to claim 6, wherein, The second region corresponds to the coding tree unit, i.e., CTU.

8. The method according to any one of the preceding claims, wherein, The size of the first region is set based on the temporal distance between the first frame and the second frame.

9. The method according to claim 8, wherein, The temporal difference is calculated based on the difference between the image sequence count (POC) of the first frame and the POC of the second frame.

10. The method according to any one of the preceding claims, wherein, Each of the plurality of frames has an associated quantization parameter, QP, and the size of the first region is set based on the difference between the QP of the first frame and the QP of the second frame.

11. The method according to any one of the preceding claims, wherein, Each of the plurality of frames has an associated temporal ID, and the size of the first region is set based on the difference between the temporal ID of the first frame and the temporal ID of the second frame.

12. The method according to any one of claims 1 to 7, wherein, The size of the first region is set based on values ​​transmitted in one of the following: a sequence parameter set, an image parameter set, an image header, and a strip header contained in the bitstream.

13. The method according to any one of the preceding claims, wherein, The size of the first region is set based on the size of the second region.

14. The method according to any one of the preceding claims, wherein, The step of deriving a value from a first region in a first frame of the plurality of frames includes: deriving a value from a block having at least one sample in the first region.

15. The method according to any one of claims 1 to 13, wherein, The step of deriving values ​​from a first region in a first frame of the plurality of frames includes: deriving values ​​from all blocks that have at least one sample in the first region.

16. The method according to claim 14 or 15, wherein, The step of deriving values ​​from a first region in a first frame of the plurality of frames includes: weighting the values ​​derived from the block or blocks based on the number of samples each block has in the first region.

17. The method according to any one of claims 14 to 16, wherein, The value is derived from the block or blocks using integer arithmetic.

18. The method according to any one of claims 1 to 13, wherein, The steps for deriving values ​​from a first region in a first frame of the plurality of frames include: deriving values ​​only from blocks that are completely contained within the first region.

19. The method according to any one of claims 1 to 13, wherein, The step of deriving values ​​from a first region in a first frame of the plurality of frames includes: deriving values ​​only from blocks located on an N×N grid within the first region, where N is an integer.

20. The method according to claim 19, wherein, N=16。 21. The method according to claim 19 or 20, wherein, The center of the grid is located at the same position as the center of the first region within the first frame.

22. The method according to any one of claims 1 to 13, wherein, The steps for deriving values ​​from a first region in a first frame of the plurality of frames include: deriving values ​​only from blocks on points of a pattern located within the first region.

23. The method according to claim 22, wherein, The dots in the pattern are spaced out in a non-uniform manner.

24. The method according to any one of the preceding claims, wherein, The step of determining, based on the value, a context increment or syntactic element or variable of a syntactic element associated with a second region in a second frame includes: determining a context increment or syntactic element or variable of a syntactic element associated with a block within the second region.

25. The method according to any one of claims 1 to 23, wherein, The step of determining a context increment or syntactic element or variable of a syntactic element associated with a second region in a second frame based on the value includes: determining a context increment or syntactic element or variable of a syntactic element associated with a block on an M×M grid located within the second region, where M is an integer.

26. The method of claim 25, wherein, The M×M grid is shifted by M / 2 in the horizontal direction and by M / 2 in the vertical direction relative to the upper left position of the second region.

27. The method according to any one of the preceding claims, wherein, The center of the first region is located at the same position within the first frame as the center of the second region within the second frame.

28. The method according to any one of claims 1 to 26, wherein, The center of the first region is located in the first frame at a position that corresponds to the position in the second frame after the center of the second region has been shifted by an amount corresponding to the motion vector derived from the region adjacent to the second region.

29. The method according to any one of claims 1 to 26, further comprising: When the center of the first region is outside the first frame, the value is derived from the upper left position of the first region.

30. The method according to any one of the preceding claims, wherein, The steps for deriving values ​​from a first region in a first frame of the plurality of frames include: accessing each block within the first region only once.

31. The method according to any one of the preceding claims, wherein, The steps for deriving a value from a first region in a first frame of the plurality of frames include: Values ​​are derived from the first region in the first frame of the plurality of frames and the third region in the third frame of the plurality of frames.

32. The method according to any one of the preceding claims, wherein, The steps for determining the context increment or syntactic element or variable of a syntactic element associated with a second region in a second frame based on the value include: determining a predictor to be added to the residual derived from the bitstream.

33. The method according to any one of claims 1 to 31, wherein, The values ​​derived from the first region include maxMttDepth, and the steps for determining, based on the values, a context increment or syntactic element or variable relating to a second region in the second frame include: determining maxMttDepth for blocks of the second region.

34. The method according to any one of claims 1 to 31, wherein, Use values ​​derived from the first region to limit the contextual increment of variables or syntactic elements associated with the second region.

35. The method according to any one of claims 1 to 31, wherein, The value derived from the first region is the minimum value minQTDepth of the quadtree depth of the block from the first region, and the step of determining the context increment or syntactic element or variable of the syntactic element associated with the second region in the second frame based on the value includes: comparing the derived minQTDepth with the quadtree depth of the block in the second region to determine whether only quadtree splitting is allowed.

36. The method according to any one of claims 1 to 31, wherein, The value derived from the first region is the maximum value of the multi-tree depth value of the block from the first region, maxMttDepth. The steps for determining the context increment or syntactic element or variable of the syntactic element associated with the second region in the second frame based on the value include: comparing the derived maxMttDepth with the maxMttDepth of the block in the second region and changing maxMttDepth based on the comparison.

37. The method according to any one of claims 1 to 31, wherein, The value derived from the first region is the average of the quadtree depth values ​​of the blocks from the first region, and the step of determining the context increment or syntactic element or variable of the syntactic element associated with the second region in the second frame based on the value includes: comparing the derived quadtree depth value with the quadtree depth of the blocks in the second region.

38. The method of claim 37, further comprising the step of: The permissible splits are determined based on the comparison.

39. The method according to claim 37 or claim 38, further comprising the step of: The maxMttDepth is changed based on the comparison.

40. An apparatus for decoding image data from a bitstream, wherein the apparatus is configured to perform the method according to any one of claims 1 to 39.

41. A method for encoding video data into a bitstream, the bitstream comprising video data corresponding to a plurality of frames arranged in decoding order, the method comprising: Derive a value for the first region in the first frame of the plurality of frames; as well as The context increment, syntactic element, or variable associated with the second region in the second frame is determined based on the value. The first frame precedes the second frame according to the decoding order.

42. The method according to claim 41, wherein, Each of the plurality of frames has an associated temporal ID, and wherein the first frame and the second frame have the same temporal ID.

43. The method according to claim 42, wherein, The first frame corresponds to the frame with the same time domain ID that is closest to the second frame in the decoding order.

44. The method according to any one of claims 41 to 43, wherein, Each of the plurality of frames has an associated quantization parameter, namely QP, and wherein the first frame and the second frame have the same QP.

45. The method according to any one of claims 41 to 44, wherein, The first frame is a reference frame.

46. ​​The method according to any one of claims 41 to 45, wherein, The first region is equal to or larger than the second region in size.

47. The method according to claim 46, wherein, The second region corresponds to the coding tree unit, i.e., CTU.

48. The method according to any one of claims 41 to 47, wherein, The size of the first region is set based on the temporal distance between the first frame and the second frame.

49. The method according to claim 48, wherein, The temporal difference is calculated based on the difference between the image sequence count (POC) of the first frame and the POC of the second frame.

50. The method according to any one of claims 41 to 49, wherein, Each of the plurality of frames has an associated quantization parameter, QP, and the size of the first region is set based on the difference between the QP of the first frame and the QP of the second frame.

51. The method according to any one of claims 41 to 50, wherein, Each of the plurality of frames has an associated temporal ID, and the size of the first region is set based on the difference between the temporal ID of the first frame and the temporal ID of the second frame.

52. The method according to any one of claims 41 to 47, wherein, The size of the first region is set based on values ​​transmitted in one of the following: a sequence parameter set, an image parameter set, an image header, and a strip header contained in the bitstream.

53. The method according to any one of claims 41 to 52, wherein, The size of the first region is set based on the size of the second region.

54. The method according to any one of claims 41 to 53, wherein, The step of deriving a value from a first region in a first frame of the plurality of frames includes: deriving a value from a block having at least one sample in the first region.

55. The method according to any one of claims 40 to 52, wherein, The step of deriving values ​​from a first region in a first frame of the plurality of frames includes: deriving values ​​from all blocks that have at least one sample in the first region.

56. The method according to claim 54 or 55, wherein, The step of deriving values ​​from a first region in a first frame of the plurality of frames includes: weighting the values ​​derived from the block or blocks based on the number of samples each block has in the first region.

57. The method according to any one of claims 54 to 56, wherein, The value is derived from the block or blocks using integer arithmetic.

58. The method according to any one of claims 41 to 53, wherein, The steps for deriving values ​​from a first region in a first frame of the plurality of frames include: deriving values ​​only from blocks that are completely contained within the first region.

59. The method according to any one of claims 41 to 53, wherein, The step of deriving values ​​from a first region in a first frame of the plurality of frames includes: deriving values ​​only from blocks located on an N×N grid within the first region, where N is an integer.

60. The method according to claim 59, wherein, N=16。 61. The method according to claim 59 or 60, wherein, The center of the grid is located at the same position as the center of the first region within the first frame.

62. The method according to any one of claims 41 to 53, wherein, The steps for deriving values ​​from a first region in a first frame of the plurality of frames include: deriving values ​​only from blocks on points of a pattern located within the first region.

63. The method according to claim 62, wherein, The dots in the pattern are spaced out in a non-uniform manner.

64. The method according to any one of claims 41 to 63, wherein, The step of determining, based on the value, a context increment or syntactic element or variable of a syntactic element associated with a second region in a second frame includes: determining a context increment or syntactic element or variable of a syntactic element associated with a block within the second region.

65. The method according to any one of claims 41 to 63, wherein, The step of determining a context increment or syntactic element or variable of a syntactic element associated with a second region in a second frame based on the value includes: determining a context increment or syntactic element or variable of a syntactic element associated with a block on an M×M grid located within the second region, where M is an integer.

66. The method according to claim 65, wherein, The M×M grid is shifted by M / 2 in the horizontal direction and by M / 2 in the vertical direction relative to the upper left position of the second region.

67. The method according to any one of claims 41 to 66, wherein, The center of the first region is located at the same position within the first frame as the center of the second region within the second frame.

68. The method according to any one of claims 41 to 66, wherein, The center of the first region is located in the first frame at a position that corresponds to the position in the second frame after the center of the second region has been shifted by an amount corresponding to the motion vector derived from the region adjacent to the second region.

69. The method according to any one of claims 41 to 66, further comprising: When the center of the first region is outside the first frame, the value is derived from the upper left position of the first region.

70. The method according to any one of claims 41 to 69, wherein, The steps for deriving values ​​from a first region in a first frame of the plurality of frames include: accessing each block within the first region only once.

71. The method according to any one of claims 41 to 70, wherein, The steps for deriving a value from a first region in a first frame of the plurality of frames include: Values ​​are derived from the first region in the first frame of the plurality of frames and the third region in the third frame of the plurality of frames.

72. The method according to any one of claims 41 to 71, wherein, The steps for determining the context increment or syntactic element or variable of a syntactic element associated with a second region in a second frame based on the value include: determining a predictor to be added to the residual derived from the bitstream.

73. The method according to any one of claims 41 to 71, wherein, The values ​​derived from the first region include maxMttDepth, and the steps for determining, based on the values, a context increment or syntactic element or variable relating to a second region in the second frame include: determining maxMttDepth for blocks of the second region.

74. The method according to any one of claims 41 to 71, wherein, Use values ​​derived from the first region to limit the contextual increment of variables or syntactic elements associated with the second region.

75. The method according to any one of claims 41 to 71, wherein, The value derived from the first region is the minimum value minQTDepth of the quadtree depth of the block from the first region, and the step of determining the context increment or syntactic element or variable of the syntactic element associated with the second region in the second frame based on the value includes: comparing the derived minQTDepth with the quadtree depth of the block in the second region to determine whether only quadtree splitting is allowed.

76. The method according to any one of claims 41 to 71, wherein, The value derived from the first region is the maximum value of the multi-tree depth value of the block from the first region, maxMttDepth. The steps for determining the context increment or syntactic element or variable of the syntactic element associated with the second region in the second frame based on the value include: comparing the derived maxMttDepth with the maxMttDepth of the block in the second region and changing maxMttDepth based on the comparison.

77. The method according to any one of claims 41 to 71, wherein, The value derived from the first region is the average of the quadtree depth values ​​of the blocks from the first region, and the step of determining the context increment or syntactic element or variable of the syntactic element associated with the second region in the second frame based on the value includes: comparing the derived quadtree depth value with the quadtree depth of the blocks in the second region.

78. The method of claim 77, further comprising the step of: The permissible splits are determined based on the comparison.

79. The method according to claim 77 or claim 78, further comprising the step of: The maxMttDepth is changed based on the comparison.

80. An apparatus for encoding image data into a bitstream, wherein the apparatus is configured to perform the method according to any one of claims 41 to 79.

81. A computer program configured to perform the method according to any one of claims 1 to 39 or claims 41 to 79 when executed.