Video Encoding and Decoding

JP2025513175A5Pending Publication Date: 2026-04-06CANON KK
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Filing Date
2023-04-14
Publication Date
2026-04-06

AI Technical Summary

Technical Problem

The complexity and code rate increase of existing video encoding standards in motion vector prediction, especially in the VVC standard, the sorting of multiple candidate motion vector predictors has a significant impact on encoding efficiency.

Method used

By introducing a threshold mechanism in video encoding, it is used to determine the threshold for decoding decisions, thereby optimizing the candidate list sorting of motion vector predictions and reducing unnecessary bit rate consumption. The method includes decoding syntax elements in the header of the image portion to determine and use thresholds to make decoding decisions.

Benefits of technology

Improved encoding performance, improved encoding efficiency by rationally selecting decoding standards, and achieved this with minimal bit rate influence through flexible signaling mechanisms.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 00000000_0000_ABST
    Figure 00000000_0000_ABST
Patent Text Reader

Abstract

A method for decoding a portion of an image is disclosed, the method including determining a threshold value from a header associated with the image portion, where a decoding criterion is based on the threshold value, and decoding the image portion based on the criterion. An encoding method and corresponding device are also disclosed.
Need to check novelty before this filing date? Find Prior Art

Description

[Technical field]

[0001] The present invention relates to video encoding and decoding. [Background technology]

[0002] The Joint Video Experts Team (JVET), a collaborative team formed by MPEG and the VCEG of ITU-T Study Group 16, has introduced a new video coding standard called VVC (Versatile Video Coding). The goal of VVC is to provide significant improvements in compression performance over the existing HEVC standard (i.e., typically twice as much as the previous one). Primary target applications and services include, but are not limited to, 360-degree and high dynamic range (HDR) video. It has shown particular effectiveness on ultra-high definition (UHD) video test material. Thus, we can expect compression efficiency improvements well beyond the 50% goal of the final standard.

[0003] Since the completion of the VVC v1 standardization, JVET has started an exploration phase by establishing an exploration software (ECM), which will collect additional tools and improvements of existing tools on top of the VVC standard to target better coding efficiency.

[0004] Among other modifications, compared to HEVC, VVC has a modified set of "merge modes" for motion vector prediction, which achieves greater coding efficiency at the expense of greater complexity. Motion vector prediction is enabled by deriving a list of "motion vector predictor candidates", and the index of the selected candidate is signaled in the bitstream. A merge candidate list is generated per coding unit (CU). However, CUs may be split into smaller blocks for Decoder-side Motion Vector Refinement (DMVR) or other methods.

[0005] This list organization and order can have a significant impact on coding efficiency, since accurate motion vector predictors reduce the size of the residual or distortion of the block predictor, and having such candidates at the top of the list reduces the number of bits required to signal the selected candidate. The present invention aims to improve upon at least one of these aspects.

[0006] Modifications incorporated into VVC v1 and ECM mean that there can be up to 10 motion vector predictor candidates, which allows for candidate diversity but can increase the bitrate if a candidate lower in the list is selected.

[0007] The present invention relates generally to methods of transmitting / receiving (or deriving) values ​​that enable a decoder to make decoding decisions, for example, to sort a list of candidates. Summary of the Invention

[0008] According to one aspect of the invention, there is provided a method of decoding a portion of an image, the method comprising determining a threshold value from a header associated with the image portion, where a decoding criterion is based on said threshold value, and decoding said image portion based on said criterion. This increases coding performance, since a proper selection of the decoding criterion increases coding efficiency. Deriving it from the header allows flexible signaling, thus having minimal impact on bitrate.

[0009] The method may include decoding a syntax element in a header that may indicate the presence of a threshold value in the header.

[0010] Optionally, the syntax element may also indicate that the threshold has a default value.

[0011] Optionally, the syntax element can also indicate that a threshold is not transmitted and that the decoding process uses an alternative method.

[0012] Optionally, to increase decoding speed / simplicity, the method may include skipping decoding of a corresponding syntax element in a hierarchically lower header if the syntax element indicates the presence of a threshold. Optionally, the syntax element is a flag. How to decide The present invention is applicable to some decoding tools, for example where the decoding includes a decision to perform motion vector refinement based on a criterion, which improves the performance of the decoding tool.

[0013] Similarly, decoding includes the decision to perform ordering or removal of motion vectors based on a criterion. Appropriate selection of the decoding criterion leads to a more accurately ordered list and thus rate reduction.

[0014] Optionally, the method further comprises determining an encoding mode, and determining the threshold is based on the determined mode.

[0015] Optionally, the criteria relate to cost.

[0016] Threshold Signaling Optionally, to reduce the bit rate, determining a threshold value comprises decoding a syntax element indicating a maximum number of bits for said threshold value.

[0017] Optionally, to reduce the bit rate, determining a threshold value comprises decoding a syntax element indicating a number of bits for said threshold value.

[0018] Optionally, to reduce the bit rate, the method includes transforming the decoded threshold value before using it in the decoding process, where the transforming includes left bit-shifting and / or squaring, inverse quantizing.

[0019] Threshold prediction To further reduce the bit rate, a threshold can be predicted.

[0020] Optionally, the method includes decoding a flag indicating whether decoding the threshold value includes predicting the threshold value or decoding the threshold value directly from the header. Similarly, determining the threshold value may include decoding the threshold value directly from the header when it is not predictable or predicting the threshold value.

[0021] Optionally, the method comprises decoding a residual and modifying the predicted threshold based on said residual.

[0022] Optionally, the predicted threshold is predicted based on a previously decoded threshold, optionally the previously decoded threshold being from a hierarchical higher level than said header.

[0023] Optionally, the predicted threshold is predicted based on a syntax element for the same header, advantageously the syntax element being indicative of the reconstructed quality of the image portion.

[0024] Optionally, the syntax element is a quantization parameter or a quantization parameter offset, which is already known by the decoder and is likely to be correlated to a threshold value.

[0025] Optionally, the method includes (a) decoding a flag indicating that the threshold was correctly predicted or (b) checking whether the threshold is associated with a syntax element for the same header. For increased flexibility, the method may include decoding a flag indicating whether to use option (a) or (b).

[0026] For ease of computation, the association between syntax elements and thresholds for the same header may be determined from a table.

[0027] Optionally, to reduce the bit rate, the tables are transmitted at a hierarchical higher level than the headers.

[0028] Optionally, for added flexibility, each syntax element may be transmitted with its associated threshold to generate a table, which is advantageously predictively coded.

[0029] Optionally, to reduce the bitrate, the tables are generated without duplicates.

[0030] Optionally, the table is used to determine the block-level threshold when the syntax element is defined at the block level.

[0031] Indexed threshold Optionally, determining the threshold value includes decoding an index indicative of the threshold value from among multiple threshold values ​​known by the decoder, which allows for signaling with a reduced number of bits.

[0032] Optionally, the method includes receiving a plurality of thresholds Optionally, the plurality of thresholds are transmitted at a hierarchically higher level than the header.

[0033] To allow for efficient decoding, the method may include decoding a variable indicative of the number of thresholds received.

[0034] To reduce the bit rate, the thresholds can be transmitted in a header at the multi-image portion level, optionally the multi-image portion level header being a sequence parameter set header.

[0035] For added flexibility, one value of the index may indicate that a threshold not in a plurality of thresholds should be decoded from the header. Optionally, the value of the index is a maximum value.

[0036] Optionally, the method further comprises adding the threshold value decoded from the header to the plurality of threshold values ​​and associating it with an index.

[0037] According to another aspect of the invention, there is provided a method for encoding a header relating to a plurality of image portions, the method comprising: determining a threshold for each image portion, the threshold being used to determine a decoding criterion for said image portion; generating an index of the thresholds, the indexes being ordered based on a frequency of occurrence of the thresholds; and encoding the index into the header. Such a method results in the most probable threshold being encoded with the smallest index, leading to rate reduction.

[0038] Optionally, the header is a sequence parameter set (SPS).

[0039] In another aspect according to the invention, there is provided a method of encoding a portion of an image, the method comprising determining a decoding criterion for the image portion, determining a threshold value used to determine the criterion, signaling the threshold value in a header associated with the image portion, and encoding the image portion based on the criterion.

[0040] According to an aspect of the present invention there is provided an apparatus for encoding image data into a bitstream, said apparatus configured to carry out any of the aspects or embodiments described above.

[0041] According to an aspect of the present invention there is provided an apparatus for decoding image data from a bitstream, said apparatus configured to perform any of the aspects or embodiments described above.

[0042] In another aspect according to the invention there is provided a (computer) program which, when executed, causes a programmable device to carry out a method according to any aspect or embodiment above. The program may be stored in a computer readable storage medium.

[0043] The program may be provided by itself or may be carried on, by or in a carrier medium. The carrier medium may be non-transitory, e.g. a storage medium, in particular a computer-readable storage medium. The carrier medium may also be transitory, e.g. a signal or other transmission medium. The signal may be transmitted over any suitable network, including the Internet.

[0044] Further features of the invention are characterized by the independent and dependent claims. Any feature in one aspect of the invention may be applied to other aspects of the invention in any appropriate combination. In particular, method aspects may be applied to apparatus aspects and vice versa. Furthermore, hardware implemented features may be implemented in software and vice versa. Any references to software and hardware features herein should be interpreted accordingly. Any apparatus feature described herein may be provided as a method feature and vice versa.

[0045] As used herein, means-plus-function features may alternatively be expressed in terms of their corresponding structure, such as a suitably programmed processor and associated memory.

[0046] It is also to be understood that specific combinations of the various features described and defined in any embodiment of the present invention can be implemented and / or provided and / or used independently. [Brief description of the drawings]

[0047] Reference is now made, by way of example, to the accompanying drawings in which: [Figure 1] FIG. 1 is a diagram for explaining the coding structure used in HEVC. [Diagram 2] FIG. 2 is a block diagram that illustrates generally a data communications system in which one or more embodiments of the present invention may be implemented. [Diagram 3]FIG. 3 is a block diagram illustrating components of a processing device in which one or more embodiments of the present invention may be implemented. [Figure 4] FIG. 4 is a flow chart illustrating steps of an encoding method according to an embodiment of the invention. [Diagram 5] FIG. 5 is a flow chart illustrating steps of a decoding method according to an embodiment of the invention. [Figure 6] FIG. 6 shows the labeling scheme used to describe blocks located relative to a current block. [Figure 7] FIG. 7 shows the labeling scheme used to describe blocks located relative to a current block. [Figure 8] 8(a) and (b) show the affine (sub-block) mode. [Figure 9] 9(a), (b), (c), and (d) show the geometric modes. [Figure 10] FIG. 10 shows the first step of deriving a merge candidate list for VVC. [Figure 11] FIG. 11 illustrates a further step in the derivation of a merge candidate list for VVC. [Figure 12] FIG. 12 shows the derivation of pairwise candidates. [Figure 13] FIG. 13 shows a template matching method based on adjacent samples. [Figure 14] FIG. 14 shows a variation of the first step of deriving the merge candidate list shown in FIG. [Figure 15] FIG. 15 illustrates a further step modification of the merge candidate list derivation shown in FIG. [Figure 16] FIG. 16 shows a modification of the derivation of pairwise candidates shown in FIG. [Figure 17] FIG. 17 illustrates the cost determination of list candidates. [Figure 18] FIG. 18 shows the process of sorting the list of merge mode candidates. [Figure 19]FIG. 19 illustrates pair-wise candidate derivation during the process of reordering the list of merge mode candidates. [Figure 20] FIG. 20 illustrates the process of ordering predictors based on distortion values. [Figure 21] FIG. 21 shows the structure of a bitstream in an exemplary coding system VVC. [Figure 22] FIG. 22 shows a system comprising an encoder or decoder and a communication network according to an embodiment of the invention. [Diagram 23] FIG. 23 is a schematic block diagram of a computing device for implementation of one or more embodiments of the present invention. [Figure 24] FIG. 24 is a diagram showing a network camera system. [Diagram 25] FIG. 25 is a diagram showing a smartphone. DETAILED DESCRIPTION OF THE PREFERRED EMBODIMENTS

[0048] Figure 1 relates to the coding structure used in the High Efficiency Video Coding (HEVC) video and Versatile Video Coding (VVC) standards. A video sequence 1 is composed of a sequence of digital images i. Each such digital image is represented by one or more matrices. The matrix coefficients represent pixels.

[0049] An image 2 of a sequence can be divided into slices 3, which in some examples may constitute the entire image. These slices are divided into non-overlapping coding tree units (CTUs). A coding tree unit (CTU) is a basic processing unit of the High Efficiency Video Coding (HEVC) video standard and conceptually corresponds structurally to the macroblock unit used in some previous video standards. A CTU is sometimes called a largest coding unit (LCU). A CTU has a luma component part and a chroma component part, each of which is called a coding tree block (CTB). These different color components are not shown in FIG. 1.

[0050] A CTU is generally sized 64 pixels by 64 pixels for HEVC, but for VVC, the size may be 128 pixels by 128 pixels. Each CTU may be iteratively divided into smaller variable-size coding units (CUs) 5 using a quadtree decomposition.

[0051] A coding unit is a basic coding element and consists of two types of subunits called prediction units (PUs) and transform units (TUs). The maximum size of a PU or TU is equal to the CU size. A prediction unit corresponds to a partition of a CU for prediction of pixel values. Various different partitions of a CU into PUs are possible as shown by 606, including a partition into four rectangular PUs and two different partitions into two rectangular PUs. A transform unit is a basic unit that undergoes spatial transformation using DCT. A CU can be partitioned into TUs based on a quadtree representation 607.

[0052] Each slice is embedded in one network abstraction layer (NAL) unit. Furthermore, the coding parameters of a video sequence are stored in a dedicated NAL unit called a parameter set. In HEVC and H.264 / AVC, two types of parameter set NAL units are used: first, the sequence parameter set (SPS) NAL unit, which collects all parameters that do not change during the entire video sequence. Typically, it handles the coding profile, the size of the video frames, and other parameters. Second, the picture parameter set (PPS) NAL unit contains parameters that can change from one picture (or frame) of the sequence to another picture (or frame). HEVC also includes the video parameter set (VPS) NAL unit, which contains parameters that describe the overall structure of the bitstream. VPS is a type of parameter set defined in HEVC and applies to all layers of the bitstream. A layer can contain multiple temporal sublayers, and all version 1 bitstreams are limited to a single layer. HEVC has certain layered extensions for scalability and multiview, which allow multiple layers with a backward-compatible version 1 base layer.

[0053] Another method of dividing an image has been introduced in VVC, which includes sub-pictures, which are independently coded groups of one or more slices.

[0054] 2 illustrates a data communication system in which one or more embodiments of the present invention may be implemented. The data communication system comprises a transmitting device, here a server 201, operable to transmit data packets of a data stream to a receiving device, here a client terminal 202, via a data communication network 200. The data communication network 200 may be a wide area network (WAN) or a local area network (LAN). Such a network may for example be a wireless network (Wifi / 802.11a or b or g), an Ethernet network, an Internet network, or a mixed network made up of several different networks. In a particular embodiment of the present invention, the data communication system may be a digital television broadcasting system in which a server 201 transmits the same data content to several clients.

[0055] The data stream 204 provided by the server 201 may be composed of multimedia data representing video and audio data. The audio and video data streams may, in some embodiments of the invention, be captured by the server 201 using a microphone and a camera, respectively. In some embodiments, the data streams may be stored in the server 201, or may be received by the server 201 from another data provider, or may be generated at the server 201. The server 201 is particularly provided with an encoder for encoding the video and audio streams to provide a compressed bitstream for transmission with a more compact representation of the data presented as input to the encoder.

[0056] In order to obtain a better ratio between the quality of the transmitted data and the amount of transmitted data, the compression of the video data may for example be according to the HEVC format or the H.264 / AVC format or the VVC format.

[0057] Client 202 receives the transmitted bitstream and decodes the reconstructed bitstream to reproduce the video images on a display device and the audio data through a loudspeaker.

[0058] In the example of FIG. 2, a streaming scenario is considered, but it will be understood that in some embodiments of the invention, data communication between the encoder and the decoder may be performed using a media storage device, such as, for example, an optical disc.

[0059] In one or more embodiments of the invention, a video image is transmitted along with data representing a compensation offset to be applied to reconstructed pixels of the image to provide filtered pixels in the final image.

[0060] 3 shows a schematic diagram of a processing device 300 configured to implement at least one embodiment of the present invention. The processing device 300 may be a device such as a microcomputer, a workstation, or a light handheld device. The device 300 may include:

[0061] - a central processing unit 311, such as a microprocessor, designated CPU; - a read-only memory 306, denoted ROM, for storing a computer program for implementing the invention; a random access memory 312, denoted RAM, storing the executable code of the method of an embodiment of the invention, as well as registers adapted to record variables and parameters necessary for implementing the method of encoding a sequence of digital images and / or the method of decoding a bitstream according to an embodiment of the invention; A communications interface 302 connected to a communications network 303 over which digital data to be processed is transmitted and received. A communication bus 313 is connected to the Optionally, the device 300 may also include the following components: - data storage means 304, such as a hard disk, for storing computer programs for implementing the methods of one or more embodiments of the invention, and data used or generated during the implementation of one or more embodiments of the invention; a disk drive 305 for a disk 306, the disk drive being adapted to read data from the disk 306 or to write data to said disk; A screen 309 for serving as a graphical interface with the user, by means of a keyboard 310 or any other pointing means, and / or for displaying data.

[0062] The device 300 may be connected to a variety of peripheral devices, such as, for example, a digital camera 320 or a microphone 308 , each connected to an input / output card (not shown) to provide multimedia data to the device 300 .

[0063] The communication bus provides communication and interoperability between the various elements included in or connected to the device 300. The representation of a bus is not limiting, in particular a central processing unit is operable to communicate instructions directly to any element of the device 300 or by another element of the device 300.

[0064] The disk 306 may be replaced by any information carrier, for example a rewritable or non-rewritable compact disk (CD-ROM), a ZIP disk or a memory card, and generally by an information storage means readable by a microcomputer or microprocessor, integrated or not integrated into the device, or removable, and adapted to store one or more programs, the execution of which enables the method of encoding a sequence of digital images and / or the method of decoding a bitstream according to the present invention to be carried out.

[0065] The executable code may be stored either in the read-only memory 306, in the hard disk 304 or on a removable digital medium as previously mentioned, such as for example the disk 306. According to a variant, the executable code of the program may be received by the communication network 303, via the interface 302, to be stored in one of the storage means of the device 300 before being executed, such as the hard disk 304.

[0066] The central processing unit 311 is adapted to control and direct the execution of instructions or parts of the software code of the program or programs according to the invention with instructions stored in one of the above mentioned storage means. On power-up, the program or programs stored in a non-volatile memory, for example the hard disk 304 or the read-only memory 306, are transferred to the random access memory 312, which contains the executable code of the program or programs, as well as registers for storing variables and parameters necessary for implementing the invention.

[0067] In this embodiment, the device is a programmable device that uses software to implement the invention, but the invention may alternatively be implemented in hardware (e.g. in the form of an application specific integrated circuit or ASIC).

[0068] 4 shows a block diagram of an encoder according to at least one embodiment of the present invention, the encoder being represented by connected modules, each module adapted to be implemented, for example, in the form of program instructions executed by the CPU 311 of the device 300, such that at least one corresponding step of a method for encoding an image in a sequence of images according to one or more embodiments of the present invention.

[0069] At 401, an original sequence of digital images i0~ is received as input by an encoder 400. Each digital image is represented by a set of samples, sometimes also called pixels (hereafter referred to as pixels).

[0070] A bitstream 410 is output by the encoder 400 after the encoding process is performed. The bitstream 410 includes multiple coding units or slices, each of which includes a slice header for transmitting coded values ​​of coding parameters used to code the slice, and the slice body includes coded video data.

[0071] An input digital image i0~ in 401 is divided by a module 402 into blocks of pixels. The blocks correspond to image portions and can be of variable size (for example 4x4, 8x8, 16x16, 32x32, 64x64, 128x128 pixels, and also several rectangular block sizes can be considered). For each input block a coding mode is selected. Two families of coding modes are provided: coding modes based on spatial predictive coding (intra prediction) and coding modes based on temporal prediction (inter coding, merge, SKIP). The possible coding modes are tested.

[0072] The module 403 implements an intra prediction process in which a given block to be coded is predicted by a predictor calculated from neighbouring pixels of said block to be coded. If intra coding is selected, the selected intra predictor and an indication of the difference between the given block and its predictor are coded to provide a residual.

[0073] Temporal prediction is performed by the motion estimation module 404 and the motion compensation module 405. First, a reference image is selected from a set of reference images 416, and a part of the reference image, also called a reference area or image part, which is the closest area (closest in terms of pixel value similarity) to a given block to be coded, is selected by the motion estimation module 404. Then, the motion compensation module 405 uses the selected area to predict the block to be coded. The difference between the selected reference area and the given block, also called a residual block, is calculated by the motion compensation module 405. The selected reference area is indicated by a motion vector.

[0074] Therefore, in both cases (spatial and temporal prediction), the residual is calculated by subtracting the predictor from the original block.

[0075] In the INTRA prediction implemented by module 403, the prediction direction is coded. In the INTER prediction implemented by modules 404, 405, 416, 418, 417, at least one motion vector or data for identifying such a motion vector is coded for temporal prediction.

[0076] If inter prediction is selected, the motion vector and information related to the residual block are coded. To further reduce the bit rate, assuming that the motion is uniform, the motion vector is coded by a difference to the motion vector predictor. The motion vector predictor from a set of motion information predictor candidates is obtained from the motion vector field 418 by the motion vector predictive coding module 417.

[0077] The encoder 400 further comprises a selection module 406 for selecting an encoding mode by applying an encoding cost criterion, such as a rate-distortion criterion. To further reduce redundancy, a transform (such as a DCT) is applied to the residual block by a transform module 407, and the resulting transformed data is then quantized by a quantization module 408 and entropy coded by an entropy coding module 409. Finally, the coded residual block of the current block being coded is inserted into a bitstream 410.

[0078] The encoder 400 also performs decoding of the encoded image to generate reference images (e.g., reference images in the reference images / pictures 416) for motion estimation of subsequent images. This allows the encoder and the decoder receiving the bitstream to have the same reference frame (reconstructed image or image portion is used). The inverse quantization ("inverse quantization") module 411 performs inverse quantization ("inverse quantization") of the quantized data, followed by inverse transformation by the inverse transformation module 412. The intra prediction module 413 uses the prediction information to decide which predictor should be used for a given block, and the motion compensation module 414 actually adds the residual obtained by module 412 to the reference area obtained from the set of reference images 416.

[0079] Then, post-filtering is applied by module 415 to filter the reconstructed frame of pixels (image or image portion). In an embodiment of the present invention, an SAO loop filter is used, and a compensation offset is added to the pixel values ​​of the reconstructed pixels of the reconstructed image. It is understood that post-filtering does not necessarily have to be performed. Also, any other type of post-filtering may be performed in addition to or instead of SAO loop filtering.

[0080] 5 shows a block diagram of a decoder 60 that may be used to receive data from an encoder according to an embodiment of the present invention. The decoder is represented by connected modules, each module adapted to carry out a corresponding step of a method implemented by the decoder 60, for example in the form of program instructions executed by the CPU 311 of the device 300.

[0081] The decoder 60 receives a bitstream 61 containing coded units (e.g. data corresponding to blocks or coding units), each unit consisting of a header containing information about coding parameters and a body containing the coded video data. As explained with respect to Fig. 4, the coded video data is entropy coded, and an index of a motion vector predictor is coded with a predefined number of bits for a given block. The received coded video data is entropy decoded by module 62. The residual data is then inverse quantized by module 63 and then an inverse transform is applied by module 64 to obtain pixel values.

[0082] Mode data indicating the coding mode is also entropy decoded, and based on the mode, INTRA type decoding or INTER type decoding is performed on the coded block (unit / set / group) of image data.

[0083] For INTRA mode, the INTRA predictor is determined by the intra prediction module 65 based on the intra prediction mode specified in the bitstream.

[0084] If the mode is INTER, motion prediction information is extracted from the bitstream to find (identify) the reference area used by the encoder. The motion prediction information includes a reference frame index and a motion vector residual. The motion vector predictor is added to the motion vector residual by the motion vector decoding module 70 to obtain a motion vector. The various motion prediction tools used in VVC are described in more detail below with reference to Figures 6 to 10.

[0085] A motion vector decoding module 70 applies motion vector decoding for each current block coded by motion prediction. Once the motion vector predictor index for the current block is obtained, the actual value of the motion vector associated with the current block may be decoded and used to apply motion compensation by module 66. The reference image portion to which the decoded motion vector points is extracted from reference image 68 and motion compensation 66 is applied. Motion vector field data 71 is updated with the decoded motion vector for use in predicting subsequent decoded motion vectors.

[0086] Finally, a decoded block is obtained. If appropriate, post-filtering is applied by a post-filtering module 67. A decoded video signal 69 is finally obtained and is supplied by the decoder 60.

[0087] FIG. 21 shows the structure of a bitstream in an exemplary coding system VVC as described in JVET-Q2001-vD.

[0088] A bitstream 2100 from a VVC coding system consists of an ordered sequence of syntax elements and coded data. The syntax elements and coded data are arranged in Network Abstraction Layer (NAL) units 2101-2108. There are different NAL unit types. The Network Abstraction Layer provides the ability to encapsulate the bitstream into different protocols such as RTP / IP, which stands for Real Time Protocol / Internet Protocol, ISO Base Media File Format, etc. The Network Abstraction Layer also provides a framework for packet loss resiliency.

[0089] The NAL units are divided into video coding layer (VCL) NAL units and non-VCL NAL units. The VCL NAL units contain the actual coded video data. The non-VCL NAL units contain additional information. This additional information may be parameters required for decoding the coded video data or supplemental data that can increase the usefulness of the decoded video data. The NAL units 2106 correspond to slices and constitute the VCL NAL units of the bitstream.

[0090] Different NAL units 2101-2105 correspond to different parameter sets, and these NAL units are non-VCL NAL units. The decoder parameter set (DPS) NAL unit 301 contains parameters that are constant for a given decoding process. The video parameter set (VPS) NAL unit 2102 contains parameters defined for the entire video, and therefore the entire bitstream. The DPS NAL unit may define more static parameters than those in the VPS. In other words, the parameters of the DPS change less frequently than the parameters of the VPS.

[0091] The sequence parameter set (SPS) NAL unit 2103 contains parameters defined for a video sequence. In particular, the SPS NAL unit can define the sub-picture layout and associated parameters of a video sequence. The parameters associated with each sub-picture specify the coding constraints that apply to the sub-picture. In particular, it comprises a flag that indicates that temporal prediction between sub-pictures is restricted to data coming from the same sub-picture. Another flag can enable or disable the loop filter across sub-picture boundaries.

[0092] The picture parameter set (PPS) NAL unit 2104, PPS contains parameters defined for a picture or a group of pictures. The adaptation parameter set (APS) NAL unit 2105 contains parameters for the loop filter of the adaptive loop filter (ALF) or reshaper model (or luma mapping with chroma scaling (LMCS) model) or scaling matrix, typically used at slice level.

[0093] The PPS syntax proposed in the current version of VVC includes syntax elements that specify the size of pictures in luma samples and the partitioning of each picture into tiles and slices.

[0094] The PPS contains syntax elements that allow to determine the slice location within a frame. Since a subpicture forms a rectangular region within a frame, it is possible to determine from the parameter set NAL unit the set of slices, parts of tiles, or tiles that belong to a subpicture. Similar to the APS, the PPS has an ID mechanism that limits the amount of transmission of the same PPS.

[0095] The main difference between the PPS and the picture header is that the PPS is transmitted, and generally, for a group of pictures, as opposed to the PH, which is transmitted systematically for each picture. Thus, in comparison to the PH, the PPS contains parameters that may be constant for several pictures.

[0096] The bitstream may also contain supplemental enhancement information (SEI) NAL units (not shown in FIG. 21). The periodicity of occurrence of these parameter sets in the bitstream is variable. A VPS defined for the entire bitstream can occur only once in the bitstream. Conversely, an APS defined for a slice may occur once for each slice in each picture. In fact, different slices may depend on the same APS, and thus, in general, there are fewer APSs than slices in each picture. In particular, the APSs are defined in the picture header. However, the ALF APSs can be refined in the slice header.

[0097] An access unit delimiter (AUD) NAL unit 2107 separates the two access units. An access unit is a set of NAL units that may contain one or more coded pictures with the same decoding timestamp. This arbitrary NAL unit contains only one syntax element in the current VVC specification: pic_type, which indicates the slice_type values ​​of all slices of the coded pictures in the AU. If pic_type is set equal to 0, the AU contains only intra slices. If it is equal to 1, it contains P and I slices. If it is equal to 2, it contains B, P or intra slices. This NAL unit contains only one syntax element of pic_type.

[0098] Reference Picture List In VVC there is the possibility to transmit a set of reference picture lists in the SPS or in the picture or slice header. The principle is to identify the POC difference between the current picture and its reference frames. Then, generally, only one set needs to be transmitted, since several frames of a GOP share the same POC difference between their reference frames. When a set has already been transmitted, it can be transmitted together with a rate-preserving index dedicated to this reference picture signaling.

[0099] Motion Estimation (INTER) Mode HEVC uses three different INTER modes: Inter mode (Advanced Motion Vector Prediction (AMVP)), "Classical" merge mode (i.e., also known as "Non-Affine Merge Mode" or "Negular" Merge Mode), and "Classical" merge skip mode (i.e., also known as "Non-Affine Merge Skip" mode or "Normal" Merge Skip Mode). The main difference between these modes is the data signaling in the bitstream. For motion vector coding, the current HEVC specification includes a contention-based scheme for motion vector prediction, which was not present in previous versions of the specification. It means that several candidates are competing with a rate-distortion criterion on the encoder side to find the best motion vector predictor or best motion information for the Inter mode or merge mode (i.e., "Classical / Regular" Merge Mode or "Classical / Regular" Merge Skip Mode), respectively. Then, an index corresponding to the best predictor or best candidate for the motion information is inserted into the bitstream along with a "residual" that represents the difference between the predicted value and the actual value. The decoder can derive the same set of predictors or candidates and use the best one according to the decoded index. Using the residual, the decoder can then recreate the original value.

[0100] In the HEVC Screen Content Extension, a new coding tool called Intra Block Copy (IBC) is signaled as one of those three INTER modes, and the difference between IBC and the equivalent INTER mode is done by checking if the reference frame is the current one. This can be done, for example, by checking the reference index of list L0 and if this is the last frame in that list, then it is presumed to be an intra block copy. Another way is to compare the picture order counts of the current and reference frames, and if they are equal, then it is an intra block copy.

[0101] The design of predictor and candidate derivation is important to achieve the best coding efficiency without disproportionately affecting the complexity. In HEVC, two motion vector derivations are used: one for inter mode (Advanced Motion Vector Prediction (AMVP)) and one for merge mode (merge derivation process-classical merge mode and classical merge skip mode). The following describes the various motion predictor modes used in VVC.

[0102] FIG. 6 illustrates the labeling scheme used herein to describe blocks located relative to a current block (ie, the block currently being coded / decoded) between frames (FIG. 6).

[0103] VVC Merge Mode In VVC, some inter modes have been added compared to HEVC, in particular new merge modes have been added to the usual merge modes of HEVC.

[0104] Affine mode (sub-block mode) In HEVC, for motion compensated prediction (MCP), only the translation motion model is applied. In the real world, there are many kinds of motions, such as zoom in / out, rotation, perspective motion, and other irregular motions.

[0105] In JEM a simplified affine transformation motion compensation prediction is applied and the general principles of the affine mode are explained below based on an extract from document JVET-G1001 presented at the JVET Conference in Turin, 13-21 July 2017, which is incorporated herein in its entirety by reference insofar as it describes other algorithms used in JEM.

[0106] As shown in FIG. 8(a), the affine motion field of a block is described by two control point motion vectors.

[0107] Affine mode is a motion compensation mode like inter mode (AMVP, "classical" merge, or "classical" merge skip). Its principle is to generate one motion information per pixel according to two or three neighboring motion information. In JEM, affine mode derives one motion information per 4x4 block as shown in Figure 8(a) (each square is a 4x4 block, and the whole block in Figure 8(a) is a 16x16 block, which is divided into 16 such square blocks of 4x4 size, and each 4x4 square block has a motion vector associated with it). Affine mode is available for AMVP mode and merge mode (i.e., conventional merge mode, also called "non-affine merge mode", and conventional merge skip mode, also called "non-affine merge skip mode") by enabling the flagged affine mode.

[0108] In the VVC specification, affine mode is also known as sub-block mode, and these terms are used interchangeably herein.

[0109] The sub-block merging mode of VVC includes a sub-block based temporal merging candidate that inherits the motion vector field of the block in the previous frame pointed to by the spatial motion vector candidate. If the neighboring block is coded in the inter-affine mode of sub-block merging, this sub-block candidate is followed by the inherited affine motion candidate, and then some of the constructed affine candidates are derived before some zero Mv candidates.

[0110] CIIP In addition to the normal merge mode and the sub-block merge mode, the VVC standard also includes a Combined Inter Merge / Intra Prediction (CIIP) merge mode, also known as the Multi-Hypothesis Intra Inter (MHII) merge mode.

[0111] The combined inter-merge / intra-prediction (CIIP) merge can be considered as a combination of the normal merge mode and the intra mode, and is described below with reference to FIG. 10. The block predictor of the current block (1001) in this mode is the average between the merge predictor block and the intra predictor block, as shown in FIG. 10. The merge predictor block is obtained in exactly the same process as in the merge mode, and is therefore a temporal block (1002) of two temporal blocks or a bi-predictor. Therefore, the merge index is signaled for this mode in the same way as in the normal merge mode. The intra predictor block is obtained based on the neighboring samples (1003) of the current block (1001). However, the amount of intra modes available for the current block is limited compared to the intra blocks. Furthermore, there is no chroma intra predictor block signaled for the CIIP block. The chroma predictor is equal to the luma predictor. As a result, 1, 2, or 3 bits are used to signal the intra predictor for the CIIP block.

[0112] The CIIP block predictor is obtained by a weighted average of the merge block predictor and the intra block predictor, where the weighting of the weighted average depends on the selected intra predictor block and / or block size.

[0113] The obtained CIIP predictor is then added to the residual of the current block to obtain the reconstructed block. Note that the CIIP mode is only valid for non-skipped blocks. In fact, the use of CIIP skip typically results in a loss of compression performance and an increase in encoder complexity. This is because the CIIP mode often has a block residual, contrary to other skip modes. As a result, its signaling for skip mode increases the bitrate. -CIIP is avoided when the current CU is skip. The consequence of this restriction is that a CIIP block cannot have a residual that contains only 0 values, since it is not possible to code a VVC block residual equal to 0. In fact, in VVC the only way to signal a block residual equal to 0 for merge mode is to use skip mode, since the CU CBF flag is inferred to be equal to true for merge mode. And when this CBF flag is true, the block residual cannot be equal to 0.

[0114] As such, CIIP should be construed herein as a mode that combines features of inter-prediction and intra-prediction, and not necessarily a label given to one particular mode.

[0115] CIIP used the same motion vector candidate list as the normal merge mode.

[0116] MMVD MMVD MERGE modes are derivations of certain normal merge mode candidates. It can be seen as an independent merge candidate list. The selected MMVD merge candidate for the current CU is obtained by adding an offset value to the motion vector component (mvx or mvy) of one of the initial normal merge candidates. The offset value is added to the motion vector of the first list L0 or to the motion vector of the second list L1 depending on the configuration of these reference frames (backward, forward or both forward and backward). The initial merge candidate is signaled by an index. The offset value is signaled by a distance index between eight possible distances (1 / 4-pel, 1 / 2-pel, 1-pel, 2-pel, 4-pel, 8-pel, 16-pel, 32-pel) and a direction index giving the x or y axis and the sign of the offset.

[0117] In VVC, typically only the first two candidates in the merge list are used for MMVD derivation and signaling via one flag.

[0118] Geometric Division Mode The geometric (GEO) MERGE mode is a specific bi-predictive mode. Figure 9 illustrates this specific block predictor generation. The block predictor includes one triangle from the first block predictor (901 or 911) and a second triangle from the second block predictor (902 or 912). However, several other possible splits of the block are possible, as shown in Figures 9(c) and 9(d). Geometric merge should be interpreted herein as a mode that combines features of two inter non-rectangular predictors, and not necessarily as a label given to one specific mode.

[0119] In the example of Figure 9(a), each partition (901 or 902) has a motion vector candidate that is a unidirectional candidate. And for each partition, an index is signaled to obtain the corresponding motion vector candidate in the list of unidirectional candidates at the decoder. Also, the first and second cannot use the same candidate. This list of candidates comes from the normal merge candidate list, with one of the two components (L0 or L1) removed for each candidate.

[0120] IBC VVC also allows the enabling of intra-block copy (IBC) merge mode, which has an independent merge candidate derivation process.

[0121] Other movement information improvements DMVR Decoder-side motion vector derivation (DMVR) improves the accuracy of MV for merge mode in VVC. For this method, bilateral matching (BM) based decoder-side motion vector refinement is applied. In this bi-predictive operation, refined MV is searched around the initial MV in reference picture list L0 and reference picture list L1. The BM method calculates the distortion between two candidate blocks in reference picture list L0 and list L1.

[0122] The DMVR method has several steps for motion refinement, which are related to several sub-pixel steps or sub-block divisions in ECM. One particularity of VVC DMVR is that the first cost is calculated between two initial blocks from L0 and L1. When the sum of absolute differences (SAD) is less than the number of samples of the current block, no refinement step is applied. This criterion is fixed and not adaptive.

[0123] BDOF VVC also integrates a bidirectional optical flow (BDOF) tool. BDOF, previously called BIO, is used to refine the bi-predictive signal of a CU at the 4x4 subblock level. BDOF is applied to a CU if it meets some conditions, in particular, if the distances (i.e., picture order count (POC) difference) from two reference pictures to the current picture are the same. As the name suggests, the BDOF mode is based on the optical flow concept, which assumes that object motion is smooth. For each 4x4 subblock, a motion refinement (v_x, v_y) is calculated by minimizing the difference between the L0 predicted sample and the L1 predicted sample. The motion refinement is then used to adjust the bi-predictive sample values ​​within the 4x4 subblock.

[0124] PROF Similarly, for the affine mode, prediction refinement by optical flow (PROF) is used.

[0125] AMVR and hpelIfIdx VVC also includes Adaptive Motion Vector Resolution (AMVR). AMVR allows the motion vector differentials of a CU to be coded with different precision. For example, for AMVP mode, quarter-luma samples, half-luma samples, integer-luma samples, or 4-luma samples are considered. The following table from the VVC specification gives the AMVR shifts based on different syntax elements:

[0126] [Table 1]

[0127] AMVR may affect the coding of modes other than those using motion vector differential coding as different merge modes. Indeed, for some candidates, the parameter hpelIfIdx, which represents an index on the luma interpolation filter for half-pel precision, is propagated for some merge candidates. For the AMVP mode, for example, hpelIfIdx is derived as follows: hpelIfIdx = AmvrShift == 3?1:0 Bi-prediction with CU-level weights (BCW) In VVC, the bi-prediction mode with CU level weighting (BCW) is extended beyond simple averaging (as performed in HEVC) to allow weighted averaging of two prediction signals P0 and P1, according to the following equation:

[0128] P bi-pred =((8-w)*P0+w*P1+4)>>3 Five weights are allowed in weighted average bi-prediction, where w∈{-2, 3, 4, 5, 10}.

[0129] For non-merged CUs, the weight index bcwIndex is signaled after the motion vector differential.

[0130] For a merged CU, the weight index is inferred from the neighboring blocks based on the merge candidate indexes.

[0131] BCW is used only for CUs with 256 or more luma samples. Furthermore, for low latency pictures, all five weights are used. For non-low latency pictures, only three weights (w∈{3,4,5}) are used.

[0132] Canonical merge list derivation In VVC, a regular merge list is derived as shown in Figures 10 and 11. First, spatial candidates B1 (1002), A1 (1006), B0 (1010), and A0 (1014) (shown in Figure 7) are added if they exist. Then, a partial redundancy check is performed to add A1 (1008) between A1 and B1 motion information (1007), B0 (1012) between B0 and B1 motion information (1011), and A0 (1016) between A0 and A1 motion information (1015).

[0133] When a merge candidate is added, the variable cnt is incremented (1015, 1009, 1013, 1017, 1023, 1027, 1115, 1108).

[0134] If the number of candidates in the list (cnt) is strictly less than four (1018), candidate B2 (1019) is added (1022) if it does not have the same motion information as A1 and B1 (1021).

[0135] Next, the time candidates are added: the bottom right candidate (1024) is added (1026) if it is available (1025), otherwise the center time candidate (1028) is added (1026) if it exists (1029).

[0136] Then, history-based (HMVP) candidates are added (1101) if they do not have the same motion information as A1 and B1 (1103). In addition, the number of history-based candidates cannot exceed the maximum number of candidates in the merge candidate list minus one (1102). Thus, after the history-based candidates, there is at least one position missing in the merge candidate list.

[0137] Then, if the number of candidates in the list is at least two, pairwise candidates are constructed (1106) and added to the merge candidate list (1107).

[0138] Next, if there are empty positions in the merge candidate list (1109), zero candidates are added (1110).

[0139] For spatial-based and history-based candidates, the parameters BCWidx and useAltHpelIf are set equal to the candidate's associated parameters. For temporal and zero candidates, they are set equal to the default value 0. These default values ​​essentially disable the method.

[0140] For a pairwise candidate, BCXidx is set equal to 0, and hpelIfIdxp is set equal to the hpelIfIdxp of the first candidate if it is equal to the hpelIfIdxp of the second candidate, otherwise it is set to 0.

[0141] Pairwise candidate derivation Pairwise candidates are constructed (1106) according to the algorithm of FIG. 12. As shown, when there are two candidates in the list (1201), hpelIfIdxp is derived as described above (1204, 1202, 1203). Then, the inter direction (interDir) is set equal to 0 (1205). If at least one reference frame is valid (different from -1) for each list L0 and L1 (1207), the parameters are set. If both are valid (1208), the mv information for this candidate is derived (1209) and set equal to the reference frame of the first candidate, the motion information is the average between the two motion vectors for this list, and the variable interDir is incremented. If only one of the candidates has motion information for this list (1210), the motion information for the pairwise candidate is set equal to this candidate (1212, 1211) and the inter direction variable interDir is incremented.

[0142] ECM Since the standardization of VVC v1 was finished, JVET started the exploration phase by establishing the Exploration Software (ECM), which collects additional tools and improvements of existing tools in addition to the VVC standard to target better coding efficiency. The different additional tools compared to VVC are described in JVET-X2025.

[0143] ECM Merge Mode Among all the tools added, some additional merge modes have been added: Affine MMVD signal offsets for merging affine candidates as MVVD encoding in normal merge mode. Similarly, we also added GEO MMVD. CIIP PDPC is an extension of CIIP. Also, two template matching merge modes have been added: normal template matching and GEO template matching.

[0144] Conventional template matching is based on template matching estimation as shown in Fig. 13. At the decoder side, for the candidate corresponding to the relevant merge index and for both available lists (L0, L1), a motion estimation based on neighboring samples of the current block (1301) and neighboring samples of multiple corresponding block positions is performed, a cost is calculated, and the motion information that minimizes the cost is selected. The motion estimation is limited by a search range, and some restrictions on this search range are also used to reduce the complexity.

[0145] In ECM, the canonical template matching candidate list is based on the canonical merge list, but some additional steps and parameters are added, which means that different merge candidate lists for the same block can be generated. Furthermore, there are only 4 candidates available for template matching in the canonical merge candidate list, compared to the 10 candidates in the canonical merge candidate list in ECM with the common test conditions defined by JVET.

[0146] Canonical merge list derivation in ECM In ECM, the regular merge list derivation has been updated. Figures 14 and 15 show this update based on Figures 10 and 11, respectively. However, for clarity, the module for history-based candidates (1101) is summarized in (1501).

[0147] In this FIG. 15, a new type of merge candidate is added: non-adjacent candidates (1540). These candidates come from blocks that are spatially located in the current frame, but not from neighboring blocks, as neighboring blocks are spatial candidates. They are selected according to distance and direction. For the history base, a list of neighboring candidates can be added so that pairwise can still be added until the list reaches the maximum number of candidates minus 1.

[0148] Zero Candidate If the list still does not reach the maximum number of candidates (Maxcand), zero candidates are added to the list. The zero candidates are added according to the possible reference frames or pairs of reference frames. The following pseudocode gives the derivation of such candidates:

[0149] [Table 2]

[0150] This pseudocode can be summarized as follows: for each reference frame index (unidirectional) or pair of reference indexes (bi-predictive), a zero candidate is added. Once all are added, only the zero candidate with reference frame index 0 is added until the number of candidates reaches its maximum value. In this way, the merge list can contain multiple zero candidates. In fact, it has been surprisingly found that this occurs frequently in real video sequences, especially at the beginning of slices or frames and sequences.

[0151] A recent modification of the derivation of merge candidates allows the number of candidates in the list to be greater than the maximum number of candidates in the final list, Maxcand. However, this number of candidates in the initial list, MaxCandInitialList, is used for the derivation. As a result, zero candidates are added until the number of candidates is MaxCandInitialList, and not until Maxcand.

[0152] BM merge mode BM merge is a merge mode dedicated to the adaptive decoder-side motion vector refinement method, which is an extension of the multi-pass DMVR in ECM. As described in JVET-X2025, this mode corresponds to two merge modes, refining MVs in only one direction. Thus, there is one merge mode for L0 and one for L1. Thus, BM merge is enabled only if the DMVR condition can be enabled. For these two merge modes, only one list of merge candidates is derived, and all candidates respect the DMVR condition.

[0153] Merge candidates for BM merge mode are derived from spatially adjacent coded blocks, TMVP, non-adjacent blocks, HMVP, and pairwise candidates in the same way as in the normal merge mode. The difference is that only those that satisfy the DMVR condition are added to the candidates. The merge index is coded in the same way as in the normal merge mode.

[0154] AMVP merge mode The AVMP merge mode, also known as the bidirectional predictor, is defined in JVET-X2025 as follows: It consists of an AMVP predictor in one direction and a merge predictor in the other direction. The mode may be enabled for a coding block when the selected merge predictor and AMVP predictor satisfy the DMVR condition, If there is at least one reference picture from the past and one reference picture from the future for the current picture, and the distances from the two reference pictures to the current picture are the same, then the bilateral matching MV refinement is applied to the merge MV candidate and the AMVP MVP as the starting point. Otherwise, if the template matching function is enabled, then the template matching MV refinement is applied to the merge predictor or the AMVP predictor with a higher template matching cost.

[0155] The AMVP part of the mode is signaled as normal unidirectional AMVP, i.e., the reference index and MVD are signaled, with a derived MVP index if template matching is used, or the MVP index is signaled if template matching is disabled.

[0156] For an AMVP direction LX, where X can be 0 or 1, the merge portion in the other direction (1 to LX) is implicitly derived by minimizing the bilateral matching cost between the AMVP predictor and the merge predictor, i.e., for the pair of AMVP motion vector and merge motion vector. For all merge candidates in the merge candidate list with other direction (1-LX) motion vectors, the bilateral matching cost is calculated using the merge candidate MV and the AMVP MV. The merge candidate with the smallest cost is selected. Bilateral matching refinement is applied to the coding block starting from the selected merge candidate MV and the AMVP MV.

[0157] The third pass of the multi-pass DMVR, which is an 8x8 sub-PU BDOF refinement of the multi-pass DMVR, is enabled for AMVP merge mode coded blocks.

[0158] The mode is indicated by a flag, and if the mode is valid, the AMVP direction LX is further indicated by a flag.

[0159] MVD Code Prediction The code prediction method is described in JVET-X0132. Motion vector differential code prediction may be applied in normal inter mode if the motion vector differential contains a non-zero component. In the current ECM version, it is applied to AMVP, affine MVD, and SMVD modes. The possible MVD code combinations are sorted according to the template matching cost, and the index corresponding to the true MVD code is derived and coded in the context model. At the decoder side, the MVD code is derived as follows: 1 / Analyze the magnitude of the MVD components. 2 / Analyze the context-coded MVD code prediction index. 3 / Build MV candidates by creating combinations between possible codes and absolute MVD values, and add it to the MV predictor. 4 / Derive the MVD code prediction cost for each derived MV based on the template matching cost and sorting. 5 / Use the MVD code prediction index to select the true MVD code. 6 / Add the true MVD to the MV predictor of the final MV.

[0160] TIM D Fusion of intra prediction, template-based intra mode derivation (TIMD) is described in JVET-X2025 as follows: For each intra prediction mode in MPM, the SATD between the template prediction sample and the reconstructed sample is calculated. The first two intra prediction modes with the minimum SATD are selected as the TIMD mode. These two TIMD modes are fused with weights after applying PDPC processing, and such weighted intra prediction is used to code the current CU. Position-dependent intra prediction combining (PDPC) is included in the derivation of TIMD modes.

[0161] Duplicate Check In Fig. 14 and Fig. 15, a duplication check for each candidate has been added (1440, 1441, 1442, 1443, 1444, 1445, and 1530). But duplication is also for non-adjacent candidates (1540) and history-based candidates (1501). It consists in comparing the motion information of the current candidate with index cnt with the motion information of each other previous candidate. If this motion information is equal, it is considered a duplication and the variable cnt is not incremented. Of course, the motion information means, for each list (L0, L1), the inter direction, the reference frame index, and the motion vector. It should be noted that zero candidates corresponding to different reference frames are not considered as duplications.

[0162] MVTH In ECM, for the overlap check, a motion vector threshold was introduced. This parameter modifies the equality check by considering two motion vectors equal if their absolute difference for each component is less than or equal to the motion vector threshold MvTh. In normal merge mode, MvTh is set equal to 1, which corresponds to the conventional overlap check without a motion vector threshold.

[0163] MvTh is equal to a value that depends on the number of luma samples in the current CU, nbSamples, for templates that match the regular merge mode, as defined as follows: if (nbSamples < 64) MvTh = 1<< MV_FRACTIONAL =16 else if (nbSamples < 256) MvTh = 2<< MV_FRACTIONAL =32 else MvTh = 4<< MV_FRACTIONAL =64 Here, MV_FRACTIONAL corresponds to the inter resolution of the codec. So, in the current ECM, 16 thSince a resolution of 1 is used, MV_FRACTIONAL is equal to 4. << is the left shift operator, where nbSamples = Height x Width (i.e. Height and Width of the current block).

[0164] For example, there is another threshold, MvThBDMVRMvdThreshold, that is used for GEO merge derivation and also for overlap checking of non-adjacent candidates, as described below.

[0165] if (nbSamples < 64) MvThBDMVRMvdThreshold = (1 << MV_ FRACTIONAL) >> 2 = 4 else if (nbSamples < 256) MvThBDMVRMvdThreshold = (1 << MV_FRACTIONAL) >> 1 = 8 else MvThBDMVRMvdThreshold = (1 << MV_FRACTIONAL) >> 0 = 16 ARMC In ECM, Adaptive Reordering of Merge Candidates Using Template Matching (ARMC) was added to reduce the number of bits in the merge index. Candidates are reordered based on the cost of each candidate according to the template matching cost calculated as in FIG. 13. In this method, only one cost is calculated per candidate. This method is applied after this list is derived and is applied only to the first five candidates in the canonical merge candidate list. It should be understood that the number five is chosen to balance the complexity of the reordering process with the potential gain, and thus a larger number (e.g., all of the candidates) could be reordered.

[0166] FIG. 18 shows an example of this method applied to a regular merge candidate list containing 10 candidates, such as the CTC.

[0167] This method is also applied to the sub-block merging mode except for the temporal candidate and to the regular TM mode for all four candidates.

[0168] In the proposal, this method was also extended to reorder and select candidates to be included in the final list of merge mode candidates. For example, in JVET-X0087, all possible non-adjacent candidates (1540) and history-based candidates (1501) are considered along with the time-non-adjacent candidates to obtain a list of candidates. This list of candidates is constructed without considering the maximum number of candidates. Then, this list of candidates is sorted. Only the correct number of candidates from this list is added to the final list of merge candidates. The correct number of candidates corresponds to the first N candidates in the list. In this example, the correct number is the maximum number of candidates minus the number of spatial and temporal candidates already in the final list. In other words, the non-adjacent candidates and the history-based candidates are processed separately from the adjacent spatial and temporal candidates. The processed list is used to generate the final merge candidate list, supplemented with the adjacent spatial and temporal merge candidates already present in the merge candidate list.

[0169] In JVET-X0091, ARMC is used to select a time candidate from three time candidates: bi-dir, L0, or L1. The selected candidate is added to the merge candidate list.

[0170] In JVET-X0133, merge time candidates are selected from among several time candidates that are sorted using ARMC. Similarly, all possible neighbor candidates are subject to ARMC, and up to nine of these candidates can be added to the list of merge candidates.

[0171] All these proposed methods use classical ARMC to sort the final list of merge candidates and then sort it. JVET X0087 reuses the costs calculated during sorting of non-adjacent and history-based candidates to avoid additional computational costs. JVET-X0133 applies systematic sorting to all candidates on the final list of merge candidates.

[0172] New ARMC Since the first implementation of ARMC, the method has been added to several other modes. ARMC applies to the regular and template matching merge modes, the subblock merge mode, as well as the IBC, MMVD, Affine MMVD, CIIP, CIIP with template matching, and BM merge modes. Additionally, the ARMC principle of sorting the list of candidates based on template matching cost is also applied to AMVP merge candidate derivation, and for the intra-method TIMD, to select the most likely predictor.

[0173] In addition, the principles of the signed residual prediction methods of the Affine MVD, AMVP, and SMVD methods have been added.

[0174] There are also additional tests of GEO merge mode and GEO with template matching, as well as the use of this reordering for reference frame index prediction.

[0175] In addition to this almost systematic use of ARMC, there have been some additional modifications of the derivation process for some candidates.

[0176] For example, ECM4.0 derivation includes a cascaded ARMC process of candidate derivation for regular merge mode, TM merge mode, and BM merge mode, as shown in Regular Merge Mode and TM Merge Mode in Figure 19 .

[0177] First, 10 temporal positions are checked compared to the previous merge candidate derivation and are added to the list of temporal candidates after a non-overlapping check. This temporal list can contain up to 9 candidates. Furthermore, this temporal list has a specific threshold since the MV threshold is always 1 and is independent of the merge mode compared to the motion threshold used for merge candidate derivation. Based on the list of up to 9 first non-overlapping positions, the ARMC process is applied and only the first temporal candidate is added to the conventional list of merge candidates if it is non-overlapping compared to the previous candidate.

[0178] Similarly, non-adjacent spatial candidates are derived among the 59 positions. A first list of non-overlapping candidates is derived, which can reach 18 candidates. However, the motion threshold is different from the one used in the temporal derivation and the rest of the list, and does not depend on the normal or template merge modes. It is set equal to mv in BDMVR. A maximum of 18 non-adjacent candidates are reordered and only the 9 first non-adjacent candidates are kept and added to the merge candidate list. Then, other candidates are added, unless the list already contains the maximum number of candidates. Furthermore, for TM merge mode, the maximum number of candidates in the list MaxCandInitialList is more than the maximum number of candidates that the final list can contain Maxcand.

[0179] Then, the ARMC process is applied to all candidates in the intermediate list, including MaxCandInitialList, as shown in Figure 19. The final list of candidates is set to the maximum number of candidates Maxcand.

[0180] The same changes as in the template matching merge mode are applied to the BM merge mode.

[0181] ARMC template cost algorithm Figure 17 shows the template cost calculation of the ARMC method. The number of candidates considered in this process, NumMergeCandInList, is equal to or greater than the maximum number (1712) that the list can contain Maxcand.

[0182] For each candidate in the list (1701), if no cost was calculated during the first ARMC process for the temporal and non-adjacent candidates (1702), the cost is set equal to 0 (1703). In the implementation, the uncalculated cost associated with candidate "i", mergeList[i].cost, is set equal to a maximum value MAXVAL. If the top template of the current block is available (1704), the distortion compared to the current block template is calculated (1705) and added to the current cost (1706). Then, or else, if the left template of the current block is available (1707), the distortion compared to the current block template is calculated (1708) and added to the current cost (1709). Then, the cost of the current merge candidate, mergeList[i].cost, is set equal to the calculated cost (1710) and the list is updated (1711). In this example, consider that the current candidate i is set to its position in terms of its cost compared to the costs of the other candidates. Once all candidates have been advanced up to the number of candidates in the list, NumMergeCandInList is set equal to the maximum number of possible candidates in the list, Maxcand.

[0183] Figure 20 shows a diagram of updating the candidate list (1710) of Figure 17. First, the variable Shift is set equal to 0. Then, the variable shift is incremented (2003) while Shift is less than the current candidate i of Figure 17 and the associated cost of the current candidate is less than the cost of the previous candidate number i-1-shift (2002). When this loop ends and the variable shift is different from 0 (2007), candidate number i is inserted at position i-shift (2010).

[0184] Multiple Hypothesis Prediction (MHP) We also added multiple hypothesis prediction (MHP) to ECM. In this method, it is possible to use up to four motion compensated prediction signals per block (instead of two, as in VVC). These individual prediction signals are superimposed to form an overall prediction signal. The motion parameters of each additional prediction hypothesis can be signaled explicitly by specifying the reference index, the motion vector predictor index, and the motion vector differential, or implicitly by specifying the merge index. A separate multiple hypothesis merge flag distinguishes between these two signaling modes.

[0185] For spatial, non-adjacent, and history-based merge candidates, a number of hypothesis parameter values ​​"addHypNeighbours" are inherited from the candidate.

[0186] For temporal candidates, as well as zero and pairwise candidates, the multiple hypotheses parameter value "addHypNeighbours" is not retained (they are distinct). LIC ECM adds Local Illumination Compensation (LIC), which is based on a linear model of illumination changes, which is calculated thanks to the neighboring samples of the current block and the neighboring samples of the previous block.

[0187] In ECM, LIC is only valid for unidirectional prediction. LIC is signaled via a flag. In merge mode, the LIC flag is not sent, but instead, the LIC flag is inherited from the merge candidate in the following way:

[0188] For spatial candidates, non-adjacent merge candidates, and history-based merge candidates, the value of the LIC flag is inherited.

[0189] For time candidates and zero candidates, the LIC flag is set equal to 0.

[0190] For pairwise candidates, the value of the LIC flag is set as shown in FIG. 16. This figure is based on FIG. 12, with the addition of modules 1620 and 1621, and the updates of modules 1609, 1612, 1611. The variable average is set equal to false (1620), and if the average for the pairwise was calculated for the current list, the LIC flag for the pairwise LICFlag[cnt] is set equal to false and the variable averageUsed is set equal to true (1609). If the candidate only has list motion information (1612, 1611), and the average was not used, the LIC flag is updated and set equal to the OR operation of its current value and the value of the candidate's LICflag.

[0191] And when the pairwise candidate is Bidir (i.e., equal to 3), LICflag is equal to false.

[0192] However, the algorithm shown in FIG. 16 only allows LICflag to be equal to something different from true if two candidates have motion information for one list, each candidate has its own list. For example, candidate 0 has motion information for L0 only, and candidate 1 has motion information for L1 only. In that case, LICflag can be equal to something different from 0, but it never happens because LIC is only for one direction. Therefore, the pairwise LICflag is always equal to false. Therefore, the pairwise candidate cannot use LIC if it is potentially needed. This therefore reduces the efficiency of the candidate and avoids the propagation of LIC for subsequent coding blocks, which in turn reduces the coding efficiency.

[0193] Furthermore, the duplicate check in the ECM software introduces some inefficiencies. As shown in Figures 14 and 15, each candidate is added to a list and the duplicate checks (1440, 1441, 1442, 1443, 1444, 1445, and 1530) only affect the increment of the variable cnt (1405, 1409, 1413, 1417, 1423, 1427, 1508). Furthermore, as explained in Figure 16, the variable BCWidx is not initialized for pairwise candidates. Thus, if the last candidate added in the list was a duplicate candidate, the value of the pairwise candidate BCWidx is the value of the previous duplicate candidate. This was not the case in VVC, since candidates are not added when they are considered duplicates.

[0194] problem One aspect to increase the coding efficiency of image or video compression is to use decoder-side decisions to reduce the rate associated with signaling. Several recent video coding standards already include techniques such as DMVR.

[0195] To make a decision, the encoder uses several parameters that can affect this decision and takes the best one according to these parameters. On the decoder side, the decision should also be adapted based on several parameters. One solution is to define the same rules for all decoders and parameters. However, this is not efficient in terms of coding efficiency and is not adaptable to all encoder implementations. At the other end of the scale, the signaling of all parameters needed to make the decision leads to an increase in the bitrate that also affects the coding efficiency (especially for short sequences or when these parameters change regularly).

[0196] The present invention aims to solve or ameliorate at least some of these problems by efficient signaling of some parameters needed to make the decision at the decoder side.

[0197] Embodiment In one example, information representative of a threshold value required to make a decoder-side decision is transmitted in the bitstream in at least one header. For example, the threshold value may be transmitted at a sequence level in an SPS, or for some pictures in a PPS, or for a picture in a picture header, or for a slice or tile in a slice header. In such a case, the bitstream includes at least one header that includes a syntax element related to a threshold value to be used when making a decoder-side decision. The threshold value may be explicitly signaled or derived, as described in detail below. In other words, the header includes a syntax element that allows to derive a threshold value, which is used to determine a criterion (such as a rate-distortion criterion) to be used in the decoding process.

[0198] The advantage of this example is the increased flexibility in making decisions at the decoder side, which increases the coding efficiency thanks to the adaptive thresholds.

[0199] Send the threshold in at least one header.

[0200] In one implementation, a syntax element indicates the transmission of the associated threshold value in one header. For example, this syntax element is a flag sps_Th_tool1_in_sps transmitted in the SPS. Alternatively or additionally, the syntax element may be pps_Th_tool1_in_pps transmitted in the PPS. Alternatively or additionally, the syntax element may be ph_Th_tool1_in_ph transmitted in the picture header. Alternatively or additionally, the syntax element may be sh_Th_tool1_in_sh transmitted in the slice header.

[0201] The advantage of this is more flexibility. In fact, in some video sequences for some video coding applications, the structure of the GOP structure is relatively simple, for example, only one reference frame in the past with the same QP, which implies that the threshold does not need to be adapted much and can only be adapted to sequence parameters in the SPS of some pictures in the PPS. On the contrary, for some applications, the GOP structure is complex, such as a QP hierarchy or some time IDs or some slices. In order to increase the coding efficiency, the threshold therefore needs more adaptation, for example, at the picture or slice level.

[0202] When transmission is enabled at a higher level, other associated syntax elements are not transmitted. In a related implementation, when a syntax element indicates transmission of the associated threshold in a previous header in the hierarchy of headers, the corresponding syntax element for the current header is not transmitted. The hierarchy of headers refers to the dependencies of headers and the order in which these headers appear in the bitstream. In general, the hierarchies are SPS, PPS, picture header, slice header.

[0203] For example, when sps_Th_tool1_in_sps is equal to 1, pps_Th_tool1_in_pps, ph_Th_tool1_in_ph, sh_Th_tool1_in_sh are not transmitted and are set equal to 0. Conversely, when the flag sps_Th_tool1_in_sps is equal to 0, the flag pps_Th_tool1_in_pps is decoded and the decoding of each ph_Th_tool1_in_ph or sh_Th_tool1_in_sh depends on this pps_Th_tool1_in_pps value.

[0204] The advantage to this is bitrate reduction, since some flags do not need to be transmitted in some cases.

[0205] Sending syntax elements to switch to default values The syntax element indicates that the threshold is to be signaled (e.g., as described above / below) or set equal to a default value. For example, a flag at slice level sh_signal_Th_tool1, when set equal to 1, indicates that another syntax element is decoded to signal the threshold, and when set equal to 0, the threshold is not transmitted and is set to a default value. This default value may be the value transmitted at a higher level.

[0206] This syntax element can be signaled at other levels such that all lower levels inherit this feature.

[0207] Sending syntax elements to switch to the default criteria Alternatively, the syntax element indicates that the decision is made on a basis that is not dependent on any threshold.

[0208] The advantage of these examples is an improvement in coding efficiency for some applications. Indeed, sometimes the adaptability of the threshold is not efficient because the adaptation requires too many bits that impact the bitrate. For example, when a picture contains several slices and the threshold should ideally be adapted to each slice, the signaling of the threshold impacts the global bitrate more than its adaptation.

[0209] How to decide The following description relates to the types of decoding decisions that the thresholds may affect. Note that several thresholds may be signaled / derived and / or one threshold may be reused for multiple decoding decisions.

[0210] The threshold is the threshold used in the DMVR method. In one category of implementations, a threshold is used in a decoder-side method to avoid additional searches. For example, VVC DVMR (Decoder-Side Motion Vector Refinement) avoids additional motion vector refinement based on distortion between two blocks from list L0 or L1. For example, the algorithm does not apply refinement when the sum of absolute differences between luma samples of these blocks is inferior to a few samples of the current luma block. Therefore, this criterion is not quality-compliant.

[0211] The advantage of transmitting a threshold for the DMVR method (which is fixed in VVC) is that it offers a possible compromise in coding efficiency at the encoder side without affecting the worst-case complexity, and therefore the availability of a decoder to decode such a bitstream. In particular, the encoder can adapt the DMVR complexity to the quality of a block, slice, frame, or sequence.

[0212] Use of a threshold for cost comparison in decoder-side reordering or selection methods of predictor lists In another category of implementations, the threshold is used for cost comparison in a decoder-side reordering or selection method of predictor lists. For example, the threshold is used for cost comparison of candidates for the purpose of selecting or reordering candidates in a merge candidate list. For example, the threshold can be a Lagrangian parameter (also called "lambda" λ). Similarly, the threshold can be used for selecting or reordering candidates in an intra-predictor list.

[0213] The advantage is an improvement in the coding efficiency of the method in case the threshold depends on several parameters and is determined by the encoder.

[0214] Thresholds for multiple modes, even with the same method.

[0215] In another category of implementations, several thresholds are transmitted for the same method, these thresholds corresponding to several modes of the set of modes for which this same decision method is applied at the decoder side.

[0216] Indeed, different modes generate different types of block predictors and the decision method can be adapted to generate the best global coding efficiency. Therefore, it may be advantageous to set different thresholds for the same method for a mode or set of modes.

[0217] Below we describe a single threshold for one tool (tool1), but this can be extended to several thresholds for multiple tools, tool2, tool3, etc.

[0218] Threshold Signaling The following description relates to how the threshold(s) are signaled in the header.

[0219] Signaling maximum number of bits The maximum number of bits required to signal the threshold is transmitted and the value of the threshold is decoded according to the number of bits.

[0220] For example, the syntax element sps_log2_bits_Th_tool1 is sent in the SPS, and then the thresholds are decoded according to the value of sps_log2_bits_Th_tool1. For example, for each threshold, N bits are decoded, where N is the value of sps_log2_bits_Th_tool1.

[0221] This maximum number of bits may be transmitted at a higher level, e.g., for the syntax element sps_log2_bits_Th_tool1, in the PPS, a threshold value is decoded and the picture header or slice is decoded based on this number of bits.

[0222] The advantage of these signaling methods is an improvement of the coding efficiency. Indeed, the threshold depends on the encoder parameters and, for a particular configuration or sequence, the encoder knows the maximum value of the threshold to be transmitted. This method has an advantage compared to the method of coding the threshold based on a unary code (see below) when the threshold is not close to zero. Indeed, the considered threshold has a high value and the unary code produces a higher rate.

[0223] Signaling the number of bits used to signal the threshold value Alternatively, the number of bits per threshold to be transmitted is signaled.

[0224] The advantage of this embodiment is improved coding efficiency when the threshold is highly variable, in which case some bits can be saved.

[0225] The threshold value is transformed or quantized.

[0226] Furthermore, the threshold value may be transformed, e.g., it may be quantized, e.g., it may be quantized at the encoder side and when received at the decoder side, it may be dequantized.

[0227] The advantage is a reduction in the number of bits required to code the thresholds, and it is efficient in any case if the thresholds are coded with bits or unary codes. The thresholds used have large values, but quantization is efficient because high precision does not lead to improved coding efficiency.

[0228] For example, the thresholds are shifted with a right shift at the encoder side and then transmitted. The decoder then applies a corresponding left shift to the decoded values ​​to obtain the thresholds to be used in the decoding process. To use the correct thresholds at the encoder side, the encoder shall also apply a left shift.

[0229] Another example is to apply a square root operation to the thresholds on the encoder side and a square operation to decode the thresholds on the decoder side, where the thresholds obtained after the square root operation are considered to be truncated to integer values.

[0230] This operation provides higher quantization than is required for some thresholds, e.g., when the threshold is high.

[0231] Threshold prediction Another way to reduce the bit rate is to derive the thresholds via prediction instead of direct signaling.

[0232] In one example, a threshold is predicted. At the encoder side, the difference between the thresholds is coded as the predicted value, which can be called the "residual." At the decoder side, the decoded residual is added to the predicted value to obtain the threshold.

[0233] The advantage is a reduction in bit rate for the coded threshold.

[0234] In one example, the predictor is a previously decoded (or encoded at the encoder side) threshold value.

[0235] The advantage is a reduction in bit rate, since the previously encoded threshold should be close to the current threshold.

[0236] A threshold value predicted by another threshold value sent at a higher level In one example, the threshold is predicted by a threshold transmitted at a higher level (if available). In particular, this predictor threshold can be transmitted for prediction only. For example, the threshold in the slice header is predicted by a threshold transmitted in the picture header. Thus, when a picture contains several slices, the threshold at the slice level should not be too different and not predicted accurately by the predictor value transmitted in the picture header, since the value of the threshold of the difference slice should be close to this picture header value. For example, the predictor of the higher level can be the average or mode of the predictors used at the lower levels, or can be specifically selected at the encoder side to minimize the total residual transmitted.

[0237] This method offers a bitrate reduction especially when the threshold is coded thanks to a unary code.

[0238] Indexed threshold In one example, the threshold value is represented by an index in the bitstream, which corresponds to a threshold value that is known to both the encoder and the decoder.

[0239] A set of values ​​corresponding to the index may be transmitted (eg, at a high level).

[0240] These variables may be transmitted by first transmitting a variable indicating the number thresholds. The thresholds are then decoded and the index corresponds to the decoded order of these decoded thresholds.

[0241] In this example, the set of thresholds is transmitted at a higher level, for example, the set of thresholds is transmitted in the SPS and can be used in the picture header or slice header.

[0242] The following syntax table of the SPS illustrates this embodiment. Thus, a set of thresholds is transmitted in the SPS. First, the flag sps_Th_tool1_enabled_flag indicates whether the associated tool1 threshold is enabled. If this flag is set equal to 1, the syntax element sps_num_Th_in_set_tool1, which represents the number of thresholds in the set, is decoded. Then, the number of bits for each threshold sps_log2_bits_Th_tool1 is decoded. In this example, this syntax element is coded with 4 bits. Thus, by taking into account that 0 bits are not possible, it means that the number of bits of the associated threshold is between 1 and 16 bits. Thus, in this example, the thresholds are between 0 and 2 to the power of 16, i.e. 65536.

[0243] Then, for i equal to 0 to sps_num_Th_in_set_tool1, the thresholds sps_Th_tool1_predictors[i] are decoded.

[0244] [Table 3]

[0245] The following syntax table shows the corresponding slice header. In this example, consider that a threshold is signaled for tool1 in the slice header. If the current method is a threshold index based method (sps_Th_tool1_based_th_set_flag equals 1), the associated index sh_Th_tool1_idx is decoded. The decoder has this value in memory thanks to the table sps_Th_tool1_predictors[]. Thus, the current decoded threshold for the current slice is set equal to sps_Th_tool1_predictors[sh_Th_tool1_idx]. In this example, sh_Th_tool1_idx is coded with a unary code since the objective of this example is to reduce the rate of threshold signaling.

[0246] [Table 4]

[0247] The advantage of this is that the number of bits required to transmit the thresholds is limited thanks to the use of the index, since the thresholds generally depend on the frame hierarchy of the GOP. Many frames share the same threshold. It is therefore preferable to transmit the thresholds at a high level and indicate the index at a low level. It should be noted that the impact of this method is only positive if the same threshold is used for several frames or slices. If at least two thresholds are used only once, the method can increase the rate.

[0248] In one example, an index represents the possibility of directly transmitting a threshold value. For example, in the previous syntax table, when the decoded index sh_Th_tool1_idx is set equal to the number of threshold values ​​decoded in the set in the SPS, sps_num_Th_in_set_tool, the threshold value sh_current_Th_tool1 is decoded for the current slice. For example, the number of bits required to decode the threshold value is equal to sps_log2_bits_Th_tool1.

[0249] The advantage is an increase in coding efficiency. In fact, when the thresholds are signaled several times, there is no need to transmit them in the SPS, especially if there are several thresholds for which index signaling is not efficient and it is preferable to signal it thanks to this available index. When the encoder knows the GOP structure and can determine the thresholds to be transmitted when it sets the SPS, it can estimate the maximum number of bits required for each threshold, based on the number of thresholds and their occurrence, which is the best compromise between index or direct coding signaling.

[0250] When a threshold is directly transmitted at a low level, it is added to the set of thresholds to improve future decoding. For example, when sh_current_Th_tool1 is decoded in the slice header, this value is added to the set of thresholds at the end, and therefore sps_Th_tool1_predictors[sps_num_Th_in_set_tool1] is set equal to sh_current_Th_tool1. Then, the number of thresholds sps_num_Th_in_set_tool1 is incremented.

[0251] In addition, thanks to this method, no threshold set needs to be transmitted with the SPS or any thresholds, since all thresholds can be added "on the fly".

[0252] This embodiment is advantageous for example for encoders that change parameters on a frame-by-frame basis to meet a target rate, in which case the encoder does not know in advance the number of thresholds, their values, etc.

[0253] An encoder that orders the thresholds according to their estimated frequency.

[0254] To further reduce the bitrate, the encoder can order the thresholds according to their occurrence. For example, the encoder knows all possible thresholds that need to be transmitted and how many times each of these thresholds needs to be signaled when defining the SPS. This can be done, for example, by analyzing the reference picture list, when the thresholds depend on the QP or the distance to the reference frame. In that case, the encoder orders the thresholds according to their number of times they have been used. The order is from most to least occurring. These thresholds are coded into the SPS thanks to this order. As a result, at the slice level, the number of bits signaling the threshold indices is minimized, since the most frequent thresholds have the least number of bits.

[0255] The advantage is a reduction in the rate.

[0256] Threshold prediction based on different syntax elements As mentioned above, the threshold is used to make the decoding decision. For example, in image or video compression, the decision is made based on a rate-distortion compromise. As a result, the threshold is statistically correlated with the parameters that affect this compromise. The following example improves the prediction of the threshold according to the values ​​of other such syntax elements.

[0257] In one example, a predictor of a threshold is obtained based on a value of another syntax element or based on a value obtained from at least one syntax element. For example, when a threshold is determined at a slice level, the decoder uses at least one syntax element or a representation thereof and obtains a value of the associated threshold according to a table associating both. For example, the table associates different QP offset values ​​with thresholds. Then, when decoding some syntax elements, the QP offset can be calculated and the associated threshold determined. In VVC, the QP of a slice or picture is signaled in the slice header or picture header according to a QP offset (sh_qp_delta or ph_qp_delta). This offset is then added to another value to obtain the QP of the current slice.

[0258] The advantage is that fewer bits are needed to determine the threshold, thus improving coding efficiency.

[0259] The following syntax table shows an example of this embodiment at the slice level.

[0260] [Table 5]

[0261] A flag signals that the threshold was predicted correctly.

[0262] In one example, a flag in the header signals that the threshold is correctly predicted. In the previous syntax element table example for a slice, this flag is sh_Th_tool1_prediction_flag, and when it is equal to 1, the decoder sets the threshold equal to the threshold of the corresponding syntax element or its representation previously obtained during decoding of the slice header.

[0263] The advantage of this is that while the encoder can determine the threshold based on syntax elements, the threshold is represented by only one bit at the slice level.

[0264] The decoder verifies whether a threshold value is associated with a syntax element or in a value obtained from at least one syntax element.

[0265] In one example, the decoder verifies whether a threshold value is associated with the syntax element or a value obtained from the at least one syntax element, and if so, the associated threshold value is set as the decoded threshold value.

[0266] The advantage compared to the previous example is that there is no rate associated with the threshold when it is possible to identify it. However, there is no possibility to signal that the predicted value is not the correct value. The implementation of this therefore depends on the flexibility of the transmitted threshold compared to the stored rate.

[0267] Flags for inference signaling Signaling to switch between signaling To provide both flexibility and rate savings, flags from higher levels signal that the prediction is correct thanks to the flag or if the prediction is imposed. For example, in the previous syntax table, the flag sps_Th_tool1_inferred_flag indicates whether the threshold is correctly predicted by the associated threshold in the table, while the flag (sh_Th_tool1_prediction_flag) signals whether the threshold is correctly predicted by the associated threshold in the table, or whether the threshold is inferred based on the associated threshold in the table. sps_Th_tool1_inferred_flag is transmitted, for example, in the SPS. In other words, the flag indicates which of the two methods described above is used (signaling that the threshold is correctly predicted or verifying whether the threshold is associated with a syntax element or for a value obtained from at least one syntax element).

[0268] The advantage is that the encoder can choose the best way to signal the threshold. This decision can be based on, for example, the number of slices per frame.

[0269] If the prediction is incorrect, the threshold is decoded directly. In one embodiment, when the flag sh_Th_tool1_prediction_flag is set equal to 0 or there is no associated threshold in the table, the threshold is decoded directly. In the previous syntax table, when sh_Th_tool1_prediction_flag is equal to 0, the syntax element sh_current_Th_tool1 is decoded and the threshold of the current slice is set equal to sh_current_Th_tool1. In this example, when the threshold is inferred by the value in the table sh_Th_tool1_prediction_flag, it is set equal to 1, and when it is not possible to infer, sh_Th_tool1_prediction_flag is set equal to 0.

[0270] The advantage is that the threshold can be set equal to another value than exists in the table, or if there is no value available.

[0271] In addition, when the threshold is decoded, the number of bits may be the maximum default possible value.

[0272] In fact, if the threshold cannot be determined by the encoder, for example, when the SPS is determined, the threshold may be higher than the maximum threshold determined, for example, by the variable sps_log2_bits_Th_tool1.

[0273] Indexed threshold Similar to the above, an index can be sent to set a different threshold for the table.

[0274] In one example, when a new threshold value is directly decoded, the value of another syntax element, or a value obtained from at least one syntax element, is added to a table with the decoded threshold value.

[0275] The advantage is for the encoder not being able to determine the syntax element values ​​or the list of values ​​previously obtained from at least one syntax element: thanks to this embodiment, the table can be updated "on the fly".

[0276] In one example, a set of thresholds is transmitted at a higher level and corresponds to a possible value identified in a set of possible values ​​related to a value of another syntax element or a value obtained from at least one syntax element, for example, a set of thresholds is transmitted in an SPS and the thresholds are determined per picture thanks to slice header information, in the picture header or per slice information.

[0277] The following syntax element table illustrates this embodiment when considering QP offsets: Thanks to the decoded reference picture list syntax element, the decoder has identified the number of QPs or QP offsets to be used in the sequence. Similarly, the list of associated QPs or QP offsets QPOffsetList[] has also been extracted, for example, from the reference picture list.

[0278] In the syntax table below, the threshold predictor sps_Th_tool1_predictors[i] is decoded from 0 to the number of QP offsets.

[0279] [Table 6]

[0280] In one example, the number of elements in the table is transmitted in the bitstream.

[0281] In that case, the table of syntax elements is similar to the table syntax elements of the embodiment relating to index based threshold prediction where the number sps_num_Th_in_set_tool1 is transmitted.

[0282] The advantage of this embodiment is flexibility, since the encoder does not need to transmit the complete table.

[0283] In one example, a set of associated values ​​of another syntax element, or a set of values ​​obtained from at least one syntax element, is transmitted along with their associated thresholds at a higher level. Compared to the previous example, the values ​​associated with the thresholds are explicitly signaled.

[0284] The following syntax table illustrates this embodiment, in which each QP or QP offset (sps_qp_offset_for_tool1[i]) is transmitted with an associated threshold (sps_Th_tool1_predictors[i]) to create a table relating both. The QP offset is encoded thanks to a signed unary maximal code.

[0285] [Table 7]

[0286] In other words, the threshold value is determined from the header by predicting the threshold value based on syntax elements associated with the same header. The association between the syntax elements and the threshold value is determined from a table. Each syntax element is transmitted with its associated threshold value to generate the table.

[0287] The advantage of this is that the decoder does not need to parse the reference picture list to identify the different possible QPs or QP offsets used for subsequent frames or slices. The impact of signaling at the SPS level can be considered small compared to the parsing complexity to determine the list of QPs or QP offsets to be used.

[0288] In a particularly advantageous implementation, a set of associated syntax element values ​​or a set of values ​​obtained from at least one syntax element is predictively coded. For example, the encoder orders the QP offsets from the smallest to the largest. This list includes only unique values ​​of QP offsets. The following table of SPS syntax elements shows the decoding process. The first QP offset is the decoded sps_qp_offset_for_tool1[0] (if i=0) followed by the associated thresholds sps_Th_tool1_predictors[i].

[0289] For other i values, if the previous qp offset, sps_qp_offset_for_tool1[i-1], is negative, the QP offset differential qp_offset is extracted by decoding it with the sign unary code (se(v)). This offset qp_offset is added to the previous QP offset sps_qp_offset_for_tool1[i-1] to get the current one sps_qp_offset_for_tool1[i-1].

[0290] Otherwise, if the previous QP offset, sps_qp_offset_for_tool1[i-1], is positive, the QP offset differential qp_offset is extracted by decoding it with the unary code (u(v)). Therefore, in that case, it is unsigned since it is certain that qp_offset is positive. This offset qp_offset is added to the previous QP offset sps_qp_offset_for_tool1[i-1] to get the current one sps_qp_offset_for_tool1[i-1].

[0291] [Table 8]

[0292] The advantage of this is that the coding efficiency is improved compared to the previous embodiment, since the rate is reduced by using this prediction.

[0293] Table Order The table is generated so that there are no duplicates. One way to achieve this is for the encoder to identify unique values ​​of QP or QP offset and order the list from smallest to largest, as described above.

[0294] Threshold prediction based on parameters that affect block quality In one example, the syntax element, or the set of values ​​obtained from at least one syntax element used to determine the threshold predictor, is a parameter that affects the reconstructed quality.

[0295] Indeed, as mentioned above, the threshold is related to a decision that depends on a rate-distortion compromise.

[0296] The parameter that has a large effect on quality is the QP or the QP offset.

[0297] In fact, this parameter is taken into account to determine the Lagrangian parameter of the rate-distortion criterion (also called lambda, λ), and the previous examples can also be based on this parameter.

[0298] The advantage is that many thresholds for decoder-side decisions depend on or are highly correlated to this parameter value, and therefore threshold prediction is better as a result.

[0299] In particular, a syntax element to be considered is the QP offset value.

[0300] Reference Frame Based Threshold Prediction In one example, the syntax element or set of values ​​obtained from at least one syntax element used to determine the threshold predictor is related to a reference frame, and in particular may take into account the time distance between the current frame and the reference frame.

[0301] The reference frame, the quality of the reference frame, and the distance affect the quality of the reconstructed frame, therefore for some thresholds of some decoder decision methods, these parameters are correlated.

[0302] In one example, when the QP signaling is per block, a similar table of QPs or QP offsets with associated thresholds is transmitted in at least one header, and the thresholds are adapted at the block level thanks to this table of thresholds.

[0303] As in the previous example, a lower rate is required to signal the thresholds, resulting in coding efficiency.

[0304] Unless otherwise stated, all of these embodiments can be combined, and in fact many combinations can be synergistic and result in efficiency gains that are greater than the sum of their parts.

[0305] Implementation of the invention Fig. 22 shows a system 191, 195 comprising at least one of the encoder 150 or the decoder 100 and a communication network 199 according to an embodiment of the present invention. According to an embodiment, the system 195 is for processing and providing content (e.g. video and audio content for displaying / outputting or streaming the video / audio content) to a user having access to the decoder 100, for example via a user interface of a user terminal comprising the decoder 100 or a user terminal capable of communicating with the decoder 100. Such a user terminal may be a computer, a mobile phone, a tablet, or any other type of device capable of providing / displaying (provided / streamed) content to a user. The system 195 obtains / receives a bitstream 101 (e.g. in the form of a continuous stream or signal while a previous video / audio is being displayed / output) via the communication network 199. According to an embodiment, the system 191 is for processing content and storing the processed content, for example the processed video and audio content for later displaying / outputting / streaming. The system 191 obtains / receives content comprising an original image sequence 151 which is received and processed by the encoder 150 (including filtering by a deblocking filter according to the invention), which generates a bitstream 101 which is communicated to the decoder 100 via the communication network 191. The bitstream 101 is then communicated to the decoder 100 in several ways, for example it may be pre-generated by the encoder 150 and stored as data in a storage device (e.g. on a server or cloud storage) in the communication network 199 until a user requests the content (i.e. bitstream data) from the storage device, at which point the data is communicated / streamed from the storage device to the decoder 100.The system 191 may also comprise a content providing device for providing / streaming to the user (e.g., by communicating data for a user interface to be displayed on the user terminal) content information (e.g., title of the content and other meta / storage location data for identifying, selecting and requesting the content) for the content stored in the storage device and receiving and processing user requests for content so that the requested content can be delivered / streamed from the storage device to the user terminal. Alternatively, the encoder 150 generates the bitstream 101 and communicates / streams it directly to the decoder 100 when the user requests the content. The decoder 100 then receives the bitstream 101 (or signal) and performs filtering using a deblocking filter according to the present invention to obtain / generate a video signal 109 and / or an audio signal, which is then used by the user terminal to provide the requested content to the user.

[0306] Any step of the method / process according to the present invention or function described herein may be implemented in hardware, software, firmware, or any combination thereof. If implemented in software, the step / function may be stored or transmitted as one or more instructions or codes or programs or computer readable media and executed by one or more hardware-based processing units such as a PC ("personal computer"), a DSP ("digital signal processor"), a circuit, a circuit element, a processor and memory, a general-purpose microprocessor or central processing unit, a microcontroller, an ASIC ("application specific integrated circuit"), a field programmable logic array (FPGA), or other equivalent integrated or discrete logic circuit, programmable computing machine, etc. Thus, the term "processor" as used herein may refer to any of the aforementioned structures, or any other structure suitable for implementing the techniques described herein.

[0307] The embodiments of the present invention may also be implemented by a wide variety of devices or apparatuses, including a wireless handset, an integrated circuit (IC), or a set of JCs (e.g., a chipset). Various components, modules, or units are described herein to illustrate functional aspects of a device / apparatus configured to perform the embodiments, but do not necessarily require implementation by different hardware units. Rather, the various modules / units may be combined in a codec hardware unit or may be provided by a collection of interoperable hardware units including one or more processors in conjunction with appropriate software / firmware.

[0308] The embodiments of the present invention can be realized by a computer of a system or apparatus including one or more processing units or circuits for reading and executing computer-executable instructions (e.g., one or more programs) recorded on a storage medium to perform one or more modules / units / functions of the above-mentioned embodiments, and / or by a method executed by the computer of the system or apparatus, e.g., by reading and executing the computer-executable instructions from the storage medium to perform one or more functions of the above-mentioned embodiments, and / or by controlling one or more processing units or circuits to perform one or more functions of the above-mentioned embodiments. The computer can include a separate computer or a network of separate processing units to read and execute the computer-executable instructions. The computer-executable instructions can be provided to the computer from a computer-readable medium, such as, for example, a network or a communication medium via a tangible storage medium. The communication medium can be a signal / bit stream / carrier wave. The tangible storage medium can be, for example, a hard disk, a random access memory (RAM), a read-only memory (ROM), a storage device of a distributed computing system, an optical disk (compact disk (CD), digital versatile disk (DVD), or Blu-ray disk (BD)). TMThe "non-transitory computer-readable storage medium" may include one or more of a memory device, a flash memory device, a memory card, etc. At least some of the steps / functions may also be implemented in hardware by machines or dedicated components such as an FPGA ("Field Programmable Gate Array") or an ASIC ("Application Specific Integrated Circuit").

[0309] Fig. 23 is a schematic block diagram of a computing device 3600 for the implementation of one or more embodiments of the present invention. The computing device 3600 can be a device such as a microcomputer, a workstation, or a light portable device. The computing device 3600 comprises a communication bus connected to: - a central processing unit (CPU) 3601, such as a microprocessor; - a random access memory (RAM) 3602 for storing executable code of the method of the embodiments of the present invention, as well as registers adapted to record variables and parameters necessary for implementing the method for encoding or decoding at least a part of an image according to the embodiments of the present invention, the memory capacity of which can be expanded, for example, by an optional RAM connected to an expansion port; - a read only memory (ROM) 3603 for storing computer programs for implementing the embodiments of the present invention; - a network interface (NET) 3604, typically connected to a communication network over which the digital data to be processed is transmitted or received. The network interface (NET) 3604 may be a single network interface or may be composed of a set of different network interfaces (e.g. wired and wireless interfaces, or different types of wired or wireless interfaces). Data packets are written to the network interface for transmission or read from the network interface for reception under the control of software applications executing in the CPU 3601;-a user interface (UI) 3605 may be used to receive input from a user or to display information to a user;-a hard disk (HD) 3606 may be provided as a mass storage device;-an input / output module (IO) 3607 may be used to receive / send data to / from external devices such as video sources or displays. Executable code may be stored in either the ROM 3603, the HD 3606, or on a removable digital medium such as a disk.According to a variant, the executable code of the program is received by the communication network via NET 3604 in order to be stored in one of the storage means of the communication device 3600, such as HD 3606, before being executed. The CPU 3601 is adapted to control and direct the execution of the instructions or parts of the software code of a program or group of programs according to an embodiment of the invention, the instructions being stored in one of the aforementioned storage means. After power-up, the CPU 3601 is able to execute instructions from the main RAM memory 3602 relating to a software application, for example after the instructions have been loaded from the program ROM 3603 or from the HD 3606. Such a software application, when executed by the CPU 3601, causes the steps of the method according to the invention to be carried out.

[0310] It will also be appreciated that according to another embodiment of the present invention, the decoder according to the above-mentioned embodiments is provided in a user terminal such as a computer, a mobile phone, a table, or any other type of device (e.g., a display device) capable of providing / displaying content to a user. According to yet another embodiment, the encoder according to the aforementioned embodiment is provided in an image capture device, also comprising a camera, video camera, or network camera (e.g., a closed circuit television or video surveillance camera), which captures and provides content for the encoder to encode. Two such examples are provided below with reference to Figures 37 and 38.

[0311] FIG. 24 is a diagram showing a network camera system 3700 including a network camera 3702 and a client device 202. As shown in FIG.

[0312] The network camera 3702 includes an imaging unit 3706 , an encoding unit 3708 , a communication unit 3710 , and a control unit 3712 .

[0313] The network camera 3702 and the client device 202 are connected to each other via the network 200 so as to be able to communicate with each other.

[0314] The image capturing unit 3706 includes a lens and an image capturing element (e.g., a charge coupled device (CCD) or a complementary metal oxide semiconductor (CMOS)) to capture an image of an object and generate image data based on the image. The image can be a still image or a video image.

[0315] The encoding unit 3708 encodes the image data using the encoding method described above, or a combination of the encoding methods described above.

[0316] The communication unit 3710 of the network camera 3702 transmits the encoded image data encoded by the encoding unit 3708 to the client device 202 .

[0317] The communication unit 3710 also receives commands from the client device 202. The commands include commands for setting parameters for the encoding unit 3708 to encode.

[0318] The control unit 3712 controls other units within the network camera 3702 according to commands received by the communication unit 3712 .

[0319] The client device 202 includes a communication unit 3714, a decryption unit 3716, and a control unit 3718.

[0320] The communication unit 3714 of the client device 202 transmits a command to the network camera 3702 .

[0321] In addition, the communication unit 3714 of the client device 202 receives the encoded image data from the network camera 3712 .

[0322] The decoder 3716 decodes the encoded image data using the decoding method described above, or a combination of the decoding methods described above.

[0323] The control unit 3718 of the client device 202 controls other units within the client device 202 in response to user operations or commands received by the communication unit 3714 .

[0324] The control unit 3718 of the client device 202 controls the display device 2120 to display the image decoded by the decoding unit 3716 .

[0325] In addition, the control unit 3718 of the client device 202 controls the display device 2120 to display a GUI (Graphical User Interface), and specifies parameter values ​​of the network camera 3702 including parameters for the encoding by the encoding unit 3708 .

[0326] Furthermore, the control unit 3718 of the client device 202 controls other units within the client device 202 in response to a user operation input to the GUI displayed by the display device 2120 .

[0327] The control unit 3718 of the client device 202 controls the communication unit 3714 of the client device 202 to send a command to the network camera 3702 that specifies parameter values ​​of the network camera 3702 in response to user input on the GUI displayed by the display device 2120.

[0328] FIG. 25 is a diagram showing a smartphone 3800.

[0329] The smartphone 3800 includes a communication unit 3802, a decoding unit 3804, a control unit 3806, and a display unit 3808.

[0330] The communication unit 3802 receives the encoded image data via the network 200 .

[0331] The decoding unit 3804 decodes the encoded image data received by the communication unit 3802 .

[0332] The decoding / encoding unit 3804 decodes / encodes the encoded image data using the above-mentioned decoding method.

[0333] The control unit 3806 controls other units within the smartphone 3800 according to user operations or commands received by the communication unit 3806.

[0334] For example, the control unit 3806 controls the display unit 3808 to display the image decoded by the decoding unit 3804. The smartphone 3800 may also include a sensor 3812 and an image recording device 3810. In this manner, the smartphone 3800 can record an image and encode the image (using the methods described above).

[0335] The smartphone 3800 can then decode the encoded images (using the methods described above) and display them via the display unit 3808, or transmit the encoded images via the communication unit 3802 and the network 200 to another device.

[0336] Substitutions and Modifications Although the present invention has been described with the embodiments, it should be understood that the present invention is not limited to the disclosed embodiments. It will be understood by those skilled in the art that various changes and modifications can be made without departing from the scope of the present invention, as defined in the appended claims. All of the features disclosed in this specification (including any accompanying claims, abstract, and drawings), and / or all of the steps of any method or process so disclosed, can be combined in any combination, except combinations in which at least some of such features and / or steps are mutually exclusive. Each feature disclosed in this specification (including any accompanying claims, abstract, and drawings) may be replaced by an alternative feature serving the same, equivalent, or similar purpose, unless expressly stated otherwise. Thus, unless otherwise stated, each feature disclosed is merely one example of a generic series of equivalent or similar functions.

[0337] It will also be understood that any result of the above comparisons, decisions, evaluations, selections, executions, performing, or considerations, e.g., selections made during an encoding or filtering process, may be indicated in or determinable / inferable from data in the bitstream, e.g., flags or data indicating the result, such that the indicated or determined / inferred result may be used in processing, e.g., during a decoding process, in lieu of actually performing the comparisons, decisions, evaluations, selections, executions, performing, or considerations.

[0338] In the claims, the word "comprising" does not exclude other elements or steps, and the indefinite articles "a" or "an" do not exclude a plurality. The mere fact that different features are recited in mutually different dependent claims does not indicate that a combination of these features cannot be used to advantage.

[0339] Reference signs appearing in the claims are by way of illustration only and shall have no limiting effect on the scope of the claims.

Claims

1. A method for decoding an image portion, The threshold is determined from the header related to the aforementioned image portion, The image portion is decoded according to the criteria based on the threshold. A method characterized by including the following.

2. The method according to claim 1, further comprising decoding a syntax element in the header that can indicate the presence of the threshold in the header.

3. The method according to 2, characterized in that the syntax element can also indicate that a threshold is not transmitted and that the decoding uses an alternative method.

4. The method according to 2, characterized in that, if the syntax element indicates the presence of the threshold, the decoding of the corresponding syntax element in a hierarchically lower header is skipped.

5. The method according to claim 1, characterized in that the decoding includes a decision to perform motion vector refinement based on the criteria.

6. The method according to claim 1, characterized in that determining the threshold includes decoding an index indicating a threshold from among a plurality of thresholds known by the decoder.

7. The method according to 6, characterized in that it includes receiving the plurality of thresholds.

8. The method according to 7, characterized in that the plurality of thresholds are transmitted at a hierarchically higher level than the header.

9. The method according to 7, characterized by comprising decoding a variable indicating the number of received thresholds.

10. The method according to claim 7, characterized in that the plurality of thresholds are transmitted in the header at a level relating to a plurality of image portions.

11. The method according to 10, characterized in that the level headers relating to multiple image portions are sequence parameter set headers.

12. The method according to 6, characterized in that one value of the index indicates that a threshold not among the plurality of thresholds should be decoded from the header.

13. The method according to 12, characterized in that the value of the index is the maximum value.

14. The method according to 12, further comprising adding the threshold decoded from the header to the plurality of thresholds and associating it with the index.

15. A method for encoding headers associated with multiple image portions, Determining a threshold for each image portion, which is used to determine a criterion for said portion, To generate an index of the aforementioned threshold, Encoding the aforementioned index into the aforementioned header A method characterized by including the following.

16. The method according to 15, characterized in that the header is a sequence parameter set (SPS).

17. A decoding device for decoding an image portion, Means for determining a threshold from a header related to the aforementioned image portion, means for decoding the image portion according to the criteria based on the threshold, A decoding device characterized by having the following features.

18. An encoding device for encoding headers associated with multiple image portions, A threshold for each image portion, and means for determining the threshold used to determine a criterion for the said portion, means for generating the threshold index, means for encoding the index into the header and An encoding device characterized by having the following features.

19. A program that causes a programmable device to execute the method described in any one of claims 1 to 16 during execution.