Video Encoding and Decoding
Patent Information
- Application Number
- JP2024504879
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2021-12-14
- Filing Date
- 2022-09-29
- Publication Date
- 2025-10-03
AI Technical Summary
The existing video coding standards, such as HEVC, face challenges in achieving efficient motion vector prediction due to the complexity and bitrate issues associated with the composition and ordering of motion vector prediction candidates in the Versatile Video Coding (VVC) standard, particularly in high dynamic range and ultra-high definition video applications.
The method involves generating and ordering motion vector prediction candidates by adding paired candidates that are likely to be selected for merge mode, ensuring they are positioned closer to the top of the list to reduce bitrate and improve encoding efficiency, while also considering conditions like merge mode type and similarity to existing candidates.
This approach enhances encoding performance by reducing bit rate and improving coding efficiency by frequently selecting pair candidates that are closer to the ideal motion vector prediction, thus optimizing the encoding process for diverse video content types.
Smart Images

Figure 00000000_0000_ABST
Abstract
Description
[Technical field]
[0001] The present invention relates to video encoding and decoding. [Background technology]
[0002] The Joint Video Experts Team (JVET), a joint team formed by MPEG and ITU-T Study Group 16 VCEG, has introduced a new video coding standard called Versatile Video Coding (VVC). The goal of VVC is to achieve significant compression improvements over the existing HEVC standard (i.e., typically twice as fast). Primary target applications and services include, but are not limited to, 360-degree video and high dynamic range (HDR) video. In particular, the test material has demonstrated its effectiveness in ultra-high definition (UHD) video. Thus, we can expect compression efficiency improvements well beyond the 50% targeted in the final standard.
[0003] After completing the standardization of VVCv1, JVET established the Exploration Software (ECM) and started the exploration phase. ECM aims to improve coding efficiency by gathering additional tools and improvements of existing tools on top of the VVC standard.
[0004] Compared to HEVC, VVC has a modified set of "merge modes" for motion vector prediction that achieves higher coding efficiency at the cost of more complexity. Motion vector prediction is enabled by deriving a list of "motion vector prediction candidates", and the index of the selected candidate is signaled in the bitstream. A merge candidate list is generated per coding unit (CU). However, CUs may be split into smaller blocks for decoder-side motion vector refinement (DMVR) or other methods.
[0005] The composition and order of this list can have a significant impact on coding efficiency, since accurate motion vector prediction reduces the size of the residual or distortion of the block prediction, and having such candidates at the top of the list reduces the number of bits required to signal the selected candidate. The present invention aims to improve on at least one of these aspects.
[0006] The modifications incorporated in VVCv1 and ECM result in a maximum of 10 motion vector prediction candidates. This allows for candidate diversity, but may increase the bit rate if candidates lower in the list are selected. The present invention broadly relates to improved derivation and ordering of one or more "pairwise" motion vector prediction candidates in a list of motion vector prediction candidates. A "pairwise" motion vector prediction candidate is a candidate that is combined or averaged from two or more other candidates in the candidate list. Summary of the Invention
[0007] According to one aspect of the present invention, there is provided a method for generating a list of motion vector prediction candidates for decoding a portion of an image encoded using a merge mode, the method comprising: determining a merge mode used to encode said image portion; and adding a pair motion vector prediction candidate to a list of motion vector prediction candidates depending on said determination.
[0008] This method improves coding performance by enabling pair candidates that are more likely to be selected for the merge mode.
[0009] Optionally, if the merge mode is template matching or GEO, the pair motion vector candidate is not added.
[0010] Optionally, if the pair candidate is an average candidate, the pair motion vector candidate is not added.
[0011] Optionally, if the merge mode is normal or CIIP merge mode, the pair motion vector candidates are added.
[0012] Optionally, to improve bitrate reduction, the method further comprises adding said candidate pair motion vectors closer to the top of said list than to the bottom.
[0013] According to another aspect of the present invention, there is provided a method for generating a list of motion vector prediction candidates for decoding a portion of an image, the method comprising: generating a pair motion vector prediction candidate; and adding the pair motion vector prediction candidate to a candidate list of motion vector prediction candidates, the position of the candidate being closer to a top of the list than to a bottom.
[0014] In this way, it turns out that the pair candidates are unexpectedly commonly selected, and the bit rate can be reduced since positions closer to the top can be coded with fewer bits than the bottom.
[0015] Optionally, the method further comprises adding the pair motion vector candidate to the list of motion vector prediction candidates at a position immediately following the motion prediction candidate used to generate the pair motion vector prediction candidate.
[0016] Optionally, to improve coding efficiency, the method further comprises adding the pair motion vector candidate in the list of motion vector prediction candidates at a position immediately following the first two spatial motion prediction candidates.
[0017] Optionally, the method further comprises adding said pair motion vector candidate to said list of motion vector prediction candidates in a second position.
[0018] According to another aspect of the present invention, there is provided a method for generating a list of motion vector prediction candidates for decoding a portion of an image, the method comprising: generating an initial list of motion vector prediction candidates; if candidate reordering is selected for the image portion, reordering at least a portion of the initial list to generate a reordered motion vector prediction candidate list; and adding pair motion vector prediction candidates to the reordered list.
[0019] This method improves coding performance by ordering multiple candidates, including pair candidates, in the most efficient order.
[0020] Optionally, to improve coding efficiency, the method further comprises determining the pair from the top two candidates in the sorted list.
[0021] Optionally, to improve coding efficiency, the method further comprises applying said permutation process to said determined candidate pairs.
[0022] Optionally, the portion of the initial sorted list is a maximum of the top N-1 candidates.
[0023] Optionally, the pair candidates are sorted as the Nth candidate.
[0024] Optionally, to improve coding efficiency, the method further comprises removing a lowest ranking candidate from the sorted list after adding the pair motion vector prediction candidates.
[0025] Optionally, all candidates in said initial list are reordered to generate said reordered motion vector prediction candidate list.
[0026] Optionally, the method further comprises determining the pair candidate using the first candidate and the i-th candidate in the sorted list, where i is an index into an initial list of motion vector prediction candidates.
[0027] Optionally, at predetermined positions one or more additional pairwise motion vector prediction candidates are included in the sorted list.
[0028] Optionally, the predetermined position is the fifth position in the sorted list.
[0029] Optionally, the predetermined position is the start of a second half of the sorted list.
[0030] Optionally, the initial list includes a first candidate pair of motion vectors, and the additional candidate pair of motion vectors is added to the reordered list immediately following the first candidate pair of motion vectors.
[0031] According to another aspect of the present invention, there is provided a method for generating a list of motion vector prediction candidates for decoding a portion of an image, the method comprising: generating an initial list of motion vector prediction candidates; and deriving at least one pair candidate from two candidates in the initial list, the two candidates including a first candidate and an i-th candidate in the list.
[0032] This method improves the relevance of the i-th candidate by combining it with the most likely candidate, improving coding performance.
[0033] Optionally, to improve coding efficiency, the i-th candidate is from the initial list unsorted.
[0034] Optionally, the method further comprises replacing the i-th candidate in the list with the determined pair candidate.
[0035] Optionally, the number of pair candidates is limited to four.
[0036] Optionally, to improve coding efficiency, the method further comprises determining whether the pair motion vector prediction candidate is similar to an existing candidate in the list before adding the pair candidate to the list. Preferably, determining whether the pair motion vector prediction candidate is similar to an existing candidate in the list comprises determining a threshold motion vector difference.
[0037] According to another aspect of the present invention, there is provided a method for generating a list of motion vector prediction candidates for decoding a portion of an image, the method comprising: generating an initial list of motion vector prediction candidates; deriving a pair candidate from two candidates in the initial list; and before adding the pair candidate to the list, determining whether the pair motion vector prediction candidate is similar to an existing candidate in the list, wherein determining whether the pair motion vector prediction candidate is similar to an existing candidate in the list comprises determining a threshold motion vector difference.
[0038] This method improves coding performance by ensuring diversity in the list of motion vector prediction candidates and narrowing it down towards the ideal candidate when appropriate.
[0039] Optionally, the method further comprises: said threshold motion vector difference being dependent on a search range of a decoder-side motion vector method.
[0040] Optionally, the threshold motion vector difference depends on whether the decoder-side motion vector method is enabled or disabled.
[0041] Optionally, the threshold motion vector difference depends on the POC distance or the POC magnitude.
[0042] Optionally, the threshold motion vector difference depends on the position in the list of the candidates used to construct the pair of candidates.
[0043] Optionally, the threshold motion vector difference is set to a first value greater than or equal to zero if the candidates used to construct the pair candidate are the first two candidates in the list, and otherwise is set to a second value, the second value being greater than the first value.
[0044] Optionally, the threshold motion vector difference depends on whether a pair candidate is inserted into the list or replaces an existing candidate.
[0045] Optionally, the threshold motion vector difference depends on whether the reference frames of the candidate pair or current frame have different orientations.
[0046] Optionally, the threshold motion vector difference depends on whether multiple reference frames of said candidate pair or current frame have the same POC distance or absolute value.
[0047] According to another aspect of the present invention, there is provided a method for generating a list of motion vector prediction candidates for decoding a portion of an image, the list including a pair motion vector prediction candidate constructed from a plurality of other motion vector prediction candidates, the method including determining at least one non-motion parameter for the pair candidate based on characteristics of at least one other candidate.
[0048] This method improves coding performance by increasing the relevance of non-motion parameters of pair candidates.
[0049] Optionally, to improve coding efficiency, the determining includes inheriting the at least one non-motion parameter from a first candidate in the list, preferably from both the first and second candidates in the list.
[0050] Optionally, said at least one other candidate includes one or both of a plurality of candidates used to construct said pair candidate.
[0051] Optionally, the non-motion parameters are inherited from one or both of the candidates used to construct the pair candidates.
[0052] Optionally, the non-motion parameters are inherited from one or both of the candidates used to construct the pair candidate if the candidates considered for the pair have the same reference frames and / or lists.
[0053] Optionally, the non-motion parameters are inherited from multiple candidates used to construct the pair candidates if the multiple candidates have the same parameter values.
[0054] Optionally, the parameters include parameters for a tool to compensate for illumination differences between the current block and a number of adjacent samples. Preferably, the parameters include weights for bi-prediction (BCWidx) or local illumination compensation (LIC).
[0055] Optionally, to improve coding efficiency, the method further comprises inheriting values of parameters associated with multiple hypotheses from one of a plurality of candidates used to construct the pair candidates.
[0056] Optionally, the method includes inheriting one or more parameters associated with the tool for compensating illumination only if said values differ from default values.
[0057] According to another aspect of the present invention, there is provided a method for generating a list of motion vector prediction candidates for decoding a portion of an image, the method comprising: generating a pair motion prediction candidate from two other motion vector prediction candidates; and adding the pair candidate to the list, wherein an average pair candidate is generated depending on characteristics of respective reference frames of a plurality of motion vector prediction candidates used to generate the pair motion prediction candidate.
[0058] In this method, average pair candidates are generated only when appropriate, thus preserving candidate diversity when needed and improving coding performance.
[0059] Optionally, said generating includes determining an average of said two candidates only if said respective reference frames are the same.
[0060] Optionally, the characteristics include positions of a number of reference frames in a list of reference frames for the current slice compared to the current frame.
[0061] Optionally, the average pair candidate is generated dependent on the position of a number of motion vector prediction candidates used to generate the pair motion prediction candidate.
[0062] In another aspect of the present invention, there is provided a method for generating a list of motion vector prediction candidates for decoding or encoding a portion of an image, the method comprising: obtaining a first list of motion vector prediction candidates; obtaining a second list of motion vector prediction candidates; and generating the list of motion vector prediction candidates for decoding or encoding the portion of an image from the first list and the second list of motion vector prediction candidates, wherein obtaining the second list comprises obtaining a plurality of motion vector prediction candidates for the second list; reordering at least a portion of the plurality of motion vector prediction candidates obtained for the second list; and adding at least one pair motion vector prediction candidate to the reordered candidates.
[0063] Optionally, the method comprises determining at least one of the pair motion vector prediction candidates from a top two candidates in the reordered motion vector prediction candidates for the second list.
[0064] Optionally, the added pair motion vector prediction candidate does not replace a motion vector prediction candidate in the reordered list.
[0065] Optionally, the added pair motion vector prediction candidate is generated from two candidates including a first candidate and an i-th candidate in the list, where i is between a second candidate and a maximum candidate in the second list.
[0066] Optionally, the method further comprises reordering the second list once the pairwise motion vector prediction candidate is added.
[0067] Optionally, if the second list includes a pair motion vector prediction candidate before reordering, the pair motion vector prediction candidate is retained together with pair motion vector prediction candidates added during candidate reordering or added to the reordered candidates.
[0068] Optionally, generating the list of motion vector prediction candidates for decoding or encoding the portion of the image from the first and second lists of motion vector prediction candidates comprises determining a difference value between the number of motion vector candidates in the first list and a maximum (or target) number, and including a number of motion vector prediction candidates from the second list (if available) equal to (or less than) the difference value.
[0069] In yet another aspect of the present invention, there is provided a method for generating a list of motion vector prediction candidates for decoding or encoding a portion of an image, the method comprising: obtaining a first list of motion vector prediction candidates; obtaining a second list of motion vector prediction candidates; and generating the list of motion vector prediction candidates for decoding or encoding the portion of an image from the first list of motion vector prediction candidates and the second list of motion vector prediction candidates, wherein obtaining the second list of motion vector prediction candidates comprises performing a first sorting process on a plurality of motion vector prediction candidates for the second list; and not adding a pair candidate when a cost of a candidate to be replaced in the first sorting process is evaluated and a second sorting process is performed following the addition of the pair candidate.
[0070] In another aspect of the present invention, there is provided a method for generating a list of motion vector prediction candidates for decoding or encoding a portion of an image, the method comprising: obtaining a cost of a motion vector prediction candidate in a first sorting process; and, if a position of the motion vector prediction candidate is not included in positions to be sorted using a further sorting process, using the cost obtained in the first sorting process in the further sorting process.
[0071] Optionally, the first list in any of the above aspects includes one or more neighboring motion vector prediction candidates (if available).
[0072] Optionally, the second list in any of the above aspects includes one or more non-adjacent motion vector prediction candidates (if available).
[0073] Optionally, the second list in any of the above aspects includes one or more history-based candidates.
[0074] Optionally, the second list in any of the above aspects includes one or more temporal candidates, for example, the second list may include three temporal candidates on which a reordering process is performed.
[0075] Optionally, the second list includes all possible adjacent candidates, all of which are subject to permutation (ARMC).
[0076] Further aspects of the invention relate to corresponding encoding methods, encoding devices, decoding devices and computer programs operable to perform the inventive decoding and / or encoding methods.
[0077] Further aspects of the invention are provided by the independent and dependent claims. [Brief description of the drawings]
[0078] Reference will now be made, by way of example only, to the accompanying drawings, in which:
[0079] [Figure 1] FIG. 1 is a diagram for explaining the coding structure used in HEVC. [Diagram 2] 1 is a block diagram that illustrates generally a data communications system in which one or more embodiments of the present invention may be implemented; [Diagram 3] 1 is a block diagram illustrating components of a processing device in which one or more embodiments of the present invention may be implemented. [Figure 4] 4 is a flow chart illustrating steps of an encoding method according to an embodiment of the present invention. [Diagram 5] 4 is a flow chart illustrating steps of a decoding method according to an embodiment of the present invention. [Figure 6] FIG. 2 illustrates a labeling scheme used to describe multiple blocks relative to a current block. [Figure 7] FIG. 2 illustrates a labeling scheme used to describe multiple blocks relative to a current block. [Figure 8]13(a) and (b) are diagrams illustrating affine (sub-block) mode. [Figure 9] 1A, 1B, 1C, and 1D are diagrams illustrating geometric modes. [Figure 10-1] FIG. 1 illustrates the first step of VVC merge candidate list derivation. [Figure 10-2] FIG. 1 illustrates the first step of VVC merge candidate list derivation. [Figure 11] FIG. 13 illustrates a further step in deriving a merge candidate list for VVC. [Figure 12] FIG. 13 is a diagram illustrating the derivation of pair candidates. [Figure 13] FIG. 1 illustrates a template matching method based on neighboring samples. [Figure 14-1] FIG. 11 is a modification of the first step of deriving the merge candidate list shown in FIG. [Figure 14-2] FIG. 11 is a modification of the first step of deriving the merge candidate list shown in FIG. [Figure 15] FIG. 12 is a modification of the further steps of the merge candidate list derivation shown in FIG. [Figure 16] FIG. 13 is a diagram in which the derivation of pair candidates shown in FIG. 12 is modified. [Figure 17-1] FIG. 1 illustrates the first step in deriving a merge candidate list. [Figure 17-2] FIG. 1 illustrates the first step in deriving a merge candidate list. [Figure 18] FIG. 13 is a diagram illustrating a process of rearranging the merge mode candidate list. [Figure 19] FIG. 13 is a diagram illustrating the derivation of pair candidates in the process of sorting the merge mode candidate list. [Figure 20a] 13A and 13B are diagrams illustrating an example of derivation of pair candidates after a sorting process of the merge mode candidate list. [Figure 20b] 13A and 13B are diagrams illustrating an example of derivation of pair candidates after a sorting process of the merge mode candidate list. [Figure 21]FIG. 1 illustrates a system including an encoder or decoder and a communication network according to an embodiment of the present invention. [Figure 22] 1 is a schematic block diagram of a computing device for implementing one or more embodiments of the present invention. [Diagram 23] FIG. 1 illustrates a network camera system. [Figure 24] FIG. 1 illustrates a smartphone. DETAILED DESCRIPTION OF THE PREFERRED EMBODIMENTS
[0080] Figure 1 relates to the coding structure used in the High Efficiency Video Coding (HEVC) video and Versatile Video Coding (VVC) standards. A video sequence 1 is composed of a succession of digital images i, each such digital image being represented by one or more matrices, whose coefficients represent pixels.
[0081] An image 2 of a sequence may be divided into multiple slices 3. A slice may constitute the entire image. These slices are divided into non-overlapping coding tree units (CTUs). A coding tree unit (CTU) is the basic processing unit of the High Efficiency Video Coding (HEVC) video standard and is a structure that conceptually corresponds to the macroblock unit used in some earlier video standards. A CTU is sometimes called a maximal coding unit (LCU). A CTU has luma and chroma component parts, each of which is called a coding tree block (CTB). These different color components are not shown in Figure 1.
[0082] A CTU is typically 64 pixels by 64 pixels in HEVC, but 128 pixels by 128 pixels in VVC. Each CTU can be iteratively divided into smaller variable-sized coding units (CUs) 5 using a quad-tree decomposition.
[0083] A coding unit is a basic coding element and consists of two types of subunits called prediction units (PUs) and transform units (TUs). The maximum size of a PU or TU is equal to the CU size. A prediction unit corresponds to a partition of a CU for predicting pixel values. Various different divisions of a CU into PUs are possible, such as a division into four square PUs and two different divisions into two rectangular PUs, as shown in 606. A transform unit is an element unit that undergoes spatial transformation using DCT. A CU can be divided into TUs based on a quadtree representation 607.
[0084] Each slice is embedded in one network abstraction layer (NAL) unit. Furthermore, the coding parameters of a video sequence are stored in a dedicated NAL unit called a parameter set. HEVC and H.264 / AVC employ two types of parameter set NAL units. First, the sequence parameter set (SPS) NAL unit collects all parameters that do not change throughout the video sequence. It typically deals with the coding profile, the size of the video frames, and other parameters. Second, the picture parameter set (PPS) NAL unit contains parameters that may change from one picture (or frame) of the sequence to another. HEVC also includes the video parameter set (VPS) NAL unit, which contains parameters that describe the overall structure of the bitstream. VPS is a type of parameter set defined in HEVC that applies to all layers of the bitstream. A layer may contain multiple temporal sublayers, while all version 1 bitstreams are limited to one layer. HEVC has specific layer extensions for scalability and multiview, which allow for a backward-compatible version 1 base layer and multiple layers.
[0085] VVC introduces another way of dividing an image: sub-pictures, which are independently coded groups of one or more slices.
[0086] 2 illustrates a data communication system in which one or more embodiments of the present invention may be implemented. The data communication system includes a transmitting device (in this case a server 201) operable to transmit data packets of a data stream to a receiving device (in this case a client terminal 202) via a data communication network 200. The data communication network 200 may be a wide area network (WAN) or a local area network (LAN). Such a network may be, for example, a wireless network (Wifi / 802.11a or b or g), an Ethernet network, an Internet network, or a mixed network including several different networks. In a particular embodiment of the present invention, the data communication system may be a digital television broadcasting system in which the server 201 transmits the same data content to several clients.
[0087] The data streams 204 provided by the server 201 may consist of multimedia data representing video and audio data. The audio and video data streams may, in some embodiments of the invention, be captured by the server 201 using a microphone and a camera, respectively. In some embodiments, the data streams are stored in the server 201, received by the server 201 from other data providers, or generated at the server 201. The server 201 comprises, inter alia, an encoder for encoding the video and audio streams to provide a compressed bitstream for transmission, which is a more compact representation of the data presented as input to the encoder.
[0088] In order to obtain a better ratio of the quality of the transmitted data to the amount of transmitted data, the compression of the video data may for example be according to the HEVC format or the H.264 / AVC format or the VVC format.
[0089] Client 202 receives the transmitted bitstream, decodes the reconstructed bitstream, and plays the video images on a display device and the audio data on a loudspeaker.
[0090] Although the example of FIG. 2 considers a streaming scenario, it will be appreciated that in some embodiments of the invention, data communication between the encoder and the decoder may be performed using a media storage device, such as an optical disc.
[0091] In one or more embodiments of the invention, a video image is transmitted along with data representative of correction offsets that are applied to reconstructed pixels of the image to provide filtered pixels in a final image.
[0092] 3 shows a schematic diagram of a processing device 300 configured to implement at least one embodiment of the present invention. The processing device 300 may be a device such as a microcomputer, a workstation or a lightweight handheld device. The device 300 comprises a communication bus 313 connected to: - a central processing unit 311 (denoted as CPU), such as a microprocessor; - a read-only memory 306 (denoted ROM) for storing a computer program for implementing the invention; a random access memory 312 (denoted RAM) for storing the executable code of the method according to an embodiment of the invention and registers suitable for recording variables and parameters necessary for implementing the method for encoding a sequence of digital images according to an embodiment of the invention and / or the method for decoding a bitstream according to an embodiment of the invention; and a communications interface 302 connected to a communications network 303 over which the digital data to be processed is sent and received;
[0093] Optionally, device 300 may also include the following components: - data storage means 304, such as a hard disk, for storing computer programs for implementing the methods of one or more embodiments of the invention, and data used or generated during the implementation of one or more embodiments of the invention; a disk drive 305 for a disk 306. The disk drive is adapted to read data from or write data to the disk 306; A screen 309 for displaying data and / or acting as a graphical interface with the user, via a keyboard 310 or other pointing means.
[0094] The device 300 can be connected to a variety of peripherals, such as a digital camera 320 and a microphone 308 , each connected to an I / O card (not shown), which provide multimedia data to the device 300 .
[0095] The communication bus provides communication and interoperability between the various elements included in or connected to the device 300. The representation of the bus is not limiting, and in particular the central processing unit is operable to communicate instructions to any element of the device 300 directly or by other elements of the device 300.
[0096] The disk 306 may be replaced by any information medium, such as, for example, a compact disk (CD-ROM), rewritable or not, a ZIP disk or a memory card, and generally speaking, by any information storage means readable by a microcomputer or microprocessor, whether integrated in the device or not, possibly removable, and suitable for storing one or more programs, the execution of which can implement the method of encoding a sequence of digital images and / or the method of decoding a bitstream according to the present invention.
[0097] The executable code may, as mentioned above, either be stored in the read-only memory 306, on the hard disk 304 or on a removable digital medium such as for example the disk 306. According to a variant, the executable code of the program may be received by the communications network 303 via the interface 302, to be stored in one of the storage means of the device 300, for example the hard disk 304, before being executed.
[0098] The central processing unit 311 is adapted to control and direct the execution of instructions or parts of the software code of a program or program according to the invention, instructions stored in one of the storage means mentioned above. On power-up, the program or programs stored in a non-volatile memory (for example the hard disk 304 or the read-only memory 306) are transferred to a random access memory 312, which contains the executable code of the program or programs, as well as registers for storing variables and parameters necessary for the implementation of the invention.
[0099] In this embodiment, the device is a programmable device that uses software to implement the invention, but the invention could alternatively be implemented in hardware (e.g., in the form of an application specific integrated circuit (ASIC)).
[0100] 4 is a block diagram of an encoder according to at least one embodiment of the present invention, the encoder being represented by a number of connected modules, each adapted to implement, for example in the form of programming instructions executed by the CPU 311 of the device 300, at least one corresponding step of a method for implementing at least one embodiment of encoding images of an image sequence according to one or more embodiments of the present invention.
[0101] The encoder 400 receives as input an original sequence of digital images i0 to in401. Each digital image is represented by a collection of samples, also called pixels (hereafter also referred to as pixels).
[0102] A bitstream 410 is output by the encoder 400 after the encoding process is performed. The bitstream 410 includes multiple coding units or slices, each of which includes a slice header for transmitting coded values of coding parameters used to code the slice, and a slice body that includes coded video data.
[0103] An input digital image i0 to 401 is divided by a module 402 into blocks of pixels. The blocks correspond to image portions and may be of variable size (for example 4x4, 8x8, 16x16, 32x32, 64x64, 128x128 pixels, and several rectangular block sizes can also be considered). A coding mode is selected for each input block. Two families of coding modes are provided: coding modes based on spatial predictive coding (intra prediction) and coding modes based on temporal predictive coding (inter coding, merge, SKIP). Possible coding modes are tested.
[0104] Module 403 performs an intra prediction process in which a given block to be coded is predicted by a prediction calculated from pixels in the neighborhood of the block to be coded. The indication of the selected intra prediction and the difference between a given block and its prediction are coded to provide a residual when intra coding is selected.
[0105] Temporal prediction is performed by a motion estimation module 404 and a motion compensation module 405. First, a reference image is selected from a set of reference images 416, and the motion estimation module 404 selects a part of the reference image (also called a reference region or image portion), which is the region closest to a given block to be coded (closest in terms of pixel value similarity). The motion compensation module 405 then uses the selected region to predict the block to be coded. The difference between the selected reference region and the given block (also called a residual block) is calculated by the motion compensation module 405. The selected reference region is indicated by a motion vector.
[0106] Therefore, in both cases (spatial and temporal prediction), the residual is calculated by subtracting the predicted value from the original block.
[0107] In the intra prediction implemented by module 403, the prediction direction is coded. In the inter prediction implemented by modules 404, 405, 416, 418, 417, at least one motion vector or data for identifying such a motion vector is coded for the temporal prediction.
[0108] If inter prediction is selected, the motion vector and information related to the residual block are coded. To further reduce the bit rate, assuming the motion is homogeneous, the motion vector is coded by the difference to the motion vector prediction. The motion vector prediction and coding module 417 obtains the motion vector prediction from a set of motion information prediction candidates from the motion vector field 418.
[0109] The encoder 400 further comprises a selection module 406 for selecting a coding mode applying a coding cost criterion, such as a rate-distortion criterion. To further reduce redundancy, a transform (such as a DCT) is applied to the residual block by a transform module 407, and the resulting transformed data is quantized by a quantization module 408 and entropy coded by an entropy coding module 409. Finally, the coded residual block of the current block being coded is inserted into a bitstream 410.
[0110] The encoder 400 also performs decoding of the encoded image to generate reference images (e.g., those of the reference image / picture 416) for motion estimation of subsequent images. This allows the encoder and decoder receiving the bitstream to have the same reference frame (reconstructed image or image portion is used). The inverse quantization ("dequantization") module 411 performs inverse quantization ("dequantization") of the quantized data, followed by inverse transformation by the inverse transformation module 412. The intra prediction module 413 uses the prediction information to decide which prediction to use for a given block, and the motion compensation module 414 actually adds the residual obtained by the module 412 to the reference area obtained from the set of reference images 416.
[0111] Post-filtering is then applied by module 415 to filter the reconstructed frame of pixels (image or image portion). In an embodiment of the invention, an SAO loop filter is used in which a compensation offset is added to the pixel values of the reconstructed pixels of the reconstructed image. It is understood that post-filtering does not always have to be performed. Also, other types of post-filtering can be performed in addition to or instead of SAO loop filtering.
[0112] 5 is a block diagram of a decoder 60 that may be used to receive data from an encoder according to an embodiment of the present invention. The decoder is represented by a number of connected modules, each adapted to implement a corresponding step of a method implemented by the decoder 60, for example in the form of programming instructions executed by the CPU 311 of the device 300.
[0113] The decoder 60 receives a bitstream 61 containing coding units (e.g. data corresponding to blocks or coding units), each including a header containing information about coding parameters and a body containing the coded video data. As explained with respect to Fig. 4, the coded video data is entropy coded, and an index of a motion vector prediction is coded with a given number of bits for a given block. The received coded video data is entropy decoded by module 62. The residual data is inverse quantized by module 63 before an inverse transform is applied by module 64 to obtain pixel values.
[0114] Mode data indicating the encoding mode is also entropy decoded, and intra-type or inter-type decoding is performed on the block (unit / set / group) of encoded image data based on the mode.
[0115] For intra modes, intra prediction is determined by intra prediction module 65 based on the intra prediction mode specified in the bitstream.
[0116] If the mode is inter, motion prediction information is extracted from the bitstream to find (identify) the reference region to be used by the encoder. The motion prediction information includes a reference frame index and a motion vector residual. The motion vector prediction is added to the motion vector residual by the motion vector decoding module 70 to obtain a motion vector. The various motion prediction tools used in VVC are described in more detail below with reference to Figures 6 to 10.
[0117] A motion vector decoding module 70 applies motion vector decoding to each current block that has been coded with motion prediction. Once the index of the motion vector prediction for the current block is obtained, the actual value of the motion vector associated with the current block can be decoded and used to apply motion compensation by module 66. The reference image portion indicated by the decoded motion vector is extracted from reference image 68 for applying motion compensation 66. Motion vector field data 71 is updated with the decoded motion vector for use in predicting subsequent decoded motion vectors.
[0118] Finally, a decoded block is obtained. If necessary, post-filtering is applied by a post-filtering module 67. A decoded video signal 69 is finally obtained and provided by the decoder 60.
[0119] Multiple motion estimation (inter) modes HEVC uses three different inter modes: Inter mode (Advanced Motion Vector Prediction: AMVP), "classical" merge mode (also called "non-affine merge mode" or "normal" merge mode), and "classical" merge Skip mode (also called "non-affine merge skip" mode or "normal" merge skip mode). The main difference between these modes is the data signal in the bitstream. Regarding motion vector coding, the current HEVC standard includes a contention-based scheme for motion vector prediction. This means that for each inter or merge mode (i.e., "classical / normal" merge mode or "classical / normal" merge skip mode), multiple candidates compete with a rate-distortion criterion on the encoder side to find the best motion vector prediction or the best motion information. Then, an index corresponding to the best candidate for the best prediction or motion information is inserted into the bitstream along with a "residual" that represents the difference between the predicted value and the actual value. The decoder can derive the same set of predictions or candidates and uses the best one according to the decoded index. Using the residual, the decoder can recreate the original value.
[0120] In the Screen Content Extension of HEVC, a new coding tool called Intra Block Copy (IBC) is signaled as one of these three inter modes, and the difference between IBC and its equivalent inter modes is that it checks if the reference frame is the current one. This can be implemented, for example, by checking the reference index of list L0 and inferring that if this is the last frame in that list, this is an intra block copy. Another way is to compare the picture order counts of the current and reference frames.
[0121] The design of prediction and candidate derivation is important to achieve the best coding efficiency without disproportionately impacting the complexity. In HEVC, two motion vector derivations are used: one for inter modes (Advanced Motion Vector Prediction (AMVP)) and one for merge modes (merge derivation processes - classical merge mode and classical merge skip mode). In the following, the different motion prediction modes used in VVC are described.
[0122] FIG. 6 illustrates the labeling scheme used herein to describe blocks located relative to the current block (ie, the block currently being coded / decoded) between frames (FIG. 6).
[0123] VVC Merge Mode VVC adds some inter modes compared to HEVC, in particular adding new merge modes to the usual merge modes of HEVC.
[0124] Affine mode (sub-block mode) In HEVC, only the translational motion model is applied for motion compensated prediction (MCP), but in the real world there are many kinds of motion, including zoom, rotation, perspective, and other irregular motions.
[0125] In JEM, a simplified affine transform motion compensation prediction is applied and the general principles of the affine mode are explained below based on an excerpt from document JVET-G1001 presented at the JVET Conference held in Turin on 13-21 July 2017. This document is incorporated herein in its entirety by reference insofar as other algorithms used in JEM are described.
[0126] As shown in FIG. 8(a), the affine motion field of a block is described by two control point motion vectors.
[0127] Affine mode is a motion compensation mode like inter modes (AMVP, "classical" merge, or "classical" merge skip). Its principle is to generate one motion information per pixel according to two or three neighboring motion information. In JEM, affine mode derives one motion information for each 4x4 block as depicted in Fig. 8(a) (each square is a 4x4 block, and the whole block in Fig. 8(a) is a 16x16 block divided into 16 blocks of 4x4 size squares - each 4x4 square block has a motion vector associated with it). Affine mode is available in AMVP and merge modes (i.e., classical merge mode, also called "non-affine merge mode", and classical merge skip mode, also called "non-affine merge skip mode") by enabling affine mode with a flag.
[0128] In the VVC specification, the affine mode is also called the sub-block mode.
[0129] The subblock merging mode of VVC includes a subblock-based temporal merging candidate that inherits the motion vector field of the block of the previous frame pointed to by the spatial motion vector candidate. This subblock candidate follows the inherited affine motion candidate if the neighboring block is coded in the inter-affine mode of subblock merging, and some constructed affine candidates are derived before some zero Mv candidates.
[0130] CIIP In addition to the normal and sub-block merge modes, the VVC standard also includes a Combined Inter-Merge / Intra Prediction (CIIP), also known as the Multi-Hypothesis Intra-Inter (MHII) merge mode.
[0131] Combined Inter Merge / Intra Prediction (CIIP) merge can be considered as a combination of the normal merge mode and the intra mode, and is described below with reference to FIG. 10 (partially illustrated in FIG. 10-1 and FIG. 10-2). The block prediction of the current block (1001) in this mode is the average between the merge prediction block and the intra prediction block, as depicted in FIG. 10. The merge prediction block is obtained by exactly the same process as in the merge mode, and is therefore bi-predictive of a temporal block (1002) or two temporal blocks. Therefore, the merge index is signaled for this mode in the same way as in the normal merge mode. The intra prediction block is obtained based on the neighboring samples (1003) of the current block (1001). However, the amount of intra modes available for the current block is limited compared to intra blocks. Furthermore, there is no chroma intra prediction block signaled for CIIP blocks. Chroma prediction is equal to luma prediction. As a result, 1, 2, or 3 bits are used for the intra prediction signal of a CIIP block.
[0132] The CIIP block prediction is obtained by a weighted average of the merge block prediction and the intra block prediction, where the weighting of the weighted average depends on the selected block size and / or intra prediction block.
[0133] Then, the obtained CIIP prediction is added to the residual of the current block to obtain the reconstructed block. It should be noted that CIIP mode is only valid for non-skip blocks. In fact, using CIIP skip usually leads to a worse compression performance and an increased encoder complexity. This is because CIIP modes often have the opposite block residual to other skip modes. As a result, signaling skip mode increases the bitrate. -CIIP is avoided if the current CU is skip. In fact, in VVC, the only way to signal a block residual equal to 0 for merge mode is to use skip mode, because the CU_CBF flag is inferred to be true for merge mode. And when this CBF flag is true, the block residual cannot be 0.
[0134] Thus, in this specification, CIIP should be interpreted as a mode that combines features of inter prediction and intra prediction, and not necessarily as a label given to one particular mode.
[0135] CIIP uses the same motion vector candidate list as the normal merge mode.
[0136] MMVD The MMVD merge mode is the derivation of a certain regular merge mode candidate. It can be seen as an independent merge candidate list. The MMVD merge candidate selected for the current CU is obtained by adding an offset value to one motion vector component (mvx or mvy) of the first regular merge candidate. The offset value is added to the motion vector of the first list L0 or to the motion vector of the second list L1 depending on the configuration of these reference frames (both backward, both forward, or forward and backward). The first merge candidate is indicated by an index. The offset value is indicated by a distance index between eight possible distances (1 / 4-pel, 1 / 2-pel, 1-pel, 2-pel, 4-pel, 8-pel, 16-pel, 32-pel) and a direction index giving the x- or y-axis and the sign of the offset.
[0137] In VVC, only the first two candidates in the regular merge list are used to derive the MMVD and are signaled with one flag.
[0138] Geometric division mode Geometric (GEO) merge mode is a particular bi-prediction mode. Figure 9 illustrates the generation of this particular block prediction. The block prediction includes one triangle from the first block prediction (901 or 911) and a second triangle from the second block prediction (902 or 912). However, several other possible divisions of the block are possible, as depicted in Figures 9(c) and 9(d). Geometric merge should be interpreted herein as a mode that combines features of two inter non-square predictions, and not necessarily as a label given to one particular mode.
[0139] In the example of Fig. 9(a), each partition (901 or 902) has a motion vector candidate that is a unidirectional candidate. And for each partition, an index is signaled to get the corresponding motion vector candidate in the list of unidirectional candidates at the decoder. And the first and second cannot be the same candidate. This list of candidates comes from the list of regular merge candidates, and for each candidate, one of the two components (L0 or L1) is removed.
[0140] IBC VVC also allows the enabling of intra-block copy (IBC) merge mode, which has an independent merge candidate derivation process.
[0141] Other movement information improvements DMVR Decoder-side motion vector derivation (DMVR) in VVC improves the accuracy of MV in merge mode. In this method, bilateral matching (BM) based decoder-side motion vector refinement is applied. In this bi-prediction operation, a refined MV is searched around the initial MV in reference picture list L0 and reference picture list L1. The BM method calculates the distortion between two candidate blocks in reference picture list L0 and list L1.
[0142] BDOF VVC also integrates a bidirectional optical flow (BDOF) tool. BDOF, previously called BIO, is used to refine the bidirectional prediction signal of a CU at the 4x4 subblock level. BDOF is applied when a CU satisfies some conditions, in particular when the distances (i.e., POC (Picture Order Count) difference) from two reference pictures to the current picture are the same. As the name suggests, the BDOF mode is based on the concept of optical flow, which assumes that object motion is smooth. For each 4x4 subblock, a motion refinement (v_x,v_y) is calculated by minimizing the difference between the L0 predicted sample and the L1 predicted sample. This motion refinement is then used to adjust the two predicted sample values of the 4x4 subblock.
[0143] PROF Similarly, prediction refinement by optical flow (PROF) is used for the affine modes.
[0144] AMVR and hpelIfIdx VVC also includes Adaptive Motion Vector Resolution (AMVR). AMVR allows the motion vector difference of a CU to be coded with different precision. For example, in AMVP mode, 1 / 4-luma samples, 1 / 2-luma samples, integer-luma samples, or 4-luma samples are considered. The following table from the VVC specification shows the AMVR shifts based on different syntax elements: TIFF2024533029000002.tif56170
[0145] AMVR may affect coding modes other than those using motion vector differential coding as different merging modes. Indeed, for some merging candidates, a parameter hpelIfIdx is propagated, which represents the index of the luma interpolation filter with half-pel precision. For example, in AMVP mode, hpelIfIdx is derived as follows: hpelIfIdx=AmvrShift==3?1:0
[0146] Bi-prediction with CU-level weights (BCW) In VVC, the Bi-prediction with CU-level weighting (BCW) mode is extended to allow not only simple averaging (as performed in HEVC) but also weighted averaging of the two prediction signals P0 and P1 according to the following formula: P bi-pred =((8-w)*P0+w*P1+4)≫3
[0147] Weighted average biprediction allows 5 weights, w∈{-2, 3, 4, 5, 10}.
[0148] For non-merged CUs, the weight index bcwIndex is signaled after the motion vector difference.
[0149] For a merged CU, the weight index is inferred from neighboring blocks based on the merge candidate indexes.
[0150] BCW is only used for CUs with 256 or more luma samples. In addition, for low-latency pictures, all five weights are used. For non-low-latency pictures, only three weights (w∈{3,4,5}) are used.
[0151] Normal merge list derivation In VVC, a normal merge list is derived as shown in Figures 10 and 11. First, if spatial candidates B1 (1002), A1 (1006), B0 (1010), and A0 (1014) (shown in Figure 7) exist, they are added. Then, partial redundancy is performed by adding A1 (1008) between the motion information of A1 and B1 (1007), adding B0 (1012) between the motion information of B0 and B1 (1011), and adding A0 (1016) between the motion information of A0 and A1 (1015).
[0152] When a merge candidate is added, the variable cnt is incremented (1015, 1009, 1013, 1017, 1023, 1027, 1115, 1108).
[0153] If the number of candidates in the list (cnt) is strictly less than four (1018), candidate B2 (1019) is added (1022) if it does not have the same motion information as A1 and B1 (1021).
[0154] Next, temporal candidates are added: if the bottom right candidate (1024) exists (1025), it is added (1026), otherwise the middle temporal candidate (1028) exists (1026), it is added (1029).
[0155] Next, a history-based (HMVP) is added (1101), but only if it has the same motion information as A1, B1 (1103). Furthermore, the number of history-based candidates must not exceed the maximum number of candidates in the merge candidate list minus one (1102). Therefore, there is at least one position missing in the merge candidate list after a history-based candidate.
[0156] Next, if the number of candidates in the list is at least two, we build pair candidates (1106) and add them to the merge candidate list (1107).
[0157] Next, if there are any empty positions in the merge candidate list (1109), a zero candidate is added (1110).
[0158] For spatial and history-based candidates, the parameters BCWidx and useAltHpelIf are set equal to the relevant parameters of the candidate. For spatial and zero candidates, they are set equal to the default value of 0. These default values essentially disable the method.
[0159] For paired candidates, BCldx is set to 0, and hpelIfIdxp is set to the hpelIfIdxp of the first candidate if it is equal to the hpelIfIdxp of the second candidate, otherwise it is set to 0.
[0160] Pair candidates The pair candidates are constructed (1106) according to the algorithm of FIG. 12. As depicted, when there are two candidates in the list (1201), hpelIfIdxp is derived as described above (1204, 1202, 1203). Then, the inter direction (interDir) is set equal to 0 (1205). For each list, L0 and L1, if at least one reference frame is valid (different from -1) (1207), the parameter is set. If both are valid (1208), the mv information of this candidate is derived (1209) and set equal to the reference frame of the first candidate, the motion information is the average between the two motion vectors of this list, and the variable interDir is incremented. If only one of the candidates has motion information for this list (1210), the motion information of the pair candidate is set equal to this candidate (1212, 1211) and the inter direction variable interDir is incremented.
[0161] ECM After completing the standardization of VVCv1, JVET launched Exploration Software (ECM) and started the exploration phase. ECM aims to improve the efficiency of encoding by adding additional tools and improving existing tools on top of the VVC standard.
[0162] ECM Merge Mode Among all the tools added, some merging modes have been added. Affine MMVD signal offset is for merging affine candidates as MVVD encoding in normal merging mode. Similarly, GEO_MMVD has also been added. CIIP_PDPC is an extension of CIIP. Also, two template matching merging modes have been added: normal template matching and GEO template matching.
[0163] Normal template matching is performed based on template matching estimation as shown in Figure 13. At the decoder side, for the candidate corresponding to the associated merge index and both available lists (L0, L1), motion estimation based on neighboring samples of the current block (1301) and motion estimation based on neighboring samples of corresponding multiple block positions is performed, the cost is calculated, and the motion information that minimizes the cost is selected. The motion estimation is limited by a search range, and some restrictions on this search range are also used to reduce the complexity.
[0164] In ECM, the normal template matching candidate list is based on the normal merge list, but some additional steps and parameters are added, which may result in different merge candidate lists for the same block. Moreover, the normal merge candidate list for template matching is only available with 4 candidates, while the normal merge candidate list for ECM with common test conditions defined in JVET is 10 candidates.
[0165] Derivation of normal merge lists in ECM In ECM, the normal merge list derivation has been updated. Figure 14 (divided into Figure 14-1 and Figure 14-2) and Figure 15 show the updates based on Figures 10 and 11, respectively. However, for clarity, the modules of the history-based candidates (1101) are collapsed into (1501).
[0166] In this Fig. 15, a new type of merge candidate is added: non-adjacent candidates (1540). These candidates come from blocks that are spatially located in the current frame, but not from adjacent blocks. They are selected by distance and direction. On a history basis, the list of adjacent candidates can be added until it reaches the maximum number of candidates minus one.
[0167] Duplicate Check In Figures 14 and 15, a duplication check for each candidate is added (1440, 1441, 1442, 1443, 1444, 1445, 1530). However, duplications also exist for non-adjacent candidates (1540) and history-based candidates (1501). This involves comparing the motion information of the current candidate with index cnt with the motion information of the other previous candidate. If this motion information is equal, it is considered a duplication and the variable cnt is not incremented. By motion information, of course, we mean the mutual orientation, the reference frame index and the motion vectors for each list (L0, L1).
[0168] MVTH In ECM, a motion vector threshold has been added for this overlap check. This parameter modifies the equality check by considering two motion vectors equal if the absolute value of their difference is less than or equal to the motion vector threshold MvTh for each component. In normal merging mode, MvTh is equal to 0, and in normal merging mode with template matching, it is set to a value that depends on the number of luma samples of the current CU.
[0169] AMRC In ECM, Adaptive Reordering of Merge Candidates Using Template Matching (AMRC) was added to reduce the number of bits in the merge index. Candidates are reordered based on the cost of each candidate according to the template matching cost calculated as in Figure 13. In this method, only one cost is calculated per candidate. This method is applied only to the first five candidates in the regular merge candidate list after this list is derived. It should be understood that the number five was chosen to balance the complexity of the reordering process with the potential benefits, and a larger number (e.g., all candidates) could be reordered.
[0170] FIG. 18 is an example of this method for a regular merge candidate list containing 10 candidates, such as CTC.
[0171] This method is also applied to the sub-block merging mode except for the spatial candidates, and to the regular TM mode for all four candidates.
[0172] In one proposal, this method was also extended to sort and select candidates to be included in the final merge mode candidate list. For example, in JVET-X0087, all possible non-adjacent candidates (1540) and history-based candidates (1501) are considered together with the temporally non-adjacent candidates to create a candidate list. This candidate list is created without considering the maximum number of candidates. This candidate list is then sorted. Only the correct number of candidates from this list are added to the final merge candidate list. The correct number of candidates corresponds to the first N candidates in the list. In this example, the correct number is the maximum number of candidates minus the number of spatial and temporal candidates already in the final list. In other words, the non-adjacent and history-based candidates are processed separately from the adjacent spatial and temporal candidates. The processed list is used to supplement the adjacent spatial and temporal merge candidates already present in the merge candidate list to generate the final merge candidate list.
[0173] In JVET-X0091, ARMC is used to select one temporal candidate from three temporal candidates: bi-dir, L0, and L1. The selected candidate is added to the merge candidate list.
[0174] In JVET-X0133, merge candidates are selected from multiple temporal candidates sorted using ARMC. Similarly, all neighboring candidates are subjected to ARMC and up to nine candidates are added to the merge candidate list.
[0175] All these proposed methods use classical ARMC sorting to sort the final list of merge candidates. JVET-X0087 reuses the costs calculated during sorting of non-adjacent and history-based candidates to avoid additional computational costs. JVET-X0133 applies a systematic sorting to all candidates on the final merge candidate list.
[0176] Multiple Hypothesis Prediction (MHP) ECM also adds multiple hypothesis prediction (MHP), which allows up to four motion compensated prediction signals per block (versus two in VVC). These individual prediction signals are superimposed to form an overall prediction signal. The motion parameters for each additional prediction hypothesis can be signaled explicitly by specifying a reference index, a motion vector prediction index, and a motion vector difference, or implicitly by specifying a merge index. A separate multiple hypothesis merge flag distinguishes between these two signaling modes.
[0177] For spatial, non-adjacent, and history-based merging candidates, the multiple hypothesis parameter value "addHypNeighbours" is inherited from the candidate.
[0178] In the case of temporal candidates, zero candidates and pair candidates, the multiple hypothesis parameter value "addHypNeighbours" is not retained (it is explicit).
[0179] LIC In ECM, Local Illumination Compensation (LIC) has been added, which is based on a linear model for illumination changes, which is calculated thanks to the neighboring samples of the current block and the previous block.
[0180] In ECM, LIC is only valid for unidirectional prediction. LIC is signaled by a flag. In merge mode, the LIC flag is not sent, but instead the LIC flag is inherited from the merge candidate in the following way:
[0181] For spatial candidates, non-adjacent merge candidates, and history-based merge candidates, the value of the LIC flag is inherited.
[0182] For temporal candidates and zero candidates, the LIC flag is set to 0.
[0183] For the pair candidates, the value of the LIC flag is set as shown in FIG. 16. This figure is based on FIG. 12, with modules 1620 and 1621 added and modules 1609, 1612, 1611 updated. The variable average is set to false (1620), and if the average of the pair was calculated for the current list, the LIC flag of the pair LICFlag[cnt] is set to false and the variable averageUsed is set to true (1609). If the candidate only has list motion information (1612-1611), and the average was not used, the LIC flag is updated and set equal to the OR operation of the current value and the value of the LIC flag of the candidate.
[0184] Also, if the pair candidate is Bidir (eg, 3), the LIC flag becomes false.
[0185] However, in the algorithm of FIG. 16, LICflag can be different from true only if two candidates have motion information of one list, and each candidate has its own list. For example, candidate 0 has motion information of only L0, and candidate 1 has motion information of only L1. In this case, the LIC flag can be non-zero, but this does not happen because LIC is only unidirectional. Therefore, the LIC flag of the pair is always false. As a result, the pair candidate cannot use LIC when it is potentially needed. This reduces the efficiency of the candidate and avoids the propagation of LIC to the next coding block, resulting in a reduction in coding efficiency.
[0186] Furthermore, the duplicate check in the ECM software introduces some inefficiencies. As depicted in Figures 14 and 15, each candidate is added to a list and the duplicate check (1440, 1441, 1442, 1443, 1444, 1445, 1530) only affects the increment of the variable cnt (1405, 1409, 1413, 1417, 1423, 1427, 1508). Also, as explained in Figure 16, the variable BCWidx is not initialized for pair candidates. As a result, if the last candidate added to the list is a duplicate candidate, the value of the pair candidate BCWidx is the value of the previous duplicate candidate. This does not happen in VVC, since if a candidate is determined to be a duplicate, the candidate is not added.
[0187] Embodiment In all of the following embodiments, pair candidates can be generated between two or more candidates. A pair candidate can represent a "compromise" position between the candidates that generated it, and as such represents an improvement towards an ideal motion vector prediction. By exploiting this (where appropriate), efficiency can be improved: a) Selecting the most suitable candidates to generate one candidate pair; b) use paired candidates only when appropriate, as doing otherwise would unseat candidates that could provide more diversity; c) placing the potential pair in the most appropriate location on the list; d) In some cases, a pair candidate may be useful, but in other cases, it may be too similar to existing candidates. e) Deriving other (non-motion) parameters of the candidate based on the multiple candidates that generated it.
[0188] Such modifications, especially when combined, lead to increased efficiency without significant complexity costs.Various embodiments relating to one or more of the above are described below.
[0189] In one embodiment, if a duplication check that only changes the number of candidates in the list is applied before a paired candidate, the non-motion parameters (e.g., BCWidx value) of the pair are set equal to a default value of 0. This ensures that non-motion parameters are not inherited from irrelevant or inappropriate candidates.
[0190] Normal merge mode In an embodiment, when generating a list of motion vector candidates, pair candidates are enabled depending on the type of merge mode. In particular, pair candidates are added only if the merge mode is a normal merge mode. This may include CIIP merge mode and MMVD. Pair candidates are efficient candidates because they are combinations or averages between the most likely candidates. Therefore, they create diversity for predictable content, and can be close to the ideal candidate. (An ideal candidate means a candidate that gives a perfect prediction of the current block, which is highly unlikely to exist in lossy codecs). Other merge modes are specialized for certain complex content, such as Geo, and exploit correlation between samples, such as template matching. In template matching, candidates need to be far apart, rather than close, to find the correct position within the search range, so pair candidates do not create enough diversity. Geo merge mode is designed to correctly split a block between two motions that exist in the neighborhood. Pair candidates create motion information that is not in this neighborhood. Therefore, this diversity is not necessary for Geo merge mode.
[0191] Template Matching Dependence on the type of merge mode also allows disabling (or not adding) a pair candidate for template matching merge mode, which can surprisingly increase coding efficiency. Indeed, if the pair is average, there will be a position between the two candidates, so the template matching merge mode will generate a region that is too close, and this candidate is very different from the other candidates, and is unlikely to generate a better prediction than the other candidates in template matching merge mode.
[0192] As a related embodiment, when the pair candidate is the average between candidates, the pair candidate for the template matching merge mode may be disabled (or not added).
[0193] GEO merge mode In a similar embodiment usefully combined with the above, in the geometric merge mode, geometric MMVD merge mode, or geometric template matching merge mode, the pair candidates are disabled (or not added). Similarly to the above, this ensures that the list has diversity among candidates of multiple merge modes, and as a result, improves the coding efficiency.
[0194] Similar to the merge mode of template matching, in all geometric merge modes, when the pair candidate is the average between candidates, the pair candidate is disabled (or not added).
[0195] Position of pair candidates in the merge list Surprisingly, in VVC, it has been found that pair candidates are very frequently selected even near the bottom of the list. Therefore, it has been found that setting the pair candidate early in the merge candidate list improves the coding efficiency. In fact, the top of the merge candidate list contains the most probable candidates. Therefore, the combinations of these most likely candidates give interesting candidates that are on average closest to the ideal candidate compared to other candidates.
[0196] In this embodiment, the loosest constraint on the pair candidate is not to place this candidate at the end of the merge candidate list. As a result, the constraint is (cnt < Maxcand - 1), where cnt is the position of the candidate (starting from zero) and Maxcand is the total number of candidates. Since non-adjacent candidates and / or history-based candidates (1102) can be deleted, these candidates can be added until the end of the list.
[0197] A stricter constraint on the position of a candidate pair is to enforce that its position be closer to the top of the list than to the bottom. In mathematical terms, this constraint can be expressed as (cnt<(Maxcand-1) / 2). In our example with Maxcand=10, cnt is 0,1,2,3,4 (i.e. the top half of the list).
[0198] In an additional embodiment, the pair is set immediately after the last candidate used to define it. Since the pair candidate is the average, or combination, of two candidates, if the original motion vector prediction candidate is not selected, a position between them may be better (and therefore selected).
[0199] An example of this implementation is to set the position of the pair motion vector prediction candidate to the third position (ie, immediately after the two candidates used to generate the pair candidate).
[0200] A similar, but alternative, additional implementation is to add a pair candidate to the list once two candidates are obtained. Figure 17 (partitioned into Figure 17-1 and Figure 17-2) illustrates this embodiment. It is based on Figure 14 with the addition of modules 1750, 1751, 1752, and 1753. This method requires minimal modification of the existing method and places the pair candidate at the appropriate top of the list.
[0201] In one embodiment, pair candidates are based solely on spatial candidates, as shown in Figure 17. The spatial candidates at the top of this list are the most likely candidates, so it makes sense that pairs should be combinations of these most likely candidates rather than candidates that would create more diversity when necessary.
[0202] In one alternative additional embodiment, the pair candidate is systematically set to the second position. The pair candidate mainly uses information from the first candidate by retaining its reference frame index (if it exists), and the first merge candidate in the list is the most selected candidate. In this sense, the pair candidate provides a compromise with the second candidate in the list being closer to the first candidate, and therefore more likely to be selected than the other candidate(s) used to generate the pair candidate.
[0203] Constructing (generating) candidate pairs The reordering process applied to the merge list provides an opportunity to improve the pair candidate generation process. The pair candidates are constructed using the first candidate applied during the reordering process. This increases the chance that the pair candidate represents a compromise between the two best candidates, as described above. This reordering process can be the Adaptive Reordering of Merge Candidates by Template Matching (AMRC) of ECM. Figure 19 illustrates this embodiment. In this figure, the construction of a pair candidate between two candidates C0, C1 is represented by the function pair(C0, C1).
[0204] This embodiment is more efficient than adding pairs as an initial step, since in this case the pair candidates are built based on the most probable candidates in the list, but it is more complicated since the pairs can be built only after the sorting process of some candidates is completed.
[0205] In an additional embodiment, where N is the number of sorted candidates, the pair candidates are constructed once N-1 candidates have been sorted in the list, thereby ensuring that the maximum number of candidates have been sorted before providing the pair candidates.
[0206] In a further embodiment, a pair candidate is inserted at position N and reordered using a reordering process, so candidate number N is not removed from the list but is set to position N+1. Similarly, other candidates after position N increment their positions, except for the last candidate in the list, which is removed.
[0207] In fact, it is preferable not to remove the candidate at the Nth position of the list, as opposed to the last candidate, since it may be more interesting than the candidate at the bottom of the list.
[0208] Before adding a pair candidate to the list, a validity check of the pair candidate can be performed. In an additional embodiment, the validity check of the pair candidate includes a duplicate check. However, this duplicate check is not a full duplicate check with the previous candidates 0 to N-1 as in the ECM, but also with all candidates from candidate N+1 to the maximum number of candidates.
[0209] Another way to construct a candidate pair is to use the first candidate and the candidate in the i-th position of the candidate list, where i>1. The first candidate is the most likely candidate, and therefore is likely to be closer to the ideal motion vector prediction than any other candidate.
[0210] In one embodiment, this pair candidate replaces the candidate at the i-th position in the candidate list.
[0211] This embodiment is efficient because in most cases the first candidate in the list should be closer to the ideal candidate since the first candidate is the most likely candidate, and the pair of the first candidate and the candidate in the i-th position of the candidate list should be closer to the ideal candidate than the candidate in the i-th position.
[0212] In one alternative, the constructed pair can be added without removing the candidate at the i-th position, but each candidate below the i-th position is incremented and the last candidate in the list is removed.
[0213] In a similar way to above, a validity check can be done before adding a pair candidate. The validity check of a pair candidate includes a duplicate check. However, this duplicate check is not a full duplicate check like ECM, which compares with the previous candidates 0 to N-1, but also with all candidates from candidate N+1 to the maximum number of candidates.
[0214] A particularly advantageous combination is to combine the method of pairing the first candidate with the candidate at the i-th position with a sorting process, where the first candidate is closer to the ideal prediction. Figures 20a and 20b show this and the next embodiment.
[0215] In a further embodiment, this process is applied to candidates that have not been reordered by the reordering process. This essentially means that the unreordered candidates are replaced with pair candidates generated from the candidate that replaces the first candidate. Since candidates lower on the list are less likely to be good predictions, replacing them with candidates closer to the first candidate may improve the prediction.
[0216] Figure 20b shows a reordering process similar to Figure 20a. In this embodiment, the pair candidate is removed from the list of normal merge candidate derivation. It is added during adaptive reordering of merge candidates by template matching ARMC (if it is not a duplicate, see details below). The pair candidate is built with the first two candidates reordered. The number of reordered candidates is kept constant, so the bottom candidate is removed. In this example, the number of calculated template matching costs (i.e., the number of reordered candidates) is 5 (4 first, followed by pair), but it can be more or less than 5 as mentioned above.
[0217] Also, pair candidates are restricted to only use the average of one list when the reference frames of the first and second permutation candidates are the same.Furthermore, inheritance of BCW index, LIC flag, and multiple hypothesis parameters is applied.
[0218] Furthermore, each merge candidate in the unsorted subgroup is replaced with a pair candidate between the first sorted candidate and this candidate if the resulting pair is non-overlapping.
[0219] If permutation (e.g. ARMC) is applied to all candidates in the list, it is also advantageous to include additional pair candidates. These candidates can be constructed using combinations of candidates from candidate number 1 to candidate number "max-1". However, since the pair candidates between the first and second candidates have already been considered, it is preferable to start with candidate number 2 (after candidate number 0 and candidate number 1).
[0220] Therefore, in one embodiment, even if all candidates have been reordered, for example by the ARMC process, additional pair candidates are added to the list.
[0221] However, sometimes it is advantageous not to place one pair candidate (or candidates) at the top of the candidate list, since these candidates produce the motion information closest to the most probable motion information, otherwise the list of candidates used may not be diverse enough to provide effective competition and therefore coding efficiency.
[0222] In one embodiment, the additional pair candidate starts at a predetermined position in the candidate list. This position may be a predetermined value. In a preferred embodiment, this value is 5, i.e., the additional pair candidate is added to the candidate list at the fifth position. In an alternative embodiment, this value (position) is equal to half the maximum number of candidates in the merge candidate list.
[0223] In one embodiment, the position (represented by a value) can be set equal to the position immediately after the first pair candidate in the merge candidate list when the first pair candidate (described above) is added. This provides a good position for ensuring diversity, but this embodiment is more complex than the previous embodiment using a predefined position because it requires tracking the position of the first pair candidate.
[0224] ARMC applied to a secondary candidate list (i.e. a subset of candidates) In the above description of ARMC, there are some implementations of the generation of the final merge list that includes the candidates of the second (or supplementary) list of candidates subject to reordering (ARMC) with the candidates of the first candidate list to form the final list of candidates used for decoding or encoding the image portion. For example, in JVET-X0087, the first list is considered to be spatial and temporal candidates, and the second (supplementary) list is considered to be non-adjacent and history-based candidates, while in JVET-X0091 and JVET-X0133, the (ARMC) reordered temporal candidates are considered to be the second list, and the other merge candidates to which the temporal candidates are added are considered to be the first list. However, it will be understood that the following embodiments are not limited to these specific proposals, and other possible permutations of the first and second lists of candidates are possible. More generally, in the following embodiments, ARMC is applied to the second (supplementary) list of candidates. The reordering allows the selection of one or more best candidates from the second list to be included in the final list of candidates. It is advantageous in terms of coding efficiency to consider ways to include pair candidates in the second list.
[0225] In one embodiment, the first two candidates in this sorted second list are used to add pair candidates to the second list in the same manner as when the ARMC process is applied to the final merge candidate list.
[0226] In an additional embodiment, the pair candidates are added to the second list without replacing the candidates, but the second list is reordered after the addition (e.g., AMRC processing), which may result in additional candidates and improved coding efficiency.
[0227] In one embodiment, some additional pair candidates are created and added to the second list of candidates. These candidates may be pair candidates between the 1st candidate and the ith candidate number, where i is between the maximum number of candidates in this subset and 2. The number of pair candidates may be determined as already described above with respect to the previous embodiment describing how to generate additional pair candidates.
[0228] These additional pair candidates can also be considered in the reordering (e.g., AMRC processing). Furthermore, the number of additional pair candidates is limited to four.
[0229] In one embodiment, if the second list contains the original pair (before reordering), any subsequent pair candidates added are not removed from the second list, which creates diversity since the original pair considers the first candidate before reordering, and the pair during ARMC processing considers two candidates that should be different (the last candidate is not added if they are the same).
[0230] In one embodiment, when a cost is calculated for a candidate and evaluated during one or more first ARMC processes and used in the final ARMC process to avoid additional computational costs, a pair candidate is not added to the second list if the candidate it could potentially replace was evaluated in the first sorting process, thereby reducing the complexity of adding new comparisons.
[0231] In an additional embodiment, the pair candidate is added to the final list at the position of the last candidate not evaluated in the first ARMC process of the second list, which provides the best chance of adding the pair candidate.
[0232] In one embodiment, during the final list sorting (ARMC) process, candidates with costs obtained during the first sorting process of the second list are considered for the final sorting process even if their position in the list is not in the group (or part of the candidates in the final list) to be sorted using the ARMC process.
[0233] Motion vector threshold for overlap check A further improvement is the overlap check process, specifically the threshold at which two candidates are considered to be overlapping. In one embodiment, the motion vector threshold for overlap check is adapted to pair candidates.
[0234] In one embodiment, the motion vector threshold for overlap check depends on the value of the search range of the decoder-side motion vector method or is based on the search range of the template matching. These search ranges basically define the possible positions of the motion vector predictions around the initial position that can be obtained by template matching. Therefore, two motion vector predictions that are within the search range should be considered as overlapping, not necessarily different predictions.
[0235] The decoder-side motion vector method is as follows: Decoder-side motion vector refinement (DMVR and BDOF in VVC) that depends on two blocks of adjacent samples from two reference frames (different from the current one). · Decoder-side motion vector refinement based on neighboring samples of the current block and neighboring samples of one or more reference frames (template matching in ECM). Decoder-side motion vector refinement based on block prediction of the current block (PDOF in VVC).
[0236] In one embodiment, the motion vector threshold for overlap check depends on enabling or disabling the decoder-side motion vector method. In particular, when the POC distance (or POC absolute value) between the current frame and the reference frames in each list is the same and the reference frames are in two different directions (one forward direction and the other backward direction), the decoder-side motion vector refinement is enabled. In this case, the motion vector threshold for overlap check depends on the value of the search range, and is otherwise set to a constant value.
[0237] In another embodiment, the motion vector threshold for the duplicate check depends on the position in the list of candidates used to construct the pair. Earlier candidates are more likely to be selected, and therefore pair candidates that represent compromise positions are more likely to be useful. In contrast, pair candidates that are similar to two candidates closer to the bottom of the list are less likely to be useful and should therefore be considered duplicates.
[0238] In one embodiment, for a pair of candidates that depend on the first and second candidates in the list (before or after the reordering process), the motion vector threshold for the overlap check is set to a value (0 or greater), and for a pair of candidates that depend on the first and i-th candidates, the motion vector threshold for the overlap check is set to a value greater than or equal to the first threshold.
[0239] The overlap check motion vector threshold may depend on whether the pair candidate is inserted into the list or added to replace a candidate in the list, for example, if the pair candidate is inserted into the list, the overlap check motion vector threshold is lower than or equal to the overlap check motion vector threshold of the pair candidate that replaces the candidate.
[0240] The motion vector threshold for the overlap check may depend on the reference frame of the pair candidate or the current frame. In particular, if these reference frames have different orientations, the value of the threshold is lower than if these reference frames have the same orientation (or the value is based on the search range). This is because if two similar motion vector predictions come from different reference frames, the pair candidate generated from these two candidates is more likely to be useful because they are independent of each other and more likely to be close to the ideal motion vector.
[0241] Similarly, the motion vector threshold for overlap checking may be based on whether the reference frame of the pair candidate and the reference frame of the current frame have the same POC distance (or the absolute value of the POC difference is the same). For example, if the reference frames have the same POC distance between the current frame and the pair candidate, the MV threshold will be lower, or the threshold will be based on the search range.
[0242] Deriving non-motion parameters of candidate pairs Non-motion parameters are parameters that are not related to motion prediction, e.g., tools that correct for lighting differences in the current image portion (e.g., block or coding unit). In one embodiment, all non-motion parameters of a pair candidate are set equal to the non-motion parameters from one of the candidates used to construct it. The non-motion parameters are hpelIfIdx, BCWidx, the multi-hypothesis parameter value "addHypNeighbours", and the LIC flag according to the ECM implementation.
[0243] In a preferred embodiment, the candidate from which the non-motion parameters are inherited is the first candidate.
[0244] In an additional embodiment, when constructing a pair from a first candidate and a second candidate, these non-motion parameters are set to those of the first candidate.
[0245] Alternatively, when constructing a pair candidate from the i-th candidate (i>1), the non-motion parameters are set to the parameters of the i-th candidate. In such an example, the i-th candidate may be significantly different from the first candidate, and the non-motion parameters of the first candidate are not appropriate. Furthermore, when multiple pair candidates are added, it is better to keep the diversity of the non-motion parameters of the different i-th candidates. As a counter example, if all i-th are replaced by pair candidates and all non-motion parameters are inherited from the first candidate, they will all have the same non-motion parameters as the first candidate, so the diversity is not enough to obtain better coding efficiency. This method is particularly relevant in combination with the previous embodiment described with reference to FIG. 20, in which the pair candidates are placed after the reordered candidates.
[0246] In an alternative embodiment, the parameters LICflag, hpelIfIdx, BCWidx are set equal to the values of the first candidate, and the multiple hypothesis parameter value "addHypNeighbours" is set equal to a default value indicating that the method is not applied to the current candidate. The advantage of this alternative embodiment is a reduced complexity, especially on the decoder side, with a minor impact on coding efficiency. Indeed, multiple hypotheses impact the coding and decoding times.
[0247] In another embodiment, the non-motion parameters depend on the first and second candidates in the list, for example BcwIdx is set equal to the value of the first and second candidates if both have the same value, otherwise it is set to a default value: BcwIdx=(BcwIdx[0]==BcwIdx[1])?BcwIdx[0]:default;
[0248] This also applies to hpelIfIdx and LICflag. For multiple hypothesis parameters, it is recommended to set them to default values since it becomes more complicated to compare all relevant parameters for two candidates. LICflag should also be set to default value in case of paired candidates.
[0249] The advantage of this embodiment is improved coding efficiency: since the first and second candidates in the list are likely to be the most promising candidates (especially when sorted), their parameters are also likely to be efficient, and comparing these parameters increases the chances that they will be useful for the current block.
[0250] In another embodiment, the non-motion parameters of a pair candidate depend on the characteristics of the candidates used to construct it.
[0251] For example, the non-motion parameters of a pair candidate are set equal to the parameters of the first candidate if the candidates considered in the pair have the same reference frame (and list), otherwise they are set to default values (default or values that disable the method). This is because in such a situation, the pair candidates should have the same non-motion parameters. In fact, for parameters related to illumination compensation, motion information close to that of the first candidate is expected to have similar illumination compensation. In the case of multiple hypothetical parameters, it is preferable to inherit the parameters of the most likely candidate rather than inheriting nothing. Also, for the half-pel accuracy index related to the accuracy of the motion vectors of the motion information, it is expected that the resolution of the motion information is related if the reference frames are the same and the motion information is close to the first candidate.
[0252] For example, BcwIdx is given by the following formula: BcwIdx=(C0_RefL0=C1_RefL0 and C0_RefL1=C1_RefL1)?BcwIdx[0]:default;
[0253] Here, C0_RefL0 is the reference index of list L0 of the first candidate, C0_RefL1 is the reference index of list L1 of the first candidate, C1_RefL0 is the reference index of list L0 of the second candidate, and C1_RefL1 is the reference index of list L1 of the second candidate. BcwIdx[0] is the BCldx of candidate 0. Also, (C?a:b) means that if the condition C is true, the value is set to a, otherwise it is set to b.
[0254] In an alternative embodiment, the non-motion parameters of a pair candidate are set equal to the parameters of the first candidate if this parameter is the same in both candidates used to construct the pair, otherwise they are set to default values. If both candidates have the same parameters, the pair is expected to have the same parameters as well.
[0255] for example: LICflag=LICflag[0]=LICflag[1]?LICflag[0]:default;
[0256] In one embodiment, the parameters of a pair candidate related to tools that correct the illumination difference between the current block and adjacent samples (LIC) or the illumination difference between block predictions (BCW) are set equal to the parameters of one of the candidates used to construct the pair if the candidates have the same reference frame (and list).
[0257] In one embodiment, the parameters of the pair candidate related to tools that correct the illumination difference between the current block and adjacent samples (LIC) or the illumination difference between block predictions (BCW) are set equal to the parameters of the candidate used to construct the pair if the candidates have parameter values different from the default values and have the same reference frame (and list).
[0258] For example, the LIC flag of a pair candidate can be calculated according to the following formula: LICflag=(C0_RefL0=C1_RefL0 and C0_RefL1=C1_RefL1)?(LICflag[0] OR LICflag[1])?:default;
[0259] In this example, LIC is set to the value "true" if any of the LIC flags of the two candidates are different from true, and if the pair is not a bidirectional candidate for the ECM's particular LIC implementation.
[0260] In an additional embodiment, if one or more parameters associated with the tool to correct illumination (LICflag or BCWidx) are different from the default values, the parameters associated with the multiple hypothesis parameter values "addHypNeighbours" of the pair candidate are set equal to those of one candidate used to construct the pair candidate.
[0261] In one alternative embodiment, the parameter associated with the multiple hypotheses parameter value "addHypNeighbours" of the candidate pair is set equal to that of the single candidate used to construct the candidate pair.
[0262] In a further embodiment, the non-motion parameters of the one candidate are the first candidates, in other words, for multiple hypotheses, the candidate selected to obtain the non-motion parameters is the first candidate in the list.
[0263] All these embodiments improve the current coding efficiency of the pair candidates.
[0264] Conditional construction (generation) of candidate pairs In one embodiment, the construction of pair candidates is restricted to some conditions. Surprisingly, it is found that while pair prediction candidates are frequently selected, some pair candidates may not be suitable. In the following embodiment, conditions are set for constructing pair candidates that are likely to be useful candidates.
[0265] In one embodiment, the average between the motion vectors of one list (L0, L1) of the first and second candidates used to construct the pair candidates is calculated only if the reference frames of the candidates are the same. Otherwise, the motion vector of the first candidate is set if available, otherwise the motion vector of the second candidate is set if available. This embodiment is configured by modifying the condition 1608 in FIG. 16 as follows: if (refLx[0]!=-1&&refLx[1]!=-1) and (refLx[0]==refLx[1])
[0266] In one embodiment, if the current frame has all these reference frames, or the first two of each list point only in one direction (backward / forward), or the reference frame of the pair candidate, the averaging between candidates is not enabled for the pair candidate. In this case, only the pair is a combination candidate for the low-latency configuration. The direction of the reference frame is obtained by checking the POC distance value of the reference frame and the current frame.
[0267] These conditions can be adaptively enabled or disabled depending on the positions of the candidates used to construct the pair. For example, the condition regarding reference frames having the same orientation can be used only when the pair candidate is constructed from the 1st and ith candidate positions in the merge candidate list.
[0268] All of these embodiments may be combined unless expressly stated otherwise, and in fact many combinations may be synergistic, resulting in efficiency gains greater than the sum of their parts.
[0269] Implementation of the Invention Fig. 21 shows a system 191, 195 including at least one of the encoder 150 or the decoder 100 and a communication network 199 according to an embodiment of the present invention. According to an embodiment, the system 195 is for processing and providing content (e.g. video and audio content for displaying / outputting or streaming the video and audio content) to a user having access to the decoder 100, for example via a user interface of a user terminal constituting the decoder 100 or capable of communicating with the decoder 100. Such a user terminal may be a computer, a mobile phone, a tablet or any other type of device capable of providing / displaying (provided / streamed) content to a user. The system 195 obtains / receives the bitstream 101 (in the form of a continuous stream or signal - for example while a previous video / audio is being displayed / output) via the communication network 199. According to an embodiment, the system 191 is for processing content and storing the processed content, for example the processed video and audio content for later displaying / outputting / streaming. The system 191 obtains / receives content including an original image sequence 151, which is received and processed (including filtered by a deblocking filter according to the invention) by an encoder 150, which generates a bitstream 101 to be communicated to the decoder 100 via a communication network 191. The bitstream 101 is then communicated to the decoder 100 in several ways, for example it may be pre-generated by the encoder 150 and stored as data in a storage device (e.g. on a server or cloud storage) in the communication network 199 until a user requests the content (i.e. bitstream data) from the storage device, at which point the data is communicated / streamed from the storage device to the decoder 100.The system 191 may also comprise a content providing device for providing / streaming to the user (e.g., by communicating data for a user interface to be displayed on the user terminal) content information (e.g., title of the content and other meta / storage location data for identifying, selecting and requesting the content) of the content stored in the storage device and for receiving and processing user requests for content so that the requested content is delivered / streamed from the storage device to the user terminal. Alternatively, the encoder 150 generates the bitstream 101 and communicates / streams it directly to the decoder 100 when the user requests the content. The decoder 100 then receives the bitstream 101 (or signal) and performs filtering with a deblocking filter according to the present invention to obtain / generate a video signal 109 and / or an audio signal, which is used by the user terminal to provide the requested content to the user.
[0270] Any step of the method / process according to the present invention or function described herein may be implemented in hardware, software, firmware, or any combination thereof. If implemented in software, the step / function is stored or transmitted as one or more instructions, codes, programs, or computer readable media and executed by one or more hardware-based processing units, such as a PC ("personal computer"), a DSP ("digital signal processor"), a programmable computing machine, which may be a circuit, processor and memory, a general-purpose microprocessor or central processing unit, a microcontroller, an ASIC ("application specific integrated circuit"), a field programmable logic array (FPGA), or other equivalent integrated or discrete logic circuit. Thus, the term "processor" as used herein may refer to any of the aforementioned structures, or other structures suitable for implementing the techniques described herein.
[0271] The embodiments of the present invention may also be implemented by a wide variety of devices or apparatuses, including a wireless handset, an integrated circuit (IC), or a set of ICs (e.g., a chipset). Various components, modules, or units are described herein to illustrate functional aspects of a device / apparatus configured to perform the embodiments, but need not necessarily be implemented by different hardware units. Rather, the various modules / units may be combined into a codec hardware unit or provided by a collection of interoperating hardware units including one or more processors in combination with appropriate software / firmware.
[0272] The embodiments of the present invention may be realized by a computer of a system or apparatus that reads and executes computer executable instructions (e.g., one or more programs) recorded on a storage medium to perform the modules / units / functions of one or more of the above-mentioned embodiments, and / or includes one or more processing units or circuits for performing one or more functions of the above-mentioned embodiments, and may also be realized by a method implemented by the computer of the system or apparatus, e.g., by reading and executing computer executable instructions from a storage medium to perform one or more functions of the above-mentioned embodiments, and / or by controlling one or more processing units or circuits for performing one or more functions of the above-mentioned embodiments. The computer may include a separate computer or a network of separate processing units for reading and executing the computer executable instructions. The computer executable instructions may be provided to the computer from a computer readable medium, e.g., a communication medium via a network or a tangible storage medium. The communication medium may be a signal / bit stream / carrier wave. The tangible storage medium is a "non-transitory computer-readable storage medium" and may include, for example, one or more of a hard disk, a random access memory (RAM), a read-only memory (ROM), a storage of a distributed computing system, an optical disk (such as a compact disk (CD), a digital versatile disk (DVD), a Blu-ray disk (BD)™), a flash memory device, a memory card, etc. At least some of the steps / functions may also be implemented in hardware by machines or dedicated components such as an FPGA ("field programmable gate array") or an ASIC ("application specific integrated circuit").
[0273] Fig. 22 is a schematic block diagram of a computing device 3600 for implementing one or more embodiments of the present invention. The computing device 3600 can be a device such as a microcomputer, a workstation, or a lightweight portable device. The computing device 3600 comprises a communication bus connected to: - a central processing unit (CPU) 3601, such as a microprocessor; - a random access memory (RAM) 3602 for storing executable code of a method according to an embodiment of the present invention and registers for recording variables and parameters required for implementing a method of encoding or decoding at least a part of an image according to an embodiment of the present invention, the memory capacity of which can be expanded, for example, by an optional RAM connected to an expansion port; - a read only memory (ROM) 3603 for storing computer programs for implementing an embodiment of the present invention; - a network interface (NET) 3604, typically connected to a communication network through which digital data to be processed is transmitted and received. The network interface (NET) 3604 can be a single network interface or can be composed of a collection of different network interfaces (for example, a wired interface and a wireless interface, or different types of wired or wireless interfaces). Data packets are written to the network interface for transmission or read from the network interface for reception under the control of a software application executed by the CPU 3601. The executable code can be stored either in the ROM 3603, in the HD 3606 or in a removable digital medium, such as a disk. According to a variant, the executable code of the program can be received by the communication network, via the NET 3604, to be stored in one of the storage means of the communication device 3600, such as the HD 3606, before being executed. The CPU 3601 is adapted to control and direct the execution of the program or parts of the program's instructions or software code according to an embodiment of the invention, these instructions being stored in one of the aforementioned storage means.After power-up, CPU 3601 can execute instructions from main RAM memory 3602 associated with software applications after those instructions have been loaded, for example, from program ROM 3603 or HD 3606. Such software applications, when executed by CPU 3601, cause the steps of the method according to the present invention to be performed.
[0274] It is also understood that according to another embodiment of the invention, the decoder according to the aforementioned embodiment is provided in a user terminal such as a computer, a mobile phone (cellular telephone), a table or any other type of device (e.g. a display device) capable of providing / displaying content to a user. According to yet another embodiment, the encoder according to the aforementioned embodiment is provided in an image capture device, also comprising a camera, a video camera or a network camera (e.g. a closed circuit television or a video surveillance camera) that captures and provides content for the encoder to encode. Two such examples are given below with reference to Figures 37 and 38.
[0275] FIG. 23 is a diagram showing a network camera system 3700 including a network camera 3702 and a client device 202 .
[0276] The network camera 3702 includes an imaging unit 3706 , an encoding unit 3708 , a communication unit 3710 , and a control unit 3712 .
[0277] The network camera 3702 and the client device 202 are connected to each other via the network 200 so that they can communicate with each other.
[0278] The imaging unit 3706 includes a lens and an image sensor (e.g., a charge-coupled device (CCD) or a complementary metal-oxide semiconductor (CMOS)) to capture an image of an object and generate image data based on the image. The image may be a still image or a moving image.
[0279] The encoding unit 3708 encodes the image data using the encoding method described above.
[0280] The communication unit 3710 of the network camera 3702 transmits the encoded image data encoded by the encoding unit 3708 to the client device 202 .
[0281] Furthermore, the communication unit 3710 receives commands from the client device 202. The commands include commands for setting parameters for encoding in the encoding unit 3708.
[0282] The control unit 3712 controls other units within the network camera 3702 according to commands received by the communication unit 3712 .
[0283] The client device 202 includes a communication unit 3714 , a decryption unit 3716 , and a control unit 3718 .
[0284] The communication unit 3714 of the client device 202 transmits a command to the network camera 3702 .
[0285] Furthermore, the communication unit 3714 of the client device 202 receives the encoded image data from the network camera 3712 .
[0286] The decoder 3716 decodes the encoded image data using any of the decoding methods described above, or a combination of the decoding methods described above.
[0287] The control unit 3718 of the client device 202 controls other units within the client device 202 in response to user operations or commands received by the communication unit 3714 .
[0288] The control unit 3718 of the client device 202 controls the display device 2120 to display the image decoded by the decoding unit 3716 .
[0289] In addition, the control unit 3718 of the client device 202 controls the display device 2120 to display a GUI (graphical user interface) for specifying parameter values of the network camera 3702, and also displays the encoding parameters of the encoding unit 3708.
[0290] Furthermore, the control unit 3718 of the client device 202 controls other units within the client device 202 in response to a user's operation input to the GUI displayed by the display device 2120 .
[0291] The control unit 3718 of the client device 202 controls the communication unit 3714 of the client device 202 to send a command specifying a parameter value for the network camera 3702 to the network camera 3702 in response to a user's operational input to the GUI displayed by the display device 2120.
[0292] FIG. 24 is a diagram showing a smartphone 3800.
[0293] The smartphone 3800 includes a communication unit 3802, a decoding unit 3804, a control unit 3806, and a display unit 3808.
[0294] The communication unit 3802 receives the encoded image data via the network 200 .
[0295] The decoding unit 3804 decodes the encoded image data received by the communication unit 3802 .
[0296] The decoding / encoding unit 3804 decodes / encodes the encoded image data using the decoding method described above.
[0297] The control unit 3806 controls other units within the smartphone 3800 in response to a user's operation or a command received by the communication unit 3806 .
[0298] For example, the control unit 3806 controls the display unit 3808 to display the image decoded by the decoding unit 3804. The smartphone 3800 may also include a sensor 3812 and an image recording device 3810. In this manner, the smartphone 3800 can record an image and encode the image (using the methods described above).
[0299] The smartphone 3800 can then decode the encoded image (as described above) and display it via the display device 3808 or transmit the encoded image to other devices via the communication device 3802 and the network 200.
[0300] Alternatives and fixes Although the present invention has been described with reference to the embodiments, it should be understood that the present invention is not limited to the disclosed embodiments. Those skilled in the art will understand that various changes and modifications can be made without departing from the scope of the present invention as defined in the appended claims. All features disclosed in this specification (including the appended claims, abstract and drawings) and / or all steps of any method or process so disclosed can be combined in any combination, except combinations in which at least some of such features and / or steps are mutually exclusive. Each feature disclosed in this specification (including the appended claims, abstract and drawings) can be replaced with an alternative feature serving the same, equivalent or similar purpose, unless expressly stated otherwise. Thus, unless expressly stated otherwise, each feature disclosed is an example of a generic series of equivalent or similar features.
[0301] It will also be understood that the result of the above-mentioned comparison, determination, evaluation, selection, execution, performing, or consideration, e.g., a selection made during an encoding or filtering process, may be indicated or determinable / inferable in data in the bitstream, e.g., a flag or data indicating the result, and the indicated or determined / inferred result may be used in processing, e.g., during a decoding process, instead of actually performing the comparison, determination, evaluation, selection, execution, performing, or consideration.
[0302] In the claims, the word "comprising" does not exclude other elements or steps, and the indefinite articles "a" or "an" do not exclude a plurality. The mere fact that different features are recited in mutually different dependent claims does not indicate that a combination of these features cannot be used to advantage.
[0303] Any reference numerals appearing in the claims are for illustrative purposes only and shall not be construed as limiting the scope of the claims.
Claims
1. 1. A method for generating a list of candidate motion vector predictions for decoding an image portion, comprising: generating an initial list of motion vector prediction candidates; if candidate reordering is selected for the image portion, reordering at least a portion of the initial list to generate a reordered list of motion vector prediction candidates; adding pair motion vector prediction candidates to the sorted list; A method comprising:
2. 2. The method of claim 1, comprising determining the pair of motion vector prediction candidates from the top two candidates in the sorted list.
3. The method of claim 2 , further comprising applying the reordering process to the determined pair of motion vector prediction candidates.
4. 2. The method of claim 1, wherein the portion of the initial sorted list is a maximum of the top N-1 candidates.
5. The method of claim 4, wherein the pair of motion vector prediction candidates is sorted as the Nth candidate.
6. The method of claim 1 , further comprising removing a lowest ranking candidate from the sorted list after adding the paired motion vector prediction candidate.
7. 2. The method of claim 1, wherein all candidates in the initial list are sorted to generate the sorted list.
8. 7. The method of claim 6, wherein one or more additional pairwise motion vector prediction candidates are included in the sorted list at predetermined positions.
9. 9. The method of claim 8, wherein the predetermined position is the fifth position in the sorted list.
10. 9. The method of claim 8, wherein the predetermined position is the beginning of a second half of the sorted list.
11. 7. The method of claim 6, wherein the initial list includes a first candidate motion vector pair, and the additional candidate motion vector pair is added to the sorted list immediately after the first candidate motion vector pair.
12. 1. A method for generating a list of candidate motion vector predictions for decoding an image portion, comprising: generating pair motion vector prediction candidates; adding the pair motion vector prediction candidate to a candidate list of motion vector prediction candidates; Including, The method of claim 1, wherein the pair of motion vector prediction candidates are located closer to the top of the list than to the bottom.
13. 13. The method of claim 12, further comprising adding the paired motion vector candidate to the list of motion vector prediction candidates immediately following the motion prediction candidate used to generate the paired motion vector prediction candidate.
14. 13. The method of claim 12, comprising adding the pair motion vector candidate in the list of motion vector prediction candidates immediately after the first two spatial motion prediction candidates.
15. 13. The method of claim 12, comprising adding the pair motion vector candidate to the list of motion vector prediction candidates in a second position.
16. 13. The method of claim 12, further comprising: before adding the paired motion vector prediction candidate to the list, determining whether the paired motion vector prediction candidate is similar to an existing candidate in the list.
17. 1. A method for generating a list of motion vector prediction candidates for encoding an image portion, comprising: generating an initial list of motion vector prediction candidates; if candidate reordering is selected for the image portion, reordering at least a portion of the initial list to generate a reordered list of motion vector prediction candidates; adding pair motion vector prediction candidates to the sorted list; A method comprising:
18. 1. A method for generating a list of motion vector prediction candidates for encoding an image portion, comprising: generating pair motion vector prediction candidates; adding the pair motion vector prediction candidate to a candidate list of motion vector prediction candidates; Including, The method of claim 1, wherein the pair of motion vector prediction candidates are located closer to the top of the list than to the bottom.
19. A decoding device for generating a list of candidate motion vector predictions for decoding an image portion, comprising: means for generating an initial list of motion vector prediction candidates; means for permuting at least a portion of the initial list to generate a permuted list of motion vector prediction candidates if permutation of candidates is selected for the image portion; means for adding pair motion vector prediction candidates to the sorted list; A decoding device comprising:
20. An encoding device for generating a list of candidate motion vector predictions for encoding an image portion, comprising: means for generating an initial list of motion vector prediction candidates; means for permuting at least a portion of the initial list to generate a permuted list of motion vector prediction candidates if permutation of candidates is selected for the image portion; means for adding pair motion vector prediction candidates to the sorted list; An encoding device comprising:
21. A decoding device for generating a list of candidate motion vector predictions for decoding an image portion, comprising: means for generating pair motion vector prediction candidates; means for adding the pair motion vector prediction candidate to a candidate list of motion vector prediction candidates; Including, The decoding device is characterized in that the position of the pair motion vector prediction candidate is closer to the top of the list than to the bottom.
22. An encoding device for generating a list of candidate motion vector predictions for encoding an image portion, comprising: means for generating pair motion vector prediction candidates; means for adding the pair motion vector prediction candidate to a candidate list of motion vector prediction candidates; Including, The encoding device is characterized in that the position of the pair motion vector prediction candidate is closer to the top of the list than to the bottom.
23. A computer program for causing a computer to execute the method according to claim 1 or 12.
24. A computer program for causing a computer to execute the method according to claim 17 or 18.