Methods and apparatus of enabling tools on scaled reference picture for video coding
By rescaling scaled reference pictures to match the current picture's resolution, the method enables bilateral and template matching processes, addressing the incompatibility issue and enhancing video coding efficiency and quality.
Patent Information
- Application Number
- PCT/CN2024/134394
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2023-12-29
- Filing Date
- 2024-11-26
- Publication Date
- 2025-07-03
AI Technical Summary
Existing video coding systems face issues when dealing with scaled reference pictures, as tools like Decoder-Side Motion Vector Refinement (DMVR) are disabled due to bilateral matching (BM) processes being incompatible with blocks of different resolutions.
The method involves rescaling scaled reference pictures to match the resolution of the current picture, enabling bilateral matching (BM) and template matching (TM) processes by applying rescaling techniques to ensure compatibility and enabling these tools for blocks with differing resolutions.
This approach allows for the effective utilization of BM and TM tools with scaled reference pictures, improving video coding efficiency and quality by enabling motion vector refinement and prediction information derivation across varying resolutions.
Smart Images

Figure CN2024134394_03072025_PF_FP_ABST
Abstract
Description
METHODS AND APPARATUS OF ENABLING TOOLS ON SCALED REFERENCE PICTURE FOR VIDEO CODINGCROSS REFERENCE TO RELATED APPLICATIONS
[0001] The present invention is a non-Provisional Application of and claims priority to U. S. Provisional Patent Application No. 63 / 615, 817, filed on December 29, 2023. The U. S. Provisional Patent Application is hereby incorporated by reference in its entirety.FIELD OF THE INVENTION
[0002] The present invention relates to video coding system. In particular, the present invention relates to situations where one or more reference pictures have different resolution from the current picture in a video coding system. BACKGROUND AND RELATED ART
[0003] Versatile video coding (VVC) is the latest international video coding standard developed by the Joint Video Experts Team (JVET) of the ITU-T Video Coding Experts Group (VCEG) and the ISO / IEC Moving Picture Experts Group (MPEG) . The standard has been published as an ISO standard: ISO / IEC 23090-3: 2021, Information technology -Coded representation of immersive media -Part 3: Versatile video coding, published Feb. 2021. VVC is developed based on its predecessor HEVC (High Efficiency Video Coding) by adding more coding tools to improve coding efficiency and also to handle various types of video sources including 3-dimensional (3D) video signals.
[0004] Fig. 1A illustrates an exemplary adaptive Inter / Intra video encoding system incorporating loop processing. For Inter / Intra Prediction 110, the prediction data is derived based on previously coded video data in the current picture (i.e., Intra prediction) or previous reconstructed reference picture (s) . The prediction data is subtracted from the input data using Adder 112 to form prediction errors, also called residues. The prediction errors are then processed by Transform (T) 114 followed by Quantization (Q) 116. The transformed and quantized residues are then coded by Entropy Encoding 120 to be included in a video bitstream corresponding to the compressed video data. The bitstream associated with the residues is then packed with side information such as motion and coding modes associated with Inter / Intra prediction, and other information such as parameters associated with loop filters applied to underlying video area. The encoder also needs reconstructed data to derive the inter / intra prediction. Accordingly, the transformed and quantized residues are processed by Inverse Quantization (IQ) 126 and Inverse Transformation (IT) 124 to reconstruct the residues. The residues are then added back to prediction data at Reconstruction (REC) 128 to form reconstructed video data. The reconstructed video data may be stored in Frame Buffer 140 and used for prediction of other frames.
[0005] As shown in Fig. 1A, incoming video data undergoes a series of processing in the encoding system. The reconstructed video data from REC 128 may be subject to various impairments due to lossy processing (e.g. quantization) . Accordingly, in-loop filter processing is often applied to the reconstructed video data before the reconstructed video data are stored in the Frame Buffer 140 in order to improve video quality. For example, deblocking filter (DF) 130, Sample Adaptive Offset (SAO) 132 and Adaptive Loop Filter (ALF) 134 may be used. The loop filter information may need to be incorporated in the bitstream so that a decoder can properly recover the required information.
[0006] The decoder, as shown in Fig. 1B, can use similar or portion of the reconstruction blocks as the encoder except for Transform 114 and Quantization 116 since the decoder only needs Inverse Quantization 126 and Inverse Transform 124. Instead of Entropy Encoding 120, the decoder uses an Entropy Decoding 160 to decode the video bitstream into quantized transform coefficients and needed coding information (e.g. in-loop filter information, Intra prediction information and Inter prediction information) . The Inter / Intra prediction 150 at the decoder side does not need to perform the Intra mode search nor motion estimation. Instead, the decoder only needs to generate Inter / Intra prediction according to Inter / Intra prediction information received from the encoder. The decoder also uses Frame Buffer to store reconstructed pictures for Inter / Intra prediction. The Inter / Intra prediction data derived by using Inter / Intra Prediction is added to the reconstructed residues from IT 124 using Adder 112. The combined data is provided to REC 128 to form reconstructed pictures.
[0007] In VVC, the reference pictures in the reference list are called scaled reference pictures if one or more of the following seven parameters are different from those of the current picture: 1) the picture width in luma samples (pps_pic_width_in_luma_samples) , 2) the picture height in luma samples (pps_pic_height_in_luma_samples) , 3) the scaling window left offset (pps_scaling_win_left_offset) , 4) the scaling window right offset (pps_scaling_win_right_offset) , 5) the scaling window top offset (pps_scaling_win_top_offset) , 6) the scaling window botton offset (pps_scaling_win_bottom_offset) , 7) the number of sub pictures -1 (sps_num_subpics_minus1) .
[0008] When one of the reference pictures is a scaled reference picture, we may encounter some problems for some video coding tools.
[0009] For example, In ECM-2.0, a multi-pass Decoder-Side Motion Vector Refinement (DMVR) method is applied in regular merge mode if the selected merge candidate meets the DMVR conditions. In the first pass, bilateral matching (BM) is applied to the coding block. In the second pass, BM is applied to each 16x16 subblock within the coding block. In the third pass, MV in each 8x8 subblock is refined by applying Bi-Directional Optical Flow (BDOF) . An example of the BM processing is shown in Fig. 2.
[0010] If one of the reference pictures is in different resolution which meets the condition of scaled reference picture, the DMVR tool is disabled due to the BM process cannot be applied to the two blocks with different resolution.
[0011] In the present invention, methods and apparatus to overcome the issue that a reference has different resolution from the current picture.
[0012] BRIEF SUMMARY OF THE INVENTION
[0013] A method and apparatus for video coding with one or more scaled reference pictures are disclosed. According to one method, input data associated with a current block in a current picture is received, wherein the input data comprise pixel data to be encoded at an encoder side or coded data associated with the current block to be decoded at a decoder side. A first reference picture in reference picture list 0 and a second reference picture in reference picture list 1 are determined. When any one of the first reference picture and the second reference picture is a scaled reference picture: BM (Bilateral Matching) process is applied to derive prediction information for the current block by using the first reference picture in reference picture list 0 and the second reference picture in reference picture list 1, wherein the scaled reference picture has at least one different resolution parameter from the current picture. The current block is encoded or decoded using coding information comprising the prediction information.
[0014] In one embodiment, when the first reference picture, the second reference picture, or both are the scaled reference picture, the first reference picture, the second reference picture, or both are rescaled respectively to a same resolution as the current picture, and the BM process is applied to rescaled first reference picture, rescaled second reference picture or both respectively.
[0015] In one embodiment, when a target reference picture corresponding to the first reference picture or the second reference picture is the scaled reference picture, a reference block in the target reference picture for the BM process is derived from the target reference picture by rescaling a corresponding block in the target reference picture. In one embodiment, when a target reference picture corresponding to the first reference picture or the second reference picture is the scaled reference picture, BM cost between two bilateral blocks associated with the first reference picture and the second reference picture is calculated for Decoder-Side Motion Vector Refinement (DMVR) to derive a refined motion vector or for Bi-Directional Optical Flow (BDOF) to derive a motion vector. In one embodiment, at least one of the two bilateral blocks is derived from a rescaled reference picture or by rescaling a corresponding block in a target scaled reference picture.
[0016] In one embodiment, when the first reference picture and the second reference picture have different resolution, at least one of the first reference picture and the second reference picture is rescaled to generate a post first reference picture and a post second reference picture so that the post first reference picture and the post second reference picture have same resolution. In one embodiment, the same resolution corresponds to minimum resolution or maximum resolution of the first reference picture and the second reference picture, or any other resolution. In one embodiment, after said at least one of the first reference picture and the second reference picture is rescaled to generate the post first reference picture and the post second reference picture, the BM process is applied to the post first reference picture and the post second reference picture to generate a temporary prediction block, and a final prediction block is generated from the temporary prediction block. In one embodiment, after said at least one of the first reference picture and the second reference picture is rescaled to generate the post first reference picture and the post second reference picture , the BM process is applied to the post first reference picture and the post second reference picture to generate a first reference block for the reference picture list 0 and a second reference block for the reference picture list 1, two rescaled reference blocks are derived by rescaling the first reference block and the second reference block respectively and a final prediction block is generated from the two rescaled reference blocks.
[0017] In one embodiment, after the BM process is applied to derive or refine an MV (Motion Vector) for the current block, a final motion compensation predictor is generated by using a high-resolution reconstructed reference picture, a reconstructed reference picture in pre-LF (Loop Filter) resolution or after-LF resolution.
[0018] According to another method, a reference picture is determined. When the reference picture is a scaled reference picture: TM (Template Matching) process is applied to derive prediction information for the current block by using the reference picture; and the current block is encoded or decoded using coding information comprising the prediction information.
[0019] In one embodiment, the reference picture is rescaled to a rescaled reference picture having a same resolution as the current picture and TM cost is determined based on the current picture and the rescaled reference picture. In one embodiment, after the TM process is applied to derive or refine an MV (Motion Vector) for the current block, a final motion compensation predictor is generated by using a high-resolution reconstructed reference picture, a reconstructed reference picture in pre-LF (Loop Filter) resolution or after-LF resolution.
[0020] According to yet another method, a reference picture is determined. When the reference picture is a scaled reference picture: the scaled reference picture is rescaled to a rescaled reference picture having a same resolution as the current picture; and a reconstruction block is generated by using the rescaled reference picture.
[0021] In one embodiment, information related to CU (Coding Unit) , PU (Prediction Unit) or TU (Transform Unit) is determined from the scaled reference picture prior to said rescaling the scaled reference picture, and wherein the information related to the CU, the PU or the TU excludes pixel value.
[0022] In one embodiment, one or more constraints on equivalent resolution between the scaled reference picture and the current picture are removed.
[0023] In one embodiment, an on / off control flag to indicate whether said rescaling the scaled reference picture and said generating the reconstruction block by using the rescaled reference picture are applied or not is signalled or parsed at a slice level, picture level, sequence level or a combination thereof.BRIEF DESCRIPTION OF THE DRAWINGS
[0024] Fig. 1A illustrates an exemplary adaptive Inter / Intra video coding system incorporating loop processing.
[0025] Fig. 1B illustrates a corresponding decoder for the encoder in Fig. 1A.
[0026] Fig. 2 illustrates an example of Bilateral Matching process.
[0027] Fig. 3 illustrates an example of rescaling a scaled reference picture according to an embodiment of the present invention, where reference picture in L0 is rescaled to have the same resolution as the current picture.
[0028] Fig. 4 illustrates an example of applying BM to a scaled reference picture according to an embodiment of the present invention, where a reference block for BM is generated from the scaled reference picture in L0.
[0029] Fig. 5 illustrates an example of calculating BM cost for DMVR on a scaled reference picture according to an embodiment of the present invention, where a rescaled block in L0 for BM cost calculation is generated according to a proposed method.
[0030] Fig. 6 illustrates an example of calculating BM cost for BDOF on a scaled reference picture according to an embodiment of the present invention, where a rescaled block in L0 for BM cost calculation is generated according to a proposed method.
[0031] Fig. 7 illustrates an example of both reference pictures in L0 / L1 being scaled pictures where both reference pictures are scaled to have same resolution.
[0032] Fig. 8 illustrates an example of generating a temporary prediction block with size (w’ x h’) using two blocks in L0 and L1, and then generating the final prediction block with size (w x h) from the temporary prediction block for BDOF.
[0033] Fig. 9 illustrates an example of calculating the BM cost by using the (w’ x h’) blocks in L0 and L1 and to generating the final prediction block from two corresponding (w x h) blocks generated from L0 and L1.
[0034] Fig. 10 illustrates an example of TM process for the case involved with a scaled reference picture.
[0035] Fig. 11 illustrates an example of motion compensation on the scaled reference picture at position (x, y) , where the position is translated to (x’, y’) according to the resolution of the current and reference picture.
[0036] Fig. 12 illustrates an example of motion compensation on the scaled reference picture at position (x, y) , where the reference picture is rescaled to a rescaled reference picture having the same resolution as the current picture and a prediction block is generated based on the rescaled reference picture.
[0037] Fig. 13 illustrates a flowchart of an exemplary video coding system that applies bilateral Matching (BM) process with scaled reference picture (s) according to an embodiment of the present invention.
[0038] Fig. 14 illustrates a flowchart of an exemplary video coding system that applies Template Matching (TM) process with scaled reference picture (s) according to an embodiment of the present invention.
[0039] Fig. 15 illustrates a flowchart of an exemplary video coding system that generates reconstruction with scaled reference picture (s) according to an embodiment of the present invention.DETAILED DESCRIPTION OF THE INVENTION
[0040] It will be readily understood that the components of the present invention, as generally described and illustrated in the figures herein, may be arranged and designed in a wide variety of different configurations. Thus, the following more detailed description of the embodiments of the systems and methods of the present invention, as represented in the figures, is not intended to limit the scope of the invention, as claimed, but is merely representative of selected embodiments of the invention. References throughout this specification to “one embodiment, ” “an embodiment, ” or similar language mean that a particular feature, structure, or characteristic described in connection with the embodiment may be included in at least one embodiment of the present invention. Thus, appearances of the phrases “in one embodiment” or “in an embodiment” in various places throughout this specification are not necessarily all referring to the same embodiment.
[0041] Furthermore, the described features, structures, or characteristics may be combined in any suitable manner in one or more embodiments. One skilled in the relevant art will recognize, however, that the invention can be practiced without one or more of the specific details, or with other methods, components, etc. In other instances, well-known structures, or operations are not shown or described in detail to avoid obscuring aspects of the invention. The illustrated embodiments of the invention will be best understood by reference to the drawings, wherein like parts are designated by like numerals throughout. The following description is intended only by way of example, and simply illustrates certain selected embodiments of apparatus and methods that are consistent with the invention as claimed herein.
[0042] As mentioned before, when one of the reference pictures is in different resolution that meets the condition of scaled reference picture, the DMVR tool is disabled due to the BM process cannot be applied on the two blocks with different resolution. In the present invention, various techniques to enable bilateral matching (TM) tools for two blocks with different resolution.
[0043] Scheme 1: Enabling BM Tools on Scaled Reference Picture
[0044] In one embodiment, the bilateral matching (BM) related tools can be enabled even when one or both reference pictures are scaled. To perform the bilateral matching on the two reference pictures, the following embodiment can be applied.
[0045] In one embodiment, the scaled reference picture is rescaled to the same resolution as the current picture, and the BM process can be the same as the non-scaled reference picture condition. When a scaled reference picture is rescaled, the scaled reference picture after rescaling is referred as a rescaled reference picture in this disclosure.
[0046] For example, the reference picture L0 is a scaled reference picture. To do bilateral matching, the reference L0 is rescaled to the same resolution as current picture as shown in Fig. 3. After generating the rescaled reference picture L0’, the bilateral matching process can be done on the block 310 in reference picture L1 and block 320 in reference picture L0’, which are the same size with current block after rescaling. Note that the reference picture L0 is a “scaled” reference picture in this example. After the reference picture is rescaled, the reference picture is referred as a “rescaled” reference picture in this disclosure.
[0047] In one embodiment, the block used for the BM process is generated from the scaled reference picture as shown in Fig. 4. For example, the reference L0 is a scaled reference picture. To do bilateral matching, the rescaled block 420 used for BM in L0 is generated from the scaled reference picture L0. To generate the rescaled block 420 with the same size as the current block, the VVC RPR inter filter can be used, or any other subsample / interpolation method can be applied. For example, the corresponding block 430 in the scaled reference picture L0 is identified and the corresponding block 430 is rescaled using the VVC RPR inter filter or any other subsample / interpolation method to generate the rescaled block 420.
[0048] After applying the above embodiment, all the BM related tools can be supported. For example, following the above example where reference picture L0 is a scaled reference picture, DMVR can be done by using the BM cost between the blocks in L0 and L1 as shown in Fig. 5, where the rescaled block (520a or 520b) in L0 and the corresponding block (510a or 510b) for the BM process are shown. The block (520a or 520b) in L0 can be generated from one of the above embodiments such as block 320 in Fig. 3 (i.e., derived from the rescaled reference picture) or block 420 in Fig. 4 (i.e., rescaling a corresponding block in the scaled reference picture) .
[0049] BDOF can also be done by using the BM cost of these two blocks as shown in Fig. 6. For example, following the above example where reference picture L0 is a scaled reference picture, the block 620 used for BDOF processing can be generated using one of the above embodiments such as block 320 in Fig. 3 (i.e., derived from the rescaled reference picture) or block 420 in Fig. 4 (i.e., rescaling a corresponding block in the scaled reference picture) .
[0050] In one embodiment, when one or both reference pictures in L0 / L1 are scaled reference pictures, both reference pictures should be at the same resolution. The resolution can be either minimum or maximum resolution among the two reference pictures or any other resolution. When the reference pictures are not at the same resolution, the reference pictures should be rescaled. After that, the BM cost can be calculated at the scaled reference pictures directly. Since the rescaling process may be applied to only one of the two reference pictures, the other reference picture may be the original reference picture without rescaling. For convenience, after the rescaling process to cause the resolution to be the same, the reference pictures are referred as post reference pictures. Accordingly, a post reference picture can be a rescaled reference picture or an original reference picture without rescaling.
[0051] In one example, both reference pictures in L0 / L1 are scaled pictures and with same resolution as shown in Fig. 7. When predicting the current block with size (w x h) , the bilateral matching can be done on block 710 in reference picture L1 and block 720 in reference picture L0 with block size (w’ x h’) . The relation between w’, w, h’a nd h can depend on the resolution of the current picture and reference picture.
[0052] In one embodiment, to generate the prediction of the current block with size (w x h) , we can generate a temporary prediction block 810 with size (w’a nd h’) using the two blocks 820 and 830 in L0 and L1, then the final prediction block with size (w x h) can be generate from the temporary prediction block 810 as shown in Fig. 8.
[0053] For example, after processing BDOF using these two blocks 820 and 830 with size (w’a nd h’) , the temporary prediction will be block 810 with size (w’a nd ‘h) as shown in Fig. 8. The final prediction with size (w x h) can be generated from block 810.
[0054] In another embodiment, the BM cost is calculated by using the (w’ x h’) blocks in L0 and L1. However, to generate the final prediction block, the corresponding (w x h) blocks 910 and 920 are generated first from L0 and L1, and the final prediction block can be generated from the two blocks (910 and 920) as shown in Fig. 9.
[0055] In one embodiment, the MV-derivation / refinement related part could use the methods proposed above, however, when generating the final motion compensation predictor, the high-resolution reconstructed reference picture is used.
[0056] In another embodiment, the MV-derivation / refinement related part could use the methods proposed above, however, when generating the final motion compensation predictor, the reference picture that reconstructed in its pre-LF resolution is used.
[0057] In another embodiment, the MV-derivation / refinement related part could use the methods proposed above, however, when generating the final motion compensation predictor, the reference picture that reconstructed in its after-LF resolution is used.
[0058] Scheme 2: Enabling TM Tools on Scaled Reference Picture
[0059] In one embodiment, the template matching (TM) related tools can be enabled even when the reference pictures are scaled. To do template matching on the reference picture, the following embodiment can be applied.
[0060] In one embodiment, the scaled reference picture is rescaled to the same resolution as the current picture, and the TM process can be the same as the non-scaled reference picture condition.
[0061] For example, to get the TM cost of current block and reference block, the reference picture is rescaled to the same resolution as the current picture, and the template cost can be calculated by using the rescaled reference picture. For example, in ECM we may want to calculate the L shape TM cost. In the following example, the L shape TM cost can be calculated between the current picture and the rescaled reference picture as shown in Fig. 10.
[0062] In one embodiment, the MV-derivation / refinement related part can use the methods proposed above. However, when generating the final motion compensation predictor, the high-resolution reconstructed reference picture is used.
[0063] In another embodiment, the MV-derivation / refinement related part can use the methods proposed above, however, when generating the final motion compensation predictor, the reference picture that is reconstructed in its pre-LF resolution is used.
[0064] In another embodiment, the MV-derivation / refinement related part can use the methods proposed above, however, when generating the final motion compensation predictor, the reference picture that is reconstructed in its after-LF resolution is used.
[0065] Scheme 3: Scaled Reference Picture with Rescaled Reconstruction
[0066] In one embodiment, when the reference picture is a scaled reference picture, we can rescale the reconstruction to the same resolution. When we want to get pixel information from the reconstruction, we can directly use the rescaled reconstruction. In this way, the complex process of position translation can be avoided. When we want to get CU / PU / TU or any other information exclude pixel value, we can still use the origin reference picture, and the position translation process is needed.
[0067] For example, in VVC, when motion compensation on the scaled reference picture at position (x, y) is performed, the position should be translated to (x’, y’) according to the resolution of the current and reference picture. After getting the position, we should use the specified filter to generate the final MC result as shown in Fig. 11.
[0068] With the embodiment applied, the motion compensation can be done in the same way as the non-scaled reference picture case since the reference picture L0’ is rescaled to the same resolution as current picture as shown in Fig. 12. For example, block 1210 is generated from the reference picture L0’ using the motion compensation process of non-scaled reference picture.
[0069] In one embodiment, when the above embodiment is applied, the constraint on equivalent resolution between the reference and current pictures can be removed.
[0070] In one embodiment, an on / off control flag is signalled at a slice level, picture level and / or sequence level to indicate whether the method is applied or not.
[0071] Any of the foregoing proposed methods of BM process on a scaled reference picture can be implemented in encoders and / or decoders. For example, any of the proposed methods can be implemented in one module of an encoders and / or decoder. Alternatively, any of the proposed methods can be implemented as a circuit coupled to one module of the encoders and / or decoder, so as to provide the information needed by the module used in encoders and / or decoder. The proposed methods of BM process on a scaled reference picture can be implemented in an encoder side or a decoder side. For example, with reference to the encoder and decoder in Fig. 1A and Fig. 1B, any of the proposed methods can be implemented in an Intra / Inter coding module (e.g. Intra Pred. 150 / MC 152 in Fig. 1B) in a decoder or an Intra / Inter coding module is an encoder (e.g. Intra Pred. 110 / Inter Pred. 112 in Fig. 1A) . However, the decoder or encoder may also use additional processing unit to implement the proposed method. While the Intra Pred. units (e.g. unit 110 / 112 in Fig. 1A and unit 150 / 152 in Fig. 1B) are shown as individual processing units, they may correspond to executable software or firmware codes stored on a media, such as hard disk or flash memory, for a CPU (Central Processing Unit) or programmable devices (e.g. DSP (Digital Signal Processor) or FPGA (Field Programmable Gate Array) ) .
[0072] Fig. 13 illustrates a flowchart of an exemplary video coding system that applies bilateral Matching (BM) process with scaled reference picture (s) according to an embodiment of the present invention. The steps shown in the flowchart may be implemented as program codes executable on one or more processors (e.g., one or more CPUs) at the encoder side. The steps shown in the flowchart may also be implemented based hardware such as one or more electronic devices or processors arranged to perform the steps in the flowchart. According to one method, input data associated with a current block in a current picture is received in step 1310, wherein the input data comprise pixel data to be encoded at an encoder side or coded data associated with the current block to be decoded at a decoder side. A first reference picture in reference picture list 0 and a second reference picture in reference picture list 1 are determined in step 1320. When any one of the first reference picture and the second reference picture is a scaled reference picture is check in step 1330. When any one of the first reference picture and the second reference picture is a scaled reference picture (i.e., the “Yes” path from step 1330) , steps 1340 and 1350 are performed. Otherwise (i.e., the “No” path from step 1330) , steps 1340 and 1350 are skipped. In step 1340, BM (Bilateral Matching) process is applied to derive prediction information for the current block by using the first reference picture in reference picture list 0 and the second reference picture in reference picture list 1, wherein the scaled reference picture has at least one different resolution parameter from the current picture. In step 1350, the current block is encoded or decoded using coding information comprising the prediction information.
[0073] Fig. 14 illustrates a flowchart of an exemplary video coding system that applies Template Matching (TM) process with scaled reference picture (s) according to an embodiment of the present invention. According to this method, input data associated with a current block in a current picture is received in step 1410, wherein the input data comprise pixel data to be encoded at an encoder side or coded data associated with the current block to be decoded at a decoder side. A reference picture is determined in step 1420. Whether the reference picture is a scaled reference picture is checked in step 1430. When the reference picture is a scaled reference picture (i.e., the “Yes” path from step 1430) , steps 1440 and 1450 are performed. Otherwise (i.e., the “No” path from step 1430) , steps 1440 and 1450 are skipped. In step 1440, TM (Template Matching) process is applied to derive prediction information for the current block by using the reference picture; and in step 1440, the current block is encoded or decoded using coding information comprising the prediction information.
[0074] Fig. 15 illustrates a flowchart of an exemplary video coding system that generates reconstruction with scaled reference picture (s) according to an embodiment of the present invention. According to this method, input data associated with a current block in a current picture is received, wherein the input data comprise pixel data to be encoded at an encoder side or coded data associated with the current block to be decoded at a decoder side. A reference picture is determined. Whether the reference picture is a scaled reference picture is checked in step 1530. When the reference picture is a scaled reference picture (i.e., the “Yes” path from step 1530) , steps 1540 and 1550 are performed. Otherwise (i.e., the “No” path from step 1530) , steps 1540 and 1550 are skipped. In step 1540 the scaled reference picture is rescaled to a rescaled reference picture having a same resolution as the current picture; and in step 1550 a reconstruction block is generated by using the rescaled reference picture.
[0075] The flowcharts shown are intended to illustrate an example of video coding according to the present invention. A person skilled in the art may modify each step, re-arranges the steps, split a step, or combine steps to practice the present invention without departing from the spirit of the present invention. In the disclosure, specific syntax and semantics have been used to illustrate examples to implement embodiments of the present invention. A skilled person may practice the present invention by substituting the syntax and semantics with equivalent syntax and semantics without departing from the spirit of the present invention.
[0076] The above description is presented to enable a person of ordinary skill in the art to practice the present invention as provided in the context of a particular application and its requirement. Various modifications to the described embodiments will be apparent to those with skill in the art, and the general principles defined herein may be applied to other embodiments. Therefore, the present invention is not intended to be limited to the particular embodiments shown and described, but is to be accorded the widest scope consistent with the principles and novel features herein disclosed. In the above detailed description, various specific details are illustrated in order to provide a thorough understanding of the present invention. Nevertheless, it will be understood by those skilled in the art that the present invention may be practiced.
[0077] Embodiment of the present invention as described above may be implemented in various hardware, software codes, or a combination of both. For example, an embodiment of the present invention can be one or more circuit circuits integrated into a video compression chip or program code integrated into video compression software to perform the processing described herein. An embodiment of the present invention may also be program code to be executed on a Digital Signal Processor (DSP) to perform the processing described herein. The invention may also involve a number of functions to be performed by a computer processor, a digital signal processor, a microprocessor, or field programmable gate array (FPGA) . These processors can be configured to perform particular tasks according to the invention, by executing machine-readable software code or firmware code that defines the particular methods embodied by the invention. The software code or firmware code may be developed in different programming languages and different formats or styles. The software code may also be compiled for different target platforms. However, different code formats, styles and languages of software codes and other means of configuring code to perform the tasks in accordance with the invention will not depart from the spirit and scope of the invention.
[0078] The invention may be embodied in other specific forms without departing from its spirit or essential characteristics. The described examples are to be considered in all respects only as illustrative and not restrictive. The scope of the invention is therefore, indicated by the appended claims rather than by the foregoing description. All changes which come within the meaning and range of equivalency of the claims are to be embraced within their scope.
Claims
1.A method of video coding, the method comprising:receiving input data associated with a current block in a current picture, wherein the input data comprise pixel data to be encoded at an encoder side or receiving coded data associated with the current block to be decoded at a decoder side;determining a first reference picture in reference picture list 0 and a second reference picture in reference picture list 1; andwhen any one of the first reference picture and the second reference picture is a scaled reference picture:applying BM (Bilateral Matching) process to derive prediction information for the current block by using the first reference picture in reference picture list 0 and the second reference picture in reference picture list 1, wherein the scaled reference picture has at least one different resolution parameter from the current picture; andencoding or decoding the current block using coding information comprising the prediction information.2.The method of Claim 1, wherein when the first reference picture, the second reference picture, or both are the scaled reference picture, the first reference picture, the second reference picture, or both are rescaled respectively to a same resolution as the current picture, and the BM process is applied to rescaled first reference picture, rescaled second reference picture or both respectively.3.The method of Claim 1, wherein when a target reference picture corresponding to the first reference picture or the second reference picture is the scaled reference picture, a reference block in the target reference picture for the BM process is derived from the target reference picture by rescaling a corresponding block in the target reference picture.4.The method of Claim 1, wherein when a target reference picture corresponding to the first reference picture or the second reference picture is the scaled reference picture, BM cost between two bilateral blocks associated with the first reference picture and the second reference picture is calculated for Decoder-Side Motion Vector Refinement (DMVR) to derive a refined motion vector or for Bi-Directional Optical Flow (BDOF) to derive a motion vector.5.The method of Claim 4, wherein at least one of the two bilateral blocks is derived from a rescaled reference picture or by rescaling a corresponding block in a target scaled reference picture.6.The method of Claim 1, wherein when the first reference picture and the second reference picture have different resolution, at least one of the first reference picture and the second reference picture is rescaled to generate a post first reference picture and a post second reference picture so that the post first reference picture and the post second reference picture have same resolution.7.The method of Claim 6, wherein the same resolution corresponds to minimum resolution or maximum resolution of the first reference picture and the second reference picture, or any other resolution.8.The method of Claim 6, wherein after said at least one of the first reference picture and the second reference picture is rescaled to generate the post first reference picture and the post second reference picture, the BM process is applied to the post first reference picture and the post second reference picture to generate a temporary prediction block, and a final prediction block is generates from the temporary prediction block.9.The method of Claim 6, wherein after said at least one of the first reference picture and the second reference picture is rescaled to generate the post first reference picture and the post second reference picture , the BM process is applied to the post first reference picture and the post second reference picture to generate a first reference block for the reference picture list 0 and a second reference block for the reference picture list 1, two rescaled reference blocks are derived by rescaling the first reference block and the second reference block respectively and a final prediction block is generated from the two rescaled reference blocks.10.The method of Claim 1, wherein after the BM process is applied to derive or refine an MV (Motion Vector) for the current block, a final motion compensation predictor is generated by using a high-resolution reconstructed reference picture, a reconstructed reference picture in pre-LF (Loop Filter) resolution or after-LF resolution.11.An apparatus for video coding, the apparatus comprising one or more electronics or processors arranged to:receive input data associated with a current block in a current picture, wherein the input data comprise pixel data to be encoded at an encoder side or receive coded data associated with the current block to be decoded at a decoder side;determine a first reference picture in reference picture list 0 and a second reference picture in reference picture list 1; andwhen any one of the first reference picture and the second reference picture is a scaled reference picture:apply BM (Bilateral Matching) process to derive prediction information for the current block by using the first reference picture in reference picture list 0 and the second reference picture in reference picture list 1, wherein the scaled reference picture has at least one different resolution parameter from the current picture; andencode or decode the current block using coding information comprising the prediction information.12.A method of video coding, the method comprising:receiving input data associated with a current block in a current picture, wherein the input data comprise pixel data to be encoded at an encoder side or coded data associated with the current block to be decoded at a decoder side;determining a reference picture; andwhen the reference picture is a scaled reference picture:applying TM (Template Matching) process to derive prediction information for the current block by using the reference picture; andencoding or decoding the current block using coding information comprising the prediction information.13.The method of Claim 12, wherein the reference picture is rescaled to a rescaled reference picture having a same resolution as the current picture and TM cost is determined based on the current picture and the rescaled reference picture.14.The method of Claim 12, wherein after the TM process is applied to derive or refine an MV (Motion Vector) for the current block, a final motion compensation predictor is generated by using a high-resolution reconstructed reference picture, a reconstructed reference picture in pre-LF (Loop Filter) resolution or after-LF resolution.15.An apparatus for video coding, the apparatus comprising one or more electronics or processors arranged to:receive input data associated with a current block in a current picture, wherein the input data comprise pixel data to be encoded at an encoder side or coded data associated with the current block to be decoded at a decoder side;determine a reference picture; andwhen the reference picture is a scaled reference picture:apply TM (Template Matching) process to derive prediction information for the current block by using the reference picture; andencode or decode the current block using coding information comprising the prediction information.16.A method of video coding, the method comprising:receiving input data associated with a current block in a current picture, wherein the input data comprise pixel data to be encoded at an encoder side or coded data associated with the current block to be decoded at a decoder side;determining a reference picture; andwhen the reference picture a scaled reference picture:rescaling the scaled reference picture to a rescaled reference picture having a same resolution as the current picture; andgenerating a reconstruction block by using the rescaled reference picture.17.The method of Claim 16, wherein information related to CU (Coding Unit) , PU (Prediction Unit) or TU (Transform Unit) is determined from the scaled reference picture prior to said rescaling the scaled reference picture, and wherein the information related to the CU, the PU or the TU excludes pixel values.18.The method of Claim 16, wherein one or more constraints on equivalent resolution between the scaled reference picture and the current picture are removed.19.The method of Claim 16, wherein an on / off control flag to indicate whether said rescaling the scaled reference picture and said generating the reconstruction block by using the rescaled reference picture are applied or not is signalled or parsed at a slice level, picture level, sequence level or a combination thereof.20.An apparatus for video coding, the apparatus comprising one or more electronics or processors arranged to:receive input data associated with a current block in a current picture, wherein the input data comprise pixel data to be encoded at an encoder side or coded data associated with the current block to be decoded at a decoder side;determine a reference picture; andwhen the reference picture is a scaled reference picture:rescale the scaled reference picture to a rescaled reference picture having a same resolution as the current picture; andgenerate a reconstruction block by using the rescaled reference picture.
Citation Information
Patent Citations
Handling of Decoder-Side Motion Vector Refinement (DMVR) Coding Tool for Reference Picture Resampling in Video Coding
US20220046271A1
Interaction between reference picture resampling and template-based inter prediction techniques in video coding
US20230199211A1
Reference picture resampling and inter-coding tools for video coding
WO2020236561A1
Video encoding and decoding using reference picture resampling
WO2023072554A1
Video encoding and decoding using reference picture resampling
WO2023194334A1