Dynamic reference conversion with different coding attributes for video coding

By employing reference regions with diverse coding attributes, the video coding system addresses inefficiencies in existing systems, achieving improved compression and decoding performance for varied video sources and display formats.

WO2025146136A1PCT designated stage expired Publication Date: 2025-07-10MEDIATEK INC
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
PCT/CN2025/070440
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2024-01-05
Filing Date
2025-01-03
Publication Date
2025-07-10

AI Technical Summary

Technical Problem

Existing video coding systems face inefficiencies due to the use of reference pictures with uniform coding attributes, which can lead to suboptimal compression and decoding performance, particularly in handling diverse video sources and display formats.

Method used

The implementation of methods and apparatus that utilize reference regions with different coding attributes, such as scaling ratios, cross-component transforms, and pixel value mappings, to adaptively adjust the coding process for improved efficiency and flexibility in video encoding and decoding.

Benefits of technology

Enhances coding efficiency by optimizing compression and decoding processes for various video sources and display formats, leading to better rate-distortion performance and adaptability to network bandwidth variations.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN2025070440_10072025_PF_FP_ABST
    Figure CN2025070440_10072025_PF_FP_ABST
Patent Text Reader

Abstract

Methods and apparatus for video coding using reference regions having a different attribute from the current block. According to one method, a first attribute associated with the current block is determined. One or more neighbouring regions for the current block are determined, wherein one second attribute is associated with each of said one or more neighbouring regions, and one or more target second attributes associated with one or more target neighbouring regions are different from the first attribute. Reference samples in said one or more target neighbouring regions are converted to converted reference samples to match the first attribute. Prediction data for the current block is derived based on the converted reference samples. The current block is encoded or decoded using the prediction data. In another method, reference samples in target neighbouring regions are converted to converted reference samples to match the first attribute.
Need to check novelty before this filing date? Find Prior Art

Description

DYNAMIC REFERENCE CONVERSION WITH DIFFERENT CODING ATTRIBUTES FOR VIDEO CODINGCROSS REFERENCE TO RELATED APPLICATIONS

[0001] The present invention is a non-Provisional Application of and claims priority to U.S. Provisional Patent Application No. 63 / 617,803, filed on January 5, 2024. The U.S. Provisional Patent Application is hereby incorporated by reference in its entirety.FIELD OF THE INVENTION

[0002] The present invention relates to video coding system. In particular, the present invention discloses using reference pictures having different coding attribute from a current block in a video coding system.BACKGROUND

[0003] Versatile video coding (VVC) is the latest international video coding standard developed by the Joint Video Experts Team (JVET) of the ITU-T Video Coding Experts Group (VCEG) and the ISO / IEC Moving Picture Experts Group (MPEG) . The standard has been published as an ISO standard: ISO / IEC 23090-3: 2021, Information technology -Coded representation of immersive media -Part 3: Versatile video coding, published Feb. 2021. VVC is developed based on its predecessor HEVC (High Efficiency Video Coding) by adding more coding tools to improve coding efficiency and also to handle various types of video sources including 3-dimensional (3D) video signals. An overview of the Versatile Video Coding (VVC) standard has been published by Bross, et. al., “Overview of the Versatile Video Coding (VVC) Standard and Its Applications” in IEEE Transactions On Circuits and Systems for Video Technology, VOL. 31, No. 10, October 2021, pp. 3736-3764.

[0004] VVC uses adaptive intra / inter prediction with quantization and transform coding for the prediction residuals similar to prior video coding standards. However, VVC also incorporates various newer coding tools. Fig. 1 illustrates an exemplary block diagram for a VVC encoder. VVC also uses block based coding, where each input picture is partitioned into CTUs (Coding Tree Units) with a typical size of 128x128 or 64x64. The input video often uses 4: 2: 0 YCbCr format, where Cb and Cr are in 1 / 4 resolution (1 / 2 in both horizontal and vertical directions) of the luma component. The input video sequence in 4: 2: 0 YCbCr format can be converted from a 4: 4: 4 RGB video through a predefined conversion process. Human eyes are less sensitive to chroma information. Therefore, a spatial down-sampling procedure is applied to the chroma channels and only 1 / 4 amount of the samples for luma are kept for chroma. After decoding, if a display device requires 4: 4: 4 RGB format for displaying the frames, the decoded 4: 2: 0 YCbCr video can be converted to the format required by display device for viewing.

[0005] As shown in Fig. 1, the Input Video Signal 102 is subtracted by prediction signal 111 using subtractor 112. The output from subtractor 112 corresponds to the residual signal. For VVC, Luma Mapping 110 is applied to input video signal 102 prior to prediction. Beside Intra-Picture Prediction 142 and Inter-Picture Prediction 146, VVC also introduces a new prediction type, named Combined Inter / Intra Prediction (CIIP, 144) . The Inter-Picture Prediction 146 is based on reconstructed pictures stored at Decoded Picture Buffer 170 and motion information from Motion Estimation 148. Since the input video signal undergoes Luma Mapping 110, a corresponding process (i.e., Luma Mapping 114) is applied to the output from Inter-Picture Prediction 146 before it can be selected and used as prediction signal 111. A selection mechanism 116 is used to adaptively select among Intra-Picture Prediction 142, Inter-Picture Prediction 146, and Combined Inter / Intra Prediction (CIIP, 144) .

[0006] Similar to HEVC, the residual signal from subtractor 112 undergoes Transform, Scaling and Quantization process 120 to generate quantized transform coefficients (labelled as “B” ) . The encoder side also needs to generate reconstructed pictures for coding process of subsequent pictures. In order to generate reconstructed pictures, Scaling and Inverse Transform process 130 is applied to the quantized transform coefficients (labelled as “B” ) to generate reconstructed residuals. VVC also introduces Chroma Scaling 132 applied to the reconstructed residuals. After chroma scaling, the reconstructed residuals are added to prediction signal 111 to form reconstructed signal. For intra prediction, the reconstructed signal can be used directly without in-loop filtering. For inter prediction, the reconstructed residuals undergo Inverse Luma Mapping, Deblocking, SAO (Sample Adaptive Offset) and ALF (Adaptive Loop Filter) 160. Since luma mapping has been applied to the input video signal, inverse luma mapping has to be applied in the reconstruction path. VVC selects filter parameters adaptively. In order to determine filter parameters, Filter Control Analysis 150 determines filter parameters based on Input Video Signal 102 and the reconstructed residuals.

[0007] In the encoder, there are various process modules that require to decide mode, parameters or other control information. For example, the prediction mode selection 116 will decide whether to use intra, inter or CIIP mode based on mode selection information from General Coder Control 190. Furthermore, General Coder Control 190 also provide control information to Intra-Picture Estimation 140, Inter-Picture Estimation 146, and Scaling and Inverse Transform 130. The quantized transform coefficients (labelled as “B” ) from Transform, Scaling and Quantization process 120 will be coded by Header Formatting and CABAC (Context-Adaptive Binary Arithmetic Coding) 180 to include in the coded bitstream. In addition, various related coder information, such as general coder control information (labelled as “A” ) , intra-prediction mode information (labelled as “C” ) from Intra-Picture Estimation 140, inter-prediction mode information (labelled as “E” ) from Inter-Picture Estimation 146, and filter control information (labelled as “D” ) from Filter Control Analysis 150 will be included in the coded bitstream using Header Formatting and CABAC 180.

[0008] Since the encoder also needs to generate reconstructed signal and uses it for prediction, the encoder comprises some key processing units as used by the decoder. The processing modules that are also used by the decoder are shown in grey colour in Fig. 1. The coded bitstream is received at the decoder side. A module to perform the reverse function of Header Formatter and CABAC 180 is used to recover the quantized transform coefficients (labelled as “B” ) along with various side information or control information (i.e., “A” , “C” , “D” , and “E” indicated in Fig. 1) . The decoder operates on the quantized transform coefficients to generate reconstructed video. At the decoder, the intra-prediction mode information and inter-prediction motion information can be generated with needed side information (i.e., “C” and “F” ) . However, for Decoder-side Intra Mode Derivation (DIMD) , the decoder will derive the mode information without the need for signalling from the encoder.

[0009] Recently, Motion Compensated Temporal Filtering (MCTF) has been disclosed by P. Wennersten, et al. ( “AHG10: Encoder-only GOP-based temporal filter” , Joint Video Experts Team (JVET) of ITU-T SG 16 WP 3 and ISO / IEC JTC 1 / SC 29 / WG 11, 15th Meeting: Gothenburg, SE, 3–12 July 2019, Document: JVET-O0549) . The filtering is done at the encoder side as a pre-processing step, where a motion estimation and motion compensation method is applied on the neighbouring pictures on 8x8 luma blocks. With reference to the encoder system block diagram in Fig. 1, the MCTF process is located prior to the input video signal 102.

[0010] Pictures are divided into a sequence of coding tree units (CTUs) . The CTU concept is same to that of the HEVC. For a picture that has three sample arrays, a CTU consists of an N×N block of luma samples together with two corresponding blocks of chroma samples. The maximum allowed size of the luma block in a CTU is specified to be 128×128 (although the maximum size of the luma transform blocks is 64×64) .

[0011] In HEVC, a CTU is split into CUs by using a quaternary-tree (QT) structure denoted as coding tree to adapt to various local characteristics. The decision whether to code a picture area using inter-picture (temporal) or intra-picture (spatial) prediction is made at the leaf CU level. Each leaf CU can be further split into one, two or four Pus according to the PU splitting type. Inside one PU, the same prediction process is applied and the relevant information is transmitted to the decoder on a PU basis. After obtaining the residual block by applying the prediction process based on the PU splitting type, a leaf CU can be partitioned into transform units (TUs) according to another quaternary-tree structure similar to the coding tree for the CU. One of key feature of the HEVC structure is that it has the multiple partition conceptions including CU, PU, and TU.

[0012] In VVC, a quadtree with nested multi-type tree using binary and ternary splits segmentation structure replaces the concepts of multiple partition unit types, i.e. it removes the separation of the CU, PU and TU concepts except as needed for CUs that have a size too large for the maximum transform length, and supports more flexibility for CU partition shapes. In the coding tree structure, a CU can have either a square or rectangular shape. A coding tree unit (CTU) is first partitioned by a quaternary tree (a. k. a. quadtree) structure. Then the quaternary tree leaf nodes can be further partitioned by multi-type tree structure. As shown in Fig. 2, there are four splitting types in multi-type tree structure, vertical binary splitting (SPLIT_BT_VER 210) , horizontal binary splitting (SPLIT_BT_HOR 220) , vertical ternary splitting (SPLIT_TT_VER 230) , and horizontal ternary splitting (SPLIT_TT_HOR 240) . The multi-type tree leaf nodes are called coding units (CUs) , and unless the CU is too large for the maximum transform length, this segmentation is used for prediction and transform processing without any further partitioning. This means that, in most cases, the CU, PU and TU have the same block size in the quadtree with nested multi-type tree coding block structure. The exception occurs when maximum supported transform length is smaller than the width or height of the colour component of the CU.

[0013] Intra Prediction

[0014] Prediction of the CU is performed based on the pixels contained in PU (prediction unit) . Prediction can be formed by neighbouring pixels or pixels in the reference frames. These pixels are available both for encoder and decoder, therefore the coding flow is valid for the codec where same prediction can be generated between the encoder and decoder. The prediction involves neighbouring pixels only is intra prediction. There are several kinds of intra prediction used in VVC, including DC, planar, angular, etc. The common step is to get the pixels from the neighbouring pixels. These neighbouring pixels may contribute from multiple CUs. For example, in Fig. 3, the neighbouring L-shape for the angular prediction is shown. Pixels from multiple neighbouring CUs are utilized in the prediction process.

[0015] Inter Prediction

[0016] Forming the prediction by referring the pixels in reference frames is the basic procedure conducted in the inter prediction. Fig. 4 is an example of bi-directional inter prediction. There are two motion vectors, MVL0 and MVL1 used to fetch the pixels.

[0017] The fetched area of the reference frame may also have multiple CUs involved. As shown in Fig. 5, the fetched area is denoted by the dark dashed area and there are nine CUs involved in this area.

[0018] Combined Inter and Intra Prediction (CIIP)

[0019] This is another type of prediction that combines inter prediction and intra prediction together. By blending the pixels from the two predictions, the final prediction is got. As mentioned in previous sections, the predictions before blending can be formed from multiple CUs.

[0020] IntraTMP (Intra Template Matching Prediction) and IBC (Intra Block Copy)

[0021] Figs. 6A-B show the concept of intraTMP, the L-shape template is used to search the area R1 to R6 to get a displacement vector (i.e., BV (block vector) ) which is similar to MV. IntraTMP is a intra coding mode and it can only search the available reconstructed area reconstructed so far. This area is both available for the encoder and the decoder while processing the current CU. As for IBC, its coding procedure is similar to intraTMP except that IBC uses original pixel of current CU to search its search region to get BV. BV of IBC should be explicitly signalled in the bitstream so that the decoder can use this information to generate the prediction for the current CU.

[0022] In the present invention, methods and apparatus to use reference regions having a different coding attribute from the current block are disclosed. BRIEF SUMMARY OF THE INVENTION

[0023] Methods and apparatus for video coding systems that use reference regions having a different coding attribute from the current block are disclosed. According to one method, input data associated with a current block in a current picture is received, wherein the input data comprise pixel data to be encoded at an encoder side or encoded data associated with the current block to be decoded at a decoder side. A first attribute associated with the current block is determined. One or more neighbouring regions for the current block are determined, wherein one second attribute is associated with each of said one or more neighbouring regions, and one or more target second attributes associated with one or more target neighbouring regions are different from the first attribute. Reference samples in said one or more target neighbouring regions are converted to converted reference samples to match the first attribute. Prediction data for the current block is derived based on the converted reference samples. The current block is encoded or decoded using the prediction data.

[0024] In one embodiment, the first attribute and the second attribute correspond to scaling ratio, cross-component transform, colour phase rotation, pixel value mapping, or a combination thereof.

[0025] In one embodiment, for intra coding, the reference samples in said one or more target neighbouring regions correspond to target samples of neighbouring CUs or target samples pointed by a block vector (BV) . In one embodiment, for inter coding, the reference samples in said one or more target neighbouring regions correspond to temporal reference areas pointed by a motion vector (MV) .

[0026] In one embodiment, the first attribute is determined according to a frame-level, region-level, CTU (Coding Tree Unit) -level, or CU (Coding Unit) -level syntax element.

[0027] In one embodiment, the reference samples in said one or more target neighbouring regions are first converted to intermediate reference samples having a third attribute, and the intermediate reference samples are then converted to the converted reference samples having the first attribute. In one embodiment, the first attribute, said one or more target second attributes, and the third attribute correspond to scaling ratio, and the third attribute is derived as a most popular scaling ratio, one or more unified scaling ratios, or a median scaling ratio among said one or more target second attributes. In one embodiment, the third attribute is selected from multiple unified scaling ratios. In one embodiment, multiple buffers are used to store reconstructed reference samples associated with the multiple unified scaling ratios and to provide the converted reference samples without on-the-fly computations. In one embodiment, the multiple unified scaling ratios are signalled or parsed in SPS (Sequence Parameter Set) , PPS (Picture Parameter Set) , PH (Picture Header) , SH (Slice Header) , or a combination thereof.

[0028] According to another method, a first attribute associated with the current block is determined. Target reference samples are selected from one or more buffers according to the first attribute. Prediction data for the current block is derived based on the target reference samples. The current block is encoded or decoded using the prediction data.

[0029] According to yet another method, a first attribute associated with one or more neighbouring regions of the current block is determined. First prediction data is derived based on reference samples in said one or more neighbouring regions. The first prediction data is converted to second prediction data, wherein the second prediction data has a same attribute as the current block. The current block is encoded or decoded using the second prediction data.BRIEF DESCRIPTION OF THE DRAWINGS

[0030] Fig. 1 illustrates an exemplary system block diagram for a VVC encoder.

[0031] Fig. 2 illustrates examples of a multi-type tree structure corresponding to vertical binary splitting (SPLIT_BT_VER) , horizontal binary splitting (SPLIT_BT_HOR) , vertical ternary splitting (SPLIT_TT_VER) , and horizontal ternary splitting (SPLIT_TT_HOR) .

[0032] Fig. 3 illustrates an example of pixels in an L-shape neighbouring region involving multiple neighbouring CUs.

[0033] Fig. 4 illustrates an example of bi-directional inter prediction.

[0034] Fig. 5 illustrates an example of pixel fetching area by motion vector and the involved CUs in the reference frame.

[0035] Fig. 6A illustrates the concept of intraTMP where the L-shape template is used to search the areas R1 to R6 to get a displacement vector (i.e., block vector) .

[0036] Fig. 6B illustrates an example of the derivation of intraTMP for a current block based on the displacement vector (i.e., block vector) .

[0037] Fig. 7 illustrates an example of candidate locations to derive the scaling ratio (SR) for prediction generation.

[0038] Fig. 8 illustrates a flowchart of an exemplary video coding system that uses one or more reference regions having one or more different attributes from the current block according to an embodiment of the present invention.

[0039] Fig. 9 illustrates a flowchart of another exemplary video coding system that uses one or more reference regions having one or more different attributes from the current block according to an embodiment of the present invention.

[0040] Fig. 10 illustrates a flowchart of yet another exemplary video coding system that uses one or more reference regions having one or more different attributes from the current block according to an embodiment of the present invention.DETAILED DESCRIPTION OF THE INVENTION

[0041] It will be readily understood that the components of the present invention, as generally described and illustrated in the figures herein, may be arranged and designed in a wide variety of different configurations. Thus, the following more detailed description of the embodiments of the systems and methods of the present invention, as represented in the figures, is not intended to limit the scope of the invention, as claimed, but is merely representative of selected embodiments of the invention. References throughout this specification to “one embodiment, ” “an embodiment, ” or similar language mean that a particular feature, structure, or characteristic described in connection with the embodiment may be included in at least one embodiment of the present invention. Thus, appearances of the phrases “in one embodiment” or “in an embodiment” in various places throughout this specification are not necessarily all referring to the same embodiment.

[0042] Furthermore, the described features, structures, or characteristics may be combined in any suitable manner in one or more embodiments. One skilled in the relevant art will recognize, however, that the invention can be practiced without one or more of the specific details, or with other methods, components, etc. In other instances, well-known structures, or operations are not shown or described in detail to avoid obscuring aspects of the invention. The illustrated embodiments of the invention will be best understood by reference to the drawings, wherein like parts are designated by like numerals throughout. The following description is intended only by way of example, and simply illustrates certain selected embodiments of apparatus and methods that are consistent with the invention as claimed herein.

[0043] In this invention, attribute adjustment is done when acquiring reference pixels for prediction generation. The attributes include scaling ratio, cross-component transform, colour phase rotation, pixel value mapping (e.g., luma mapping, gamma curve mapping or HDR tone mapping) .

[0044] Intra Coding with Scaled Neighbourhood

[0045] Reference Picture Resampling (RPR) is a coding tool that makes the resolution change during coding the whole video sequence. For displaying the video, frames with resolution change are restored to display resolution. For certain sequences with carefully selected scaling ratio, the RPR coded bitstream can show better rate distortion (RD) performance than the non-RPR coded version. Another use case for RPR to be turned on is the adaptation of network bandwidth. When the network bandwidth is insufficient, the encoding process of the video by turning on RPR and setting a scaling ratio to code the current picture in lower resolution is an effective and practical method for overcoming the variation of the bandwidth for video transmission. The same scaling concept can be applied to regions (where multiple CTUs are involved) , CTU-level or CU-level coding. For example, a CTU or CU can be coded with scaling ratio other than 1.0. A block with less texture can be better coded with scaled down pixel samples and provides better rate-distortion (RD) performance. As shown in Fig. 3, there are eight CUs denoted by CU_T0, CU_T1, CU_T2, CU_L0, CU_L1, CU_TL, CU_BL and CU_TR. These CUs can be coded with multiple scaling ratio (SR) options, such as, SR = 1.0, 1.25, 1.5 and 2.0. SR larger than 1.0 means the CU is coded with down-scaled processing. Possible processes are: 1. A At the encoder side, down-scale the original pixels and generate a down-scaled prediction.  Get residue from the down-scaled original and prediction pixels. Applied transform coding on this residue. The decoder side receives the residue, generates the down-scaled prediction and adds received residue to get the down-scaled reconstruction. Scale up the reconstruction to SR = 1.0 to get final reconstruction. 1. B At the encoder side, get residue by subtracting prediction (SR = 1.0) pixels from the  original pixels. Generate down-scaled residue. Apply transform coding on this residue. The decoder side receives the down-scaled residue and generates corresponding up-scaled residue. The prediction of SR 1.0 is generated and added with the up-scaled residue to get reconstruction for SR = 1.0.

[0046] Process 1. A (type A) applies scaling on both original pixels and prediction pixels while process 1. B (type B) applies scaling on the residue. Both processes require to generate a prediction of the current CU with certain SR. Process A requires prediction with target SR of the current CU while process B requires prediction with SR = 1.0. However, neighbouring CUs may not be coded with this required target SR. We propose several methods to deal with the prediction generation process where the SR used in prediction generation is based on the SR of the coding CU itself or based on the SRs from the neighbouring CUs.

[0047] For one embodiment (Mi. 1) , the reconstruction pixels of neighbouring CUs are scaled to the target SR for coding the current CU. The prediction is generated with these scaled pixels. For another embodiment (Mi. 2.1) , the most popular SR among the neighbouring CUs is selected as the target SR. If there are two most popular SRs with equal occurrences of usage to encode neighbouring CUs, the SR closer to 1.0 is selected as the target SR. Reconstruction pixels of neighbouring CUs are scaled to the target SR and the prediction of the current CU is generated. If the target SR of neighbouring prediction generation (denoted as SRP) is different from the SR for encoding the CU (denoted as SRC) , the prediction should be scaled to the required resolution for coding. Table 1 - Exemplary SR of CUs

[0048] For example, from Table 1, the most popular SRs in the neighbourhood involved for prediction generation are 1.5 and 1.25 because they are utilized in coding three times. SR 1.25 is closer to 1.0, therefore it is selected as target SR for prediction generation. SR 1.25 is different from the SR 1.5 for coding the current CU. Therefore, the generated prediction is scaled to 1.5 for type A residue coding process. If type B residue coding process is applied, the prediction is scaled to SR 1.0 accordingly. In another embodiment, if there are two most popular SRs with equal occurrence of usage to encode neighbouring CUs, the SR closer to the SR of the current CU (which is signalled by explicit syntax element) is selected as the target SR. However, for some situation, the two most popular SRs are with the same difference from the SR of the current CU. For example, the current CU is going to be coded with SR = 1.5 and the two most popular SRs are 2.0 and 1.0. The difference is 0.5 for the two popular SRs. A predefined policy can be applied to determine the SR for prediction generation.

[0049] The policy can be to choose the SR closer to 1.0 if the difference is the same for multiple neighbouring SRs. Another policy can consider the areas of the neighbouring CUs or distances between the neighbouring CUs and the current CU. Popular SR with larger accumulated CU area wins or popular SR with shorter accumulated CU distance wins. Therefore, SR for prediction generation is determined by multiple policies. In another implementation, the adopted SR resolving policy can be signalled by SPS (Sequence Parameter Set)  / PPS (Picture Parameter Set)  / PH (Picture Header)  / SH (Slice Header) , etc. In another implementation, for counting the occurrences of usage of SR around the neighbourhood, the SR used for coding the current CU (SRC) is also considered. SRC is known to the encoder and explicitly signalled in the bitstream for the decoder. Therefore, for counting the SR usage around the neighbourhood, SRC can participate in this process.

[0050] In another embodiment (Mi. 2.2) , SR from one designated location around the neighbourhood is used for prediction generation. As depicted in Fig. 7, one SR from location from (1) or (2) or (3) can be used as the SR for prediction generation.

[0051] In another embodiment (Mi. 2.3) , median SR from the SRs of the neighbouring CUs is used for prediction generation. Take Fig. 7 as an example, median SR from (1) , (2) and (3) is chosen as the SR for prediction generation. If the median SR is different from the SR selected for coding the current CU, scaling on the generated prediction is required for calculating the residue. Similar to counting most popular SR around neighbourhood, SRC can also be considered for median SR generation.

[0052] SR for prediction determined by the neighbourhood can be viewed as an on-the-fly method for prediction generation. However, the on-the-fly method requires extra computations. For example, if the SR for coding the current CU changes, the SR for prediction generation can be altered and scaling of the neighbourhood is required to be done again with the changed SR. We propose a unified SR for scaling the neighbouring reconstruction and the prediction generation is done accordingly.

[0053] For one embodiment (Mi. 3.1) , the unified SR for scaling the reconstruction samples of the neighbourhood for prediction generation for future CUs can be specified by SPS / PPS / PH / SH, etc. In another embodiment the unified SR can be a default value of 1.0 (or 1.5 or 2.0) . In another implementation (Mi. 3.2) , the default SR can be updated with syntax elements in SPS / PPS / PH / SH, etc. Instead of single unified SR, multiple unified SRs can be utilized in coding. For example, SRs 1.0 / 1.5 / 2.0 are utilized at the same time. Three reconstruction buffers are maintained for prediction generation and the best prediction from a certain SR is chosen for prediction generation according to the rate-distortion optimization result in video encoder. The chosen SR for coding a CU is signalled with region / CTU / CU-level syntax element.

[0054] In one embodiment (Mi. 4.1) , the multiple unified SRs are from a predefined set of SRs. In another embodiment (Mi. 4.2) , the multiple unified SRs are signalled in SPS / PPS / PH / SH, etc. With the limited number of allowed SRs chosen from these multiple SRs, extra required frame-level buffers for storing reconstruction samples can be pre-allocated for each SR. After coding a CU, the reconstructions of the CU are generated for all the unified SRs. Then reconstructions for all the multiple unified SRs are stored to the corresponding buffers to be used for prediction generation for coding future CUs. Possible multiple unified SRs are listed in Table 2. Table 2 - Possible multiple unified SRs

[0055] Table 2 is not an exhaust list of multiple unified SRs and only serves as an example of possible choices for embodiments. It should be mentioned that the SR used in coding the CU may or may not be the SR specified by the single or multiple unified SRs. Though for the case of multiple unified SRs, coding the CU with one of the SR from the multiple unified SR is practical and can avoid scaling the prediction for coding current CU. If the SR used for coding the current CU is not the same as any unified SR, first prediction is generated based on the unified SR. Then first prediction is scaled to align with SRC as a second prediction for coding. If the SR used for coding the current CU is not as any of the unified SR in the multiple unified SRs, policy, such as choosing the SR from multiple unified SR with value closer to SRC or choosing the SR from multiple unified SR with value closer to 1.0 can be applied to determine the SR used for prediction. First prediction is generated based on the determined SR. Then first prediction is scaled to SRC for coding the current CU as a second prediction used in the coding flow.

[0056] On-the-fly method though requires extra computations, it requires less storage. On-the-fly methods does not require picture-level buffers for prediction generation. It only requires storage for scaling the neighbourhood and the storage is generally at a less amount of memory than a picture buffer.

[0057] In another embodiment (Mi. 5) , a region / CTU / CU-level flag is signalled to fallback the SR to 1.0 for prediction generation. Region means a sub-area of picture normally larger than a CTU and may or may not have its boundary aligned to the CTU boundary. This syntax element introduces another choice for rate distortion optimization (RDO) , which may be beneficial for certain sequences to generate extra coding gain. This additional flag changes the default behaviour as to select the SR for prediction generation according to the neighbourhood (e.g. most popular SR, SR from a pre-defined location, median SR from CUs, etc. ) and provides a short cut to select SR 1.0. This syntax element exists when the derived SR is not 1.0. In another embodiment, the fallback SR is signalled in SPS / PPS / PH / SH. A region / CTU / CU-level flag is signalled to determine whether to fallback the SR for prediction generation as the SR signalled in SPS / PPS / PH / SH when the SR derived from the neighbourhood is different from the signalled SR by SPS / PPS / PH / SH.

[0058] In another embodiment (Mi. 6) , a region / CTU / CU-level syntax element is signalled as an index to specify the SR used for prediction generation. The index can be an indicator of a pre-defined order of the SRs as shown in Table 3. Otherwise, the index can be an indicator for selecting the SR according to popularity of the SR around the neighbourhood as shown in Table 4. To know the popularity, usage of the SRs from the neighbourhood is gathered and sorted according to the usage number. Table 3 - Index of SRs with pre-defined order Table 4 - Index of SRs with order of popularity

[0059] Inter Coding with Scaled Reference

[0060] For inter coding, motion search is invoked to get the motion vector (MV) . If region, CTU-level or CU-level coding is applied in the reference picture, the search area needs to be converted to a unified scaling ratio (SR) for motion search. For one embodiment (M. 1) , on-the-fly conversion of the searched area (for motion estimation, ME) or prediction reference pixels (for motion compensation, MC) is conducted. Pixels to be used belong to CUs with different SRs are converted to current CU’s SR. One-the-fly method normally can be implemented with less memory. However, on-the-fly conversion of the search area may impose intensive computations. Therefore, it is more practical for every frame to generate a representative frame with a certain SR for motion search. Besides, the current CU is coded with another SR, which may or may not be identical to the SR of the representative frame. We propose several methods to complete the coding task.

[0061] For one embodiment (M. 2) , one SR for the representative frame is signalled in SPS / PPS / PH / SH, etc. After encoding or decoding a frame, the representative frame is generated according to this SR for this picture to be referred by other frame while doing inter coding. For another embodiment, if no syntax element is used to specify the SR, the default SR is 1.0 for the representative frame. If the SR of the current CU is different from the SR of the selected representative frame, for motion search, one of the following procedures can be applied: 2. A Scaling the CU’s original pixels to align the SR of the representative frame (denoted as  SRU) before motion search. Perform motion search to get CU’s MV and use it for coding. 2. B Scaling the CU’s original pixels and the fetched search range to current CU’s SR (SRC)  before motion search. Get the MV and use it for coding. 2. C Scaling the fetched search range to SR 1.0 before motion search and use unscaled original  pixels to do ME against this scaled search range. Get the MV and use it for coding.

[0062] If encoding time is relaxed, all the procedures 2. A, 2. B, and 2. C can be tried to get the minimal RDO cost. Which procedure is adopted for coding can be signalled by a CU-level syntax element.

[0063] For prediction generation as the motion compensation process for inter coding, if SRU of the representative frame is not the same as the SR of the current CU (SRC) . For method 2. A, fetch the pixels from the representative frame and scale the prediction to align SRC to complete the MC process. For method 2. B, the search range is scaled to SRC first, therefore prediction obtained by motion compensation can be used directly for coding. However, for method 2. A and 2. C, after generating the prediction by the motion compensation process with MV, the prediction may require to be scaled to SRC before calculating the residue. Besides, for calculating the residue of the methods 2. A and 2. B in the encoder side, the original pixels are also required to be scaled to SRU and SRC.

[0064] In another embodiment (M. 3) , multiple SRs are specified in SPS / PPS / PH / SH, etc. or specified from a predefine set of SRs. After encoding or decoding a frame, the representative frames, involving multiple frame-level buffers with scaled reconstructed samples, are generated according to these SRs. While performing inter coding of a CU with SRC (may or may not one from the multiple SRs) , the SR of the reference frame to be used is selected based on the difference of SRC and SRs of the representative frames. The representative frame with minimal difference (could be zero) is chosen for referring to generate prediction. If two representative frames with different SRs have the same difference measurement relating to SRC, other resolution policies can be applied to determine the final representative frame. For example, just use the representative frame with SR closer to 1.0.

[0065] In another embodiment (M. 4) , the SRC is selected from one of the multiple SRs adopted for the representative frames. The selection can be done with one or more extra region / CU / CT-level syntax elements. Therefore, for P slice, CU can always find the representative frame with identical SR as SRC for referring. For B slice, there are two reference frames, the SRC can be selected from the intersection of the SRs specified for each reference frame. In another embodiment for B slice, the SR for coding the CU is chosen from one SR from one reference frame. If another reference frame does not have that selected SR, reference samples of another reference frame can be scaled to the selected SR for ME and MC for bi-directional coding.

[0066] In another embodiment, for I frame (frame with only I slice) to be referred as the reference frame for inter coding, the SR for the representative frame is aligned to the unified SR selected for intra prediction as described in the previous section.

[0067] Coding of Cross-Component Transform (CCT)

[0068] For boosting the coding gain, cross-component transform (CCT) is applied before coding. For example, with input Cb and Cr, a cross-component chroma linear transformation can be applied to get Cp and Cq with the following formulas: Forward transformation: Cp = a*Cb + b*Cr + c                 (Eq. 1) Cq = d*Cb + e*Cr + f               (Eq. 2) Backward transformation: Cb = [e* (Cp-c) -b* (Cq-f) ]  /  (a*e-b*d)       (Eq. 3) Cr = [d* (Cp-c) -a* (Cq-f) ]  /  (b*d-a*e)              (Eq. 4)

[0069] An exemplary case for the coefficients for a to f and are used to encode 10-bit video: a = 1 / 2, b = 1 / 2, c = 0, d = 1 / 2, e = -1 / 2, f = 512 Cp = 1 / 2 *Cb + 1 / 2 *Cr + 0 = (Cb + Cr)  / 2 Cq = 1 / 2 *Cb -1 / 2 *Cr + 512 = (Cb -Cr + 1024)  / 2 Cb = [-1 / 2 *Cp -1 / 2 * (Cq-512) ]  /  (-1 / 2) = Cp + Cq -512 Cr = [1 / 2 *Cp -1 / 2 * (Cq-512) ]  /  (1 / 2) = Cp -Cq + 512

[0070] In the above equations, after the CCT, Cp and Cq are then encoded with the coding flow for chroma. The transform can be applied to frame, region, CTU or CU-level coding. For intra coding, the neighbourhood consists of multiple CUs. Some CUs may be coded with CCT, but others are not. Therefore, to generate a prediction for coding current CU, the neighbourhood should be converted to the same domain (i.e., CCT or non-CCT) . In one embodiment (Mc. 1) , the conversion of the neighbourhood is aligned with the coding mode chosen by the current CU. CU’s coding mode choice may depend on frame-level, region-level, CTU-level or CU-level setting. For example, if CCT is applied for the whole CTU, the CU contained by the CTU shall be coded by CCT. Frame-level or region-level CCT setting can be signalled in SPS / PPS / PH / SH. For frame-level intra coding, CCT coding does not require the neighbourhood to be transformed to the same domain because the whole frame is coded with identical coding mode (i.e., either CCT or non-CCT) .

[0071] In another embodiment (Mc. 2.1) , for region / CTU / CU-level CCT coding, pre-defined location, (e.g. the location as shown in Fig. 7) is chosen as the unified coding attribute (i.e., CCT or non-CCT) for prediction generation. In another embodiment (Mc. 2.2) , for region / CTU / CU-level CCT coding, if most of the CUs in the neighbourhood are coded with CCT applied, current CU is inferred to be coded with CCT by default. Otherwise, non-CCT mode for coding current CU is inferred. Just like counting for popular scaling factors around the neighbourhood, the coding mode of the current CU may or may not be considered for counting the popularity of CCT attribute. In yet another embodiment (Mc. 3) , by signalling an extra flag to decide whether to code the CU with the inferred coding mode. If the flag indicates not to take the default inferred mode, the other mode is taken for coding.

[0072] For inter coding, reference conversion is required according to the CCT domain required for coding. In another embodiment (Mc. 4) , if the reference frame is not coded with the same mode as the current CU, the reference frame should be converted to the same domain for motion estimation and prediction generation of MC, etc.

[0073] In another embodiment for inter coding (Mc. 5) , the CCT attribute from the reference frame is taken first for prediction generation. Then if the coding attribute of current CU is different from the attribute of reference frame, forward CCT or backward CCT is applied on the prediction to get another prediction with aligned CCT attribute to predict current CU.

[0074] In another embodiment (Mc. 6) , to realize picture-level CCT. We propose to maintain two reconstruction buffers. If forward CCT is applied to the current picture, we save the reconstruction samples after ALF process in one of the reconstruction buffers. Then, we apply backward CCT transformation on reconstruction samples and save the result in the other reconstruction buffer. If forward CCT is not applied to the current picture, we save the reconstruction sample after ALF process in one of the reconstruction buffers. Then, we apply forward CCT transformation on reconstruction samples and save the result in the other reconstruction buffer. These dual reconstruction buffers can be used for picture-level coding for prediction generation in the next frame according to current CU’s attribute of CCT.

[0075] Coding of Colour Phase Rotation

[0076] For region / CTU / CU-level intra coding, if CUs in the referred area are coded with different attributes from the current CU, conversion should be applied in inferred area. Just like CCT, colour phase rotation (CPHR) can be applied to chroma components during coding. Procedure of forward colour phase rotation is described as follows: 1. Convert Cb, Cr to U, V (for 10-bit video, subtract 512 from Cb, Cr to get U, V) 2. Rotate [U, V] T by rotation matrix with phase θ to get [U’, V’] T 3. Convert U’, V’ to Cb’, Cr’ for coding (for 10-bit video, add 512 to U’, V’ to get Cb’, Cr’)

[0077] Backward colour phase rotation is like forward procedure but with negated phase (-θ) for vector rotation.

[0078] Cb’ and Cr’ from forward colour phase rotation are used in the coding and after reconstruction, backward colour phase rotation is applied to colour components to get the final reconstruction in phase with the original video. CPHR can be applied to frame-level, region-level, CTU-level or CU-level coding. For intra coding, the neighbourhood consists of multiple CUs. Some CUs are coded with CPHR but others are not. Therefore, to generate a prediction for coding current CU, the neighbourhood should be converted to the same domain (CPHR or non-CPHR) . In one embodiment (Mr. 1) , the conversion of the neighbourhood is aligned with the coding mode chosen by the current CU. CU’s coding mode choice may depend on frame-level, region-level, CTU-level or CU-level setting. For example, if CPHR is applied for the whole CTU, the CU contained by the CTU shall be coded by CPHR. For frame-level intra coding, CPHR coding does not require the neighbourhood to be transformed to the same domain because the whole frame is coded with an identical domain.

[0079] In another embodiment (Mr. 2.1) , for region / CTU / CU-level CPHR coding, pre-defined location, (e.g. the location as shown in Fig. 7) is chosen as the unified coding attribute (i.e., CPHR or non-CPHR) for prediction generation. In another embodiment (Mr. 2.2) , for region / CTU / CU-level CPHR coding, if most of the CUs in the neighbourhood are coded with CPHR applied, the current CU is inferred to be coded with CPHR by default. Just like counting for popular scaling factors around the neighbourhood, the coding mode of the current CU may or may not be considered for counting the popularity of CPHR attribute. In yet another embodiment (Mr. 3) , an extra flag is signalled to indicate whether to code the CU with the inferred coding mode. If the flag represents not to take the default inferred mode, the other mode is taken for coding.

[0080] For inter coding, reference conversion is required according to the CPHR domain required for coding. In another embodiment (Mr. 4) , if the reference frame is not coded with the same mode as current CU, the reference frame should be converted to the same domain for motion estimation and prediction generation, etc.

[0081] In another embodiment for inter coding (Mr. 5) , the CPHR attribute from the reference frame is taken first for prediction generation. Then, if the coding attribute of the current CU is different from the attribute of the reference frame, forward CPHR or backward CPHR is applied on the prediction to get another prediction with aligned CPHR attribute to predict the current CU.

[0082] In another embodiment (Mr. 6) , like that for CCT, dual reconstruction buffers with non-CPHR and CPHR formats can be used for picture-level coding for prediction generation according to current CU’s attribute of CPHR.

[0083] Coding of Pixel Value Mapping

[0084] Just like CCT, pixel value mapping (PVM) can be applied to the components during coding, as described by the following formulas: Forward PVM: p’ = LUTF [p] (LUTF: forward look-up table) Backward PVM: p” = LUTB [p’] (LUTB: backward look-up table)

[0085] PVM can be applied to frame, region, CTU or CU-level coding. For intra coding, the neighbourhood consists of multiple CUs. Some CUs may be coded with PVM but others are not. Therefore, to generate a prediction for coding the current CU, the neighbourhood should be converted to the same domain (i.e., PVM or non-PVM) . In one embodiment (Mp. 1) , the conversion of the neighbourhood is aligned to the coding mode chosen by the current CU. CU’s coding mode choice may depend on frame, region, CTU or CU-level setting. For example, if PVM is applied for the whole CTU, the CU contained by the CTU shall be coded by PVM. For frame-level intra coding, PVM coding does not require the neighbourhood to be transformed to the same domain because the whole frame is coded with an identical domain.

[0086] In another embodiment (Mp. 2.1) , for region / CTU / CU-level PVM coding, pre-defined location (e.g. the location as shown in Fig. 7) is chosen as the unified coding attribute (PVM or non-PVM) for prediction generation. In another embodiment (Mp. 2.2) , for region / CTU / CU-level PVM coding, if most of the CUs in the neighbourhood are coded with PVM applied, the current CU is inferred to be coded with PVM by default. In yet another embodiment (Mp. 3) , an extra flag is signalled to indicate whether to code the CU with the inferred coding mode. If the flag represents not to take the default inferred mode, the other mode is taken for coding.

[0087] For frame-level inter coding, reference conversion is required according to the PVM domain required for coding. In another embodiment (Mp. 4) , if the reference frame is not coded with the same mode as the current CU, the reference frame should be converted to the same domain for motion estimation and prediction generation, etc.

[0088] In another embodiment for inter coding (Mp. 5) , the PVM attribute from the reference frame is taken first for prediction generation. Then, if the coding attribute of the current CU is different from the attribute of the reference frame, forward PVM or backward PVM is applied on the prediction to get another prediction with aligned PVM attribute to predict the current CU.

[0089] In another embodiment (Mp. 6) , like that for CCT, dual reconstruction buffers with non-PVM and PVM formats can be used for picture-level coding for prediction generation according to current CU’s attribute of PVM.

[0090] Intra Coding with BV

[0091] Intra coding with BV (IntraTMP and IBC) faces similar situation as the inter coding with MV. The methods proposed for dealing with inter coding with MV can be extended to intra coding with BV for unifying the coding attributes for prediction generation and the down-stream coding processes. Assume there are multiple reconstruction buffers for reference (e.g. one buffer for one scaling ratio, one buffer for one attribute value) where reconstruction samples with a certain attribute are stored. In one embodiment, for intra coding with BV, aligning the attribute of reference areas to current CU’s attribute. The reconstructed samples are converted to attribute, including scaling, CCT, CPHR, and PVM, as that used in coding current CU.

[0092] In another embodiment, for intra coding with BV, adopting the attribute of the reference for first prediction generation and then adapt the prediction to second prediction according to current CU’s attribute.

[0093] Any of the foregoing proposed methods of using one or more reference regions having one or more different attributes from the current block can be implemented in encoders and / or decoders. For example, any of the proposed methods of using one or more reference regions having one or more different attributes from the current block can be implemented in an inter / intra / prediction module of an encoder, and / or an inter / intra / prediction module of a decoder. Alternatively, any of the proposed methods can be implemented as a circuit coupled to the inter / intra / prediction module of the encoder and / or the inter / intra / prediction module of the decoder, so as to provide the information needed by the inter / intra / prediction module.

[0094] Fig. 8 illustrates a flowchart of an exemplary video coding system that uses one or more reference regions having one or more different attributes from the current block according to an embodiment of the present invention. The steps shown in the flowchart may be implemented as program codes executable on one or more processors (e.g., one or more CPUs) at the decoder side. The steps shown in the flowchart may also be implemented by hardware such as one or more electronic devices or processors arranged to perform the steps in the flowchart. According to one method, input data associated with a current block in a current picture is received in step 810, wherein the input data comprise pixel data to be encoded at an encoder side or encoded data associated with the current block to be decoded at a decoder side. A first attribute associated with the current block is determined in step 820. One or more neighbouring regions for the current block are determined in step 830, wherein one second attribute is associated with each of said one or more neighbouring regions, and one or more target second attributes associated with one or more target neighbouring regions are different from the first attribute. Reference samples in said one or more target neighbouring regions are converted to converted reference samples to match the first attribute in step 840. Prediction data for the current block is derived based on the converted reference samples in step 850. The current block is encoded or decoded using the prediction data in step 860.

[0095] Fig. 9 illustrates a flowchart of another exemplary video coding system that uses one or more reference regions having one or more different attributes from the current block according to an embodiment of the present invention. According to this method, input data associated with a current block in a current picture is received in step 910, wherein the input data comprise pixel data to be encoded at an encoder side or encoded data associated with the current block to be decoded at a decoder side. A first attribute associated with the current block is determined in step 920. Target reference samples are selected from one or more buffers according to the first attribute in step 930. Prediction data for the current block is derived based on the target reference samples in step 940. The current block is encoded or decoded using the prediction data in step 950.

[0096] Fig. 10 illustrates a flowchart of another exemplary video coding system that uses one or more reference regions having one or more different attributes from the current block according to an embodiment of the present invention. According to yet another method, input data associated with a current block in a current picture is received in step 1010, wherein the input data comprise pixel data to be encoded at an encoder side or encoded data associated with the current block to be decoded at a decoder side. A first attribute associated with one or more neighbouring regions of the current block is determined in step 1020. First prediction data is derived based on reference samples in said one or more neighbouring regions in step 1030. The first prediction data is converted to second prediction data, wherein the second prediction data has a same first attribute as the current block in step 1040. The current block is encoded or decoded using the second prediction data in step 1050.

[0097] The flowcharts shown are intended to illustrate an example of video coding according to the present invention. A person skilled in the art may modify each step, re-arranges the steps, split a step, or combine steps to practice the present invention without departing from the spirit of the present invention. In the disclosure, specific syntax and semantics have been used to illustrate examples to implement embodiments of the present invention. A skilled person may practice the present invention by substituting the syntax and semantics with equivalent syntax and semantics without departing from the spirit of the present invention.

[0098] The above description is presented to enable a person of ordinary skill in the art to practice the present invention as provided in the context of a particular application and its requirement. Various modifications to the described embodiments will be apparent to those with skill in the art, and the general principles defined herein may be applied to other embodiments. Therefore, the present invention is not intended to be limited to the particular embodiments shown and described, but is to be accorded the widest scope consistent with the principles and novel features herein disclosed. In the above detailed description, various specific details are illustrated in order to provide a thorough understanding of the present invention. Nevertheless, it will be understood by those skilled in the art that the present invention may be practiced.

[0099] Embodiment of the present invention as described above may be implemented in various hardware, software codes, or a combination of both. For example, an embodiment of the present invention can be one or more circuit circuits integrated into a video compression chip or program code integrated into video compression software to perform the processing described herein. An embodiment of the present invention may also be program code to be executed on a Digital Signal Processor (DSP) to perform the processing described herein. The invention may also involve a number of functions to be performed by a computer processor, a digital signal processor, a microprocessor, or field programmable gate array (FPGA) . These processors can be configured to perform particular tasks according to the invention, by executing machine-readable software code or firmware code that defines the particular methods embodied by the invention. The software code or firmware code may be developed in different programming languages and different formats or styles. The software code may also be compiled for different target platforms. However, different code formats, styles and languages of software codes and other means of configuring code to perform the tasks in accordance with the invention will not depart from the spirit and scope of the invention.

[0100] The invention may be embodied in other specific forms without departing from its spirit or essential characteristics. The described examples are to be considered in all respects only as illustrative and not restrictive. The scope of the invention is therefore, indicated by the appended claims rather than by the foregoing description. All changes which come within the meaning and range of equivalency of the claims are to be embraced within their scope.

Claims

1.A method of video coding, the method comprising:receiving input data associated with a current block in a current picture, wherein the input data comprise pixel data to be encoded at an encoder side or encoded data associated with the current block to be decoded at a decoder side;determining a first attribute associated with the current block;determining one or more neighbouring regions for the current block, wherein one second attribute is associated with each of said one or more neighbouring regions, and one or more target second attributes associated with one or more target neighbouring regions are different from the first attribute;converting reference samples in said one or more target neighbouring regions to converted reference samples to match the first attribute;deriving prediction data for the current block based on the converted reference samples; andencoding or decoding the current block using the prediction data.2.The method of Claim 1, wherein the first attribute and the second attribute correspond to scaling ratio, cross-component transform, colour phase rotation, pixel value mapping, or a combination thereof.3.The method of Claim 1, wherein for intra coding, the reference samples in said one or more target neighbouring regions correspond to target samples of neighbouring CUs or target samples pointed by a block vector (BV) .4.The method of Claim 1, wherein for inter coding, the reference samples in said one or more target neighbouring regions correspond to temporal reference areas pointed by a motion vector (MV) .5.The method of Claim 1, wherein the first attribute is determined according to a frame-level, region-level, CTU (Coding Tree Unit) -level, or CU (Coding Unit) -level syntax element.6.The method of Claim 1, wherein the reference samples in said one or more target neighbouring regions are first converted to intermediate reference samples having a third attribute, and the intermediate reference samples are then converted to the converted reference samples having the first attribute.7.The method of Claim 6, wherein the first attribute, said one or more target second attributes, and the third attribute correspond to scaling ratio, and the third attribute is derived as a most popular scaling ratio, one or more unified scaling ratios, or a median scaling ratio among said one or more target second attributes.8.The method of Claim 7, wherein the third attribute is selected from multiple unified scaling ratios.9.The method of Claim 8, wherein multiple buffers are used to store reconstructed reference samples associated with the multiple unified scaling ratios and to provide the converted reference samples without on-the-flight computations.10.The method of Claim 8, wherein the multiple unified scaling ratios are signalled or parsed in SPS (Sequence Parameter Set) , PPS (Picture Parameter Set) , PH (Picture Header) , SH (Slice Header) , or a combination thereof.11.An apparatus for video coding, the apparatus comprising one or more electronics or processors arranged to:receive input data associated with a current block in a current picture, wherein the input data comprise pixel data to be encoded at an encoder side or encoded data associated with the current block to be decoded at a decoder side;determine a first attribute associated with the current block;determine one or more neighbouring regions for the current block, wherein one second attribute is associated with each of said one or more neighbouring regions, and one or more target second attributes associated with one or more target neighbouring regions are different from the first attribute;convert reference samples in said one or more target neighbouring regions to converted reference samples to match the first attribute;derive prediction data for the current block based on the converted reference samples; andencode or decode the current block using the prediction data.12.A method of video coding, the method comprising:receiving input data associated with a current block in a current picture, wherein the input data comprise pixel data to be encoded at an encoder side or encoded data associated with the current block to be decoded at a decoder side;determining a first attribute associated with the current block;selecting target reference samples from one or more buffers according to the first attribute;deriving prediction data for the current block based on the target reference samples; andencoding or decoding the current block using the prediction data.13.The method of Claim 12, wherein the first attribute corresponds to scaling ratio, cross-component transform, colour phase rotation, pixel value mapping, or a combination thereof.14.The method of Claim 12, wherein multiple buffers are allocated to store reference samples corresponding to different attributes.15.The method of Claim 12, wherein the first attribute is determined according to a frame-level, region-level, CTU (Coding Tree Unit) -level, or CU (Coding Unit) -level syntax element.16.An apparatus for video coding, the apparatus comprising one or more electronics or processors arranged to:receive input data associated with a current block in a current picture, wherein the input data comprise pixel data to be encoded at an encoder side or encoded data associated with the current block to be decoded at a decoder side;determine a first attribute associated with the current block;select target reference samples from one or more buffers according to the first attribute;derive prediction data for the current block based on the target reference samples; andencode or decode the current block using the prediction data.17.A method of video coding, the method comprising:receiving input data associated with a current block in a current picture, wherein the input data comprise pixel data to be encoded at an encoder side or encoded data associated with the current block to be decoded at a decoder side;determining a first attribute associated with one or more neighbouring regions of the current block;deriving first prediction data based on reference samples in said one or more neighbouring regions;converting the first prediction data to second prediction data, wherein the second prediction data has a same attribute as the current block; andencoding or decoding the current block using the second prediction data.18.The method of Claim 17, wherein the first attribute corresponds to scaling ratio, cross-component transform, colour phase rotation, pixel value mapping, or a combination thereof.19.The method of Claim 17, wherein for intra coding, the reference samples in said one or more target neighbouring regions correspond to target samples of neighbouring CUs or target samples pointed by a block vector (BV) .20.The method of Claim 17, wherein for inter coding, the reference samples in said one or more target neighbouring regions correspond to temporal reference areas pointed by a motion vector (MV) .21.The method of Claim 17, wherein the first attribute is determined according to a frame-level, region-level, CTU (Coding Tree Unit) -level, or CU (Coding Unit) -level syntax element.22.An apparatus for video coding, the apparatus comprising one or more electronics or processors arranged to:receive input data associated with a current block in a current picture, wherein the input data comprise pixel data to be encoded at an encoder side or encoded data associated with the current block to be decoded at a decoder side;determine a first attribute associated with one or more neighbouring regions of the current block;derive first prediction data based on reference samples in said one or more neighbouring regions;convert the first prediction data to second prediction data, wherein the second prediction data has a same attribute as the current block; andencode or decode the current block using the second prediction data.

Citation Information

Patent Citations

  • Reference subgraph scaling ratio for subgraphs in video coding

    CN114846794A

  • System and method for anomaly detection of submarines

    KR1020250061043A

  • Coding apparatus, coding method, decoding apparatus, decoding method, transmitting apparatus, and receiving apparatus

    US20200068222A1

  • Image decoding method and apparatus therefor

    WO2022216051A1

  • Method, apparatus, and recording medium for image encoding / decoding

    WO2023090894A1