Methods and apparatus of temporal partition prediction with reference picture resampling
By disabling temporal partition prediction when resolution differences occur and recalculating partition information based on refined positions and specific scaling ratios, the method ensures efficient and adaptive video coding despite reference picture resampling.
Patent Information
- Application Number
- PCT/CN2025/107017
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-07-05
- Filing Date
- 2025-07-04
- Publication Date
- 2026-01-08
AI Technical Summary
The resolution differences between current and reference pictures due to reference picture resampling in video coding systems break the assumption of consistent spatial relationships, rendering existing temporal partition prediction methods less effective.
Implement methods to maintain temporal partition prediction by disabling it when resolution differences are significant, recalculating partition prediction information based on refined positions, and using multiple sets of information tailored to specific scaling ratios, ensuring efficient partitioning decisions.
This approach maintains computational efficiency and coding quality by adapting partition predictions to resolution changes, avoiding suboptimal coding decisions and unnecessary complexity.
Smart Images

Figure CN2025107017_08012026_PF_FP_ABST
Abstract
Description
METHODS AND APPARATUS OF TEMPORAL PARTITION PREDICTION WITH REFERENCE PICTURE RESAMPLING
[0001] CROSS REFERENCE TO RELATED APPLICATION
[0002] This application claims the benefit of U.S. Provisional Application No. 63 / 667,854, filed on July 5th, 2024. The content of the application is incorporated herein by reference.BACKGROUND OF THE INVENTION1. FIELD OF THE INVENTION
[0003] The invention relates to video coding technologies, and more particularly to methods for enabling temporal partition prediction in the presence of reference picture resampling.
[0004] 2. DESCRIPTION OF THE PRIOR ART
[0005] Video compression technologies continually evolve to meet the growing demand for efficient digital video transmission and storage. The Versatile Video Coding (VVC) standard, also known as H. 266, represents a significant advancement over previous standards such as High Efficiency Video Coding (HEVC / H. 265) . One aspect of VVC is its sophisticated block partitioning system, which divides video frames into coding units (CUs) using a combination of quad-tree (QT) and multi-type tree (MTT) structures.
[0006] In HEVC, a coding tree unit (CTU) was split into CUs using a quaternary-tree structure. VVC enhances this approach by implementing a more flexible system where a CTU is first partitioned by a quad-tree structure, and then the leaf nodes can be further partitioned using binary or ternary splitting methods through the MTT structure.
[0007] The recursive partitioning process is computationally intensive but crucial for achieving optimal compression efficiency. To reduce encoding time while maintaining compression quality, temporal partition prediction techniques have been developed. These techniques leverage information from previously coded frames to predict partitioning decisions for the current frame, significantly reducing the computational burden of exhaustive partitioning searches. VVC also introduced reference picture resampling (RPR) . RPR is particularly valuable for sequences with zoom-in or zoom-out camera movements and for adapting to varying network conditions.
[0008] However, when RPR is enabled, the resolution of the current picture may differ from that of one or more reference pictures. This resolution difference breaks the assumption of temporal partition prediction, which is that spatial relationships between current and reference frames remain consistent. As a result standard methods become less effective and requiring new approaches to handle these mismatches.SUMMARY OF THE INVENTION
[0009] An embodiment provides a method of video coding in a video system for encoding or decoding video pictures. The method comprises receiving input data associated with a current block in a current picture, wherein the input data comprise pixel data to be encoded at an encoder side or data associated with the current block to be decoded at a decoder side, determining a temporal partition prediction based on a relation between current picture and a reference picture, partitioning the current block according to the temporal partition prediction, and encoding or decoding the current block according to a block partitioning structure of the current block.
[0010] An embodiment provides an apparatus of video processing in a video coding system for encoding or decoding video pictures. The apparatus comprises one or more electronic circuits configured to receive input data associated with a current block in a current picture, wherein the input data comprise pixel data to be encoded at an encoder side or data associated with the current block to be decoded at a decoder side, determine a temporal partition prediction based on a relation between the current picture and a reference picture, partition the current block according to the temporal partition prediction, and encode or decode the current block according to a block partitioning structure of the current block.
[0011] In certain aspects, the method further comprises determining a scaling ratio between the current picture and the reference picture, and selectively disabling the temporal partition prediction when said scaling ratio deviates from unity.
[0012] In certain aspects, the method comprises calculating a refined position in the reference picture based on a scaling ratio between the current picture and the reference picture, and deriving partition prediction information from said refined position. The refined position calculation implements a series of mathematical operations translating coordinates from the current picture space to the reference picture space while accounting for scaling window offsets.
[0013] In certain aspects, the method provides for deriving a plurality sets of partition prediction information, each set corresponding to a different scaling ratio between the current picture and reference picture, and selecting an appropriate set for use in partitioning. Each set may be characterized by specific grid resolution and area size parameters optimized for particular scaling ratio combinations.
[0014] To the accomplishment of the foregoing and related ends, certain embodiments comprise the features hereinafter fully described and particularly pointed out in the claims. The following description and accompanying drawings set forth in detail certain illustrative aspects of the embodiments. These aspects are indicative, however, of but a few of the various ways in which the principles of the embodiments may be employed, and the present disclosure is intended to include all such aspects and their equivalents. These and other objectives of the present invention will no doubt become obvious to those of ordinary skill in the art after reading the following detailed description of the preferred embodiment that is illustrated in the various figures and drawings.BRIEF DESCRIPTION OF THE DRAWINGS
[0015] FIG. 1 is a diagram illustrating different types of multi-type tree splitting modes according to the embodiments.
[0016] FIG. 2 is a diagram illustrating a grid point and its neighborhood used for deriving partition information according to the embodiments.
[0017] FIG. 3 is a diagram illustrating reference picture resampling relationships among a current picture and reference pictures according to the embodiments.
[0018] FIG. 4 is a flow diagram illustrating a method video coding in a video system for encoding or decoding video pictures according to the embodiments.
[0019] FIG. 5A is a block diagram illustrating an exemplary encoder structure in an adaptive inter / intra video coding system according to the embodiments.
[0020] FIG. 5B is a block diagram illustrating an exemplary decoder structure in an adaptive inter / intra video coding system according to the embodiments.DETAILED DESCRIPTION
[0021] The present disclosure relates to methods and apparatus for enabling temporal partition prediction with reference picture resampling in video coding systems. In particular, the disclosure provides techniques for maintaining effective temporal partition prediction when current and reference pictures have different resolutions due to reference picture resampling (RPR) .
[0022] Various examples of the disclosure are described below. It will be apparent that the examples described herein may be embodied in many different forms and should not be construed as limited to the examples set forth herein. Rather, these examples are provided so that this disclosure will be thorough and complete and will convey the scope of the disclosure to those skilled in the art.
[0023] The video coding standards referred to in this disclosure include but are not limited to H. 266 / Versatile Video Coding (VVC) , H. 265 / High Efficiency Video Coding (HEVC) , H. 264 / Advanced Video Coding (AVC) , and other existing or future video coding standards.
[0024] The examples provided throughout this disclosure may be applied to both video encoder and decoder implementations. While specific reference may be made to encoder operations in particular examples, those skilled in the art will recognize that corresponding decoder operations are implied to maintain bitstream compatibility, and vice versa.
[0025] The following description sets forth numerous specific details to provide a thorough understanding of the examples. However, it will be apparent to those skilled in the art that the examples may be practiced without these specific details. In other instances, well-known methods, procedures, components, and circuits have not been described in detail so as not to obscure the examples unnecessarily.
[0026] The acronyms and abbreviations used in this disclosure are listed herein:
[0027] AVC: Advanced Video Coding (H. 264)
[0028] BT: Binary Tree (a type of partition in VVC)
[0029] CU: Coding Unit
[0030] CTU: Coding Tree Unit
[0031] ECM: Enhanced Compression Model
[0032] HEVC: High Efficiency Video Coding (H. 265)
[0033] ID: Identifier
[0034] MTT: Multi-Type Tree
[0035] PPS: Picture Parameter Set
[0036] PU: Prediction Unit
[0037] QT: Quad Tree (or Quaternary Tree)
[0038] RPR: Reference Picture Resampling
[0039] TT: Ternary Tree (a type of partition in VVC)
[0040] TU: Transform Unit
[0041] VVC: Versatile Video Coding (H. 266)
[0042] 1 Partitioning
[0043] 1.1 Partitioning of the CTUs using a tree structure
[0044] Please refer to FIG. 1. In HEVC, a CTU (coding tree unit) is split into CUs (coding units) by using a quaternary-tree structure to adapt to various local characteristics. The decision whether to code a picture area using inter-picture (temporal) or intra-picture (spatial) prediction is made at the leaf CU level. Each leaf CU can be further split into one, two or four PUs (prediction units) according to the PU splitting type. Inside one PU, the prediction process is applied and the relevant information is transmitted to the decoder on a PU basis. After obtaining the residual block by applying the prediction process based on the PU splitting type, a leaf CU can be partitioned into TUs (transform units) according to another quaternary-tree structure similar to the coding tree for the CU. One of the features of the HEVC structure is that it has the multiple partition conceptions including CU, PU, and TU. For example, a 64×64 CTU containing complex texture might be split into multiple CUs of varying sizes (32×32, 16×16, 8×8) , with some CUs using inter prediction for areas matching previous frames and others using intra prediction for content predicted well by neighborhoods of the CUs from current frame. A moving object might be encoded with a 16×8 PU splitting type to better match its motion, while transform units might use 8×8 or 4×4 sizes to efficiently represent the residual data.
[0045] In VVC, a quadtree with nested multi-type tree (MTT) using binary and ternary split segmentation structure replaces the concept of multiple partition unit types, i.e., it removes the separation of the CU, PU and TU concept except as needed for CUs that have a size too large for the maximum transform length and supports more flexibility for CU partition shapes. In the coding tree structure, a CU can have either a square or rectangular shape. A coding tree unit (CTU) is first partitioned by a quaternary tree (a. k. a. quadtree) structure. Then the quaternary tree leaf nodes can be further partitioned by a multi-type tree structure. As shown in FIG. 1, there are four splitting types in multi-type tree structure, vertical binary splitting (SPLIT_BT_VER) , horizontal binary splitting (SPLIT_BT_HOR) , vertical ternary splitting (SPLIT_TT_VER) , and horizontal ternary splitting (SPLIT_TT_HOR) . The multi-type tree leaf nodes are called coding units (CUs) . For instance, a scene with a horizontal boundary (like a horizon line) might benefit from SPLIT_BT_HOR, dividing the CTU into two rectangles with the boundary aligned to the horizon. A centered object might be efficiently coded using SPLIT_TT_VER or SPLIT_TT_HOR, creating three partitions with the object isolated in the middle section.
[0046] Since the CTU is partitioned recursively, larger depth for partitioning results in longer encoding coding time and better compression result. In this innovation, depth of partitioning is estimated by temporal information and better trade-off between encoding time and compression result can be attained. By intelligently predicting these partition patterns based on temporal information, the encoder can avoid exhaustive recursive partition searches while maintaining compression efficiency.
[0047] 1.2 Estimate partition prediction information temporally
[0048] In current ECM (Enhanced Compression Model) software, temporal partition prediction is utilized to adaptively increase or decrease the MTT depth while coding a CU. Partition prediction information is stored for each picture after encoding or decoding to be utilized for coding future frames. This partition prediction information can include:
[0049] a. minimum temporal QT depth
[0050] b. maximum temporal QT depth
[0051] c. average temporal QT depth
[0052] d. minimum temporal MTT depth
[0053] e. maximum temporal MTT depth
[0054] f. average temporal MTT depth
[0055] g. minimum temporal BT depth
[0056] h. maximum temporal BT depth
[0057] i. average temporal BT depth
[0058] j. minimum temporal TT depth
[0059] k. maximum temporal TT depth
[0060] l. average temporal TT depth
[0061] This information is calculated with a certain grid resolution. Assume the grid resolution is N by N and for each grid point the information is calculated upon an area of neighborhood of 2M by 2M as shown in FIG. 2. All the CUs with their upper-left corner located within this area are the target CUs involved in the calculation of the partition prediction information. For instance, in a 4K video with N=16, each grid point might be 16 pixels apart from its nearest another grid point, and with M=128, the neighborhood area would be 256×256 pixels. In a scene with a stationary camera showing a person talking, the grid points covering the person's face might have deeper partition depths to capture facial details, while grid points in the background would have shallower depths.
[0062] Minimum temporal QT depth and maximum temporal MTT depth are derived by searching all the CUs in the 2M-by-2M area. Average temporal QT depth and average temporal MTT depth are calculated by accumulating the area-weighted depth information (i.e., CU area × depth and the area covered by one CU is only calculated once) for the two partition types (QT and MTT calculated separately) among all the CUs in 2M-by-2M area. Then the area-weighted depth information is divided by the accumulated CU area to obtain the average partition depth. This information is stored after the whole picture is encoded or decoded and to be used in the coding process for the other frame. For example, the area covering a certain object might have an average temporal QT depth of 1 and an average temporal MTT depth of 2, indicating moderate partitioning complexity, while the background might have minimum temporal QT depth of 0 and maximum temporal MTT depth of 0, showing no further splitting is necessary.
[0063] For processing the other frame, a collocated picture is decided and its partition prediction information is referenced while coding current frame. The partition information from the collocated frame is used to constrain the block partition when splitting the current block. For current implementation in ECM, several constraints are checked to increase or decrease the MTT depth for CU split. For example, if QT depth is smaller than the QT depth in collocated picture and also the MTT depth is smaller than the MTT depth in collocated picture, the MTT depth can be increased for CU split. Another constraint like if the MTT depth of the collocated picture is smaller than current MTT depth, MTT split is not allowed even the MTT depth is still smaller then than the maximal value specified in the configuration file. These extra constraints are applied with higher precedence compared with the original block partition constraints set by the coding parameters (for example, MaxMTTHierarchyDepth, MaxBTNonISlice, MaxTTNonISlice and MaxMTTHierarchyDepthByTid, etc. ) . Therefore, maximal MTT depth can be increased or decreased adaptively for seeking better coding gain for every CU. As an illustration, when processing an object that appeared in previous frames, the encoder might restrict MTT splitting to match the depth used when that object was previously coded, saving computation time while maintaining consistent quality. Conversely, when new objects enter the frame, the encoder might allow deeper MTT splitting to properly capture these previously unseen details.
[0064] 1.3 Reference picture resampling
[0065] In H. 266 / VVC, a coding tool called reference picture resampling (RPR) is introduced. For coding a picture, a scaling ratio can be determined. Common scaling ratios are 1.25, 1.5 and 2.0. These ratios are used to scale current picture for the current picture to become a smaller picture. For example, if scaling ratio is 2.0, a picture with resolution 3840×2160 is converted to 1920×1080 while encoding this picture. For coding with inter slices, coded pictures become the reference pictures for coding current frame. The scaling ratios of the reference pictures can be different from that of the current picture. For motion compensation can work properly, the referred samples from the reference picture are scaled according to the size relationship between the current picture and the reference picture. As shown in Fig. 3, Ref 0 is smaller than the current picture, therefore for Ref 0 to be used as the reference picture for motion compensation, Ref 0 should be scaled up to be used as reference while coding current picture. Ref 1 is larger than current picture, therefore for Ref 1 to be used as the reference picture for motion compensation, Ref 1 should be scaled down to be used as reference while coding current picture. RPR is especially useful while coding sequences with zoom-in or zoom-out camera operations. The scaling ratio for RPR is in the range from 1 / 8 to 2.0.
[0066] This flexibility in handling different resolution pictures offers advantages for various video content types. For instance, during a video conference where bandwidth fluctuates, the encoder might dynamically adjust the resolution of frames to maintain smooth playback. A frame might be encoded at full resolution (scaling ratio 1.0) when bandwidth is sufficient, then scaled down (ratio 1.5 or 2.0) during network congestion periods, with seamless transitions between these resolutions. Furthermore, RPR enables adaptive resolution techniques for content-aware encoding. Scenes with limited detail or slow motion can be encoded at lower resolutions to save bandwidth, while complex scenes can be encoded at higher resolutions.
[0067] 2. Methods for temporal partition prediction while enabling reference picture resampling
[0068] 2.1 Turn off temporal partition prediction if picture sizes differ between the current picture and the reference picture
[0069] When RPR is turned on, the scaling ratios between the current picture and the reference picture can be different. This makes the picture resolution of the current picture be different from the picture resolution of the reference picture. In the meanwhile, partition prediction information calculated for temporal partition prediction from the reference picture cannot be used directly for doing temporal partition prediction. As described in section 1.2, the partition prediction information is calculated from a selected region in the collocated picture. The basic assumption is that there is a certain relationship between the partition situation between the current picture and the collocated picture within this selected area. With RPR enabled and the scaling ratios of the current picture and the collocated picture can be set with different values, hence the relationship is broken. The partition prediction information from the collocated picture is no longer suitable for prediction the partition of the current picture. For one embodiment when scaling ratios set to the current picture and the collocated picture are different (or called deviates from unity) , the temporal partition prediction for current picture is turned off (or called disabled) . In VVC, the scaling ratio is calculated based on a scaling window with four offset parameters and the signaled picture size. Take parameters in PPS (picture parameter set) as example, these parameters are:
[0070] pps_pic_width_in_luma_samples, pps_pic_height_in_luma_samples
[0071] pps_scaling_win_left_offset, pps_scaling_win_right_offset,
[0072] pps_scaling_win_top_offset, pps_scaling_win_bottom_offset
[0073] For common cases, the four offsets are set to zero. Effective frame width and height are calculated as follows (SubWidthC, SubHeightC are assigned according to chroma format) :
[0074] CurrPicScalWinWidthL = pps_pic_width_in_luma_samples -SubWidthC* (pps_scaling_win_right_offset + pps_scaling_win_left_offset)
[0075] CurrPicScalWinHeightL = pps_pic_height_in_luma_samples -SubHeightC* (pps_scaling_win_bottom_offset + pps_scaling_win_top_offset)
[0076] With reference frame’s width and height, fRefWidth and fRefHeight, the scaling ratios RefPicScaleW and RefPicScaleH are calculated with CurrPicScalWinWidthL and CurrPicScalWinHeightL being considered as follows:
[0077] RefPicScaleW = ( (fRefWidth << 14) + (CurrPicScalWinWidthL >> 1) ) / CurrPicScalWinWidthL
[0078] RefPicScaleH = ( (fRefHeight << 14) + (CurrPicScalWinHeightL >> 1) ) / CurrPicScalWinHeightL
[0079] With RPR tuned on, if RefPicScaleW or RefPicScaleH is not 16384 (which corresponding to scaling ratio 1.0 (for example, in 14-bit precision) ) , temporal partition prediction is turned off for current picture.
[0080] The syntax is explained below:
[0081] Picture dimension parameters:
[0082] pps_pic_width_in_luma_samples: The width of the picture in luma samples
[0083] pps_pic_height_in_luma_samples: The height of the picture in luma samplesScaling window offset parameters:
[0084] pps_scaling_win_left_offset: Left offset of the scaling window
[0085] pps_scaling_win_right_offset: Right offset of the scaling window
[0086] pps_scaling_win_top_offset: Top offset of the scaling window
[0087] pps_scaling_win_bottom_offset: Bottom offset of the scaling window
[0088] These offset parameters allow for adjusting the active picture area, which is particularly useful in zoom-in / zoom-out scenarios.
[0089] Effective frame dimensions calculation:
[0090] CurrPicScalWinWidthL: Effective frame width calculated by subtracting the scaled horizontal offsets from the full picture width.
[0091] CurrPicScalWinHeightL: Effective frame height calculated by subtracting the scaled vertical offsets from the full picture height.
[0092] The calculation uses SubWidthC and SubHeightC, which are chroma subsampling factors assigned based on the video's chroma format.
[0093] YUV 420: SubWidthC = 2, SubHeightC = 2
[0094] YUV 422: SubWidthC = 2, SubHeightC = 1
[0095] YUV 444: SubWidthC = 1, SubHeightC = 1
[0096] Scaling ratio calculation:
[0097] RefPicScaleW: Horizontal scaling factor between reference picture and current picture
[0098] RefPicScaleH: Vertical scaling factor between reference picture and current picture
[0099] These ratios are calculated with a formula that performs shifting the reference picture width / height left by 14 bits (multiplying by 214 or 16384) to create fixed-point precision, adding half the effective width / height of the current picture for rounding, and dividing by the effective width / height of the current picture.
[0100] With RPR tuned on, if RefPicScaleW or RefPicScaleH is not 16384 (which corresponding to scaling ratio 1.0) , temporal partition prediction is turned off for current picture. In an illustration, consider a video sequence where the current picture has a resolution of 1920×1080 and the reference picture with a resolution of 3840×2160. Assuming no scaling window offsets for simplicity (all offsets set to 0) , the calculation would proceed as follows:
[0101] RefPicScaleW = ( (3840 << 14) + (1920 >> 1) ) / 1920 = (62914560 + 960) / 1920 = 32768
[0102] RefPicScaleH = ( (2160 << 14) + (1080 >> 1) ) / 1080 = (35389440 + 540) / 1080 = 32768
[0103] Since both RefPicScaleW and RefPicScaleH equal 32768 (corresponding to a scaling ratio of 2.0) , and neither equals 16384, temporal partition prediction would be disabled when encoding current picture. This prevents the encoder from using potentially inappropriate partition information from a frame with different resolution, avoiding suboptimal coding decisions that could result in either reduced compression efficiency or unnecessary computational complexity.
[0104] 2.2 Re-positioning to get temporal partition prediction information if picture sizes differ between the current picture and the reference picture
[0105] In section 2.1, if scaling ratios differ between current picture and reference picture, temporal partition prediction is turned off. However, if position for getting the partition prediction information is refined according to the scaling ratios of current picture and reference picture, the information used for partition prediction can be found. In one embodiment, re-positioning according to scaling ratio is conducted and the refined position is used to retrieve the partition prediction information. For example, assume the representative CU position (for one embodiment, this position could be the center location of a CU) of current picture is (xc, yc) and the scaling ratios for the current picture by referring the reference picture are (RefPicScaleW, RefPicScaleH) . The updated / refined position, (xr, yr) , to fetch the partition prediction information is calculated as follows:
[0106] refx = ( (xc - (SubWidthC *pps_scaling_win_left_offset) ) << 4) *RefPicScaleW
[0107] xrpre = ( (Sign (refx) * ( (Abs (refx) + 128) >> 8) ) + fRefLeftOffset + 32) >> 6
[0108] refy = ( (yc - (SubHeightC *pps_scaling_win_top_offset) ) << 4) *RefPicScaleH
[0109] yrpre = ( (Sign (refy) * ( (Abs (refy) + 128) >> 8) ) + fRefTopOffset + 32) >> 6
[0110] where:
[0111] fRefLeftOffset = ( (SubWidthC *pps_scaling_win_left_offset) << 10)
[0112] fRefTopOffset = ( (SubHeightC *pps_scaling_win_top_offset) << 10)
[0113] xr = xrpre >> 4
[0114] yr = yrpre >> 4
[0115] The process begins with a position (xc, yc) in the current picture, which represents the representative position of a CU (often its center) . RefPicScaleW and RefPicScaleH are the horizontal and vertical scaling ratios between current and reference pictures.
[0116] The formula:
[0117] refx = ( (xc - (SubWidthC *pps_scaling_win_left_offset) ) << 4) *RefPicScaleW
[0118] refy = ( (yc - (SubHeightC *pps_scaling_win_top_offset) ) << 4) *RefPicScaleH
[0119] converts the current position to a scaled position, accounting for scaling window offsets and multiplying by the scaling ratio. The << 4 shifts left by 4 bits for precision.
[0120] The formula:
[0121] xrpre = ( (Sign (refx) * ( (Abs (refx) + 128) >> 8) ) + fRefLeftOffset + 32) >> 6
[0122] yrpre = ( (Sign (refy) * ( (Abs (refy) + 128) >> 8) ) + fRefTopOffset + 32) >> 6
[0123] applies rounding by adding 128 and right-shifting by 8, preserving the sign, then adds offset values and performs another right-shift for further scaling.
[0124] xr = xrpre >> 4 and yr = yrpre >> 4 produce final coordinates (xr, yr) in the reference picture.
[0125] The offset variables:
[0126] fRefLeftOffset = ( (SubWidthC *pps_scaling_win_left_offset) << 10)
[0127] fRefTopOffset = ( (SubHeightC *pps_scaling_win_top_offset) << 10)
[0128] are used in the calculations above.
[0129] In certain embodiments, alternative formulas can also be used, since the partition prediction information is generated based on a relatively large area (2M×2M, typically the size of a CTU) , so small differences in the rounding mechanism may not significantly affect the quality of the fetched information.
[0130] Once the refined position (xr, yr) is calculated, the encoder or decoder can determine the corresponding grid point in the reference frame and use its partition prediction information to guide the partitioning of the current CU, even when the reference and current pictures have different resolutions.
[0131] Based on the updated position (xr, yr) , the corresponding grid point in the reference frame is determined and the partition prediction information from the grid point is used to predict the partition of current frame while coding the CU whose center position is (xc, yc) . For example, consider a case where a CU in the current picture has its center at a particular position and the reference picture has a different resolution according to the specified scaling ratios. After applying the re-positioning formulas, the encoder calculates a corresponding position in the reference picture that accounts for the resolution difference. The encoder would then fetch the partition prediction information from the nearest grid point in the reference picture. This information might indicate that the corresponding area in the reference picture benefited from certain QT and MTT depths, suggesting similar partitioning might be appropriate for the current CU. By doing this coordinate transformation, the encoder can maintain the benefits of temporal prediction for partitioning decisions despite the resolution differences between pictures. This approach strikes a balance between computational efficiency and coding performance, avoiding both the exhaustive partition searches and the potential quality loss that would occur if temporal prediction were simply disabled.
[0132] 2.3 Multiple temporal partition prediction information according to the scaling ratios of the current picture and the reference picture
[0133] In addition to re-position the location to fetch the partition prediction information, the grid resolution and the region used to calculate the partition prediction information can be chosen adaptively according to the scaling ratios of the current picture and the reference picture. The scaling ratios in the following description refers to the scaling ratio calculated by referring the maximum resolution picture of a sequence.
[0134] As shown in FIG. 2, the grid resolution is N by N and for every grid point a 2M-by-2M area is used to calculate the partition prediction information to be used for future picture. Therefore, N and M determine the granularity and scope of information aggregation when extracting the partition prediction information from a reference frame. For one embodiment, the value of N and M is adaptively determined by the scaling ratios of the current picture and reference picture. One picture can be referred multiple times by different pictures of the sequence. Therefore, multiple sets of N and M values are required for dealing with all possible scaling ratio differences between current picture and reference picture. Table 1 and Table 2 shows two examples for determining the value of N and M according to different scaling ratios. For different ratios, the calculated partition prediction information is stored to different information buffers with distinct identifications (IDs) . The required storage size is determined by the grid resolution N. For a frame to be referred multiple times with different scaling ratio relationships, multiple partition prediction information buffers are required to store the calculated information. Besides, in Table 1 and Table 2, the picture is assumed to be scaled with equal factor vertically and horizontally. Therefore, only one scaling ratio for reference picture or current picture is listed in the table. For cases, where vertical scaling factor is different from horizontal scaling factor, more combinations of N and M could also be determined with similar concept disclosed here.
[0135] Table 1
[0136] Table 2
[0137] In other words, this approach recognizes that the relationship between partition structures at different resolutions is complex and depends on the specific scaling ratios involved. By adaptively selecting the grid resolution (N) and the size of the area (M) used for calculating partition information, the method can better account for how content details and optimal partitioning change across different scaling ratios.
[0138] For example, when a reference picture with scaling ratio 1.0 is used to predict a current picture with scaling ratio 2.0, this represents a downscaling where fine details in the reference picture may be compressed in the current picture. In this case, using the same grid resolution but adjusting the area size (as shown in Table 1) allows the algorithm to aggregate information from an appropriately larger area of the reference picture.
[0139] Conversely, when a reference picture with scaling ratio 1.5 is used to predict pictures at various other scaling ratios (as shown in Table 2) , both the grid resolution and area size are adjusted to maintain an appropriate balance between precision and coverage. For instance, when predicting a picture with scaling ratio 2.0 from a reference with scaling ratio 1.5, a smaller grid resolution (N=10) and larger relative area size are used to account for the increased downscaling.
[0140] Each combination of reference and current picture scaling ratios is assigned a unique buffer ID, allowing the system to store and retrieve the appropriate partition prediction information for each scenario. By maintaining these multiple sets of information, the encoder can make better-informed partitioning decisions that account for how content characteristics and optimal partitioning strategies change with different scaling ratios, ultimately achieving better compression efficiency while still benefiting from the computational savings of temporal partition prediction.
[0141] 2.4 Methods for generating multiple temporal partition prediction information
[0142] There are several methods to obtain the multiple partition prediction information. These methods are listed below:
[0143] Method 1: Preserve CU partition information of the whole picture (e.g., coding structure of a picture) . Recalculate partition prediction information on demand.
[0144] In method 1, the system preserves the complete CU partition information of the entire picture-essentially storing the full coding structure after a picture has been encoded or decoded. This includes detailed information about every CU in the picture, such as the exact position and size of each CU, the partitioning decisions made (QT, BT, or TT splits at each level) , the depth values for each type of partitioning (QT depth, MTT depth, etc. ) , and the hierarchical relationship between parent and child nodes in the partition tree.
[0145] When temporal partition prediction is needed for a new picture with a different scaling ratio, the system doesn't rely on pre-calculated summary information. Instead, it accesses this stored detailed structure and performs "on-demand" calculation of the partition prediction information specifically tailored to the current scaling ratio relationship.
[0146] This approach offers maximum flexibility and precision, as it can dynamically generate exactly the right partition prediction information for any scaling ratio combination. The system can customize the grid resolution (N) and area size (M) parameters based on the specific scaling relationship between the current and reference pictures.
[0147] Method 2: Calculate all possible combinations of the scaling ratio relationships after encoding or decoding a new frame. Table 3 shows an example for scaling ratio of current picture equaling to 1.25.
[0148] Table 3
[0149] Method 2 takes a more systematic approach to generating multiple temporal partition prediction information sets by calculating all possible scaling ratio combinations after encoding or decoding each frame. Rather than storing the detailed coding structure of each picture (as in method 1) , method 2 pre-computes and stores the derived partition prediction information for all potential scaling ratio relationships that might be needed in future frames. The example in Table 3 illustrates this approach for a current picture with scaling ratio 1.25.
[0150] The video coding system can identify all possible reference picture scaling ratios that might be used in combination with the current picture (1.0, 1.25, 1.33, 1.5, 1.75, and 2.0 in the example) . For each combination, it can calculate the appropriate grid resolution (N) and area size (M) parameters based on the relative scaling relationship. For instance, when both reference and current pictures have the same scaling ratio (1.25) , the grid resolution remains at 16 and the area size is calculated as (CTU size*1.0) / 2. In another example, when the reference picture has a higher scaling ratio (e.g., 2.0) compared to the current picture (1.25) , the grid resolution can be reduced to 12 and the area size is adjusted to (CTU size*0.625) / 2.
[0151] The method 2 can generate and store partition prediction information for each combination, assigning different buffer IDs (0-5 in the example) to each set. When encoding or decoding future frames, the video coding system can simply retrieve the appropriate pre-calculated information based on the actual scaling ratio relationship.
[0152] The advantage of method 2 is that it ensures partition prediction information is immediately available for any scaling ratio combination, without requiring on-demand calculation. This improves processing speed during the actual encoding or decoding of video frames.
[0153] Method 3: Generating multiple temporal partition prediction information according to signaled picture-level syntax elements. Explicitly signal which information IDs are required. Encoder knows the best scaling ratios therefore the encoder can send this information in PPS or picture header.
[0154] The advantages of explicit signaling the IDs are:
[0155] 1. Traverse once of the whole frame to get the partition prediction information for different grid resolutions and information collection areas.
[0156] 2. No need to keep partition information of every frame as in method 1. For getting the partition information, CU information like location, size, QT depth, MTT depth, etc. should be kept. Keeping partition prediction information requires less storage space because it has already bee sub-sampled by the grid stride.
[0157] These explicit signaled IDs help the decoder to manage the buffers for storing partition prediction information. For example, using the ID assigned as in Table 3. Required IDs can be signaled in PPS or picture header. If only ID 1 and ID 5 are signaled as necessary, decoder only needs to generate two combinations of partition prediction information for future decoding process.
[0158] In other word, method 3 introduces a more selective approach by explicitly signaling which partition prediction information sets are actually needed. Rather than calculating all possible combinations (Method 2) or storing detailed structures for on-demand calculation (Method 1) , method 3 uses picture-level syntax elements to communicate which specific information sets should be generated and stored.
[0159] The encoder, having a broader view of the encoding process, can determine which scaling ratio combinations will actually be used in the video sequence. By explicitly signaling this information through syntax elements in the Picture Parameter Set (PPS) or picture header, the encoder guides the system to generate only the necessary information sets.
[0160] For example, using ID system in Table 3, if the encoder determines that only the combinations corresponding to ID 1 (reference 1.25, current 1.25) and ID 5 (reference 2.0, current 1.25) will be needed, it can signal just these two IDs. This means the decoder only needs to generate and store these specific sets of partition prediction information, reducing both computational load and storage requirements.
[0161] This approach can enhance efficiency in processing and reduce storage requirements. The video coding system needs to traverse the frame only once to generate partition information for different grid resolutions and area sizes that are actually needed, rather than calculating unused combinations. Also, by storing only the necessary partition prediction information (which is already sub-sampled by the grid stride) , method 3 requires less memory than storing full CU structures or all possible combinations.
[0162] Any of the foregoing proposed methods can be implemented in encoders and / or decoders. These methods offer flexible integration options within existing video coding architectures. For example, any of the proposed methods can be implemented directly within an inter / intra / prediction module of an encoder, where they would operate alongside existing prediction mechanisms to guide block partitioning decisions during the encoding process. Similarly, they can be integrated into the corresponding inter / intra / prediction module of a decoder to ensure that the same partitioning decisions are applied during the decoding process.
[0163] Alternatively, these methods can be implemented as separate specialized circuits or processing units that operate in conjunction with existing modules. In this configuration, a dedicated circuit would be coupled to the inter / intra / prediction module of the encoder and / or decoder, functioning as an auxiliary component that calculates and provides temporal partition prediction information to the primary module when needed. This modular approach allows for easier integration into existing hardware designs without requiring modifications to core processing elements. The auxiliary circuit would analyze the scaling relationship between current and reference pictures, apply the appropriate method (turning off prediction, re-positioning, or using multiple information sets) , and feed the resulting partition prediction information to the main inter / intra / prediction module to guide its partitioning decisions.
[0164] This implementation flexibility ensures that the proposed methods can be adopted across various hardware and software platforms, from dedicated video processing chips to software encoders, allowing manufacturers to select the integration approach that best suits their specific architecture and performance requirements.
[0165] 3. Flow Diagram
[0166] FIG. 4 is a flow diagram illustrating a method 400 video coding in a video system for encoding or decoding video pictures according to the embodiments. The method 400 comprises the following steps:
[0167] S402: Receive input data associated with a current block in a current picture;
[0168] S404: Determine a temporal partition prediction based on a relation between current picture and a reference picture;
[0169] S406: Partition the current block according to the temporal partition prediction; and
[0170] S408: Encode or decode the current block according to a block partitioning structure of the current block.
[0171] The input data comprise pixel data to be encoded at an encoder side or data associated with the current block to be decoded at a decoder side.
[0172] In this step S402, the video coding system obtains the data necessary for processing the current block. At the encoder side, this input data comprises original pixel data (such as RGB or YUV values) representing the visual content to be compressed. The encoder receives this uncompressed pixel data, typically organized into blocks of predetermined dimensions (e.g., 64×64, 32×32, or 16×16 pixels) that serve as the basic units for subsequent coding operations. At the decoder side, the input data comprises encoded bitstream elements associated with the current block, including syntax elements that specify coding parameters, prediction modes, transform coefficients, and other information necessary to reconstruct the block. This step establishes the foundation for all subsequent processing operations in the video coding pipeline.
[0173] In step S404, the video coding system analyzes the relationship between the current picture and a previously coded reference picture to predict appropriate partitioning structures. This step may involve examining already-coded blocks in collocated or nearby regions of the reference picture and extracting partition information such as quad tree (QT) depth, multi-type tree (MTT) depth, binary tree (BT) depth, and ternary tree (TT) depth values. When reference picture resampling is enabled, the video coding system can account for potential resolution differences between the current and reference pictures. Depending on implementation, this step may involve one of the three approaches described in the invention: (1) determining whether to disable temporal prediction based on scaling ratios, (2) calculating refined positions in the reference picture to obtain relevant partition information, or (3) selecting from multiple sets of partition prediction information based on the specific scaling ratio relationship.
[0174] Based on the temporal partition prediction determined in the previous step, in step S406, the video coding system applies appropriate partitioning decisions to the current block. This step involves recursively splitting the block using combinations of quad tree (QT) , binary tree (BT) , and ternary tree (TT) partitioning according to the predicted structure. For example, the partitioning process can begin with quad tree splitting and may continue with multi-type tree splitting to various depths as guided by the temporal prediction information. The video coding system may apply constraints to increase or decrease partition depths based on the relationship between current and temporal partition values, optimizing the trade-off between encoding complexity and compression efficiency. Thus, the initial coding unit is transformed into a hierarchical structure of smaller coding units that better adapt to the local characteristics of the video content.
[0175] In step S408, the video coding system performs the encoding or decoding operations using the determined block partitioning structure. At the encoder side, the video coding system may perform mode decision, prediction (inter or intra) , transform, quantization, and entropy coding for each leaf node in the partition structure. The partitioning information itself can also be encoded into the bitstream to enable proper decoding. At the decoder side, the video coding system can interpret the received partition information, and then it can perform inverse processes including entropy decoding, inverse quantization, inverse transform, and prediction to reconstruct the pixel values of the current block. The optimized partitioning structure can enable more efficient compression by allowing different prediction and transform operations to be applied to sub-blocks of appropriate sizes, resulting in improved rate-distortion performance of the video coding system.
[0176] 4. Inter / Intra Video Coding System
[0177] FIGs. 5A and 5B illustrates an exemplary adaptive inter / intra video coding system for performing the above-described video coding techniques. For intra-prediction module 110, the prediction data is derived based on previous coded video data in the current picture. For inter-prediction module 112, motion estimation (ME) is performed at the encoder side and motion compensation (MC) is performed based on the result of motion estimation to provide prediction data derived from other pictures and motion data. A selection switch 114 selects between intra-prediction module 110 or inter-prediction module 112, and the selected prediction data is supplied to an adder 116 to form prediction errors, also called residues. The residues are then processed by transform module (T) 118 followed by quantization module (Q) 120. The transformed and quantized residues are then coded by an entropy encoder 122 to be included in a video bitstream corresponding to the compressed video data. The bitstream associated with the transform coefficients is then packed with side information such as motion and coding modes associated with Intra-prediction and Inter-prediction, and other information such as parameters associated with loop filters applied to the underlying image area. The side information associated with Intra-prediction module 110, Inter-prediction module 112 and in-loop filter (ILPF) 130, are provided to the entropy encoder 122 as shown in FIG. 5A. When an inter-prediction mode is used, a reference picture or pictures have to be reconstructed at the encoder end as well. Consequently, the transformed and quantized residues are processed by inverse quantization module (IQ) 124 and inverse transform module (IT) 126 to recover the residues. The residues are then added back to prediction data 136 at a reconstruction module (REC) 128 to reconstruct video data. The reconstructed video data may be stored in a reference picture buffer 134 and used for prediction of other frames.
[0178] As shown in FIG. 5A, incoming video data undergoes a series of encoding operations in the encoding system. The reconstructed video data from REC 128 may be subject to various impairments due to these encoding operations. To improve video quality, an in-loop filter 130 is applied to the reconstructed video data before it is stored in the Reference Picture Buffer 134. The in-loop filter 130 may include multiple filtering operations such as a deblocking filter (DF) , a sample adaptive offset (SAO) , and an adaptive loop filter (ALF) . Since a decoder needs to apply identical filtering operations, the loop filter information must be incorporated into the bitstream. Therefore, this loop filter information is provided to the entropy encoder 122 for incorporation into the encoded bitstream. As illustrated in FIG. 5A, the in-loop filter 130 processes the reconstructed video data before the filtered samples are stored in the reference picture buffer 134. This encoding system architecture shown in FIG. 5A represents an exemplary structure of a typical video encoder, which may be implemented in various video coding standards such as High Efficiency Video Coding (HEVC) , VP8, VP9, Advanced Video Coding (H. 264) , or Versatile Video Coding (VVC) .
[0179] The decoder architecture, as illustrated in FIG. 5B, shares several functional similarities with the encoder but operates in a complementary manner to reconstruct the original video data. Unlike the encoder which requires both transform module 118 and quantization module 120 for compression, the decoder only needs inverse quantization module 124 and inverse transform module 126 to reverse the compression process. In the decoder, the entropy decoder 140, replaces the encoder's entropy encoder 122. This entropy decoder 140 performs the task of interpreting the received video bitstream, extracting both the quantized transform coefficients and essential coding information, including ILPF information, Intra-prediction information, and Inter-prediction information.
[0180] The intra-prediction module 150 of the decoder operates more efficiently than its encoder counterpart since it does not need to perform the computationally intensive mode search process. Instead, it directly generates the Intra-prediction signal by applying the intra-prediction information received from the entropy decoder 140. This information precisely specifies which prediction mode to use, eliminating the need for the extensive mode evaluation process required at the encoder side.
[0181] Similarly, the Inter-prediction process at the decoder is streamlined compared to the encoder. The motion compensation module (MC) 152 only needs to execute the motion compensation operation based on the motion vectors and reference picture information received through the entropy decoder 140. This is simpler than the encoder's Inter-prediction process, which must perform both motion estimation to find the best motion vectors and motion compensation to generate the prediction signal. The decoder can apply the received motion information to reconstruct the Inter-predicted blocks, accessing the necessary reference picture data from its reference picture buffer 134.
[0182] 5. Additional Note
[0183] The terminology employed in the description of the various embodiments herein is intended for the purpose of describing particular embodiments and should not be construed as limiting. In the context of this description and the appended claims, the singular forms "a" , "an" , and "the" are intended to encompass plural forms as well, unless the context clearly indicates otherwise.
[0184] It should be understood that the term "and / or" as used herein is intended to encompass any and all possible combinations of one or more of the associated listed items. Furthermore, it should be noted that the terms "includes, " "including, " "comprises, " and / or "comprising, " when used in this specification, indicate the presence of stated features, integers, steps, operations, elements, and / or components, but do not exclude the presence or addition of one or more other features, integers, steps, operations, elements, components, and / or groups thereof.
[0185] In the context of this disclosure, the terms "coupled, " "connected, " "connecting, " "electrically connected, " and similar expressions are used interchangeably to broadly denote the state of being electrically or electronically connected. Furthermore, an entity is deemed to be in "communication" with another entity (or entities) when it electrically transmits and / or receives information signals to / from the other entity, irrespective of whether these signals contain image / voice information or data / control information, and regardless of the signal type (analog or digital) . It is important to note that this communication can occur through either wired or wireless means. The use of these terms is intended to encompass all forms of electrical or electronic connectivity relevant to the described embodiments.
[0186] The use of ordinal designators like "first, " "second, " and so forth in the specification and claims serves to differentiate between multiple instances of similarly named elements. These designators do not imply any inherent sequence, priority, or chronological order in the manufacturing process or functional relationship between elements. Rather, they are employed solely as a means of uniquely identifying and distinguishing between separate instances of elements that share a common name or description.
[0187] The directional terms used in the embodiments such as up, down, left, right, upper-side, down-side, in front of or behind are just the directions referring to the attached figures. Thus, the direction terms used in the present disclosure are for illustration, and are not intended to limit the scope of the present disclosure. It should be noted that the elements which are specifically described or labeled may exist in various forms for those skilled in the art.
[0188] As may be used throughout this specification and the appended claims, terms of approximation and degree such as "substantially, " "approximately, " "generally, " "essentially, " "nearly, " "about, " and similar expressions are used to account for variations in precision, manufacturing tolerances, measurement accuracy, environmental conditions, and inherent material properties that may affect the described features or characteristics. Such variations may range from ±20%in broader applications to progressively tighter tolerances of ±10%, ±5%, ±3%, ±2%, ±1%, or ±0.5%in more precise implementations. The specific degree of variation encompassed by these terms of approximation in any given context is informed by the nature of the component, relationship, or parameter being described, the technical requirements of the particular embodiment, and the understanding of one skilled in the relevant art.
[0189] This interpretation of terminology is provided to ensure clarity and consistency throughout the specification and claims, and should not be construed as restricting the scope of the disclosed embodiments or the appended claims.
[0190] The various illustrative components, logic, logical blocks, modules, circuits, operations and algorithm processes described in connection with the embodiments disclosed herein may be implemented as electronic hardware, firmware, software, or combinations of hardware, firmware or software, including the structures disclosed in this specification and the structural equivalents thereof. The interchangeability of hardware, firmware and software has been described generally, in terms of functionality, and illustrated in the various illustrative components, blocks, modules, circuits and processes described above. Whether such functionality is implemented in hardware, firmware or software depends upon the particular application and design constraints imposed on the overall system.
[0191] The hardware and data processing apparatus utilized to implement the various illustrative components, logics, logical blocks, modules, and circuits described herein may comprise, without limitation, one or more of the following: a general-purpose single-chip or multi-chip processor, a digital signal processor (DSP) , an application specific integrated circuit (ASIC) , a field programmable gate array (FPGA) , other programmable logic devices (PLDs) , discrete gate or transistor logic, discrete hardware components, or any suitable combination thereof. Such hardware and apparatus shall be configured to perform the functions described herein.
[0192] A general-purpose processor may include, but is not limited to, a microprocessor, or alternatively, any conventional processor, controller, microcontroller, or state machine. In certain implementations, a processor may be realized as a combination of computing devices. Such combinations may include, for example, a DSP and a microprocessor, a plurality of microprocessors, one or more microprocessors in conjunction with a DSP core, or any other such configuration as may be suitable for the intended application.
[0193] It is to be understood that in some embodiments, particular processes, operations, or methods may be executed by circuitry specifically designed for a given function. Such function-specific circuitry may be optimized to enhance performance, efficiency, or other relevant metrics for the particular task at hand. The selection of specific hardware implementation shall be determined based on the particular requirements of the application, which may include, inter alia, performance specifications, power consumption constraints, cost considerations, and size limitations.
[0194] In certain aspects, the subject matter described herein may be implemented as software. Specifically, various functions of the disclosed components, or steps of the methods, operations, processes, or algorithms described herein, may be realized as one or more modules within one or more computer programs. These computer programs may comprise non-transitory processor-executable or computer-executable instructions, encoded on one or more tangible processor-readable or computer-readable storage media. Such instructions are configured for execution by, or to control the operation of, data processing apparatus, including the components of the devices described herein. The aforementioned storage media may include, but are not limited to, Random Access Memory (RAM) , Read Only Memory (ROM) , Electrically Erasable Programmable Read-Only Memory (EEPROM) , Compact Disc Read-Only Memory (CD-ROM) or other optical disk storage, magnetic disk storage or other magnetic storage devices, or any other medium capable of storing program code in the form of instructions or data structures. It should be understood that combinations of the above-mentioned storage media are also contemplated within the scope of computer-readable storage media for the purposes of this disclosure.
[0195] Various modifications to the embodiments described in this disclosure may be readily apparent to persons having ordinary skill in the art, and the generic principles defined herein may be applied to other embodiments without departing from the spirit or scope of this disclosure. Thus, the claims are not intended to be limited to the embodiments shown herein, but are to be accorded the widest scope consistent with this disclosure, the principles and the novel features disclosed herein.
[0196] In certain implementations, the embodiments may comprise the disclosed features and may optionally include additional features not explicitly described herein. Conversely, alternative implementations may be characterized by the substantial or complete absence of non-disclosed elements. For the avoidance of doubt, it should be understood that in some embodiments, non-disclosed elements may be intentionally omitted, either partially or entirely, without departing from the scope of the invention. Such omissions of non-disclosed elements shall not be construed as limiting the breadth of the claimed subject matter, provided that the explicitly disclosed features are present in the embodiment.
[0197] Additionally, various features that are described in this specification in the context of separate embodiments also can be implemented in combination in a single implementation. Conversely, various features that are described in the context of a single implementation also can be implemented in multiple embodiments separately or in any suitable subcombination. As such, although features may be described above as acting in particular combinations, and even initially claimed as such, one or more features from a claimed combination can in some cases be excised from the combination, and the claimed combination may be directed to a subcombination or variation of a subcombination.
[0198] The depiction of operations in a particular sequence in the drawings should not be construed as a requirement for strict adherence to that order in practice, nor should it imply that all illustrated operations must be performed to achieve the desired results. The schematic flow diagrams may represent example processes, but it should be understood that additional, unillustrated operations may be incorporated at various points within the depicted sequence. Such additional operations may occur before, after, simultaneously with, or between any of the illustrated operations.
[0199] Additionally, it should be understood that the various figures and component diagrams presented and discussed within this document are provided for illustrative purposes only and are not drawn to scale. These visual representations are intended to facilitate understanding of the described embodiments and should not be construed as precise technical drawings or limiting the scope of the invention to the specific arrangements depicted.
[0200] In certain implementations, multitasking and parallel processing may prove advantageous. Furthermore, while various system components are described as separate entities in some embodiments, this separation should not be interpreted as mandatory for all embodiments. It is contemplated that the described program components and systems may be integrated into a single software package or distributed across multiple software packages, as dictated by the specific implementation requirements.
[0201] It should be noted that other embodiments, beyond those explicitly described, fall within the scope of the appended claims. The actions specified in the claims may, in some instances, be performed in an order different from that in which they are presented, while still achieving the desired outcomes. This flexibility in execution order is an inherent aspect of the claimed processes and should be considered within the scope of the invention.
[0202] While the invention has been described in connection with certain embodiments, it will be understood by those skilled in the art that various modifications and adaptations can be made without departing from the scope of the invention. The specific embodiments presented are intended to illustrate the invention and not to limit its application or construction. Those skilled in the art will readily observe that numerous modifications and alterations of the device and method may be made while retaining the teachings of the invention. Accordingly, the above disclosure should be construed as limited only by the metes and bounds of the appended claims.
Claims
1.A method of video coding in a video system for encoding or decoding video pictures, comprising:receiving input data associated with a current block in a current picture, wherein the input data comprise pixel data to be encoded at an encoder side or data associated with the current block to be decoded at a decoder side;determining a temporal partition prediction based on a relation between the current picture and a reference picture;partitioning the current block according to the temporal partition prediction; andencoding or decoding the current block according to a block partitioning structure of the current block.2.The method of claim 1, wherein determining the temporal partition prediction comprises:determining a scaling ratio between the current picture and the reference picture; anddisabling the temporal partition prediction when the scaling ratio deviates from unity.3.The method of claim 1, wherein determining the temporal partition prediction comprises:calculating a refined position in the reference picture based on a scaling ratio between the current picture and the reference picture; andderiving a set of partition prediction information associated with the temporal partition prediction from the refined position.4.The method of claim 3, wherein calculating the refined position from an original position in the current picture comprises:calculating a horizontal reference value by multiplying a horizontal coordinate difference between the original position and a scaling window left offset by the horizontal scaling ratio;calculating a vertical reference value by multiplying a vertical coordinate difference between the original position and a scaling window top offset by the vertical scaling ratio;deriving a preliminary horizontal refined position by applying a rounding operation to the horizontal reference value and adjusting by a reference left offset value;deriving a preliminary vertical refined position by applying a rounding operation to the vertical reference value and adjusting by a reference top offset value; anddetermining the final refined horizontal and vertical positions by applying bit-shifting operations to the preliminary positions,wherein the horizontal scaling ratio and the vertical scaling ratio represent relative size differences between the current picture and the reference picture, and wherein the scaling window offsets and reference offsets are determined based on chroma subsampling parameters.5.The method of claim 1, wherein determining the temporal partition prediction comprises:deriving one or more sets of partition prediction information associated with the temporal partition prediction, each set corresponding to a scaling ratio between the current picture and reference picture; andselecting a set of partition prediction information associated with the temporal partition prediction.6.The method of claim 5, wherein the one or more sets of partition prediction information are characterized by a grid resolution and an area size.7.The method of claim 5, wherein deriving the one or more sets of partition prediction information comprises signaling or parsing, via picture level syntax elements, a set of partition prediction information.8.The method of claim 5, wherein deriving the one or more sets of partition prediction information comprises:storing coding unit (CU) partition information for the reference picture; andderiving the one or more set of partition prediction information from the stored CU partition information for the reference picture.9.The method of claim 5, wherein deriving the one or more sets of partition prediction information comprises:generating one or more sets of partition prediction information for a plurality of scaling ratio combinations between the reference picture and the current picture.10.The method of claim 5, wherein deriving the one or more sets of partition prediction information comprises:signaling or parsing, via picture level syntax elements, the one or more sets of partition prediction information.11.The method of claim 1, wherein determining the temporal partition prediction comprises:calculating one or more partition depth values from a collocated region in the reference picture, wherein the one or more partition depth values include at least one of:minimum temporal quad tree (QT) depth;maximum temporal QT depth;average temporal QT depth;minimum temporal multi-type tree (MTT) depth;maximum temporal MTT depth;average temporal MTT depth;minimum temporal binary tree (BT) depth;maximum temporal BT depth;average temporal BT depth;minimum temporal ternary tree (TT) depth;maximum temporal TT depth; andaverage temporal TT depth; andapplying at least one constraint to block partitioning of the current block based on the one or more partition depth values.12.The method of claim 11, wherein calculating the one or more partition depth values comprises:determining a calculation area centered at a grid point in the collocated region of the reference picture;determining one or more of coding units (CUs) having upper-left corners located within the calculation area;determining the minimum temporal QT depth by finding a lowest QT depth value among all the identified CUs; anddetermining the maximum temporal MTT depth by finding a highest MTT depth value among the CUs.13.The method of claim 11, wherein partitioning the current block comprises using the QT with nested MTT segmentation structure comprising binary and ternary splits.14.An apparatus of video processing in a video coding system for encoding or decoding video pictures, the apparatus comprising one or more electronic circuits configured to:receive input data associated with a current block in a current picture, wherein the input data comprise pixel data to be encoded at an encoder side or data associated with the current block to be decoded at a decoder side;determine a temporal partition prediction based on a relation between the current picture and a reference picture;partition the current block according to the temporal partition prediction; andencode or decode the current block according to a block partitioning structure of the current block.
Citation Information
Patent Citations
Method and apparatus for encoding and decoding image using object boundary based partition
CN101682778A
Video encoding or decoding method and device related to high-level information signaling
CN114902660A
Encoder, decoder, encoding method, and decoding method
US20190342550A1
Method and apparatus for video coding
US20210051346A1