Video encoding or decoding method and apparatus with scaling ratio constraints

By introducing proportional constraints in the video encoding and decoding system, ensuring that the scale window ratio of the reference image to the current image is within a specific range, the problems of video data adaptability and jump effect are solved, and the video encoding efficiency and quality are improved.

CN120475174APending Publication Date: 2025-08-12HFI INNOVATION INC
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510591023.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2019-12-18
Filing Date
2020-12-10
Publication Date
2025-08-12

AI Technical Summary

Technical Problem

In video encoding and decoding systems, the prior art is difficult to effectively constrain the scaling ratio between the reference image and the current image, resulting in the video data being unable to seamlessly adapt to dynamic channel conditions and user preferences, and may cause a jump effect.

Method used

By introducing a scale constraint in the video processing method, ensuring that the ratio between the reference image and the scaling window width, height or size of the current image is within a specific range, this scale constraint is used to generate a reference block for motion compensation and encode or decode the current block.

Benefits of technology

It realizes seamless adaptation of video data at different resolutions and sample bit depths, improves video encoding efficiency, reduces jumping effect, and enhances the flexibility and quality of video transmission.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120475174A_ABST
    Figure CN120475174A_ABST
Patent Text Reader

Abstract

A video processing method and apparatus for processing a current block in a current image by reference image re-sampling, includes receiving input video data of the current block, determining a zoom window of the current image and a zoom window of a reference image. The current image may have a different zoom window size from the reference image. The ratio between the zoom window width, height, or size of the reference image and the zoom window width, height, or size of the current image is constrained within a ratio constraint. A reference block is generated from the reference picture according to the ratio and used to encode or decode the current block.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] Cross-references This application claims priority to U.S. Provisional Patent Applications Serial No. 62 / 946,540, entitled “Method of Scaling Ratio Constraint,” filed on December 11, 2019, and Serial No. 62 / 949,506, entitled “Method of Scaling Window Constraint,” filed on December 18, 2019, respectively. The U.S. Provisional Patent Applications are hereby incorporated by reference in their entirety. Technical Field

[0002] The present invention relates to a video processing method and apparatus in a video encoding and decoding system, and more particularly to scaling constraints for reference picture resampling. Background Art

[0003] The Versatile Video Codec (VVC) standard is an upcoming video codec standard that builds upon the previous High Efficiency Video Coding (HEVC) standard by enhancing existing codec tools and introducing several new codec tools in various codec building blocks. The VVC standard improves compression performance as well as transmission and storage efficiency, and supports new formats such as High Dynamic Range and omni-directional 360 video. The VVC standard makes video transmission over mobile networks more efficient by allowing systems or locations with poor data rates to receive larger files faster. VVC supports layer coding and spatial or temporal scalability of the signal-to-noise ratio (SNR).

[0004] Reference picture resampling (RPR) In the VVC standard, fast presentation switching for adaptive streaming services is desirable to deliver multiple presentations of the same video content at the same time, each with different characteristics. The different characteristics involve different spatial resolutions or different sample bit depths. In real-time video communication, by allowing the resolution to be changed in the encoded video sequence without inserting I-pictures, not only can the video data be seamlessly adapted to dynamic channel conditions and user preferences, but the beating effect caused by I-pictures can also be removed. Reference picture resampling (RPR) allows pictures with different resolutions to reference each other for inter-frame prediction. Figure 1This diagram illustrates an example of applying reference picture resampling to encode or decode a current picture, where inter-coded blocks of the current picture are predicted from a reference picture of the same or different size. Spatial scalability is beneficial in streaming applications. When spatial scalability is supported, the reference picture can have a different size than the current picture. The VVC standard uses RPR to support on-the-fly upsampling and downsampling motion compensation.

[0005] Table 1 shows an example of signaling the RPR enable flag and maximum picture size in a Sequence Parameter Set (SPS). The RPR enable flag, sps_ref_pic_resampling_enabled_flag, signaled in the SPS is used to indicate whether RPR is enabled for pictures referencing the SPS. When this RPR enable flag is equal to 1, the current picture referencing the SPS may have slices that reference a reference picture in the active entry of the reference picture layer, where the reference picture layer has one or more of the following seven parameters that are different from those of the current picture. The seven parameters include syntax elements associated with the following: picture width pps_pic_width_in_luma_samples, picture height pps_pic_height_in_luma_samples, left scaling window offset pps_scaling_win_left_offset, right scaling window offset pps_scaling_win_right_offset, upper scaling window offset pps_scaling_win_top_offset, lower scaling window offset pps_scaling_win_bottom_offset, and number of sub-pictures sps_num_subpics_minus1. For a current picture that refers to a reference picture (the reference picture has one or more of the seven parameters different from the current picture), the reference picture may belong to the same layer as the layer containing the current picture or a different layer. The syntax element sps_res_change_in_clvs_allowed_flag equal to 1 indicates that the picture spatial resolution may change within the Coded Layer Video Sequence (CLVS) referencing the SPS, and the syntax element equal to 0 indicates that the picture spatial resolution does not change within any CLVS referencing the SPS. The maximum picture size is signaled in the SPS via the syntax elements sps_pic_width_max_in_luma_samples and sps_pic_height_max_in_luma_samples and must not be larger than the picture size of the Decoded Picture Buffer (DPB) of the Output Layer Set (OLS) signaled in the corresponding Video Parameter Set (VPS).

[0006] Table 1

[0007] When RPR is used to predict a current picture, a picture size ratio is derived from the reference picture width or height and the current picture width or height. The picture size ratio is constrained to be in the range between 1 / 8 and 2. For example, the picture width and height measured in luma samples are derived from the syntax elements pic_width_in_luma_samples and pic_height_in_luma_samples signaled in a Picture Parameter Set (PPS). The syntax element pic_width_in_luma_samples specifies the width (in luma samples) of each decoded picture referenced to the PPS. This syntax element must not be equal to 0 and should be an integer multiple of Max(8,MinCbSizeY) and constrained to be less than or equal to pic_width_max_in_luma_samples. When the subpics_present_flag is equal to 1 or when the RPR enable flag ref_pic_resampling_enabled_flag is equal to 0, the value of the syntax element pic_width_in_luma_samples shall be equal to pic_width_max_in_luma_samples. The syntax element pic_height_in_luma_samples specifies the height (in luma samples) of each decoded picture that references the PPS. This syntax element must not be equal to 0 and must be an integer multiple of Max(8, MinCbSizeY) and must be less than or equal to pic_height_max_in_luma_samples. When the subpics_present_flag is equal to 1 or when the RPR enable flag ref_pic_resampling_enabled_flag is equal to 0, the value of the syntax element pic_height_in_luma_samples is set to pic_height_max_in_luma_samples.

[0008] In the current design of RPR in VVC draft 6, when the image sizes of a current picture and a reference picture are specified, the following constraint must be satisfied. This constraint limits the image size ratio of the reference picture to the current picture to be in the range [1 / 8, 2]. Assume that the variables refPicWidthInLumaSamples and refPicHeightInLumaSamples are the image width and image height of a reference picture referenced by a current picture. The bitstream specification requires that all of the following conditions are met: the image width pic_width_in_luma_samples of the current image multiplied by two should be greater than or equal to the image width refPicWidthInLumaSamples of the reference image, the image height pic_height_in_luma_samples of the current image multiplied by two should be greater than or equal to the image height refPicHeightInLumaSamples of the reference image, the image width pic_width_in_luma_samples of the current image should be less than or equal to the image width refPicWidthInLumaSample of the reference image multiplied by eight, and the image height pic_height_in_luma_samples of the current image should be less than or equal to the image height refPicHeightInLumaSamples of the reference image multiplied by eight.

[0009] The picture size scaling ratio between a reference picture and a current picture is derived from the syntax elements pic_width_in_luma_samples and pic_height_in_luma_samples signaled in a PPS associated with the reference picture, and pic_width_in_luma_samples and pic_height_in_luma_samples signaled in a PPS associated with the current picture. The scaling window offset for RPR is also derived from the syntax elements signaled in the PPS. Table 2 shows these syntax elements signaled in the PPS and their corresponding semantics.

[0010] Table 2

[0011] The syntax element scaling_window_flag equal to 1 indicates that the scaling window offset parameter is present in the PPS, and scaling_window_flag equal to 0 indicates that the scaling window offset parameter is not present in the PPS. When the RPR enable flag ref_pic_resampling_enabled_flag is equal to 0, the value of this syntax element scaling_window_flag shall be equal to 0. The syntax elements scaling_win_left_offset, scaling_win_right_offset, scaling_win_top_offset, and scaling_win_bottom_offset specify the scaling offsets (in units of luma samples). These scaling offsets are applied to the picture size for scaling calculations. Scaling offsets can be negative. When a scaling window flag scaling_window_flag is equal to 0, the values of the four scaling offset syntax elements scaling_win_left_offset, scaling_win_right_offset, scaling_win_top_offset, and scaling_win_bottom_offset are inferred to be equal to 0.

[0012] The sum of the left and right offsets, scaling_win_left_offset and scaling_win_right_offset, should be less than the image width, pic_width_in_luma_samples, and the sum of the top and bottom offsets, scaling_win_top_offset and scaling_win_bottom_offset, should be less than the image height, pic_height_in_luma_samples. The variable PicOutputWidthL, representing the width of the scaling window, is derived by subtracting the left and right offsets from the image width. PicOutputWidthL = pic_width_in_luma_samples – (scaling_win_right_offset + scaling_win_left_offset). The variable PicOutputHeightL, representing the height of the scaling window, is derived by subtracting the top and bottom offsets from the image height. PicOutputHeightL = pic_height_in_luma_samples – (scaling_win_bottom_offset + scaling_win_top_offset) .

[0013] A variable fRefWidth is set equal to PicOutputWidthL of luma samples in a reference picture RefPicList[i][j], and a variable fRefHight is set equal to PicOutputHeightL of luma samples in the reference picture RefPicList[i][j]. A derived reference picture scale for the horizontal direction RefPicScale[i][j][0] is calculated by ((fRefWidth<<14) + (PicOutputWidthL>>1)) / PicOutputWidthL, and a derived reference picture scale for the vertical direction RefPicScale[i][j][1] is calculated by ((fRefHeight<<14) + (PicOutputHeightL>>1)) / PicOutputHeightL. Therefore, the derived reference image scale is RefPicIsScaled[i][j] = (RefPicScale [i][j][0] != (1<<14)) || (RefPicScale[i][j][1] != (1<<14).

[0014] In a recent proposal of the VVC standard, the scaling window offsets are measured in chroma samples, and when these scaling window offset syntax elements are not present in the PPS, the values of the four scaling offset syntax elements, scaling_win_left_offset, scaling_win_right_offset, scaling_win_top_offset, and scale_win_bottom_offset, are inferred to be equal to conf_win_left_offset, conf_win_right_offset, conf_win_top_offset, and conf_win_bottom_offset, respectively. A variable indicating the scaling window width, CurrPicScalWinWidthL, is derived from the picture width, SubWidthC, the left scaling offset, and the right scaling offset, and a variable indicating the scaling window height, CurrPicScalWinHeightL, is derived from the picture height, SubHeightC, the top scaling offset, and the bottom scaling offset, as shown below. CurrPicScalWinWidthL = pic_width_in_luma_samples – SubWidthC * ( scaling_win_right_offset + scaling_win_left_offset ); and CurrPicScalWinHeightL = pic_height_in_luma_samples – SubHeightC * ( scaling_win_bottom_offset + scaling_win_top_offset ). Summary of the Invention

[0015] In an exemplary embodiment of a video processing method for processing a current block in a current image, a video encoding or decoding system implementing the video processing method receives input video data associated with the current block; determines a scaling window width, height, or size of the current image; determines a scaling window width, height, or size of a reference image; generates a reference block based on a ratio between the scaling window width, height, or size of the reference image and the scaling window width, height, or size of the current image; uses the reference block to perform motion compensation for the current block; and encodes or decodes the current block in the current image. The ratio between the scaling window width, height, or size of the reference image and the scaling window width, height, or size of the current image is constrained within a ratio constraint.

[0016] In some exemplary embodiments, the ratio constraint is between 1 / M and N, where M and N are positive integers. To make the ratio of the scaling window width of the reference image to the scaling window width of the current image within the ratio constraint, N times the scaling window width of the current image is greater than or equal to the scaling window width of the reference image, and the scaling window width of the current image is less than or equal to M times the scaling window width of the reference image. To make the ratio of the scaling window height of the current image to the scaling window height of the reference image within the ratio constraint, N times the scaling window height of the current image is greater than or equal to the scaling window height of the reference image, and the scaling window height of the current image is less than or equal to M times the scaling window height of the reference image. In one embodiment, the scaling window size includes the scaling window width and the scaling window height. To ensure that the ratio of the zoom window size of the current image to the zoom window size of the reference image is within the ratio constraint, N times the zoom window width of the current image is greater than or equal to the zoom window width of the reference image, N times the zoom window height of the current image is greater than or equal to the zoom window height of the reference image, the zoom window width of the current image is less than or equal to M times the zoom window height of the reference image, and the zoom window height of the current image is less than or equal to M times the zoom window height of the reference image. For example, the ratio constraint is between 1 / 8 and 2. When the zoom window size of the current image is smaller than the zoom window size of the reference image, twice the zoom window width of the current image is greater than or equal to the zoom window width of the reference image, and twice the zoom window height of the current image is greater than or equal to the zoom window height of the reference image. When the scaling window size of the current image is larger than the scaling window size of the reference image, the scaling window width of the current image is less than or equal to eight times the scaling window width of the reference image, and the scaling window height of the current image is less than or equal to eight times the scaling window height of the reference image.

[0017] In some embodiments, the zoom window width of the current image is derived from an image width, a left zoom window offset, and a right zoom window offset of the current image, and the zoom window height of the current image is derived from an image height, an upper zoom window offset, and a lower zoom window offset of the current image. The image width, left zoom window offset, right zoom window offset, image height, upper zoom window offset, and lower zoom window offset of the current image are signaled in a PPS associated with the current image.

[0018] In one embodiment, the zoom window offset is measured in luma samples, the zoom window width of the current image is derived by subtracting the left and right zoom window offsets from the image width of the current image; and the zoom window height of the current image is derived by subtracting the top and bottom zoom window offsets from the image height of the current image. In another embodiment, the zoom window offset is measured in chroma samples, the zoom window width of the current image is derived from the image width, left and right zoom window offsets, and a variable SubWidthC, and the zoom window height of the current image is derived from the image height, top and bottom zoom window offsets, and a variable SubHeightC. These variables SubWidthC and SubHeightC indicate the downsampling ratios associated with the chroma bit planes in the horizontal and vertical dimensions. The zoom window width of the current image is derived by multiplying the variable SubWidthC by the sum of the left zoom window offset and the right zoom window offset, and then subtracting it from the image width of the current image; and the zoom window height of the current image is derived by multiplying the variable SubHeightC by the sum of the upper zoom window offset and the lower zoom window offset, and then subtracting it from the image height of the current image.

[0019] In one embodiment, a reference image scaling ratio for motion compensation is derived from the scaling window width, height, or size of the current image and the scaling window width, height, or size of the reference image; and the reference image scaling ratio is constrained to be in the range of [2048, 32768].

[0020] In one embodiment, for generating a bit stream of encoded data corresponding to a video sequence on an encoder side, or receiving a bit stream of encoded data corresponding to a video sequence on a decoder side, the following is a bit stream specification: twice the scaling window width of the current image is greater than or equal to the scaling window width of the reference image, twice the scaling window height of the current image is greater than or equal to the scaling window height of the reference image, the scaling window width of the current image is less than or equal to eight times the scaling window width of the reference image, and the scaling window height of the current image is less than or equal to eight times the scaling window height of the reference image.

[0021] Aspects of the present disclosure further provide a video processing device in a video encoding or decoding system, the device comprising one or more electronic circuits configured to: receive input video data of a current block in a current picture; determine a scaling window width, height, or size of the current picture; determine a scaling window width, height, or size of a reference picture; generate a reference block from the reference picture; perform motion compensation for the current block using the reference block; and encode or decode the current block in the current picture, wherein a ratio between the scaling window width, height, or size of the current picture and the scaling window width, height, or size of the reference picture is within a ratio constraint.

[0022] Aspects of the present disclosure further provide a non-transitory computer-readable medium for storing program instructions that cause a processing circuit of a device to perform a video processing method to encode or decode a current block in a current image. The video processing method determines a scaling window width, height, or size of the current image; determines a scaling window width, height, or size of a reference image; generates a reference block from the reference image; and encodes or decodes the current block based on the reference block. A ratio between the scaling window width, height, or size of the current image and the scaling window width, height, or size of the reference image is constrained within a ratio constraint. Other aspects and features of the present invention will become apparent to those skilled in the art from the following description of specific embodiments. BRIEF DESCRIPTION OF THE DRAWINGS

[0023] Various embodiments in which the present disclosure is presented as examples will be explained in more detail with reference to the following drawings, in which: Figure 1 A hypothetical example of enabling reference image resampling is shown.

[0024] Figure 2 An example of enabling reference picture resampling considering the scaling window size of each picture is shown.

[0025] Figure 3 An exemplary flow chart of a video encoding or decoding system is shown according to an embodiment of the present invention to check a scaling window ratio between a current picture and a reference picture.

[0026] Figure 4 An embodiment of a video processing method is shown as a flow chart to encode or decode a current block by enabling reference picture resampling in a video encoding or decoding system.

[0027] Figure 5An exemplary system block diagram is shown for a video encoding system embodying a video processing method according to an embodiment of the present invention.

[0028] Figure 6 An exemplary system block diagram is shown for a video decoding system embodying a video processing method according to an embodiment of the present invention. DETAILED DESCRIPTION

[0029] It will be readily understood that the components of the present invention, as generally described and illustrated in the figures herein, may be arranged and designed in a variety of different configurations. Accordingly, the following more detailed description of embodiments of the systems and methods of the present invention, as illustrated in the figures, is not intended to limit the scope of the claimed invention, but is merely representative of selected embodiments of the invention.

[0030] In VVC draft 6, a bitstream specification requirement applies to constrain the image size ratio between a reference image and a current image to be within [1 / 8, 2]. The image size ratio is derived from the width / height / size of a reference image and the width / height / size of a current image. The image size ratio constraint is specified within [1 / 8, 2] because the interpolation filter only supports scaling ratios between 1 / 8 and 2. Some embodiments of the present invention apply the [1 / 8, 2] ratio constraint to the scaling ratio between a reference scaling window width, height, or size and a current scaling window width, height, or size. The scaling ratio is calculated using the scaling window width, height, or size (rather than the image width, height, or size). Figure 2 An example of performing motion compensation by referring to two reference images with different image sizes and different scaling window sizes is shown. Figure 2 A current image 20 is shown having a zoom window 202. Although a first reference image 22 (reference image 0) is smaller than the current image 20, a zoom window 222 of the first reference image 22 is larger than the zoom window 202 of the current image. This indicates that a scaling factor of less than 1 is applied to downscale the zoom window 222 to be referenced by the current image. A second reference image 24 (reference image 1) is larger than the current image 20. However, a zoom window 242 of the second reference image 24 is smaller than the zoom window 202 of the current image. Therefore, a scaling factor of greater than 1 is applied to upscale the zoom window 242 to be referenced by the current image.

[0031] In one embodiment, a scaling window width PicOutputWidthL of a current picture is derived from a picture width pic_width_in_luma_samples, a left scaling window offset scaling_win_left_offset, and a right scaling window offset scaling_win_right_offset signaled in the PPS associated with the current picture, i.e., PicOutputWidthL = pic_width_in_luma_samples - ( scaling_win_right_offset + scaling_win_left_offset ); and a scaling window height PicOutputHeightL of the current picture is derived from a picture height pic_height_in_luma_samples, an upper scaling window offset scaling_win_top_offset, and a lower scaling window offset scaling_win_bottom_offset, i.e., PicOutputHeightL = pic_height_in_luma_samples - ( scaling_win_bottom_offset + scaling_win_top_offset ). When scaling_window_flag is equal to 1, it is assumed that refPicOutputWidthL and refPicOutputHiehgtL are a scaling window width and a scaling window height of a reference picture, respectively. A reference block in the reference picture is determined to be referenced by a current block of the current picture. For example, a video encoding system determines the reference block by motion estimation, and a video decoding system determines the reference block by analyzing motion information of the current block signaled in a video bitstream. When the ratio between the scaling window size of the reference picture and the scaling window size of the current picture is within the ratio constraint [1 / 8, 2], bitstream conformance requires that all of the following four conditions are met. Double the zoom window width of the current image is greater than or equal to the zoom window width of the reference image, double the zoom window height of the current image is greater than or equal to the zoom window height of the reference image, the zoom window width of the current image is less than or equal to eight times the zoom window width of the reference image, and the zoom window height of the current image is less than or equal to eight times the zoom window height of the reference image.That is to say, PicOutputWidthL * 2 ≥ refPicOutputWidthL, PicOutputHeightL * 2 ≥ refPicOutputHeightL, PicOutputWidthL ≤ refPicOutputWidthL * 8, andPicOutputHeightL ≤ refPicOutputHeightL * 8.

[0032] Generalizing the above embodiment and constraining the scaling window width and scaling window height of the current image based on the scaling window width and scaling window height of the reference image, the bitstream specification requires that all of the following conditions be met. N times the scaling window width of the current image is greater than or equal to the scaling window width of the reference image, N times the scaling window height of the current image is greater than or equal to the scaling window height of the reference image, the scaling window width of the current image is less than or equal to M times the scaling window width of the reference image, and the scaling window height of the current image is less than or equal to M times the scaling window height of the reference image. The ratio between the scaling window size of the current image and the scaling window size of the reference image is within a ratio constraint [1 / M, N], where N and M are positive integers. For example, in the previous embodiment, N is 2 and M is 8. PicOutputWidthL *N ≥ refPicOutputWidthL, PicOutputHeight * N ≥ refPicOutputHeight, PicOutputWidthL ≤ refPicOutputWidthL * M, and PicOutputHeightL ≤ refPicOutputHeightL * M.

[0033] In one embodiment, a scale constraint [1 / M, N] is determined to encode or decode a current picture. An encoder or decoder checks whether one or more reference pictures satisfy the scale constraint by determining a scaling window width, height, or size for the current picture and a scaling window width, height, or size for the reference picture. Only reference pictures having a scaling window width, height, or size that satisfies the scale constraint can be referenced by the current picture. Figure 3 FIG. 1 is a flow chart illustrating an example of this embodiment.

[0034] In some other embodiments, a ratio constraint [1 / M, N] is determined, and an encoder or decoder determines a scaling window width, height, or size of a current image based on a scaling window width, height, or size of a reference image to satisfy the ratio constraint. In one embodiment, the same ratio constraint may constrain both the scaling window ratio and the image size ratio, and the encoder or decoder also determines an image size of the current image based on an image size of the reference image to comply with the ratio constraint.

[0035] In another embodiment, the scaling window offset signaled in the PPS is measured in chroma samples. A scaling window width, PicOutputWidthL, for the current picture is derived from the picture width, pic_width_in_luma_samples, signaled in the PPS; a left scaling window offset, scaling_win_left_offset; and a right scaling window offset, scaling_win_right_offset; and a variable, SubWidthC. The value of the variable, SubWidthC, is defined based on the color sampling format of the video data; for example, when the color sampling format is 4:2:0, SubWidthC is equal to 2. PicOutputWidthL = pic_width_in_luma_samples – SubWidthC * (scaling_win_right_offset + scaling_win_left_offset). Similarly, the scaling window height (PicOutputHeightL) of the current image is derived from the image height (pic_height_in_luma_samples), the top scaling window offset (scaling_win_top_offset), the bottom scaling window offset (scaling_win_bottom_offset), and the variable (SubHeightC). The value of the variable (SubHeightC) is defined based on the color sampling format of the video data; for example, when the color sampling format is 4:2:0, SubHeightC is equal to 2. PicOutputHeightL = pic_height_in_luma_samples – SubHeightC * (scaling_win_bottom_offset + scaling_win_top_offset). The variables (SubWidthC and SubHeightC) indicate the downsampling ratios associated with the chroma bit planes in the horizontal and vertical dimensions, respectively. The variables (SubWidthC and SubHeightC) indicate the downsampling ratios associated with the chroma bit planes in the horizontal and vertical dimensions, respectively.

[0036] Assume that refPicOutputWidthL and refPicOutputHeightL are a scaling window width and a scaling window height of a reference picture referenced by a current block of the current picture, where refPicOutputWidthL and refPicOutputHeightL are derived from the picture width and height, the scaling window offset, and the variables SubWidthC and SubHeightC. Bitstream conformance requires that all four of the following conditions be met: twice the scaling window width of the current picture is greater than or equal to the scaling window width of the reference picture, twice the scaling window height of the current picture is greater than or equal to the scaling window height of the reference picture, the scaling window width of the current picture is less than or equal to eight times the scaling window width of the reference picture, and the scaling window height of the current picture is less than or equal to eight times the scaling window height of the reference picture. PicOutputWidthL * 2 ≥ refPicOutputWidthL, PicOutputHeightL * 2 ≥ refPicOutputHeightL, PicOutputWidthL ≤ refPicOutputWidthL * 8, PicOutputHeightL ≤ refPicOutputHeightL * 8.

[0037] A reference picture scaling ratio RefPicScale[i][j][0], RefPicScale[i][j][1] is derived from the scaling window size, width, or height specified in the PPS for motion compensation. This reference picture scaling ratio affects which filters are used in the motion compensation stage and also affects the memory bandwidth used for the motion compensation stage. In addition to constraining the picture size ratio, embodiments of the present invention also constrain the reference picture scaling ratio. For example, the reference picture scaling ratios RefPicScale[i][j][0] and RefPicScale[i][j][1] should be constrained to be in the range [2048, 32768], which is equivalent to a scaling ratio of [1 / 8, 2]. The bitstream specification requires that all of the following conditions be met: RefPicScale[i][j][0] should be greater than or equal to 2048 and should be less than or equal to 32768, and RefPicScale[i][j][1] should be greater than or equal to 2048 and should be less than or equal to 32768.

[0038] For example, depending on the scaling, three different interpolation filter sets can be selected for motion compensation. A first interpolation filter set (Set 0) includes an 8-tap DCT-IF filter, an affine 6-tap DCT-IF filter, and a 6-tap half-pixel IF filter; a second interpolation filter set (Set 1) includes an 8-tap RPR filter and a corresponding 6-tap affine filter (scaled 1.5 times); and a third interpolation filter set (Set 2) includes an 8-tap RPR filter and a corresponding 6-tap affine filter (scaled 2.0 times). To process a current block associated with a scaling between 1 / 8 and 1.25, filters in Set 0 are selected; to process a current block associated with a scaling between 1.25 and 1.75, filters in Set 1 are selected; and to process a current block associated with a scaling between 1.75 and 2, filters in Set 2 are selected.

[0039] Exemplary Flowchart Figure 3According to one embodiment of the present invention, an exemplary flow chart of a video encoding or decoding system is depicted for checking a scaling window ratio between a current picture and a reference picture. In step S302, the video encoding or decoding system receives input video data associated with a current picture and, in step S304, determines a scaling window width, height, or size for the current picture. For example, the scaling window size includes both a scaling window width and a scaling window height. In this embodiment, the scaling window width of the current picture is derived from a picture width, a left scaling window offset, and a right scaling window offset of the current picture, and the scaling window height of the current picture is derived from a picture height, an upper scaling window offset, and a lower scaling window offset of the current picture. Syntax elements associated with these scaling window offsets and the picture width and height are signaled in a PPS corresponding to the current picture. In step S306, a scaling window width, height, or size for a reference picture is determined. Similarly, the scaling window width of the reference picture is derived from a picture width, a left scaling window offset, and a right scaling window offset of the reference picture, and the scaling window height of the reference picture is derived from a picture height, a top scaling window offset, and a bottom scaling window offset of the reference picture. Syntax elements associated with these scaling window offsets and picture widths and heights of the reference picture are signaled in a PPS corresponding to the reference picture. In step S308, the video encoding or decoding system checks whether a ratio between the scaling window width, height, or size of the current picture and the scaling window width, height, or size of the reference picture is within a ratio constraint [1 / M, N]. For example, a ratio constraint of [1 / 8, 2] indicates that the scaling window width / height of the current picture is greater than or equal to the scaling window width / height of the reference picture when two times the scaling window width / height of the current picture is greater than or equal to the scaling window width / height of the reference picture, and when the scaling window width / height of the current picture is less than or equal to eight times the scaling window width / height of the reference picture. When the ratio is within the ratio constraint, in step S310, the reference image is included in a reference image list of one or more blocks in the current image, so that the reference image can be referenced by the blocks in the current image. In step S312, when the ratio is not within the ratio constraint, the reference image is excluded from the reference image list because it cannot be referenced by any block in the current image. In step S314, the video encoding or decoding system further encodes or decodes the current image.

[0040] Figure 4According to one embodiment of the present invention, an exemplary flow chart of a video encoding or decoding system is depicted for encoding or decoding a current block by enabling reference image resampling. In step S402, the video encoding or decoding system receives input video data of a current block in a current image. In step S404, a reference block in a reference image is determined for prediction or motion compensation of the current block. A ratio between a scaling window width, height, or size of the reference image and a scaling window width, height, or size of the current image is within a ratio constraint [1 / M, N]. In step S406, the video encoding or decoding system generates a reference block from a reference area in the reference image based on the ratio; and in step S408, the reference block is used to encode or decode the current block.

[0041] The proposed video processing methods for reference picture resampling can be implemented in a video encoder or decoder. For example, the proposed video processing methods can be implemented in an inter-frame prediction module of an encoder and / or an inter-frame prediction module of a decoder. Alternatively, any of the proposed methods can be implemented in one or a combination of inter-frame prediction modules of a decoder and / or in a circuit coupled to one or a combination of inter-frame prediction modules to provide information required by the inter-frame prediction module. Figure 5An exemplary system block diagram of a video encoder 500 implementing various embodiments of the present invention is shown. The intra-frame prediction module 510 provides an intra-frame predictor based on reconstructed video data of a current picture. The inter-frame prediction module 512 performs motion estimation (ME) and motion compensation (MC) to provide an inter-frame predictor based on video data from one or more other pictures. According to some embodiments of the present invention, to encode a current block in a current picture, a reference region in a valid reference picture is determined, and a scaling ratio between any valid reference picture and the current picture is within a scaling constraint [1 / M, N]. The reference block is generated from the reference region and used for motion compensation of the current block. The scaling constraint is defined based on an interpolation filter used for motion compensation; for example, the scaling constraint is between 1 / 8 and 2. In another embodiment, the intra-frame prediction module 510 determines a scaling window width, height, or size for the current picture based on the scaling constraint and the scaling window width, height, or size of one or more reference pictures of the current picture. A switch 514 selects one of the intra prediction module 510 or the inter prediction module 512 to provide the selected predictor to the addition module 516 to form a prediction error, also known as a prediction residual. The prediction residual of the current block is further processed by the transform module (T) 518, followed by the quantization module (Q) 520. The transformed and quantized residual signal is then encoded by the entropy encoder 532 to form a video bitstream. The video bitstream is then packaged together with the side information. The transformed and quantized residual signal of the current block is then processed by the inverse quantization module (IQ) 522 and the inverse transform module (IT) 524 to restore the prediction residual. Figure 5 As shown, the prediction residual is restored by adding back the selected predictor at reconstruction module (REC) 526 to generate reconstructed video data. The reconstructed video data can be stored in a reference picture buffer (Ref. Pict. Buffer) 530 and used for prediction of other pictures. The reconstructed video data restored from REC module 526 may have been subject to various impairments due to the encoding process. Therefore, before being stored in reference picture buffer 530, an in-loop processing filter 528 is applied to the reconstructed video data to further improve image quality.

[0042] Used to decode from Figure 5 The video encoder 500 generates a video bit stream corresponding to a video decoder 600. Figure 6As shown. The video bitstream is input to the video decoder 600 and decoded by the entropy decoder 610 to parse and restore the transformed and quantized residual signal and other system information. The decoding process of the decoder 600 is similar to the reconstruction loop of the encoder 500, except that the decoder 600 only requires motion-compensated prediction in an inter-frame prediction module 614. Each block is decoded by either the intra-frame prediction module 612 or the inter-frame prediction module 614. According to some embodiments of the present invention, to determine a current block in a current image, the inter-frame prediction module 614 determines a reference region in a reference image. A ratio between a scaling window width, height, or size of the reference image and a scaling window width, height, or size of the current image is within a ratio constraint [1 / M, N]. A reference block is then generated from the reference region based on the ratio, and the reference block is used by the inter-frame prediction module 614 to perform motion compensation for the current block. Based on the decoded mode information, a switch 616 selects an intra predictor from the intra prediction module 612 or an inter predictor from the inter prediction module 614. The transformed and quantized residual signal associated with each block is restored by an inverse quantization (IQ) module 620 and an inverse transformation (IT) module 622. The restored residual signal is reconstructed by adding back the predictor in a reconstruction REC module 618 to produce reconstructed video. The reconstructed video is further processed by an in-loop filter 624 to produce the final decoded video. If the currently decoded picture is a reference picture for a subsequent picture in decoding order, the reconstructed video of the currently decoded picture is also stored in a reference picture buffer (Ref. Pict. Buffer) 626.

[0043] Figure 5 and Figure 6The various components of the video encoder 500 and the video decoder 600 may be implemented by hardware components, one or more processors configured to execute program instructions stored in a memory, or a combination of hardware and processors. For example, a processor executes program instructions to control the reception of input data associated with a current image. The processor is equipped with a single or multiple processing cores. In some examples, the processor executes program instructions to perform functions in some components in the encoder 500 and the decoder 600, and the memory electrically coupled to the processor is used to store program instructions, information corresponding to the reconstructed image of the block and / or intermediate data in the encoding or decoding process. The memory in some embodiments includes a non-temporary computer-readable medium, such as a semiconductor or solid-state memory, a random access memory (RAM), a read-only memory (ROM), a hard disk, an optical disk, or other suitable storage medium. The memory may also be a combination of two or more of the non-temporary computer-readable media listed above. As Figure 5 and Figure 6 As shown, the encoder 500 and the decoder 600 may be implemented in the same electronic device, and thus various functional components of the encoder 500 and the decoder 600 may be shared or reused if implemented in the same electronic device.

[0044] Embodiments of the processing method performed in a video codec system can be implemented in circuitry integrated into a video compression chip, or in program code integrated into video compression software to perform the processing described above. For example, determining a current block within a current image can be implemented in program code executed on a computer processor, such as a digital signal processor (DSP), a microprocessor, or a field programmable gate array (FPGA). These processors can be configured to perform specific tasks according to the present invention by executing machine-readable software code or firmware defining specific methods embodied in the present invention.

[0045] References in this specification to "an embodiment," "some embodiments," or similar language mean that the specific features, structures, or characteristics described in conjunction with the embodiment may be included in at least one embodiment of the present invention. Therefore, the phrases "in an embodiment" or "in some embodiments" appearing in various places throughout this specification do not necessarily refer to the same embodiment, which may be implemented alone or in combination with one or more other embodiments. In addition, the described features, structures, or characteristics may be combined in any suitable manner in one or more embodiments. However, those skilled in the relevant art will recognize that the present invention may be practiced without one or more specific details or with other methods, components, etc. In other cases, well-known structures or operations are not shown or described in detail to avoid obscuring aspects of the present invention.

[0046] Without departing from the spirit or essential characteristics of the present invention, the present invention may be implemented in other specific forms. The examples described are considered to be illustrative and not restrictive in all aspects only. Therefore, the scope of the present invention is indicated by the appended claims rather than the preceding description. All changes within the meaning and scope of equivalents belonging to the claims are intended to be included within their scope.

[0047] The present invention may be embodied in other specific forms without departing from its spirit or essential characteristics. The described examples are to be considered in all respects as illustrative and not restrictive. The scope of the present invention is therefore indicated by the appended claims rather than by the foregoing description. All variations coming within the meaning and scope of the claims are intended to be included within their scope.

Claims

1. A video processing method, used in a video encoding or decoding system, comprising: receiving input video data for a current block in a current image; Determine the width, height, or size of the zoom window of the current image; determining a scaling window width, height, or size of a reference image, wherein a ratio between the scaling window width, height, or size of the reference image and the scaling window width, height, or size of the current image is within a ratio constraint, wherein the scaling window size includes both the scaling window width and the scaling window height, and the ratio between the scaling window size of the current image and the scaling window size of the reference image is within the ratio constraint when twice the scaling window width of the current image is greater than or equal to the scaling window width of the reference image, when twice the scaling window height of the current image is greater than or equal to the scaling window height of the reference image, when the scaling window width of the current image is less than or equal to eight times the scaling window width of the reference image, and when the scaling window height of the current image is less than or equal to eight times the scaling window height of the reference image; generating a reference block from the reference image according to the ratio between the scaling window width, height, or size of the reference image and the scaling window width, height, or size of the current image; Using the reference block to perform motion compensation for the current block; and The current block in the current image is encoded or decoded.

2. The video processing method according to claim 1, wherein: The ratio constraint is between 1 / 8 and 2.

3. The video processing method according to claim 1, wherein: The zoom window width of the current image is derived by the image width of the current image, the left zoom window offset and the right zoom window offset, and the zoom window height of the current image is derived by the image height of the current image, the upper zoom window offset and the lower zoom window offset.

4. The video processing method according to claim 3, wherein: The zoom window width of the current image is derived by subtracting the left zoom window offset and the right zoom window offset from the image width of the current image; and the zoom window height of the current image is derived by subtracting the upper zoom window offset and the lower zoom window offset from the image height of the current image.

5. The video processing method according to claim 3, wherein: The image width, left zoom window offset, right zoom window offset, image height, top zoom window offset, and bottom zoom window offset of the current image are signaled in an image parameter set associated with the current image.

6. The video processing method according to claim 3, wherein: The left scaling window offset, the right scaling window offset, the top scaling window offset, and the bottom scaling window offset are measured in chroma samples.

7. The video processing method according to claim 6, characterized in that: The scaling window width of the current image is further derived by the variable SubWidthC, and the scaling window height of the current image is further derived by the variable SubHeightC, where the variables SubWidthC and SubHeightC indicate the downsampling ratios associated with the chroma bit planes in the horizontal and vertical dimensions.

8. The video processing method according to claim 7, wherein: The zoom window width of the current image is derived by multiplying the variable SubWidthC by the sum of the left zoom window offset and the right zoom window offset, and then subtracting it from the image width of the current image; and the zoom window height of the current image is derived by multiplying the variable SubHeightC by the sum of the upper zoom window offset and the lower zoom window offset, and then subtracting it from the image height of the current image.

9. The video processing method according to claim 1, wherein: The reference image scaling ratio for motion compensation is derived from the scaling window width, height, or size of the current image and the scaling window width, height, or size of the reference image; and the reference image scaling ratio is constrained to be in the range of [2048, 32768].

10. The video processing method according to claim 1, wherein: Further including: A bitstream of encoded data corresponding to a video sequence is generated at an encoder side or received at a decoder side, wherein the bitstream complies with a bitstream specification that: twice the scaling window width of the current image is greater than or equal to the scaling window width of the reference image, twice the scaling window height of the current image is greater than or equal to the scaling window height of the reference image, the scaling window width of the current image is less than or equal to eight times the scaling window width of the reference image, and the scaling window height of the current image is less than or equal to eight times the scaling window height of the reference image.

11. A video data processing apparatus in a video encoding or decoding system, the apparatus comprising one or more electronic circuits configured to: receiving input video data for a current block in a current image; Determine the width, height, or size of the zoom window of the current image; determining a scaling window width, height, or size of a reference image, wherein a ratio between the scaling window width, height, or size of the reference image and the scaling window width, height, or size of the current image is within a ratio constraint, wherein the scaling window size includes both the scaling window width and the scaling window height, and the ratio between the scaling window size of the current image and the scaling window size of the reference image is within the ratio constraint when two times the scaling window width of the current image is greater than or equal to the scaling window width of the reference image, when two times the scaling window height of the current image is greater than or equal to the scaling window height of the reference image, when the scaling window width of the current image is less than or equal to eight times the scaling window width of the reference image, and when the scaling window height of the current image is less than or equal to eight times the scaling window height of the reference image; generating a reference block from the reference image according to the ratio between the scaling window width, height, or size of the reference image and the scaling window width, height, or size of the current image; Using the reference block to perform motion compensation for the current block; and The current block in the current image is encoded or decoded.

12. A non-transitory computer-readable medium for storing program instructions, the program instructions causing a processing circuit of a device to perform a video processing method for video data, the method comprising: receiving input video data for a current block in a current image; Determine the width, height, or size of the zoom window of the current image; determining a scaling window width, height, or size of a reference image, wherein a ratio between the scaling window width, height, or size of the reference image and the scaling window width, height, or size of the current image is within a ratio constraint, wherein the scaling window size includes both the scaling window width and the scaling window height, and the ratio between the scaling window size of the current image and the scaling window size of the reference image is within the ratio constraint when two times the scaling window width of the current image is greater than or equal to the scaling window width of the reference image, when two times the scaling window height of the current image is greater than or equal to the scaling window height of the reference image, when the scaling window width of the current image is less than or equal to eight times the scaling window width of the reference image, and when the scaling window height of the current image is less than or equal to eight times the scaling window height of the reference image; generating a reference block from the reference image according to the ratio between the scaling window width, height, or size of the reference image and the scaling window width, height, or size of the current image; Using the reference block to perform motion compensation for the current block; and The current block in the current image is encoded or decoded.