Video encoding or decoding method and device with zoom ratio constraint

By introducing scaling window offset and proportional constraints in the video encoding or decoding system, the problem of lax proportional constraints in the prior art is solved, and more efficient video processing and better video quality are achieved.

CN114788285BActive Publication Date: 2025-05-16HFI INNOVATION INC
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202080085721.0
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Priority Date
2019-12-18
Filing Date
2020-12-10
Publication Date
2025-05-16
Estimated Expiration
2040-12-10

AI Technical Summary

Technical Problem

In video encoding and decoding systems, it is difficult for the prior art to effectively constrain the range of image size ratios during reference image resampling (RPR), resulting in difficult to ensure video processing efficiency and quality.

Method used

By introducing scaling window offset and proportional constraints in the video encoding or decoding system, it is ensured that the scaling window width, height or size ratio between the current image and the reference image is within the range of [1/M, N], specifically, such as between 1/8 and 2.

Benefits of technology

It realizes more efficient video processing in video encoding and decoding systems, improves the efficiency of video transmission and storage, reduces the data rate requirements, and improves video quality.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114788285B_ABST
    Figure CN114788285B_ABST
Patent Text Reader

Abstract

A video processing method and apparatus, for processing a current block in a current image by reference picture resampling, comprises: receiving input video data of the current block, determining a scaling window of the current image and a scaling window of a reference image. The current image and the reference image may have different scaling window sizes. The ratio between the scaling window width, height, or size of the current image and the scaling window width, height, or size of the reference image is constrained within a ratio constraint. Based on the ratio, a reference block is generated from the reference image and used to encode or decode the current block.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] Cross-references

[0002] The present invention claims priority to U.S. Provisional Patent Applications No. 62 / 946,540, entitled “Method of Scaling Ratio Constraint,” filed on December 11, 2019, and No. 62 / 949,506, entitled “Method of Scaling Window Constraint,” filed on December 18, 2019, respectively. The U.S. Provisional Patent Applications are hereby incorporated by reference in their entirety. Technical Field

[0003] The present invention relates to a video processing method and apparatus in a video encoding and decoding system. In particular, the present invention relates to scaling constraints for reference picture resampling. Background Art

[0004] The Versatile Video Codec (VVC) standard is an upcoming emerging video codec standard that is gradually evolving from the previous High Efficiency Video Coding (HEVC) standard by enhancing existing codec tools and introducing multiple new codec tools in various building blocks of the codec. The VVC standard improves compression performance and transmission and storage efficiency, and supports new formats such as High Dynamic Range and omni-directional 360 video. The VVC standard makes video transmission in mobile networks more efficient because it allows systems or locations with poor data rates to receive larger files faster. VVC supports layer coding, spatial or temporal scalability of signal-to-noise ratio (SNR).

[0005] Reference picture resampling (RPR) In the VVC standard, fast presentation switching for adaptive streaming services is desirable to deliver multiple presentations of the same video content at the same time, each with different characteristics. The different characteristics involve different spatial resolutions or different sample bit depths. In real-time video communication, by allowing the resolution to be changed in the encoded video sequence without inserting I-pictures, not only can the video data be seamlessly adapted to dynamic channel conditions and user preferences, but the beating effect caused by I-pictures can also be removed. Reference picture resampling (RPR) allows pictures with different resolutions to reference each other for inter-frame prediction. Figure 1 An example of applying reference picture resampling to encode or decode a current picture is shown, where inter-coded blocks of the current picture are predicted from reference pictures of the same or different size. Spatial scalability is beneficial in streaming applications. When spatial scalability is supported, the picture size of the reference picture can be different from the current picture. RPR is used in the VVC standard to support on-the-fly up-sampling and down-sampling motion compensation.

[0006] Table 1 shows an example of signaling the RPR enable flag and the maximum picture size in a sequence parameter set (SPS). The RPR enable flag sps_ref_pic_resampling_enabled_flag signaled in a sequence parameter set (SPS) is used to indicate whether RPR is enabled in a picture that references the SPS. When this RPR enable flag is equal to 1, the current picture that references the SPS may have slices that reference a reference picture in the active entry of the reference picture layer, and the reference picture layer has one or more of the following seven parameters that are different from the current picture. The seven parameters include syntax elements associated with: picture width pps_pic_width_in_luma_samples, picture height pps_pic_height_in_luma_samples, left scaling window offset pps_scaling_win_left_offset, right scaling window offset pps_scaling_win_right_offset, top scaling window offset pps_scaling_win_top_offset, bottom scaling window offset pps_scaling_win_bottom_offset, number of sub-pictures sps_num_subpics_minus1. For a current picture that refers to a reference picture (the reference picture has one or more of the seven parameters different from the current picture), the reference picture may belong to the same layer as the layer containing the current picture or a different layer. The syntax element sps_res_change_in_clvs_allowed_flag equal to 1 indicates that the picture spatial resolution may change within the Coded Layer Video Sequence (CLVS) referencing the SPS, and this syntax element equal to 0 indicates that the picture spatial resolution does not change within any CLVS referencing the SPS. The maximum picture size is signaled in the SPS via the syntax elements sps_pic_width_max_in_luma_samples and sps_pic_height_max_in_luma_samples and must not be larger than the picture size of the Output Layer Set (OLS) Decoded Picture Buffer (DPB) signaled in the corresponding Video Parameter Set (VPS).

[0007] Table 1

[0008]

[0009] When RPR is used to predict a current picture, a picture size ratio is derived from the reference picture width or height and the current picture width or height. The picture size ratio is constrained to be in a range between 1 / 8 and 2. For example, the picture width and height measured in luma samples are derived by the syntax elements pic_width_in_luma_samples and pic_height_in_luma_samples signaled in a Picture Parameter Set (PPS). The syntax element pic_width_in_luma_samples indicates the width (in luma samples) of each decoded picture of the reference PPS. This syntax element must not be equal to 0 and should be an integer multiple of Max(8,MinCbSizeY) and is constrained to be less than or equal to pic_width_max_in_luma_samples. When the subpics_present_flag is equal to 1 or when the RPR enable flag ref_pic_resampling_enabled_flag is equal to 0, the value of this syntax element pic_width_in_luma_samples shall be equal to pic_width_max_in_luma_samples. The syntax element pic_height_in_luma_samples specifies the height of each decoded picture of the reference PPS (in units of luma samples). This syntax element shall not be equal to 0 and shall be an integer multiple of Max(8,MinCbSizeY) and shall be less than or equal to pic_height_max_in_luma_samples. When the subpics_present_flag is equal to 1 or when the RPR enable flag ref_pic_resampling_enabled_flag is equal to 0, the value of this syntax element pic_height_in_luma_samples is set equal to pic_height_max_in_luma_samples.

[0010] In the current design of RPR in VVC draft 6, when the image sizes of a current image and a reference image are specified, the following constraint must be satisfied. This constraint limits the image size ratio of the reference image to the current image to be in the range [1 / 8, 2]. Assume that the variables refPicWidthInLumaSamples and refPicHeightInLumaSamples are the image width and image height of a reference image referenced by a current image. The bitstream specification requires that all of the following conditions are met: the image width pic_width_in_luma_samples of the current image multiplied by two should be greater than or equal to the image width refPicWidthInLumaSamples of the reference image, the image height pic_height_in_luma_samples of the current image multiplied by two should be greater than or equal to the image height refPicHeightInLumaSamples of the reference image, the image width pic_width_in_luma_samples of the current image should be less than or equal to the image width refPicWidthInLumaSample of the reference image multiplied by eight, and the image height pic_height_in_luma_samples of the current image should be less than or equal to the image height refPicHeightInLumaSamples of the reference image multiplied by eight.

[0011] The picture size scaling ratio between a reference picture and a current picture is derived from the syntax elements pic_width_in_luma_samples and pic_height_in_luma_samples signaled in a PPS associated with the reference picture, and pic_width_in_luma_samples and pic_height_in_luma_samples signaled in a PPS associated with the current picture. The scaling window offset for the RPR is also derived from the syntax elements signaled in the PPS. Table 2 shows these syntax elements signaled in the PPS and the corresponding semantics.

[0012] Table 2

[0013]

[0014] The syntax element scaling_window_flag equal to 1 indicates that the scaling window offset parameter is present in the PPS, and scaling_window_flag equal to 0 indicates that the scaling window offset parameter is not present in the PPS. When an RPR enable flag ref_pic_resampling_enabled_flag is equal to 0, the value of this syntax element scaling_window_flag shall be equal to 0. The syntax elements scaling_win_left_offset, scaling_win_right_offset, scaling_win_top_offset, and scaling_win_bottom_offset specify the scaling offsets (in units of luma samples). These scaling offsets are applied to the picture size for scaling calculations. Scaling offsets can be negative. When a scaling window flag scaling_window_flag is equal to 0, the values ​​of the four scaling offset syntax elements scaling_win_left_offset, scaling_win_right_offset, scaling_win_top_offset, and scaling_win_bottom_offset are inferred to be equal to 0.

[0015] The sum of the left and right offsets scaling_win_left_offset and scaling_win_right_offset should be less than the image width pic_width_in_luma_samples, and the sum of the top and bottom offsets scaling_win_top_offset and scaling_win_bottom_offset should be less than the image height pic_height_in_luma_samples. A variable PicOutputWidthL representing a scaling window width is derived by subtracting the left and right offsets from the image width. PicOutputWidthL = pic_width_in_luma_samples - (scaling_win_right_offset + scaling_win_left_offset). A variable PicOutputHeightL representing a scaling window height is derived by subtracting the top and bottom offsets from the image height. PicOutputHeightL=pic_height_in_luma_samples–(scaling_win_bottom_offset+scaling_win_top_offset).

[0016] A variable fRefWidth is set equal to PicOutputWidthL of luma samples in a reference picture RefPicList[i][j], and a variable fRefHight is set equal to PicOutputHeightL of luma samples in the reference picture RefPicList[i][j]. A derived reference picture scale for the horizontal direction RefPicScale[i][j][0] is calculated by ((fRefWidth<<14)+(PicOutputWidthL>>1)) / PicOutputWidthL, and a derived reference picture scale for the vertical direction RefPicScale[i][j][1] is calculated by ((fRefHeight<<14)+(PicOutputHeightL>>1)) / PicOutputHeightL. Therefore, the derived reference image scaling factor is RefPicIsScaled[i][j]=(RefPicScale[i][j][0]!=(1<<14))||(RefPicScale[i][j][1]!=(1<<14).

[0017] In a recent proposal of the VVC standard, the scaling window offsets are measured in chroma samples, and when the scaling window offset syntax elements are not present in the PPS, the values ​​of the four scaling offset syntax elements scaling_win_left_offset, scaling_win_right_offset, scaling_win_top_offset, and scale_win_bottom_offset are inferred to be equal to conf_win_left_offset, conf_win_right_offset, conf_win_top_offset, and conf_win_bottom_offset, respectively. A variable CurrPicScalWinWidthL indicating the scaling window width is derived from the picture width, SubWidthC, the left scaling offset, and the right scaling offset, and a variable CurrPicScalWinHeightL indicating the scaling window height is derived from the picture height, SubHeightC, the top scaling offset, and the bottom scaling offset, as shown below. CurrPicScalWinWidthL=pic_width_in_luma_samples–SubWidthC*(scaling_win_right_offset+scaling_win_left_offset); and CurrPicScalWinHeightL=pic_height_in_luma_samples–SubHeightC*(scaling_win_bottom_offset+scaling_win_top_offset). Summary of the invention

[0018] In an exemplary embodiment of a video processing method for processing a current block in a current image, a video encoding or decoding system implementing the video processing method: receives input video data associated with the current block; determines a scaling window width, height, or size of the current image; determines a scaling window width, height, or size of a reference image; generates a reference block by a ratio between the scaling window width, height, or size of the current image and the scaling window width, height, or size of the reference image; uses the reference block to perform motion compensation for the current block; and encodes or decodes the current block in the current image. A ratio between the scaling window width, height, or size of the current image and the scaling window width, height, or size of the reference image is constrained within a ratio constraint.

[0019] In some exemplary embodiments, the ratio constraint is between 1 / M and N, where M and N are positive integers. To make the ratio of the scaling window width of the current image to the scaling window width of the reference block between the ratio constraint, N times the scaling window width of the current image is greater than or equal to the scaling window width of the reference image, and the scaling window width of the current image is less than or equal to M times the scaling window width of the reference image. To make the ratio of the scaling window height of the current image to the scaling window height of the reference image between the ratio constraint, N times the scaling window height of the current image is greater than or equal to the scaling window height of the reference image, and the scaling window height of the current image is less than or equal to M times the scaling window height of the reference image. In one embodiment, the scaling window size includes the scaling window width and the scaling window height. To make the ratio of the scaling window size of the current image to the scaling window size of the reference image within the ratio constraint, N times the scaling window width of the current image is greater than or equal to the scaling window width of the reference image, N times the scaling window height of the current image is greater than or equal to the scaling window height of the reference image, the scaling window width of the current image is less than or equal to M times the scaling window height of the reference image, and the scaling window height of the current image is less than or equal to M times the scaling window height of the reference image. For example, the ratio constraint is between 1 / 8 and 2. When the scaling window size of the current image is smaller than the scaling window size of the reference image, twice the scaling window width of the current image is greater than or equal to the scaling window width of the reference image, and twice the scaling window height of the current image is greater than or equal to the scaling window height of the reference image. When the zoom window size of the current image is larger than the zoom window size of the reference image, the zoom window width of the current image is less than or equal to eight times the zoom window width of the reference image, and the zoom window height of the current image is less than or equal to eight times the zoom window height of the reference image.

[0020] In some embodiments, the zoom window width of the current image is derived by an image width, a left zoom window offset, and a right zoom window offset of the current image, and the zoom window height of the current image is derived by an image height, an upper zoom window offset, and a lower zoom window offset of the current image. The image width, left zoom window offset, right zoom window offset, image height, upper zoom window offset, and lower zoom window offset of the current image are signaled in a PPS associated with the current image.

[0021] In one embodiment, the scaling window offset is measured in luma samples, the scaling window width of the current image is derived by subtracting the left scaling window offset and the right scaling window offset from the image width of the current image; and the scaling window height of the current image is derived by subtracting the top scaling window offset and the bottom scaling window offset from the image height of the current image. In another embodiment, the scaling window offset is measured in chroma samples, the scaling window width of the current image is derived by the image width, left and right scaling window offsets, and a variable SubWidthC, and the scaling window height of the current image is derived by the image height, top and bottom scaling window offsets, and a variable SubHeightC. The variables SubWidthC and SubHeightC indicate the downsampling ratios associated with the chroma bit planes in the horizontal and vertical dimensions. The zoom window width of the current image is derived by multiplying the variable SubWidthC by the sum of the left zoom window offset and the right zoom window offset, and then subtracting it from the image width of the current image; and the zoom window height of the current image is derived by multiplying the variable SubHeightC by the sum of the upper zoom window offset and the lower zoom window offset, and then subtracting it from the image height of the current image.

[0022] In one embodiment, a reference image scaling ratio for motion compensation is derived from: the scaling window width, height, or size of the current image, and the scaling window width, height, or size of the reference image; and the reference image scaling ratio is constrained to be within a range of [2048, 32768].

[0023] In one embodiment, for generating a bit stream of encoded data corresponding to a video sequence at an encoder side, or receiving a bit stream of encoded data corresponding to a video sequence at a decoder side, the following is a bit stream specification: twice the scaling window width of the current image is greater than or equal to the scaling window width of the reference image, twice the scaling window height of the current image is greater than or equal to the scaling window height of the reference image, the scaling window width of the current image is less than or equal to eight times the scaling window width of the reference image, and the scaling window height of the current image is less than or equal to eight times the scaling window height of the reference image.

[0024] The present disclosure further provides a video processing device in a video encoding or decoding system, the device comprising one or more electronic circuits configured to: receive input video data of a current block in a current image; determine a scaling window width, height, or size of the current image; determine a scaling window width, height, or size of a reference image; generate a reference block from the reference image; use the reference block to perform motion compensation for the current block; and encode or decode the current block in the current image. A ratio between the scaling window width, height, or size of the current image and the scaling window width, height, or size of the reference image is within a ratio constraint.

[0025] Aspects of the present disclosure further provide a non-transitory computer-readable medium for storing program instructions that cause a processing circuit of a device to perform a video processing method to encode or decode a current block in a current image. The video processing method determines a scaling window width, height, or size of the current image; determines a scaling window width, height, or size of a reference image; generates a reference block from the reference image; and encodes or decodes the current block based on the reference block. A ratio between the scaling window width, height, or size of the current image and the scaling window width, height, or size of the reference image is constrained within a ratio constraint. Other aspects and features of the present invention will become apparent to those of ordinary skill in the art through the following description of specific embodiments. BRIEF DESCRIPTION OF THE DRAWINGS

[0026] Various embodiments in which examples are presented in the present disclosure will be explained in more detail with reference to the following drawings, in which:

[0027] Figure 1 A hypothetical example of enabling reference image resampling is shown.

[0028] Figure 2 An example of enabling reference image resampling taking into account the scaling window size of each image is shown.

[0029] Figure 3 An exemplary flow chart of a video encoding or decoding system is shown according to an embodiment of the present invention to check a scaling window ratio between a current picture and a reference picture.

[0030] Figure 4 An embodiment of a video processing method is shown as a flow chart to encode or decode a current block by enabling reference picture resampling in a video encoding or decoding system.

[0031] Figure 5An exemplary system block diagram is shown for a video encoding system embodying a video processing method according to an embodiment of the present invention.

[0032] Figure 6 An exemplary system block diagram is shown for a video decoding system embodying a video processing method according to an embodiment of the present invention. DETAILED DESCRIPTION

[0033] It is readily understood that the components of the present invention as generally described and illustrated in the drawings herein may be arranged and designed in a variety of different configurations. Therefore, as shown in the drawings, the following more detailed description of the embodiments of the systems and methods of the present invention is not intended to limit the scope of the claimed invention, but merely represents selected embodiments of the present invention.

[0034] Constraining reference image scaling In VVC draft 6, a bitstream specification requirement is applied to constrain an image size ratio of a reference image to a current image to be within [1 / 8, 2]. The image size ratio is derived from the width / height / size of a reference image and the width / height / size of a current image. The image size ratio constraint is specified to be within [1 / 8, 2] because the interpolation filter only supports scaling ratios between 1 / 8 and 2. Some embodiments of the present invention apply a [1 / 8, 2] ratio constraint to a scaling ratio between a current scaling window width, height, or size and a reference scaling window width, height, or size. The scaling ratio is calculated by the width, height, or size of the scaling window (not the width, height, or size of the image). Figure 2 An example of performing motion compensation by referring to two reference images with different image sizes and different scaling window sizes is shown. Figure 2 A current image 20 is shown having a zoom window 202, and although a first reference image 22 (reference image 0) is smaller than the current image 20, a zoom window 222 of the first reference image 22 is larger than the zoom window 202 of the current image, which means that a zoom ratio of less than 1 is applied to downscale the zoom window 222 to be referenced by the current image. A second reference image 24 (reference image 1) is larger than the current image 20, however, a zoom window 242 of the second reference image 24 is smaller than the zoom window 202 of the current image, so a zoom ratio of greater than 1 is applied to upscale the zoom window 242 to be referenced by the current image.

[0035] In one embodiment, a scaling window width PicOutputWidthL of a current picture is derived by a picture width pic_width_in_luma_samples, a left scaling window offset scaling_win_left_offset, and a right scaling window offset scaling_win_right_offset signaled in the PPS associated with the current picture, i.e., PicOutputWidthL=pic_width_in_luma_samples−(scaling_win_right_offset+scaling_win_left_offset). t_offset); and a scaling window height PicOutputHeightL of the current picture is derived by a picture height pic_height_in_luma_samples, an upper scaling window offset scaling_win_top_offset, and a lower scaling window offset scaling_win_bottom_offset, i.e. PicOutputHeightL = pic_height_in_luma_samples - (scaling_win_bottom_offset + scaling_win_top_offset). When scaling_window_flag is equal to 1, it is assumed that refPicOutputWidthL and refPicOutputHiehgtL are a scaling window width of a reference picture and a scaling window height of the reference picture, respectively. A reference block in the reference picture is determined to be referenced by a current block of the current picture. For example, a video encoding system determines the reference block by motion estimation, and a video decoding system determines the reference block by analyzing motion information of the current block signaled in a video bitstream. When the ratio between the scaling window size of the current image and the scaling window size of the reference image is within the ratio constraint [1 / 8, 2], the bitstream conformance requires that all four of the following conditions are met: twice the scaling window width of the current image is greater than or equal to the scaling window width of the reference image, twice the scaling window height of the current image is greater than or equal to the scaling window height of the reference image, the scaling window width of the current image is less than or equal to eight times the scaling window width of the reference image, and the scaling window height of the current image is less than or equal to eight times the scaling window height of the reference image.That is, PicOutputWidthL*2≥refPicOutputWidthL, PicOutputHeightL*2≥refPicOutputHeightL, PicOutputWidthL≤refPicOutputWidthL*8, and PicOutputHeightL≤refPicOutputHeightL*8.

[0036] Generalizing the above embodiment and constraining the scaling window width and scaling window height of the current image based on the scaling window width and scaling window height of the reference image, the bitstream specification requires that all of the following conditions are met. N times the scaling window width of the current image is greater than or equal to the scaling window width of the reference image, N times the scaling window height of the current image is greater than or equal to the scaling window height of the reference image, the scaling window width of the current image is less than or equal to M times the scaling window width of the reference image, and the scaling window height of the current image is less than or equal to M times the scaling window height of the reference image. The ratio between the scaling window size of the current image and the scaling window size of the reference image is within a ratio constraint [1 / M, N], where N and M are positive integers. For example, in the previous embodiment, N is 2 and M is 8. PicOutputWidthL*N≥refPicOutputWidthL, PicOutputHeight*N≥refPicOutputHeight, PicOutputWidthL≤refPicOutputWidthL*M, and PicOutputHeightL≤refPicOutputHeightL*M.

[0037] In one embodiment, a ratio constraint [1 / M, N] is determined to encode or decode a current image, and an encoder or decoder checks whether one or more reference images satisfy the ratio constraint by determining a scaling window width, height, or size of the current image and a scaling window width, height, or size of the reference image. Only reference images having a scaling window width, height, or size that satisfies the ratio constraint can be referenced by the current image. Figure 3 is a flow chart illustrating an example of this embodiment.

[0038] In some other embodiments, a ratio constraint [1 / M, N] is determined, and an encoder or decoder determines a scaling window width, height, or size of a current image based on a scaling window width, height, or size of a reference image so as to satisfy the ratio constraint. In one embodiment, the same ratio constraint can constrain both the scaling window ratio and the image size ratio, and the encoder or decoder also determines an image size of the current image based on an image size of the reference image so as to comply with the ratio constraint.

[0039] In another embodiment, the scaling window offset signaled in the PPS is measured in chroma samples, and a scaling window width PicOutputWidthL of a current picture is derived by a picture width pic_width_in_luma_samples, a left scaling window offset scaling_win_left_offset, and a right scaling window offset scaling_win_right_offset signaled in the PPS, and a variable SubWidthC. The value of the variable SubWidthC is defined according to the color sampling format of the video data; for example, when the color sampling format is 4:2:0, SubWidthC is equal to 2. PicOutputWidthL = pic_width_in_luma_samples - SubWidthC*(scaling_win_right_offset + scaling_win_left_offset). Similarly, a scaling window height PicOutputHeightL of the current image is derived from an image height pic_height_in_luma_samples, an upper scaling window offset scaling_win_top_offset, a lower scaling window offset scaling_win_bottom_offset, and a variable SubHeightC. The value of the variable SubHeightC is defined according to the color sampling format of the video data; for example, when the color sampling format is 4:2:0, SubHeightC is equal to 2. PicOutputHeightL = pic_height_in_luma_samples - SubHeightC*(scaling_win_bottom_offset + scaling_win_top_offset). The variables SubWidthC and SubHeightC indicate the downsampling ratios associated with the chroma bit planes in the horizontal and vertical directions, respectively. The variables SubWidthC and SubHeightC indicate the downsampling ratios associated with the chroma bit planes in the horizontal and vertical dimensions, respectively.

[0040] Assume that refPicOutputWidthL and refPicOutputHeightL are a scaling window width and a scaling window height of a reference picture to which a current block of the current picture refers, where refPicOutputWidthL and refPicOutputHeightL are derived from the picture width and height, the scaling window offset, and the variables SubWidthC and SubHeightC. Bitstream conformance requires that all four of the following conditions be satisfied: twice the scaling window width of the current picture is greater than or equal to the scaling window width of the reference picture, twice the scaling window height of the current picture is greater than or equal to the scaling window height of the reference picture, the scaling window width of the current picture is less than or equal to eight times the scaling window width of the reference picture, and the scaling window height of the current picture is less than or equal to eight times the scaling window height of the reference picture. PicOutputWidthL*2≥refPicOutputWidthL,PicOutputHeightL*2≥refPicOutputHeightL,PicOutputWidthL≤refPicOutputWidthL*8,PicOutputHeightL≤refPicOutputHeightL*8.

[0041] A reference image scaling ratio RefPicScale[i][j][0], RefPicScale[i][j][1] is derived from the scaling window size, width, or height specified in the PPS for motion compensation. This reference image scaling ratio affects which filters are used in the motion compensation stage and also affects the memory bandwidth used for the motion compensation stage. In addition to constraining the image size ratio, embodiments of the present invention also constrain the reference image scaling ratio. For example, the reference image scaling ratios RefPicScale[i][j][0] and RefPicScale[i][j][1] should be constrained to be within the range of [2048,32768], equivalent to a scaling ratio of [1 / 8,2]. The bitstream specification requires that all of the following conditions be met: RefPicScale[i][j][0] should be greater than or equal to 2048 and should be less than or equal to 32768, and RefPicScale[i][j][1] should be greater than or equal to 2048 and should be less than or equal to 32768.

[0042] For example, depending on the scaling, three different interpolation filter groups may be selected in motion compensation. A first interpolation filter group (group 0) includes an 8-tap DCT-IF filter, an affine 6-tap DCT-IF filter, and a 6-tap half-pixel IF filter, and a second interpolation filter group (group 1) includes an 8-tap RPR filter, and a corresponding 6-tap affine filter (scaled 1.5 times), and a third interpolation filter group (group 2) includes an 8-tap RPR filter and a corresponding 6-tap affine filter (scaled 2.0 times). To process a current block associated with a scaling between 1 / 8 and 1.25, filters in group 0 are selected, to process a current block associated with a scaling between 1.25 and 1.75, filters in group 1 are selected, and to process a current block associated with a scaling between 1.75 and 2, filters in group 2 are selected.

[0043] Exemplary Flowchart Figure 3According to an embodiment of the present invention, an exemplary flow chart of a video encoding or decoding system is depicted to check a scaling window ratio between a current image and a reference image. In step S302, the video encoding or decoding system receives input video data associated with a current image, and in step S304, determines the scaling window width, height, or size of the current image. For example, the scaling window size includes both scaling window width and scaling window height. In this embodiment, the scaling window width of the current image is derived by a picture width, a left scaling window offset, and a right scaling window offset of the current image, and the scaling window height of the current image is derived by a picture height, an upper scaling window offset, and a lower scaling window offset of the current image. Syntax elements associated with these scaling window offsets and picture width and height are signaled in a PPS corresponding to the current image. In step S306, a scaling window width, height, or size of a reference image is determined. Similarly, the scaling window width of the reference image is derived by a picture width, a left scaling window offset and a right scaling window offset of the reference image, and the scaling window height of the reference image is derived by a picture height, an upper scaling window offset and a lower scaling window offset of the reference image. Syntax elements associated with these scaling window offsets and picture widths and heights of the reference image are signaled in a PPS corresponding to the reference image. In step S308, the video encoding or decoding system checks whether a ratio between the scaling window width, height, or size of the current image and the scaling window width, height, or size of the reference image is within a ratio constraint [1 / M, N]. For example, the ratio constraint of [1 / 8, 2] indicates that when twice the scaling window width / height of the current image is greater than or equal to the scaling window width / height of the reference image, and when the scaling window width / height of the current image is less than or equal to eight times the scaling window width / height of the reference image. When the ratio is within the ratio constraint, in step S310, the reference image is included in a reference image list of one or more blocks in the current image, so that the reference image can be referenced by the blocks in the current image. In step S312, when the ratio is not within the ratio constraint, the reference image is excluded from a reference image list because the reference image cannot be referenced by any block in the current image. In step S314, the video encoding or decoding system further encodes or decodes the current image.

[0044] Figure 4According to an embodiment of the present invention, an exemplary flow chart of a video encoding or decoding system is depicted to encode or decode a current block by enabling reference image resampling. In step S402, the video encoding or decoding system receives input video data of a current block in a current image. In step S404, a reference block in a reference image is determined for prediction or motion compensation of the current block. A ratio between a scaling window width, height, or size of the current image and a scaling window width, height, or size of the reference image is within a ratio constraint [1 / M, N]. In step S406, the video encoding or decoding system generates a reference block from a reference area in the reference image based on the ratio; and in step S408, the reference block is used to encode or decode the current block.

[0045] Implementation of Video Encoders and Decoders The above-mentioned proposed video processing methods for reference picture resampling can be implemented in a video encoder or a decoder. For example, the proposed video processing methods can be implemented on an inter-frame prediction module of an encoder and / or an inter-frame prediction module of a decoder. Alternatively, any of the proposed methods can be implemented on one or a combination of inter-frame prediction modules of a decoder and / or a circuit coupled to one or a combination of inter-frame prediction modules to provide the information required by the inter-frame prediction module. Figure 5An exemplary system block diagram of a video encoder 500 implementing various embodiments of the present invention is shown. An intra prediction module 510 provides an intra predictor based on reconstructed video data of a current picture. An inter prediction module 512 performs motion estimation (ME) and motion compensation (MC) to provide an inter predictor based on video data from one or more other pictures. According to some embodiments of the present invention, to encode a current block in a current picture, a reference region in an active reference picture is determined, and a scaling ratio between the current picture and any active reference picture is between a scaling constraint [1 / M, N]. The reference block is generated from the reference region and is used for motion compensation of the current block. The scaling constraint is defined based on an interpolation filter used for motion compensation, for example, the scaling constraint is between 1 / 8 and 2. In another embodiment, the intra prediction module 510 determines a scaling window width, height, or size of the current picture based on the scaling constraint and a scaling window width, height, or size of one or more reference pictures of the current picture. A switch 514 selects one of the intra-frame prediction module 510 or the inter-frame prediction module 512 to provide the selected predictor to the addition module 516 to form a prediction error, also known as a prediction residual. The prediction residual of the current block is further processed by the transformation module (T) 518, followed by the quantization module (Q) 520. The transformed and quantized residual signal is then encoded by the entropy encoder 532 to form a video bitstream. The video bitstream is then packaged together with side information. The transformed and quantized residual signal of the current block is then processed by the inverse quantization module (IQ) 522 and the inverse transformation module (IT) 524 to restore the prediction residual. Figure 5 As shown, the prediction residual is restored by adding back the selected predictor at the reconstruction module (REC) 526 to generate reconstructed video data. The reconstructed video data can be stored in a reference picture buffer (Ref.Pict.Buffer) 530 and used for prediction of other pictures. The reconstructed video data restored from the REC module 526 may be subjected to various impairments due to the encoding process, so before being stored in the reference picture buffer 530, an in-loop processing filter (In-loop Processing Filter) 528 is applied to the reconstructed video data to further improve the image quality.

[0046] Used to decode from Figure 5 The video encoder 500 generates a video bit stream corresponding to a video decoder 600. Figure 6As shown. The video bitstream is the input of the video decoder 600 and is decoded by the entropy decoder 610 to parse and restore the transformed and quantized residual signal and other system information. The decoding process of the decoder 600 is similar to the reconstruction loop at the encoder 500, except that the decoder 600 only requires motion compensation prediction in an inter-frame prediction module 614. Each block is decoded by the intra-frame prediction module 612 or the inter-frame prediction module 614. According to some embodiments of the present invention, a current block in a current image is determined, and the inter-frame prediction module 614 determines a reference area in a reference image. A ratio between a scaling window width, height, or size of the current image and a scaling window width, height, or size of the reference image is within a ratio constraint [1 / M, N]. A reference block is then generated from the reference area based on the ratio, and the reference block is used to perform motion compensation for the current block in the inter-frame prediction module 614. Based on the decoded mode information, a switch 616 selects an intra predictor from the intra prediction module 612 or an inter predictor from the inter prediction module 614. The transformed and quantized residual signal associated with each block is restored by an inverse quantization (IQ) module 620 and an inverse transformation (IT) module 622. The restored residual signal is reconstructed by adding back the predictor in a reconstruction REC module 618 to produce a reconstructed video. The reconstructed video is further processed by an in-loop filter 624 to produce the final decoded video. If the current decoded picture is a reference picture for a subsequent picture in decoding order, the reconstructed video of the current decoded picture is also stored in a reference picture buffer (Ref.Pict.Buffer) 626.

[0047] Figure 5 and Figure 6The various components of the video encoder 500 and the video decoder 600 in the embodiment may be implemented by hardware components, one or more processors configured to execute program instructions stored in a memory, or a combination of hardware and processors. For example, a processor executes program instructions to control the reception of input data associated with a current image. The processor is equipped with a single or multiple processing cores. In some examples, the processor executes program instructions to perform functions in some components in the encoder 500 and the decoder 600, and the memory electrically coupled to the processor is used to store program instructions, information corresponding to the reconstructed image of the block and / or intermediate data in the encoding or decoding process. The memory in some embodiments includes a non-temporary computer-readable medium, such as a semiconductor or solid-state memory, a random access memory (RAM), a read-only memory (ROM), a hard disk, an optical disk, or other suitable storage medium. The memory may also be a combination of two or more of the non-temporary computer-readable media listed above. As Figure 5 and Figure 6 As shown, the encoder 500 and the decoder 600 may be implemented in the same electronic device, so if implemented in the same electronic device, various functional components of the encoder 500 and the decoder 600 may be shared or reused.

[0048] Embodiments of the processing method in a video codec system may be implemented in a circuit integrated in a video compression chip, or in a program code integrated in video compression software to perform the processing described above. For example, determining a current block in a current image may be implemented in program code executed on a computer processor, a digital signal processor (DSP), a microprocessor, or a field programmable gate array (FPGA). These processors may be configured to perform specific tasks according to the present invention by executing machine-readable software code or firmware code that defines specific methods embodied in the present invention.

[0049] References in this specification to "an embodiment", "some embodiments" or similar language mean that the specific features, structures or characteristics described in conjunction with the embodiment may be included in at least one embodiment of the present invention. Therefore, the phrases "in an embodiment" or "in some embodiments" appearing in various places throughout this specification do not necessarily refer to the same embodiment, which may be implemented alone or in combination with one or more other embodiments. In addition, the described features, structures or characteristics may be combined in any suitable manner in one or more embodiments. However, those skilled in the art in the relevant art will recognize that the present invention may be practiced without one or more specific details or using other methods, components, etc. In other cases, well-known structures or operations are not shown or described in detail to avoid obscuring various aspects of the present invention.

[0050] Without departing from the spirit or essential features of the present invention, the present invention may be implemented in other specific forms. The described examples are considered to be illustrative and not restrictive in all aspects only. Therefore, the scope of the present invention is indicated by the appended claims rather than the previous description. All changes within the meaning and scope of the equivalents belonging to the claims will be included within their scope.

[0051] The present invention may be implemented in other specific forms without departing from its spirit or essential characteristics. The examples described are illustrative in all respects and not restrictive. Therefore, the scope of the present invention is indicated by the appended claims rather than the foregoing description. All changes within the meaning of the claims and the same scope should be included within their scope.

Claims

1. A video processing method, used in a video encoding or decoding system, the method comprising: receiving input video data for a current block in a current image; Determine the current image's zoom window width, height, or size; determining a scaling window width, height, or size of a reference image, wherein a ratio between the scaling window width, height, or size of the current image and the scaling window width, height, or size of the reference image is within a ratio constraint, wherein the ratio constraint is between 1 / 8 and 2; generating a reference block from the reference image according to the ratio between the scaling window width, height, or size of the current image and the scaling window width, height, or size of the reference image; Using the reference block to perform motion compensation for the current block; and Encode or decode the current block in the current image.

2. The video processing method according to claim 1, characterized in that: The scaling window size includes: the scaling window width and the scaling window height, and when twice the scaling window width of the current image is greater than or equal to the scaling window width of the reference image, when twice the scaling window height of the current image is greater than or equal to the scaling window height of the reference image, when the scaling window width of the current image is less than or equal to eight times the scaling window width of the reference image, and when the scaling window height of the current image is less than or equal to eight times the scaling window height of the reference image, the ratio between the scaling window size of the current image and the scaling window size of the reference image is within the ratio constraint.

3. The video processing method according to claim 1, characterized in that: The zoom window width of the current image is derived by the image width of the current image, the left zoom window offset and the right zoom window offset, and the zoom window height of the current image is derived by the image height of the current image, the upper zoom window offset and the lower zoom window offset.

4. The video processing method according to claim 3, characterized in that: The zoom window width of the current image is derived by subtracting the left zoom window offset and the right zoom window offset from the image width of the current image; and the zoom window height of the current image is derived by subtracting the upper zoom window offset and the lower zoom window offset from the image height of the current image.

5. The video processing method according to claim 3, characterized in that: The image width, left zoom window offset, right zoom window offset, image height, top zoom window offset, and bottom zoom window offset of the current image are signaled in an image parameter set associated with the current image.

6. The video processing method according to claim 3, characterized in that: The left scaling window offset, the right scaling window offset, the top scaling window offset, and the bottom scaling window offset are measured in chroma samples.

7. The video processing method according to claim 6, characterized in that: The scaling window width of the current image is further derived by a variable SubWidthC, and the scaling window height of the current image is further derived by a variable SubHeightC, wherein the variables SubWidthC and SubHeightC indicate the downsampling ratios associated with the chroma bit planes in the horizontal and vertical dimensions.

8. The video processing method according to claim 7, characterized in that: The zoom window width of the current image is derived by multiplying the variable SubWidthC by the sum of the left zoom window offset and the right zoom window offset, and then subtracting it from the image width of the current image; and the zoom window height of the current image is derived by multiplying the variable SubHeightC by the sum of the upper zoom window offset and the lower zoom window offset, and then subtracting it from the image height of the current image.

9. The video processing method according to claim 1, characterized in that: The reference image scaling ratio for motion compensation is derived from: the scaling window width, height, or size of the current image, and the scaling window width, height, or size of the reference image; and the reference image scaling ratio is constrained to be in the range of [2048,32768].

10. The video processing method according to claim 1, characterized in that: Further including: A bit stream of encoded data corresponding to a video sequence is generated at an encoder side or received at a decoder side, wherein the bit stream complies with a bit stream specification that: twice the scaling window width of the current image is greater than or equal to the scaling window width of the reference image, twice the scaling window height of the current image is greater than or equal to the scaling window height of the reference image, the scaling window width of the current image is less than or equal to eight times the scaling window width of the reference image, and the scaling window height of the current image is less than or equal to eight times the scaling window height of the reference image.

11. A video data processing device in a video encoding or decoding system, the device comprising one or more electronic circuits configured to: receiving input video data for a current block in a current image; Determine the current image's zoom window width, height, or size; determining a scaling window width, height, or size of a reference image, wherein a ratio between the scaling window width, height, or size of the current image and the scaling window width, height, or size of the reference image is within a ratio constraint, wherein the ratio constraint is between 1 / 8 and 2; generating a reference block from the reference image according to the ratio between the scaling window width, height, or size of the current image and the scaling window width, height, or size of the reference image; Using the reference block to perform motion compensation for the current block; and Encode or decode the current block in the current image.

12. A non-transitory computer readable medium for storing program instructions, the program instructions causing a processing circuit of a device to perform a video processing method for video data, the method comprising: receiving input video data for a current block in a current image; Determine the current image's zoom window width, height, or size; determining a scaling window width, height, or size of a reference image, wherein a ratio between the scaling window width, height, or size of the current image and the scaling window width, height, or size of the reference image is within a ratio constraint, wherein the ratio constraint is between 1 / 8 and 2; generating a reference block from the reference image according to the ratio between the scaling window width, height, or size of the current image and the scaling window width, height, or size of the reference image; Using the reference block to perform motion compensation for the current block; and Encode or decode the current block in the current image.

Citation Information

Patent Citations

  • Video encoding and decoding method using space zoom prediction

    CN102752588A

  • Reference layer sample position derivation for scalable video coding

    CN105900431A