Method and apparatus for encoding and decoding a video sequence

By constraining RPR-related parameters, the problem of increased computing complexity and memory bandwidth caused by RPR in video encoding and decoding is solved, and efficient encoding and decoding of video sequences is achieved, reducing the memory bandwidth pressure in the worst-case scenario.

CN114982236BActive Publication Date: 2025-06-06HFI INNOVATION INC
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202080092706.9
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Priority Date
2020-12-10
Filing Date
2020-12-11
Publication Date
2025-06-06
Estimated Expiration
2040-12-11

AI Technical Summary

Technical Problem

In video encoding and decoding, the computational complexity and memory bandwidth increase caused by reference image resampling (RPR) is especially in the worst case, and the pressure on memory bandwidth is needed to be effectively relieved.

Method used

Effective encoding and decoding of video sequences is achieved in the encoder and decoder by constraining RPR-related parameters such as the scaling window size, image width and height of the current image and reference image, and the specified maximum image size.

Benefits of technology

Effectively reduces the worst-case memory bandwidth consumption and improves the efficiency and performance of video encoding and decoding, especially when dealing with high resolution and dynamic channel conditions.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114982236B_ABST
    Figure CN114982236B_ABST
Patent Text Reader

Abstract

A method and device for encoding and decoding a video sequence are disclosed. According to the method, a bitstream corresponding to the encoded data of the video sequence is generated on the encoder side or received on the decoder side, wherein the bitstream conforms to bitstream consistency: one or more constraints are satisfied. The one or more constraints are related to a set of reference picture resampling (RPR) parameters, including a scaling window width or height of the current image, a scaling window width or height of the reference image, a width or height of the current image, and a maximum image width or height specified for the video sequence. Scaling information of the RPR mode can be derived using the set of RPR parameters. Then, when the RPR mode is enabled for a target image, the target image of the video sequence is encoded on the encoder side or decoded on the decoder side by using the scaling information.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] Related references

[0002] This application claims priority to U.S. Provisional Application No. 62 / 946,542 filed on December 11, 2019, U.S. Provisional Application No. 62 / 953,232 filed on December 24, 2019, and U.S. Provisional Application No. 62 / 954,020 filed on December 27, 2019, respectively. The entire contents of the above are incorporated herein by reference. Technical Field

[0003] The present invention relates to video codecs including Reference Picture Resampling (RPR) codec tools. In particular, the present invention relates to constraining RPR parameters to mitigate worst-case memory bandwidth. Background Art

[0004] The High Efficiency Video Coding (HEVC) standard has been developed under a joint video project of the ITU-T Video Coding Experts Group (VCEG) and the ISO / IEC Moving Picture Experts Group (MPEG) standardization organizations, especially in collaboration with a group called the Joint Collaborative Team on Video Coding (JCT-VC). The emerging video codec standard development is called Versatile Video Coding (VVC), which has been carried out in recent years as the next generation of video codecs beyond HEVC. VVC supports Reference Picture Resampling (RPR) as a tool for adaptive streaming services to support upsampling and downsampling motion compensation in real-time (on-the-fly). The technologies related to adaptive streaming services are as follows.

[0005] Reference Picture Resampling (RPR)

[0006] During the development of VVC, according to the "Requirements for a Future Video Coding Standard", the standard should support fast representation switching in the case where an adaptive streaming service provides multiple representations of the same content, each with different properties (such as spatial resolution or sample bit depth). In real-time video communication, allowing the resolution to be changed in a coded video sequence without inserting I pictures can not only make the video data seamlessly adapt to dynamic channel conditions or user preferences, but also eliminate the jitter effect caused by I pictures. A hypothetical example of adaptive resolution change (ARC) with reference picture resampling (RPR) is shown in Figure 1 In which the current image (110) is predicted based on reference images (Ref0 120 and Ref1 130) of different sizes. Figure 1 As shown, the reference image Ref0 (120) has a lower resolution than the current image (110). In order to use the reference image Ref0 as a reference, Ref0 must be scaled up to the same resolution as the current image. The reference image Ref1 (130) has a higher resolution than the current image (110). In order to use the reference image Ref1 as a reference, Ref1 must be scaled down to the same resolution as the current image.

[0007] To support spatial scalability, the image size of the reference image can be different from the current image, which is useful for streaming applications. Methods for supporting reference picture resampling (RPR), also known as adaptive resolution change (ARC), have been studied for inclusion in the VVC specification. At the 14th JVET meeting in Geneva, some contributions on RPR were submitted for discussion.

[0008] When RPR is used, the picture size ratio is derived from the width and height of the reference picture and the width and height of the current picture. The picture size ratio is constrained to be in the range [1 / 8 to 2]. In other words, the picture size ratio is between 1 / 8 and 2. The picture width / height in units of luma samples can be sent in the bitstream, such as PPS, with the following semantics.

[0009] pic_width_in_luma_samples specifies the width of each decoded picture referencing the PPS in units of luma samples. pic_width_in_luma_samples shall not be equal to 0, shall be an integer multiple of Max(8, MinCbSizeY), and shall be less than or equal to pic_width_max_in_luma_samples.

[0010] When subpics_present_flag is equal to 1 or ref_pic_resampling_enabled_flag is equal to 0, the value of pic_width_in_luma_samples shall be equal to pic_width_in_luma_samples.

[0011] pic_height_in_luma_samples specifies the height of each decoded picture referencing the PPS in units of luma samples. pic_height_in_luma_samples shall not be equal to 0, shall be an integer multiple of Max(8, MinCbSizeY), and shall be less than or equal to pic_height_max_in_luma_samples.

[0012] When subpics_present_flag is equal to 1 or ref_pic_resampling_enabled_flag is equal to 0, the value of pic_height_in_luma_samples shall be equal to pic_height_in_luma_samples.

[0013] When the image sizes of the current image and the reference image are specified, the following constraint should be satisfied. This constraint limits the image size ratio of the reference image to the current image to be in the range [1 / 8, 2].

[0014] Let refPicWidthInLumaSamples and RefPicHeightInLumaSamples be the pic_width_in_luma_samples and pic_height_in_luma_samples, respectively, of the reference picture that references the current picture of this PPS. It is a requirement for bitstream conformance that all of the following conditions are met:

[0015] –pic_width_in_luma_samples*2 should be greater than or equal to refPicWidthInLumaSamples.

[0016] –pic_height_in_luma_samples*2 should be greater than or equal to refPicHeightInLumaSamples.

[0017] –pic_width_in_luma_samples should be less than or equal to refPicWidthInLumaSamples*8.

[0018] –pic_height_in_luma_samples should be less than or equal to refPicHeightInLumaSamples*8.

[0019] In VVC, the scaling factor and scaling offset of RPR are derived from the syntax information sent in PPS. The PPS syntax is shown in the following table.

[0020] Table 1. Image parameter set RBSP syntax for scaling and scaling offset

[0021]

[0022]

[0023] The semantics of the grammar is described below.

[0024] scaling_window_flag equal to 1 indicates that the scaling window offset parameter is present in the PPS. scaling_window_flag equal to 0 indicates that the scaling window offset parameter is not present in the PPS. When ref_pic_resampling_enabled_flag is equal to 0, the value of scaling_window_flag shall be equal to 0.

[0025] scaling_win_left_offset, scaling_win_right_offset, scaling_win_top_offset, and scaling_win_bottom_offset specify an offset in units of luma samples that is applied to the image size for scaling calculations. When scaling_window_flag is equal to 0, the values ​​of scaling_win_left_offset, scaling_win_right_offset, scaling_win_top_offset, and scaling_win_bottom_offset are inferred to be equal to 0.

[0026] The value of scaling_win_left_offset+scaling_win_right_offset should be less than pic_width_in_luma_samples, and the value of scaling_win_top_offset+scaling_win_bottom_offset should be less than pic_height_in_luma_samples.

[0027] The variables PicOutputWidthL and PicOutputHeightL are exported as follows:

[0028] PicOutputWidthL=pic_width_in_luma_samples-

[0029] (scaling_win_right_offset+scaling_win_left_offset).

[0030] PicOutputHeightL=pic_height_in_luma_samples-(scaling_win_bottom_offset+scaleing_win_top_offset).

[0031] fRefWidth is set equal to PicOutputWidthL (in luma samples) of the reference image RefPicList[i][j], and fRefHeight is set equal to PicOutputHeightL (in luma samples) of the reference image RefPicList[i][j].

[0032] RefPicScale[i][j][0]=((fRefWidth<<14)+(PicOutputWidthL>>1)) / PicOutputWidthL.

[0033] RefPicScale[i][j][1]=((fRefHeight<<14)+(PicOutputHeightL>>1)) / PicOutputHeightL.

[0034] RefPicIsScaled[i][j]=(RefPicScale[i][j][0]!=(1<<14))||(RefPicScale[i][j][1]!=(1<<14)).

[0035] Although RPR adds flexibility to encoding and decoding video bitstreams, motion compensation associated with RPR scaling leads to increased computational complexity and memory bandwidth. In order to mitigate the worst-case memory bandwidth, the present invention discloses a method and apparatus for constraining parameters related to RPR. Summary of the invention

[0036] A method and apparatus for encoding a video sequence are disclosed, wherein a reference picture resampling (RPR) mode is disclosed. According to the method, a bitstream corresponding to the encoded data of the video sequence is generated on the encoder side or received on the decoder side, wherein the bitstream conforms to the bitstream consistency: one or more constraints are satisfied. The constraints are related to a set of RPR parameters, which include a scaling window width or height of the current image, a scaling window width or height of the reference image, a current image width or height, and a maximum image width or height specified for the video sequence. The scaling information of the RPR mode is derived using the set of RPR parameters. Then, when the RPR mode is enabled for a target image, the target image of the video sequence can be encoded on the encoder side or decoded on the decoder side by using the scaling information.

[0037] In one embodiment, the constraint is not affected by any interpolation filter associated with the RPR process. For example, the constraint may include a first value less than or equal to (SpsMaxPicWidth*PicOutputWidthL), and wherein the first value is determined by CurPicWidth and refPicOutputWidthL. CurPicWidth corresponds to the current image width, refPicOutputWidthL corresponds to the scaling window width of the reference image, PicOutputWidthL corresponds to the scaling window width of the current image, and SpsMaxPicWidth corresponds to the maximum image width specified for the video sequence. In another example, the constraint may include a second value less than or equal to (SpsMaxPicHeight*PicOutputHeight), and wherein the second value is determined by CurPicHeight and refPicOutputHeight. CurPicHeight corresponds to the current image height, refPicOutputHeight corresponds to the scaling window height of the reference image, PicOutputHeight corresponds to the scaling window height of the current image, and SpsMaxPicHeight corresponds to the maximum image height specified for the video sequence. In another example, the constraints may include (CurPicWidth*refPicOutputWidthL) is less than or equal to (SpsMaxPicWidth*PicOutputWidthL). CurPicWidth corresponds to the current image width, refPicOutputWidthL corresponds to the scaling window width of the reference image, PicOutputWidthL corresponds to the scaling window width of the current image, and SpsMaxPicWidth corresponds to the maximum image width specified for the video sequence. In another example, the constraints may include (CurPicHeight*refPicOutputHeight) is less than or equal to (SpsMaxPicHeight*PicOutputHeight). CurPicHeight corresponds to the current image height, refPicOutputHeight corresponds to the scaling window height of the reference image, PicOutputHeight corresponds to the scaling window height of the current image, and SpsMaxPicHeight corresponds to the maximum image height specified for the video sequence.

[0038] In an embodiment, the constraint includes: (pic_width_in_luma_samples*pic_height_in_luma_samples*(16*refPicOutputWidthL+7*PicOutputWidthL)*(4*refPicOutputHeightL+7*PicOutputHeightL)) is less than or equal to (pic_width_max_in_luma_samples*pic_height_max_in_luma_samples height*253*PicOutputWidthL*PicOutputHeightL). In an embodiment, the constraint includes: (pic_width_in_luma_samples*pic_height_in_luma_samples*(4*refPicOutputWidthL+7*PicOutputWidthL)*(16*refPicOutputHeightL+7*PicOutputHeightL)) is less than or equal to (pic_width_max_in_luma_samples*pic_height_max_in_luma_samples height*253*PicOutputWidthL*PicOutputHeightL). pic_width_in_luma_samples and pic_height_in_luma_samples correspond to the current image width and height in luma samples, respectively, refPicOutputWidthL and refPicOutputHeightL correspond to the scaled window width and height of the reference image, respectively, PicOutputWidthL and PicOutputHeightL correspond to the scaled image width and height of the current image, respectively, and pic_width_max_in_luma_samples and pic_height_max_in_luma_samples correspond to the maximum image width and height in luma samples, respectively.

[0039] In one embodiment, the constraints include (CurPicWidth*CurPicHeight*refPicOutputWidthL*refPicOutputHeightL) is less than or equal to (SpsMaxPicWidth*SpsMaxPicHeight*PicOutputWidthL*PicOutputHeightL). CurPicWidth and CurPicHeight correspond to the current image width and height, respectively, refPicOutputWidthL and refPicOutputHeightL correspond to the scaling window width and height of the reference image, respectively, and PicOutputWidthL and PicOutputHeightL correspond to the scaling window width and height of the current image, respectively. SpsMaxPicWidth and SpsMaxPicHeight correspond to the maximum image width and height specified for the video sequence, respectively.

[0040] In an embodiment, the constraint includes: (pic_width_in_luma_samples*pic_height_in_luma_samples*(8*refPicOutputWidthL+7*PicOutputWidthL)*(8*refPicOutputHeightL+7*PicOutputHeightL)) is less than or equal to (pic_width_max_in_luma_samples*pic_height_max_in_luma_samples height*225*PicOutputWidthL*PicOutputHeightL). pic_width_in_luma_samples and pic_height_in_luma_samples correspond to the current image width and height in luma samples, refPicOutputWidthL and refPicOutputHeightL correspond to the scaling window width and height of the reference image, PicOutputWidthL and PicOutputHeightL correspond to the scaling window width and height of the current image, and pic_width_max_in_luma_samples and pic_height_max_in_luma_samples correspond to the maximum image width and height in luma samples. BRIEF DESCRIPTION OF THE DRAWINGS

[0041] Figure 1A hypothetical example of Adaptive Resolution Change (ARC) with Reference Picture Resampling (RPR) is shown, where the current picture is predicted from reference pictures (Ref0 and Ref1) of different sizes.

[0042] Figure 2 An exemplary block diagram of a system for introducing constrained reference picture resampling (RPR) parameters according to an embodiment of the present invention is shown. DETAILED DESCRIPTION

[0043] The following description is the best mode of implementing the present invention. This description is for the purpose of illustrating the general principles of the present invention and should not be considered as limiting. The scope of the present invention is best determined by reference to the attached claims.

[0044] It is readily understood that the elements of the present invention as generally described and illustrated in the drawings herein may be arranged and designed in a variety of different configurations. Therefore, as shown in the drawings, the following more detailed description of the embodiments of the systems and methods of the present invention is not intended to limit the scope of the claimed invention, but merely represents selected embodiments of the present invention.

[0045] References in this specification to "an embodiment", "some embodiments" or similar language mean that a particular feature, structure or characteristic described in conjunction with the embodiment may be included in at least one embodiment of the present invention. Therefore, the phrases "in an embodiment" or "in some embodiments" appearing in various places throughout this specification do not necessarily refer to the same embodiment.

[0046] Furthermore, the described features, structures, or characteristics may be combined in any suitable manner in one or more embodiments. However, those skilled in the art will recognize that the invention may be practiced without one or more of the specific details or with other methods, elements, etc. In other cases, well-known structures or operations are not shown or described in detail to avoid obscuring aspects of the invention.

[0047] The illustrated embodiments of the present invention will be best understood by referring to the drawings, wherein like parts are represented by like numerals throughout. The following description is intended to be exemplary only and simply illustrates certain selected embodiments of apparatus and methods consistent with the invention claimed herein.

[0048] In the specification, like reference numerals throughout the drawings and the specification designate corresponding or similar elements between the different views.

[0049] According to VVC, derived reference picture scales (RefPicScale[i][j][0], RefPicScale[i][j][1]) are used for motion compensation. RefPicScale is derived from the scale window size / width / height specified in the PPS. It affects which filters will be used during the motion compensation stage and also affects the memory bandwidth used for motion compensation. For example, for a 16x16 block, when performing motion compensation, if the scale is equal to 1 (e.g. both RefPicScale[i][j][0] and RefPicScale[i][j][1] are equal to 16384), a (16+L–1)x(16+L–1) reference block is required, where L is the filter tap length for motion compensation. In the SPS, the maximum picture size of the sequence is specified. Based on the maximum picture size, the worst-case bandwidth can be calculated and constrained. However, when the scaling factor is equal to 2 (e.g., RefPicScale[i][j][0] and RefPicScale[i][j][1] are both equal to 32768), then (32+L–1)x(32+L–1) reference blocks are required. Taking into account the effect of the filter tap length, the bandwidth is almost quadrupled. The required bandwidth is affected by the scaling factor. For example, if the size of the current image is equal to the maximum image size, and the scaling factor of one of the reference images is greater than 1, the required bandwidth of this current image may be greater than the bandwidth expected by the system. Several methods have been proposed to constrain the worst-case bandwidth.

[0050] Method-1: Constrained Scaling

[0051] In the present invention, the scaling ratio (of the reference image size to the current image size) is constrained to be no greater than the ratio (of the maximum image size in the SPS to the current image size). For example, scaling_ratio_x*current_picture_width*scaling_ratio_y*current_picture_height should be less than or equal to max_picture_width*max_picture_height. current_picture_width or current_picture_height can be the width or height of the image (e.g., expressed or derived in PPS), or a consistency window (e.g., expressed or derived in PPS), or a scaling window (e.g., expressed or derived in PPS). max_picture_width or max_picture_height height can be the maximum image width or height of the current sequence (e.g., sent or derived in SPS). In one embodiment, scaling_ratio_x and scaling_ratio_y can be RefPicIsScaled[][][0] and RefPicIsScaled[][][1]. For example, RefPicIsScaled[][][0]*current_picture_width*RefPicIsScaled[][][1]*current_picture_height should be less than or equal to max_picture_width*max_picture_height*2^K, or (((RefPicIsScaled[][][0]*RefPicIsScaled[][][1]*current_picture_width*current_picture_height)>>K) should be less than or equal to max_picture_width*max_picture_height. K can be 28.

[0052] In another embodiment, the horizontal scaling ratio and the vertical scaling can be constrained separately. For example, the horizontal and vertical scaling ratios (of the reference image size to the current image size) are constrained to be no greater than the ratio of (the maximum image width in the SPS to the current image width) and the ratio of (the maximum image height in the SPS to the current image height), respectively. For example, scaling_ratio_x*current_picture_width should be less than or equal to max_picture_width. scaling_ratio_y*current_picture_height should be less than or equal to max_picture_height. current_picture_width or current_picture_height can be the width or height of the image (e.g., expressed or derived in PPS, or a consistency window (e.g., expressed or derived in PPS), or a scaling window (e.g., expressed or derived in PPS). max_picture_width or max_picture_height height may be the maximum image width or height of the current sequence (e.g., sent or derived in the SPS). In one embodiment, scaling_ratio_x and scaling_ratio_y may be RefPicIsScaled[][][0] and RefPicIsScaled[][][1]. For example, RefPicIsScaled[][][0]*current_picture_width should be less than or equal to max_picture_2^K, or (((RefPicIsScaled[][0]*current_picture_width)>>K) should be less than or equal to max_picture_width. Similarly, RefPicIsScaled[][][1]*current_picture_height should be less than or equal to max_picture_height*2^K, or (((RefPicIsScaled[][][1]*current_picture_height)>>K)>>max_picture_height. K may be 14.

[0053] In another embodiment, the size of the scaling window, the current picture size, the reference picture size and / or the maximum picture size in the current sequence are constrained. For example, let refPicOutputWidthL and refPicOutputHeightL be the PicOutputWidthL and PicOutputHeightL of the reference picture of the current picture that references this PPS, respectively. A requirement for bitstream conformance is that the value determined by all of refPicOutputWidthL, current_picture_width, refPicOutputHeightL and current_picture_height shall be less than or equal to PicOutputWidthL*max_picture_width*PicOutputHeightL*max_picture_height. For example, satisfying all of the following conditions is a requirement for bitstream conformance:

[0054] –refPicOutputWidthL*current_picture_width*refPicOutputHeightL*current_picture_height should be less than or equal to PicOutputWidthL*max_picture_width*PicOutputHeightL*max_picture_height.

[0055] In another embodiment, the width and height of the scaling window, the width and height of the current image, the width and height of the reference image, and / or the maximum width and height of the images in the current sequence are constrained separately. For example, refPicOutputWidthL and refPicOutputHeightL are respectively PicOutputWidthL and PicOutputHeightL of the reference image of the current image that references the PPS. The requirement for bitstream consistency is that the first value determined by all refPicOutputWidthL and current_picture_width, refPicOutputHeightL should be less than or equal to PicOutputWidthL*max_picture_width; the second value determined by refPicOutputHeightL and current_picture_height should be less than or equal to PicOutputHeightL*max_picture_height. For example, satisfying all of the following conditions is a requirement for bitstream consistency:

[0056] –refPicOutputWidthL*current_picture_width should be less than or equal to PicOutputWidthL*max_picture_width,

[0057] –refPicOutputHeightL*current_picture_height should be less than or equal to PicOutputHeightL*max_picture_height.

[0058] current_picture_width and current_picture_height may be the image width and height sent in the SPS or PPS. For example, the width and height of the image may be pic_width_in_luma_samples and pic_height_in_luma_samples, or may be the width and height of the consistent cropping window, or may be the width and height of the scaling window.

[0059] In another embodiment, interpolation filters are considered. The worst case MC memory bandwidth of the current image is equal to CurPicWidth*CurPicHeight*WorstCaseBlockBW / WorstCaseBlockSize. Therefore, the worst case MC memory bandwidth will not increase if the following conditions are met.

[0060] CurPicWidth*CurPicHeight*(8*ScalingRatioX+7)*(8*ScalingRatioY+7) / (8*8)<=SpsMaxPicWidth*SpsMaxPicHeight*(8+7)*(8+7) / (8*8).

[0061] In the above equation, ScalingRatioX is equal to (refPicOutputWidthL / PicOutputWidthL), ScalingRatioY is equal to (refPicOutputHeightL / PicOutputHeightL), SpsMaxPicWidth is the maximum image width sent in SPS, and SpsMaxPicHeight is the maximum image height sent in SPS.

[0062] After formula simplification, the following equation can be rewritten into the following form.

[0063] CurPicWidth*CurPicHeight*(8*refPicOutputWidthL+7*PicOutputWidthL)*(8*refPicOutputHeightL+7*PicOutputHeightL)<=SpsMaxPicWidth*SpsMaxPicHeight*225*PicOutputWidthL*PicOutputHeightL.

[0064] Similarly, for the chroma components, the constraints are as follows:

[0065] CurPicWidth*CurPicHeight*(8 / SubWidthC*refPicOutputWidthL+3*PicOutputWidthL)*(8 / SubHeightC*refPicOutputHeightL+3*Pi cOutputHeightL)<=SpsMaxPicWidth*SpsMaxPicHeight*(8 / SubWidthC+3)*(8 / SubHeightC+3)*PicOutputWidthL*PicOutputHeightL.

[0066] In the above equations, for 4:2:0, 4:2:2, and 4:4:4, (SubWidthC, SubHeightC) are (2, 2), (2, 1), and (1, 1), respectively.

[0067] The suggested text for resizing the window is as follows.

[0068] scaling_win_left_offset, scaling_win_right_offset, scaling_win_top_offset, and scaling_win_bottom_offset specify an offset in units of luma samples that is applied to the image size for scaling calculations. When scaling_window_flag is equal to 0, the values ​​of scaling_win_left_offset, scaling_win_right_offset, scaling_win_top_offset, and scaling_win_bottom_offset are inferred to be equal to 0.

[0069] The value of scaling_win_left_offset+scaling_win_right_offset should be less than pic_width_in_luma_samples, and the value of scaling_win_top_offset+scaling_win_bottom_offset should be less than pic_height_in_luma_samples.

[0070] The variables PicOutputWidthL and PicOutputHeightL are derived as follows:

[0071] PicOutputWidthL=pic_width_in_luma_samples-

[0072] (scaling_win_right_offset+scaling_win_left_offset).

[0073] PicOutputHeightL=pic_height_in_luma_samples-

[0074] (scaling_win_bottom_offset+scaling_win_top_offset).

[0075] In another example, scaling_win_left_offset, scaling_win_right_offset, scaling_win_top_offset, and scaling_win_bottom_offset are sent in units of chroma samples. The variables PicOutputWidthL and PicOutputHeightL are derived as follows:

[0076] PicOutputWidthL=pic_width_in_luma_samples-

[0077] SubWidthC*(scaling_win_right_offset+scaling_win_left_offset).

[0078] PicOutputHeightL=pic_height_in_luma_samples-

[0079] SubHeightC*(scaling_win_bottom_offset+scaling_win_top_offset).

[0080] In the above equation, SubWidthC and SubHeightC specify the sampling ratio of luma samples to chroma samples in the horizontal and vertical directions, respectively.

[0081] According to an embodiment, the following constraints are imposed. When scaling_window_flag is equal to 1, let refPicOutputWidthL and refPicOutputHeightL be the PicOutputWidthL and PicOutputHeightL of the reference picture of the current picture that refers to the PPS, respectively. It is a requirement for bitstream consistency that the following conditions are met:

[0082] –pic_width_in_luma_samples*pic_height_in_luma_samples*(8*refPicOutputWidthL+7*PicOutputWidthL)*(8*refPicOutputHeightL+7*PicOutputHeightL) should be less than or equal to pic_width_max_in_luma_samples*pic_height_max_in_lum_samples height*225*PicOutputWidthL*PicOutputHeightL.

[0083] –pic_width_in_luma_samples*pic_height_in_luma_samples*(8 / SubWidthC*refPicOutputWidthL+3*PicOutputWidthL)*(8 / SubH eightC*refPicOutputHeightL+3*PicOutputHeightL) should be less than or equal to pic_width_max_in_luma_samples*pic_height_max_in_luma_samples height*(8 / SubWidthC+3)*(8 / SubHeightC+3)*PicOutputWidthL*PicOutputHeightL.

[0084] In another approach, we can consider only the luminance component. The suggested text for scaling the window size is as follows.

[0085] scaling_win_left_offset, scaling_win_right_offset, scaling_win_top_offset, and scaling_win_bottom_offset specify an offset in units of luma samples that is applied to the image size for scaling calculations. When scaling_window_flag is equal to 0, the values ​​of scaling_win_left_offset, scaling_win_right_offset, scaling_win_top_offset, and scaling_win_bottom_offset are inferred to be equal to 0.

[0086] The value of scaling_win_left_offset+scaling_win_right_offset should be less than pic_width_in_luma_samples, and the value of scaling_win_top_offset+scaling_win_bottom_offset should be less than pic_height_in_luma_samples.

[0087] The variables PicOutputWidthL and PicOutputHeightL are derived as follows:

[0088] PicOutputWidthL=pic_width_in_luma_samples-

[0089] (scaling_win_right_offset+scaling_win_left_offset).

[0090] PicOutputHeightL=pic_height_in_luma_samples-

[0091] (scaling_win_bottom_offset+scaling_win_top_offset).

[0092] According to an embodiment, the following constraints are imposed. When scaling_window_flag is equal to 1, let refPicOutputWidthL and refPicOutputHeightL be the PicOutputWidthL and PicOutputHeightL of the reference picture of the current picture that refers to the PPS, respectively. It is a requirement for bitstream consistency that the following conditions are met:

[0093] –pic_width_in_luma_samples*pic_height_in_luma_samples*(8*refPicOutputWidthL+7*PicOutputWidthL)*(8*refPicOutputHeightL+7*PicOutputHeightL) should be less than or equal to pic_width_max_in_luma_samples*pic_height_max_in_lum_samples height*225*PicOutputWidthL*PicOutputHeightL.

[0094] Note that the numbers 8, 7, and 3 above can be replaced by other numbers. For example, after the formula is simplified, it can be rewritten as follows.

[0095] CurPicWidth*CurPicHeight*(M*refPicOutputWidthL+N*PicOutputWidthL)*(O*refPicOutputHeightL+P*PicOutputHeightL)<=SpsMaxPicWidth*SpsMaxPicHeight*(M+N)*(O+P)*PicOutputWidthL*PicOutputHeightL.

[0096] Similarly, for the chroma components, the constraints are as follows.

[0097] CurPicWidth*CurPicHeight*(A / SubWidthC*refPicOutputWidthL+B*PicOutputWidthL)*(C / SubHeightC*refPicOutputHeightL+D*Pi cOutputHeightL)<=SpsMaxPicWidth*SpsMaxPicHeight*(A / SubWidthC+B)*(C / SubHeightC+D)*PicOutputWidthL*PicOutputHeightL.

[0098] In the above equations, for 4:2:0, 4:2:2, and 4:4:4, (SubWidthC, SubHeightC) are (2, 2), (2, 1), and (1, 1), respectively.

[0099] In one embodiment, M and O may be 1, N and P may be 0, A and C may be 1 or 2, and B and D may be 0.

[0100] In one embodiment, 16x4 and 4x16 are used as block sizes to calculate the worst case bandwidth. The constraints can be rewritten as follows.

[0101] After the formula is simplified, it is rewritten as follows:

[0102] CurPicWidth*CurPicHeight*(16*refPicOutputWidthL+7*PicOutputWidthL)*(4*refPicOutputHeightL+7*PicOutputHeightL)<=SpsMaxPicWidth*SpsMaxPicHeight*253*PicOutputWidthL*PicOutputHeightL,

[0103] CurPicWidth*CurPicHeight*(4*refPicOutputWidthL+7*PicOutputWidthL)*(16*refPicOutputHeightL+7*PicOutputHeightL)<=SpsMaxPicWidth*SpsMaxPicHeight*253*PicOutputWidthL*PicOutputHeightL.

[0104] Similarly, for the chroma components, the constraints are as follows:

[0105] CurPicWidth*CurPicHeight*(16 / SubWidthC*refPicOutputWidthL+3*PicOutputWidthL)*(4 / SubHeightC*refPicOutputHeightL+3*PicOutputHeightL)<=

[0106] SpsMaxPicWidth*SpsMaxPicHeight*(16 / SubWidthC+3)*(4 / SubHeightC+3)*PicOutputWidthL*PicOutputHeightL,

[0107] CurPicWidth*CurPicHeight*(4 / SubWidthC*refPicOutputWidthL+3*PicOutputWidthL)*(16 / SubHeightC*refPicOutputHeightL+3*Pi cOutputHeightL)<=SpsMaxPicWidth*SpsMaxPicHeight*(4 / SubWidthC+3)*(16 / SubHeightC+3)*PicOutputWidthL*PicOutputHeightL.

[0108] In the above equations, (SubWidthC, SubHeightC) are (2, 2), (2, 1) and (1, 1) for 4:2:0, 4:2:2 and 4:4:4, respectively.

[0109] The suggested text for scaling the window size is as follows.

[0110] scaling_win_left_offset, scaling_win_right_offset, scaling_win_top_offset, and scaling_win_bottom_offset specify an offset in units of luma samples that is applied to the image size for scaling calculations. When scaling_window_flag is equal to 0, the values ​​of scaling_win_left_offset, scaling_win_right_offset, scaling_win_top_offset, and scaling_win_bottom_offset are inferred to be equal to 0.

[0111] The value of scaling_win_left_offset+scaling_win_right_offset should be less than pic_width_in_luma_samples, and the value of scaling_win_top_offset+scaling_win_bottom_offset should be less than pic_height_in_luma_samples.

[0112] The variables PicOutputWidthL and PicOutputHeightL are derived as follows:

[0113] PicOutputWidthL=pic_width_in_luma_samples-(scaling_win_right_offset+scaling_win_left_offset).

[0114] PicOutputHeightL=pic_height_in_luma_samples-(scaling_win_bottom_offset+scaling_win_top_offset).

[0115] According to an embodiment, the following constraints are imposed. When scaling_window_flag is equal to 1, let refPicOutputWidthL and refPicOutputHeightL be the PicOutputWidthL and PicOutputHeightL of the reference picture of the current picture that references this PPS, respectively. It is a requirement for bitstream consistency that the following conditions are met:

[0116] –pic_width_in_luma_samples*pic_height_in_luma_samples*(16*refPicOutputWidthL+7*PicOutputWidthL)*(4*refPicOutputHeightL+7*PicOutputHeightL) should be less than or equal to pic_width_max_in_luma_samples*pic_height_max_in_lum_samples height*253*PicOutputWidthL*PicOutputHeightL.

[0117] –pic_width_in_luma_samples*pic_height_in_luma_samples*(4*refPicOutputWidthL+7*PicOutputWidthL)*(16*refPicOutputHeightL+7*PicOutputHeightL) should be less than or equal to pic_width_max_in_luma_samples*pic_height_Width_sample height*253*PicOutputWidthL*PicOutputHeightL.

[0118] –pic_width_in_luma_samples*pic_height_in_luma_samples*(16 / SubWidthC*refPicOutputWidthL+3*PicOutputWidthL)*(4 / Sub HeightC*refPicOutputHeightL+3*PicOutputHeightL) should be less than or equal to pic_width_max_in_luma_samples*pic_height_max_in_luma_samples height*(16 / SubWidthC+3)*(4 / SubHeightC+3)*PicOutputWidthL*PicOutputHeightL.

[0119] –pic_width_in_luma_samples*pic_height_in_luma_samples*(4 / SubWidthC*refPicOutputWidthL+3*PicOutputWidthL)*(16 / Sub HeightC*refPicOutputHeightL+3*PicOutputHeightL) should be less than or equal to pic_width_max_in_luma_samples*pic_height_max_in_luma_samples height*(4 / SubWidthC+3)*(16 / SubHeightC+3)*PicOutputWidthL*PicOutputHeightL.

[0120] Additionally, in another example, we can consider only the luminance component.

[0121] In another embodiment, a tolerance ratio may be called. For example, Ty may be a tolerance value for brightness. Tc may be a tolerance value for chromaticity. Ty and Tc may be integers or real numbers (e.g., 1.0, 1.2, 1.1, 1.5, etc.).

[0122] For example, after formula simplification, the constraints can be rewritten as follows.

[0123] CurPicWidth*CurPicHeight*(M*refPicOutputWidthL+N*PicOutputWidthL)*(O*refPicOutputHeightL+P*PicOutputHeightL)<=SpsMaxPicWidth*SpsMaxPicHeight*(M+N)*(O+P)*PicOutputWidthL*PicOutputHeightL*Ty.

[0124] Similarly, for the chroma components, the constraints are as follows.

[0125] CurPicWidth*CurPicHeight*(A / SubWidthC*refPicOutputWidthL+B*PicOutputWidthL)*(C / SubHeightC*refPicOutputHeightL+D*Pic OutputHeightL)<=SpsMaxPicWidth*SpsMaxPicHeight*(A / SubWidthC+B)*(C / SubHeightC+D)*PicOutputWidthL*PicOutputHeightL*Tc,

[0126] In the equation, for 4:2:0, 4:2:2, and 4:4:4, (SubWidthC, SubHeightC) are (2, 2), (2, 1), and (1, 1), respectively.

[0127] In one example, M and O may be 1, N and P may be 0, A and C may be 1 or 2, and B and D may be 0. In one example, M and O may be 8, N and P may be 7 or 8, A and C may be 8, and B and D may be 3 or 4. In one example, M may be 16, O may be 4, N and P may be 7 or 8, A may be 16, C may be 4, and B and D may be 3 or 4. In one example, M may be 4, O may be 16, N and P may be 7 or 8, A may be 4 and C may be 16, and B and D may be 3 or 4.

[0128] In another embodiment, the vertical constraint and the horizontal constraint can be separated. For example, after the formula is simplified, the constraint can be rewritten as follows.

[0129] CurPicWidth*(M*refPicOutputWidthL+N*PicOutputWidthL)<=SpsMaxPicWidth*(M+N)*PicOutputWidthL*Ty_x.

[0130] CurPicHeight*(O*refPicOutputHeightL+P*PicOutputHeightL)<=SpsMaxPicHeight*(O+P)*PicOutputHeightL*Ty_y.

[0131] Similarly, for the chroma components, the constraints are as follows.

[0132] CurPicWidth*(A / SubWidthC*refPicOutputWidthL+B*PicOutputWidthL)<=SpsMaxPicWidth*(A / SubWidthC+B)*PicOutputWidthL*Tc_x.

[0133] CurPicHeight*(C / SubHeightC*refPicOutputHeightL+D*PicOutputHeightL)<=SpsMaxPicHeight*(C / SubHeightC+D)*PicOutputHeightL*Tc_y.

[0134] In the above equations, (SubWidthC, SubHeightC) are (2, 2), (2, 1) and (1, 1) for 4:2:0, 4:2:2 and 4:4:4, respectively.

[0135] In one example, after formula simplification, the constraints can be rewritten as follows.

[0136] CurPicWidth*(16*refPicOutputWidthL+7*PicOutputWidthL)<=SpsMaxPicWidth*(16+7)*PicOutputWidthL*Ty_x.

[0137] CurPicHeight*(4*refPicOutputHeightL+7*PicOutputHeightL)<=SpsMaxPicHeight*(4+7)*PicOutputHeightL*Ty_y.

[0138] CurPicWidth*(4*refPicOutputWidthL+7*PicOutputWidthL)<=SpsMaxPicWidth*(4+7)*PicOutputWidthL*Ty_x.

[0139] CurPicHeight*(16*refPicOutputHeightL+7*PicOutputHeightL)<=SpsMaxPicHeight*(16+7)*PicOutputHeightL*Ty_y.

[0140] In another example, the constraint can be rewritten as follows:

[0141] CurPicWidth*(4*refPicOutputWidthL+7*PicOutputWidthL)<=SpsMaxPicWidth*(4+7)*PicOutputWidthL*Ty_x.

[0142] CurPicHeight*(4*refPicOutputHeightL+7*PicOutputHeightL)<=SpsMaxPicHeight*(4+7)*PicOutputHeightL*Ty_y.

[0143] In another example, the constraint can be rewritten as follows:

[0144] CurPicWidth*(16*refPicOutputWidthL+7*PicOutputWidthL)<=SpsMaxPicWidth*(16+7)*PicOutputWidthL*Ty_x.

[0145] CurPicHeight*(16*refPicOutputHeightL+7*PicOutputHeightL)<=SpsMaxPicHeight*(16+7)*PicOutputHeightL*Ty_y.

[0146] Similarly, for the chroma components, the constraints are as follows:

[0147] CurPicWidth*(16 / SubWidthC*refPicOutputWidthL+3*PicOutputWidthL)<=SpsMaxPicWidth*(16 / SubWidthC+3)*PicOutputWidthL*Tc_x.

[0148] CurPicHeight*(4 / SubHeightC*refPicOutputHeightL + 3*PicOutputHeightL) <= SpsMaxPicHeight*(4 / SubHeightC + 3)*PicOutputHeightL*Tc_y。

[0149] CurPicWidth*(4 / SubWidthC*refPicOutputWidthL + 3*PicOutputWidthL) <= SpsMaxPicWidth*(4 / SubWidthC + 3)*PicOutputWidthL*Tc_x。

[0150] CurPicHeight*(16 / SubHeightC*refPicOutputHeightL + 3*PicOutputHeightL) <= SpsMaxPicHeight*(16 / SubHeightC + 3)*PicOutputHeightL*Tc_y。

[0151] In another example, the constraint can be rewritten as follows:

[0152] CurPicWidth*(4 / SubWidthC*refPicOutputWidthL + 3*PicOutputWidthL) <= SpsMaxPicWidth*(4 / SubWidthC + 3)*PicOutputWidthL*Tc_x。

[0153] CurPicHeight*(4 / SubHeightC*refPicOutputHeightL + 3*PicOutputHeightL) <= SpsMaxPicHeight*(4 / SubHeightC + 3)*PicOutputHeightL*Tc_y。

[0154] In another example, the constraint can be rewritten as follows:

[0155] CurPicWidth*(16 / SubWidthC*refPicOutputWidthL + 3*PicOutputWidthL) <= SpsMaxPicWidth*(16 / SubWidthC + 3)*PicOutputWidthL*Tc_x。

[0156] CurPicHeight*(16 / SubHeightC*refPicOutputHeightL+3*PicOutputHeightL)<=SpsMaxPicHeight*(16 / SubHeightC+3)*PicOutputHeightL*Tc_y.

[0157] In the above equations, (SubWidthC, SubHeightC) are (2, 2), (2, 1) and (1, 1) for 4:2:0, 4:2:2 and 4:4:4, respectively.

[0158] In another embodiment, the filter tap length can be ignored. For example, after formula simplification, the constraint can be rewritten as follows:

[0159] CurPicWidth*CurPicHeight*refPicOutputWidthL*refPicOutputHeightL<=SpsMaxPicWidth*SpsMaxPicHeight*PicOutputWidthL*PicOutputHeightL*T.

[0160] In another example, the constraint can be rewritten as follows:

[0161] CurPicWidth*refPicOutputWidthL<=SpsMaxPicWidth*PicOutputWidthL*Tx.

[0162] CurPicHeight*refPicOutputHeightL<=SpsMaxPicHeight*PicOutputHeightL*Ty.

[0163] In another embodiment, in the above method, the width or height of the image or zoom window may be replaced by the number of CTUs in the width / height of the image / zoom window.

[0164] Method 2: When the scaling ratio is greater than a threshold, small inter-frame blocks are disabled.

[0165] For motion compensation, smaller blocks require larger bandwidth per sample due to the interpolation filter. To reduce memory bandwidth, it is recommended to disable small inter-frame blocks when the scaling ratio is greater than a threshold. For example, when the scaling ratio is greater than K or greater than or equal to K, inter-frame blocks with a height and / or height less than N are disabled. In another example, when the scaling ratio is greater than K, inter-frame blocks with an area less than M are disabled. K can be 1, 1.25, 1.5, 1.75, or 2.0. N can be 8, 16, or 32. M can be 32, 64, 128, 256, 512, or 1024.

[0166] In one embodiment, RefPicIsScaled[][][0] and / or RefPicIsScaled[][][1] may be used. For example, when RefPicIsScaled[][][0] and / or RefPicIsScaled[][][1] is greater than K or greater than or equal to K, inter-frame blocks with a height and / or height less than N are disabled. In another example, when RefPicIsScaled[][][0] and / or RefPicIsScaled[][][1] is greater than K, inter-frame blocks with an area less than M are disabled. K may be (1, 1.25, 1.5, 1.75, or 2.0)*2^P, where P may be 14. N may be 8, 16, or 32. M may be 32, 64, 128, 256, 512, or 1024.

[0167] Method-3: No inter prediction outside the current zoom window

[0168] In the proposed method, if the current CU is outside the current image scaling window, the encoder should not use the inter-frame mode for the frame CU. Therefore, the bitstream should have such consistency that all CUs outside the scaling window should not be in inter-frame mode. In another embodiment, the relevant syntax can be saved. For example, any CU outside the scaling window will only need to send syntax related to intra-frame mode, so that all syntax related to inter-frame mode can be saved. This method will also guarantee the worst-case bandwidth.

[0169] In the method disclosed above, the constraints on the scaling of the RPR are shown in various equations or formulas. These equations or formulas are not intended to provide an exhaustive list of possible implementations based on embodiments of the present invention. As known in the art, by rearranging or recombining the terms in these equations or formulas, these equations or formulas may have various forms. Embodiments of the present invention encompass all equivalent equations or formulas.

[0170] Implementation of any of the above-mentioned proposed methods may be in an encoder and / or a decoder. For example, any of the proposed methods may be implemented in a scaling or motion compensation module or a parameter determination module of an encoder, and / or a scaling or motion compensation module or a parameter determination module of a decoder. Alternatively, any of the proposed methods is implemented as one or more circuits coupled to a scaling or motion compensation module or a parameter determination module of an encoder and / or a scaling or motion compensation module or a parameter determination module of a decoder to provide information required by the scaling or motion compensation module or the parameter determination module.

[0171] The video encoder must follow the above syntax design in order to generate a legal bitstream, and the video decoder can correctly decode the bitstream only if the parsing process complies with the above syntax design. When the syntax is skipped in the bitstream, the encoder and decoder should set the syntax value to the inferred value to ensure that the encoding and decoding results match.

[0172] Figure 2 An exemplary block diagram of a system for introducing constrained reference image resampling (RPR) parameters according to an embodiment of the present invention is shown. The steps shown in the flowchart and other subsequent flowcharts in the present disclosure can be implemented as program codes that can be executed on one or more processors (e.g., one or more CPUs) on the encoder side and / or the decoder side. The steps shown in the flowchart can also be implemented based on hardware such as one or more electronic devices or processors arranged to perform the steps in the flowchart. According to the method, in step 210, a bitstream corresponding to the encoded data of a video sequence is generated on the encoder side or received on the decoder side, and the bitstream conforms to the bitstream consistency: one or more constraints are satisfied, and the one or more constraints are related to a set of RPR parameters, the set of RPR parameters including the scaling window width or height of the current image, the scaling window width or height of the reference image, the current image width or height, and the maximum image width or height specified for the video sequence, wherein the scaling information of the RPR mode is derived using the set of RPR parameters. In step 220, when the RPR mode is enabled for the target image, the target image of the video sequence is encoded on the encoder side or decoded on the decoder side by using the scaling information.

[0173] The flowchart shown is intended to illustrate an example of an embodiment according to the present invention. Those skilled in the art may modify each step, rearrange the steps, split the steps or combine the steps to practice the present invention without departing from the spirit of the present invention.

[0174] The above description is given to enable those skilled in the art to practice the present invention provided in the context of a specific application and its requirements. Various modifications to the described embodiments will be apparent to those skilled in the art, and the general principles defined herein may be applied to other embodiments. Therefore, the present invention is not intended to be limited to the specific embodiments shown and described, but is consistent with the widest range of principles and novel features disclosed herein. In the above detailed description, various specific details are shown in order to provide a thorough understanding of the present invention. However, those skilled in the art will appreciate that the present invention can be implemented.

[0175] The embodiments of the present invention as described above can be implemented in various hardware, software codes or a combination of the two. For example, an embodiment of the present invention can be a program code integrated into one or more circuits in a video compression chip or integrated into video compression software to perform the processing described herein. An embodiment of the present invention can also be a program code executed on a digital signal processor (Digital Signal Processor, DSP) to perform the processing described herein. The present invention may also be related to many functions performed by a computer processor, a digital signal processor, a microprocessor or a field programmable gate array (fieldprogrammable gate arragy, referred to as FPGA). These processors can be configured to perform specific tasks according to the present invention by executing machine-readable software code or firmware code that defines the specific method embodied in the present invention. The software code or firmware code can be developed in different programming languages ​​and different formats or styles. The software code can also be compiled for different target platforms. However, different code formats, styles and languages ​​of software codes and other means of configuring codes to perform tasks according to the present invention will not depart from the spirit and scope of the present invention.

[0176] The present invention may be implemented in other specific forms without departing from the spirit or essential features of the present invention. The described examples should be considered in all respects as illustrative only and not restrictive. Therefore, the scope of the present invention is indicated by the appended claims rather than the above description. All changes falling within the equivalent meaning and scope of the claims should be included within their scope.

Claims

1. A method for encoding and decoding a video sequence, wherein a reference picture resampling mode is supported, the method include: A bitstream of encoded data corresponding to the video sequence is generated at an encoder side or received at a decoder side, wherein the bitstream complies with bitstream conformance: one or more constraints are satisfied, wherein the one or more constraints are related to a set of reference picture resampling parameters, the set of reference picture resampling parameters including a scaling window width or height of a current picture, a scaling window width or height of a reference picture, a current picture width or height, and a maximum picture width or height specified for the video sequence, wherein scaling information of the reference picture resampling mode is derived using the set of reference picture resampling parameters, the one or more constraints including: a first value is less than or equal to (SpsMaxPicWidth*PicOutputWidthL), and wherein the first value is determined based on CurPicWidth and refPicOutputWidthL, and wherein CurPicWidth corresponds to the current picture width, refPicOutputWidthL corresponds to the scaling window width of the reference picture, PicOutputWidthL corresponds to the scaling window width of the current picture, and SpsMaxPicWidth corresponds to the maximum picture width specified for the video sequence; and Using the scaling information, a target picture of the video sequence is encoded at the encoder side or decoded at the decoder side.

2. The method for encoding and decoding a video sequence according to claim 1, It is characterized in that The one or more constraints include: (CurPicWidth*refPicOutputWidthL) is less than or equal to (SpsMaxPicWidth*PicOutputWidthL), and wherein CurPicWidth corresponds to the current image width, refPicOutputWidthL corresponds to the scaling window width of the reference image, PicOutputWidthL corresponds to the scaling window width of the current image, and SpsMaxPicWidth corresponds to the maximum image width specified for the video sequence.

3. The method for encoding and decoding a video sequence according to claim 1, It is characterized in that The one or more constraints include: a second value being less than or equal to (SpsMaxPicHeight*PicOutputHeight), and wherein the second value is determined based on CurPicHeight and refPicOutputHeight, and wherein CurPicHeight corresponds to the current image height, refPicOutputHeight corresponds to the scaling window height of the reference image, PicOutputHeight corresponds to the scaling window height of the current image, and SpsMaxPicHeight corresponds to the maximum image height specified for the video sequence.

4. The method for encoding and decoding a video sequence according to claim 1, It is characterized in that The one or more constraints include: (CurPicHeight*refPicOutputHeight) is less than or equal to (SpsMaxPicHeight*PicOutputHeight), and wherein CurPicHeight corresponds to the current image height, refPicOutputHeight corresponds to the scaling window height of the reference image, PicOutputHeight corresponds to the scaling window height of the current image, and SpsMaxPicHeight corresponds to the maximum image height specified for the video sequence.

5. The method for encoding and decoding a video sequence according to claim 1, It is characterized in that The one or more constraints include (pic_width_in_luma_samples*pic_height_in_luma_samples*(16*refPicOutputWidthL+7*PicOutputWidthL)*(4*refPicOutputHeightL+7*PicOutputHeightL)) is less than or equal to (pic_width_max_in_luma_samples*pic_height_max_in_luma_samples) height*253*PicOutputWidthL*PicOutputHeightL), and wherein pic_width_in_luma_samples and pic_height_in_luma_samples correspond to the current image width and height in units of luma samples, respectively, refPicOutputWidthL and refPicOutputHeightL correspond to the scaling window width and height of the reference image, respectively, PicOutputWidthL and PicOutputHeightL correspond to the scaling window width and height of the current image, pic_width_max_in_luma_samples and pic_height_max_in_luma_samples correspond to the maximum image width and height in units of luma samples, respectively.

6. The method for encoding and decoding a video sequence according to claim 1, It is characterized in that The one or more constraints include (pic_width_in_luma_samples*pic_height_in_luma_samples*(4*refPicOutputWidthL+7*PicOutputWidthL)*(16*refPicOutputHeightL+7*PicOutputHeightL)) is less than or equal to (pic_width_max_in_luma_samples*pic_height_max_in_luma_samples) height*253*PicOutputWidthL*PicOutputHeightL), where pic_width_in_luma_samples and pic_height_in_luma_samples correspond to the current image width and height in units of luma samples, respectively, refPicOutputWidthL and refPicOutputHeightL correspond to the scaling window width and height of the reference image, respectively, PicOutputWidthL and PicOutputHeightL correspond to the scaling window width and height of the current image, and pic_width_max_in_luma_samples and pic_height_max_in_luma_samples correspond to the maximum image width and height in units of luma samples, respectively.

7. The method for encoding and decoding a video sequence according to claim 1, It is characterized in that The one or more constraints include (CurPicWidth*CurPicHeight*refPicOutputWidthL*refPicOutputHeightL) is less than or equal to (SpsMaxPicWidth*SpsMaxPicHeight*PicOutputWidthL*PicOutputHeightL), and wherein CurPicWidth and CurPicHeight correspond to the current image width and height, respectively, refPicOutputWidthL and refPicOutputHeightL correspond to the scaling window width and height of the reference image, respectively, PicOutputWidthL and PicOutputHeightL correspond to the scaling window width and height of the current image, respectively, and SpsMaxPicWidth and SpsMaxPicHeight correspond to the maximum image width and height specified for the video sequence, respectively.

8. The method for encoding and decoding a video sequence according to claim 1, It is characterized in that The one or more constraints include: (pic_width_in_luma_samples*pic_height_in_luma_samples*(8*refPicOutputWidthL+7*PicOutputWidthL)*(8*refPicOutputHeightL+7*PicOutputHeightL)) is less than or equal to (pic_width_max_in_luma_samples*pic_height_max_in_luma_samples) height*225*PicOutputWidthL*PicOutputHeightL), where pic_width_in_luma_samples and pic_height_in_luma_samples correspond to the current image width and height in units of luma samples, respectively, refPicOutputWidthL and refPicOutputHeightL correspond to the scaling window width and height of the reference image, respectively, PicOutputWidthL and PicOutputHeightL correspond to the scaling window width and height of the current image, pic_width_max_in_luma_samples and pic_height_max_in_luma_samples correspond to the maximum image width and height in units of luma samples, respectively.

9. An apparatus for encoding or decoding a video sequence, wherein a reference picture resampling mode is supported, the apparatus comprising one or more electronic circuits configured to: A bitstream of encoded data corresponding to the video sequence is generated at an encoder side or received at a decoder side, wherein the bitstream complies with bitstream conformance: one or more constraints are satisfied, wherein the one or more constraints are related to a set of reference picture resampling parameters, the set of reference picture resampling parameters comprising a scaling window width or height of a current picture, a scaling window width or height of a reference picture, a current picture width or height, and a maximum picture width or height specified for the video sequence, wherein scaling information of the reference picture resampling mode is derived using the set of reference picture resampling parameters, the one or more constraints include: A first value is less than or equal to (SpsMaxPicWidth*PicOutputWidthL), and wherein the first value is determined based on CurPicWidth and refPicOutputWidthL, and wherein CurPicWidth corresponds to the current image width, refPicOutputWidthL corresponds to the scaling window width of the reference image, PicOutputWidthL corresponds to the scaling window width of the current image, and SpsMaxPicWidth corresponds to the maximum image width specified for the video sequence; and Using the scaling information, a target picture of the video sequence is encoded at the encoder side or decoded at the decoder side.

10. The apparatus for encoding and decoding a video sequence according to claim 9, It is characterized in that The one or more constraints include: (CurPicWidth*refPicOutputWidthL) is less than or equal to (SpsMaxPicWidth*PicOutputWidthL), and wherein CurPicWidth corresponds to the current image width, refPicOutputWidthL corresponds to the scaling window width of the reference image, PicOutputWidthL corresponds to the scaling window width of the current image, and SpsMaxPicWidth corresponds to the maximum image width specified for the video sequence.

11. The apparatus for encoding and decoding a video sequence according to claim 10, It is characterized in that The one or more constraints include: a second value being less than or equal to (SpsMaxPicHeight*PicOutputHeight), and wherein the second value is determined based on CurPicHeight and refPicOutputHeight, and wherein CurPicHeight corresponds to the current image height, refPicOutputHeight corresponds to the scaling window height of the reference image, PicOutputHeight corresponds to the scaling window height of the current image, and SpsMaxPicHeight corresponds to the maximum image height specified for the video sequence.

12. The apparatus for encoding and decoding a video sequence according to claim 10, It is characterized in that The one or more constraints include: (CurPicHeight*refPicOutputHeight) is less than or equal to (SpsMaxPicHeight*PicOutputHeight), and wherein CurPicHeight corresponds to the current image height, refPicOutputHeight corresponds to the scaling window height of the reference image, PicOutputHeight corresponds to the scaling window height of the current image, and SpsMaxPicHeight corresponds to the maximum image height specified for the video sequence.

13. The apparatus for encoding and decoding a video sequence according to claim 10, It is characterized in that The one or more constraints include (pic_width_in_luma_samples*pic_height_in_luma_samples*(16*refPicOutputWidthL+7*PicOutputWidthL)*(4*refPicOutputHeightL+7*PicOutputHeightL)) is less than or equal to (pic_width_max_in_luma_samples*pic_height_max_in_luma_samples) height*253*PicOutputWidthL*PicOutputHeightL), and wherein pic_width_in_luma_samples and pic_height_in_luma_samples correspond to the current image width and height in units of luma samples, respectively, refPicOutputWidthL and refPicOutputHeightL correspond to the scaling window width and height of the reference image, respectively, PicOutputWidthL and PicOutputHeightL correspond to the scaling window width and height of the current image, pic_width_max_in_luma_samples and pic_height_max_in_luma_samples correspond to the maximum image width and height in units of luma samples, respectively.

14. The apparatus for encoding and decoding a video sequence according to claim 10, It is characterized in that The one or more constraints include (pic_width_in_luma_samples*pic_height_in_luma_samples*(4*refPicOutputWidthL+7*PicOutputWidthL)*(16*refPicOutputHeightL+7*PicOutputHeightL)) is less than or equal to (pic_width_max_in_luma_samples*pic_height_max_in_luma_samples) height*253*PicOutputWidthL*PicOutputHeightL), where pic_width_in_luma_samples and pic_height_in_luma_samples correspond to the current image width and height in units of luma samples, respectively, refPicOutputWidthL and refPicOutputHeightL correspond to the scaling window width and height of the reference image, respectively, PicOutputWidthL and PicOutputHeightL correspond to the scaling window width and height of the current image, and pic_width_max_in_luma_samples and pic_height_max_in_luma_samples correspond to the maximum image width and height in units of luma samples, respectively.

15. The apparatus for encoding and decoding a video sequence according to claim 10, It is characterized in that The one or more constraints include (CurPicWidth*CurPicHeight*refPicOutputWidthL*refPicOutputHeightL) is less than or equal to (SpsMaxPicWidth*SpsMaxPicHeight*PicOutputWidthL*PicOutputHeightL), and wherein CurPicWidth and CurPicHeight correspond to the current image width and height, respectively, refPicOutputWidthL and refPicOutputHeightL correspond to the scaling window width and height of the reference image, respectively, PicOutputWidthL and PicOutputHeightL correspond to the scaling window width and height of the current image, respectively, and SpsMaxPicWidth and SpsMaxPicHeight correspond to the maximum image width and height specified for the video sequence, respectively.

16. The apparatus for encoding and decoding a video sequence according to claim 10, It is characterized in that The one or more constraints include: (pic_width_in_luma_samples*pic_height_in_luma_samples*(8*refPicOutputWidthL+7*PicOutputWidthL)*(8*refPicOutputHeightL+7*PicOutputHeightL)) is less than or equal to (pic_width_max_in_luma_samples*pic_height_max_in_luma_samples) height*225*PicOutputWidthL*PicOutputHeightL), where pic_width_in_luma_samples and pic_height_in_luma_samples correspond to the current image width and height in units of luma samples, respectively, refPicOutputWidthL and refPicOutputHeightL correspond to the scaling window width and height of the reference image, respectively, PicOutputWidthL and PicOutputHeightL correspond to the scaling window width and height of the current image, pic_width_max_in_luma_samples and pic_height_max_in_luma_samples correspond to the maximum image width and height in units of luma samples, respectively.