Video encoding or decoding method and apparatus related to high-level information signaling

By using high-level information signaling in video encoding or decoding systems to control RPR and sub-image segmentation, the problems of image resolution changes and sub-image segmentation processing in the prior art are solved, and more efficient video encoding and bitstream consistency are achieved.

CN114902660BActive Publication Date: 2025-05-27HFI INNOVATION INC
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202080090562.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Priority Date
2020-12-28
Filing Date
2020-12-30
Publication Date
2025-05-27
Estimated Expiration
2040-12-30

AI Technical Summary

Technical Problem

When processing video data, existing video encoding standards are difficult to effectively process image resolution changes and sub-image segmentation, resulting in low encoding efficiency and inconsistent bitstream.

Method used

By introducing high-level information signaling in the video encoding or decoding system, it is determined that the reference image resampling (RPR) and sub-image segmentation are enabled or disabled, ensuring that images with different resolutions are allowed to be referenced in intra prediction, or only images with the same resolution are allowed to be referenced in inter prediction, and whether to split the image into sub-images is determined based on the bitstream consistency requirements.

Benefits of technology

It improves the encoding efficiency and bitstream consistency of video encoding, supports dynamic upsampling and downsampling motion compensation, and adapts to the needs of different video application scenarios.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114902660B_ABST
    Figure CN114902660B_ABST
Patent Text Reader

Abstract

A video processing method and apparatus for processing video images with reference to a high-level syntax set, including receiving input data, determining a first syntax element indicating whether reference picture resampling is disabled or constrained, determining a second syntax element indicating whether sub-picture segmentation is disabled or constrained, and encoding or decoding the video image. The first and second syntax elements are restricted such that sub-picture segmentation is disabled or constrained when reference picture resampling is enabled, or reference picture resampling is disabled or constrained when the sub-picture segmentation is enabled. The first syntax element and the second element are syntax elements signaled in the high-level syntax set.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] Cross-reference

[0002] This application claims priority to U.S. Provisional Patent Application No. 62 / 955,364, filed on December 30, 2019, entitled "Methods and apparatus related to high-level information signaling for coding image and video data", and U.S. Provisional Patent Application No. 62 / 955,539, filed on December 31, 2019, entitled "Methods and apparatus related to high-level information signaling for coding image and video data", the entire disclosures of which are hereby incorporated by reference. Technical Field

[0003] The present invention relates to video processing methods and apparatuses in video coding and decoding systems. Specifically, the present invention relates to high-level information signaling for reference image resampling and sub-image segmentation. Background Art

[0004] The Versatile Video Coding (VVC) standard is an emerging video coding standard that has been gradually developed based on the previous High Efficiency Video Coding (HEVC) standard by enhancing existing coding tools and introducing a variety of new coding tools in various building blocks of the encoder structure. Figure 1A block diagram of an HEVC encoding system is shown. The input video signal 101 is predicted by an intra or inter prediction module 102 to form a prediction signal 103, which is derived from the decoded picture buffer. The prediction residual signal 105 between the input video signal 101 and the prediction signal 103 is processed by a linear transformation in a transform and quantization module 106. The transform coefficients are quantized in the transform and quantization module 106 and entropy coded together with other side information in an entropy coding module 108 to form a video bitstream 107. A reconstructed signal 109 is generated from the prediction signal 103 and the quantized residual signal 111 by a reconstruction module 112, where the quantized residual signal 111 is generated by an inverse quantization and inverse transform module 110 after inverse transforming the dequantized transform coefficients. The reconstructed signal 109 is further processed by loop filtering in a deblocking filter 114 and a non-deblocking filter 116 to remove coding artifacts. The decoded picture 113 is stored in a frame buffer 118 for predicting future pictures in the input video signal 101. In the HEVC standard, an encoded picture is partitioned into non-overlapping square block regions represented by associated coding tree units (CTUs). An encoded picture may be represented by a number of slices, each including an integer number of CTUs. Individual CTUs within a slice are processed in raster scan order. A bi-predictive (B) slice may be decoded using either intra prediction or inter prediction, by predicting the sample values of each block using up to two motion vectors and reference indices. A predictive (P) slice is decoded using intra prediction or inter prediction, by referring to up to one motion vector and reference index to predict the sample values of each block. An intra (I) slice is decoded using only intra prediction.

[0005] A recursive quadtree (QT) structure may be used to partition a CTU into multiple non-overlapping coding units (CUs) to accommodate various local motion and texture characteristics. One or more prediction units (PUs) are designated for each CU. The PU together with the associated CU syntax serves as a basic unit for signaling prediction sub-information. The specified prediction process is employed to predict the values of the relevant syntax samples within the PU. A CU may be further partitioned using a residual quadtree (RQT) structure to represent the associated prediction residual signal. The leaf nodes of the RQT correspond to transform units (TUs). A TU consists of a transform block of 8x8, 16x16, or 32x32 luma samples, or four TBs of 4x4 luma samples and two corresponding TBs of chroma samples of the picture in 4:2:0 color format. An integer transform is applied to the TBs and the hierarchical values of the quantized coefficients and other side information are entropy coded in the video bitstream. Figure 2A shows the block partitioning of an exemplary CTU and Figure 2B shows its corresponding QT representation. Figure 2A and Figure 2BThe solid lines therein indicate the CU boundaries within a CTU and the dashed lines indicate the TU boundaries within a CTU. The terms Coding Tree Block (CTB), Coding Block (CB), Prediction Block (PB), and Transform Block (TB) are defined to specify a 2-D sample array of one color component associated with a CTU, a CU, a PU, and a TU, respectively. A CTU consists of a luma CTB, two chroma CTBs, and associated syntax elements. Similar relationships are valid for CUs, PUs, and TUs. Tree segmentation is typically applied to both luma and chroma components, but exceptions can occur when the chroma components reach certain minimum sizes.

[0006] The VVC standard improves compression performance and the efficiency of transmission and storage, and supports new formats such as high dynamic range and panoramic 360 video. The video iconography encoded by the VVC standard is segmented into non-overlapping square block regions represented by CTUs, similar to the HEVC standard. Each CTU is partitioned into one or more CUs of smaller sizes by a quadtree with embedded multi-type trees using binary and ternary splits. Each resulting CU partition is square or rectangular in shape. The VVC standard makes video transmission in mobile networks more efficient because it allows systems or locations with poor data rates to receive larger files faster. The VVC standard supports layer coding, spatial or signal-to-noise ratio (SNR) temporal scalability.

[0007] Reference Picture Resampling (RPR) In the VVC standard, fast representation switching for adaptive streaming services is expected to deliver multiple representations of the same video content simultaneously, each with different performance. Different performance involves different spatial resolutions or different sample bit depths. In real-time video communication, by allowing resolution changes within an encoded video sequence without inserting I pictures, not only can video data seamlessly adapt to dynamic channel conditions and user preferences, but also the jerking effects caused by I pictures are removed. Reference Picture Resampling (RPR) allows pictures with different resolutions to be referenced from each other in inter-frame prediction. In traditional and streaming applications, RPR provides higher codec efficiency for adaptation of spatial resolution and bit rate. RPR can also be used in application scenarios when scaling the entire video region or some regions of interest is needed. Figure 3 Examples of applying reference picture resampling to encode or decode a current picture are shown, where an inter-frame coded block of the current picture is predicted from a reference picture with the same or different size. When spatial scalability is supported because it is beneficial in streaming applications, the picture size of the reference picture can be different from that of the current picture. RPR is adopted in the VVC standard to support dynamic upsampling and downsampling motion compensation.

[0008] The current picture size and resolution scaling parameters are signaled in the sequence parameter set (SPS), and the maximum picture size in terms of picture width and height for encoding a coded layer video sequence (CLVS) is specified in the SPS. A video picture in a CLVS is a picture associated with the same SPS and in the same layer. Table 1 shows an example of signaling the RPR enable flag and the maximum picture size in the SPS. The RPR enable flag sps_ref_pic_resampling_enabled_flag signaled in the SPS is used to indicate whether RPR is enabled for pictures that refer to the SPS. When this RPR enable flag is equal to 1, the current picture that refers to the SPS may have stripes that refer to a reference picture in an active entry of a reference picture layer, and the reference picture layer has one or more of the following 7 parameters that are different from the current picture. The 7 parameters include the syntax element pps_pic_width_in_luma_samples related to the picture width, the syntax element pps_pic_height_in_luma_samples related to the picture height, the left scaling window offset pps_scaling_win_left_offset, the right scaling window offset pps_scaling_win_right_offset, the top scaling window offset pps_scaling_win_top_offset, the bottom scaling window offset pps_scaling_win_bottom_offset, and the number of sub-pictures sps_num_subpics_minus1. For a current picture that refers to a reference picture, if the reference picture has one or more of these seven parameters different from the current picture, the reference picture may belong to the same layer or a different layer that contains the current picture. The syntax element sps_res_change_in_clvs_allowed_flag equal to 1 specifies that the picture spatial resolution may change within the CLVS that refers to the SPS, and this syntax element equal to 0 specifies that the picture spatial resolution does not change within any CLVS that refers to the SPS. When this syntax element sps_res_change_in_clvs_allowed_flag is not present in the SPS, the value is inferred to be equal to 0. The maximum picture size is signaled in the SPS by the syntax elements sps_pic_width_max_in_luma_samples and sps_pic_height_max_in_luma_samples, and the maximum picture size shall not be greater than the output layer set (OLS) decoded picture buffer (DPB) picture size signaled in the corresponding video parameter set (VPS). Table 1 shows the syntax elements related to RPR signaled in the SPS.

[0009] Table 1

[0010]

[0011] The RPR-related syntax elements for signaling in the PPS are shown in Table 2. The syntax element pic_width_in_luma_samples specifies the width of each decoded picture in the PPS in units of luma samples. The syntax element shall not be equal to 0 and shall be an integer multiple of Max(8, MinCbSizeY), and shall be constrained to be less than or equal to sps_pic_width_max_in_luma_samples signaled in the corresponding SPS. When the sub-picture presence flag subpics_present_flag is equal to 1 or when the RPR enable flag ref_pic_resampling_enabled_flag is equal to 0, the value of this syntax element pic_width_in_luma_samples shall be equal to sps_pic_width_max_in_luma_samples. The syntax element pic_height_in_luma_samples specifies the height of each decoded picture in the PPS in units of luma samples. This syntax element shall not be equal to 0 and shall be an integer multiple of Max(8, MinCbSizeY), and shall be less than or equal to sps_pic_height_max_in_luma_samples. When the sub-picture presence flag subpics_present_flag is equal to 1 or when the RPR enable flag ref_pic_resampling_enabled_flag is equal to 0, the value of the syntax element pic_height_in_luma_samples is set to be equal to sps_pic_height_max_in_luma_samples.

[0012] Table 2

[0013]

[0014] The syntax element scaling_window_flag equal to 1 specifies that the scaling window offset parameter exists in the PPS, and scaling_window_flag equal to 0 specifies that the scaling window offset parameter does not exist in the PPS. When the RPR enable flag ref_pic_resampling_enabled_flag is equal to 0, the value of this syntax element scaling_window_flag will be equal to 0. The syntax elements scaling_win_left_offset, scaling_win_right_offset, scaling_win_top_offset, and scaling_win_bottom_offset specify the scaling offsets in units of luma samples. These scaling offsets are applied to the image dimensions for scale ratio calculation. The scaling offsets can be negative. When the scaling window flag scaling_window_flag is equal to 0, the values of the four scaling offset syntax elements scaling_win_left_offset, scaling_win_right_offset, scaling_win_top_offset, and scaling_win_bottom_offse are inferred to be 0.

[0015] When using RPR to predict the current block in the current picture, the scale ratio is derived from the scaling window dimensions specified by these PPS syntax elements, including scaling_win_left_offset, scaling_win_right_offset, scaling_win_top_offset, and scaling_win_bottom_offset. The variables PicOutputWidthL and PicOutputHeightL representing the scaling window width and the scaling window height are derived as follows:

[0016] PicOutputWidthL = pic_width_in_luma_samples – (scaling_win_right_offset + scaling_win_left_offset).

[0017] PicOutputHeightL = pic_height_in_luma_samples – (scaling_win_bottom_offset + scaling_win_top_offset)

[0018] Derive a variable PicOutputWidthL representing the scaled window width by subtracting the right and left offsets from the image width, and derive the scaled window height by subtracting the top and bottom offsets from the image height. The sum of the values of the left and right offsets scaling_win_left_offset and scaling_win_right_offset will be less than the image width pic_width_in_luma_samples, and the sum of the values of the top and bottom offsets scaling_win_top_offset and scaling_win_bottom_offset will be less than the image height pic_height_in_luma_samples.

[0019] In the latest proposal of the VVC standard, the scaled window offsets are measured in chroma samples, and when these scaled window offset syntax elements are not present in the PPS, the values of the four scaled offset syntax elements scaling_win_left_offset, scaling_win_right_offset, scaling_win_top_offset, and scaling_win_bottom_offset are inferred to be equal to conf_win_left_offset, conf_win_right_offset, conf_win_top_offset, and conf_win_bottom_offset, respectively. As shown below, derive a variable CurrPicScalWinWidthL indicating the scaled window width from the image width, the variable SubWidthC, the left scaled offset, and the right scaled offset. Derive a variable CurrPicScalWinHeightL indicating the scaled window height from the image height, the variable SubHeightC, the top scaled offset, and the bottom scaled offset. For example, derive the scaled window width and the scaled window height by the following equations. CurrPicScalWinWidthL = pic_width_in_luma_samples – SubWidthC * (scaling_win_right_offset + scaling_win_left_offset); and CurrPicScalWinHeightL = pic_height_in_luma_samples – SubHeightC * (scaling_win_bottom_offset + scaling_win_top_offset).

[0020] Sub - picture segmentation In the VVC standard, an image can be further split into one or more sub - pictures for encoding or decoding. The sub - pictures proposed in the VVC standard are similar to the Motion Constrained Tile Sets (MCTS) in the HEVC standard. For applications such as viewport - related 360 - degree video stream optimization and region - of - interest applications, this coding tool allows independent coding and extraction of rectangular subsets of an encoded image sequence. Even when the sub - pictures are extractable, the sub - pictures allow motion vectors that point to coding blocks outside the sub - pictures, thus allowing padding of the sub - picture boundaries in this case. The layout of the sub - pictures of the VVC standard is signaled in the SPS, and thus the layout of the sub - pictures is constrained within the CLVS. Each sub - picture includes one or more complete rectangular stripes. A sub - picture identifier (ID) can optionally be assigned to each sub - picture. Table 3 shows exemplary syntax elements related to sub - picture segmentation signaled in the SPS.

[0021] Table 3

[0022] Summary of the Invention

[0023] In an exemplary embodiment of a video processing method for processing video data, a video encoding or decoding system implementing the video processing method receives input video data related to a video image that references a high - level syntax set, determines a first syntax element signaled or to be signaled in the high - level syntax set, the first syntax element indicating whether the RPR is disabled or constrained for the video image related to the high - level syntax set, determines a second syntax element, the second syntax element indicating whether sub - picture segmentation is disabled or constrained for the video image related to the high - level syntax set, encodes or decodes the video image by allowing reference to images with different resolutions in intra - prediction when the RPR is enabled or allowing reference to only images with the same resolution in inter - prediction when the RPR is disabled, and encodes or decodes the video image by splitting each image into one or more sub - pictures when sub - picture segmentation is enabled or processing each image without splitting into sub - pictures when sub - picture segmentation is disabled. In some embodiments, the first syntax element and the second syntax element are restricted to disabling or constraining sub - picture segmentation when the RPR is enabled. An example of the high - level syntax set is the Sequence Parameter Set. In some embodiments, since the second syntax element is signaled conditionally in the high - level syntax set, i.e., sub - picture segmentation is allowed only when the RPR is disabled, the second syntax element is signaled in the high - level syntax set only when the RPR is disabled.

[0024] In some embodiments, when the sub - picture segmentation is enabled, all video pictures referring to the high - level syntax set have the same derived value for a scaling window width and the same derived value for a scaling window height. The scaling window width of a picture is derived from a picture width signaled in a picture parameter set referred to by the picture, a left - hand scaling window offset, and a right - hand scaling window offset, and the scaling window height of the picture is derived from a picture height, an upper - hand scaling window offset, and a bottom - hand scaling window offset. In the case where the scaling window offset is measured in chroma samples, the scaling window width and height are further derived by variables SubWidthC and SubHeightC respectively.

[0025] According to some embodiments, the first and second syntax elements are restricted to disable or constrain sub - picture segmentation when RPR is enabled or the first and second syntax elements are restricted to disable or constrain RPR when sub - picture segmentation is enabled. In one embodiment, the first syntax element is signaled conditionally in the high - level syntax set and the first syntax element is signaled only when sub - picture segmentation is disabled. In some particular cases, the first syntax element is an RPR - enable flag that specifies whether RPR is enabled, and the second syntax element is a sub - picture - segmentation - presence flag that specifies whether sub - picture parameters are present in the high - level syntax set. In one embodiment, when the sub - picture - segmentation - presence flag is equal to 1, the RPR - enable flag is inferred to be equal to 0 or not signaled, and the RPR - enable flag is signaled in the high - level syntax set only when the sub - picture - segmentation - presence flag is equal to 0. When the RPR - enable flag is not signaled, the RPR - enable flag is inferred to be equal to 0. In another embodiment, when the RPR - enable flag is equal to 1, the sub - picture - segmentation - presence flag is inferred to be equal to 0 or not signaled, and the sub - picture - segmentation - presence flag is signaled in the high - level syntax set only when the RPR - enable flag is equal to 0. When the sub - picture - presence flag is not signaled, the sub - picture - segmentation - presence flag is inferred to be equal to 0.

[0026] In one embodiment, the first syntax element is a resolution - change flag that specifies whether the picture - space resolution is variable in CLVS, and the second syntax element is a sub - picture - segmentation - presence flag that specifies whether sub - picture information is present for CLVS in the high - level syntax set. In this embodiment, when the resolution - transform flag is equal to 1, the sub - picture - segmentation - presence flag is inferred to be equal to 0, and when the picture - space resolution is variable within CLVS referring to the high - level syntax set, sub - picture information is not present for CLVS in the high - level syntax set.

[0027] An embodiment of the video processing method further includes determining a third syntax element according to the second syntax element, where the second syntax element is a sub - picture segmentation presence flag that specifies whether sub - picture information exists for a coded layer video sequence in a high - level syntax set, and the third syntax element is a sub - picture identifier flag that specifies whether a sub - picture identifier mapping exists in the high - level syntax element. The video processing method determines a relevant sub - picture layout when the sub - picture segmentation presence flag is equal to 1. When the sub - picture segmentation presence flag is equal to 0, the sub - picture identifier flag is not coded and is inferred to be 0, and when the sub - picture segmentation presence flag is equal to 1, the sub - picture identifier flag is only signaled in the high - level syntax set.

[0028] One aspect of the present invention further provides a video processing apparatus in a video encoding or decoding system. The apparatus includes one or more electronic circuits for receiving input video data of a video image that references a high - level syntax set, determining whether a first syntax element signaled or to be signaled in the high - level syntax set indicates whether RPR is disabled or constrained, determining a second syntax element indicating whether sub - picture segmentation is disabled or constrained, and encoding or decoding the video image by allowing reference to images with different resolutions in inter - frame prediction when RPR is enabled, or by splitting each image into one or more sub - pictures when sub - picture segmentation is enabled. Embodiments of the video encoding or decoding system encode or decode video images according to bit - stream consistency requirements, which disable or constrain sub - pictures when RPR is enabled or disable or constrain RPR when sub - picture segmentation is enabled.

[0029] One aspect of the present invention further provides a non - transitory computer - readable medium storing program instructions such that a processing circuit of a device executes a video processing method to encode or decode a video image that references a high - level syntax set. The video processing method encodes or decodes the video image by disabling or constraining sub - picture segmentation when RPR is enabled or by disabling or constraining RPR when sub - picture segmentation is enabled. After reading the following description of specific embodiments, other aspects and features of the present invention will be apparent to those of ordinary skill in the art. BRIEF DESCRIPTION OF THE DRAWINGS

[0030] Various embodiments of the present invention presented by way of example will be described in detail with reference to the following drawings, and in which:

[0031] Figure 1 A system block diagram of a video encoding system showing various encoding tools in the implementation of the HEVC standard is shown.

[0032] Figure 2A An example of block partitioning within a coding tree unit according to the HEVC standard is shown.

[0033] Figure 2Bshows a quadtree - based coding tree representation corresponding to the Figure 2A coding tree units shown in

[0034] Figure 3 shows an exemplary example enabling reference picture resampling.

[0035] Figure 4 shows an exemplary flowchart of a video encoding or decoding system of an embodiment disabling sub - picture segmentation when reference picture resampling is enabled.

[0036] Figure 5 shows an exemplary flowchart of a video encoding or decoding system of an embodiment disabling reference picture resampling when sub - picture segmentation is enabled.

[0037] Figure 6 shows an exemplary system block diagram of a video encoding system for a merging video processing method according to an embodiment of the present invention.

[0038] Figure 7 shows an exemplary system block diagram of a video decoding system for a merging video processing method according to an embodiment of the present invention. Detailed Description

[0039] It will be readily understood that the elements of the present invention as described and shown in the figures herein can be arranged and designed in a variety of different configurations. Thus, the following detailed description of the present invention as represented in the figures is not intended to limit the scope of the present invention, but is merely representative of selected embodiments of the present invention.

[0040] For the same picture size and the same scaling window, when sub - picture segmentation is enabled, it requires bitstream consistency. All pictures referring to the same SPS have the same picture size as the maximum picture size specified in the SPS. However, when all pictures referring to the same SPS have the same picture size but different spatial scaling parameters, the coding parameters related to the signaling spatial position and segmentation (sub - picture layout specified in the SPS) may not be consistent with the pictures in the encoded video sequence. In some embodiments of the present invention, when the two pictures have the same picture size, the value of the scaling parameter is constrained to be the same for any two pictures referring to the same SPS. In this way, all pictures referring to the same SPS and having the same picture size will have the same derived values of the scaling window width PicOutputWidthL and the scaling window height PicOutputHeightL. The proposed modification related to the semantics of the picture parameter set RBSP according to an embodiment is provided as follows.

[0041] The syntax element scaling_window_flag being equal to 1 specifies that the scaling window offset parameters are present in the PPS, and scaling_window_flag being equal to 0 specifies that the scaling window offset parameters are not present in the PPS. When the syntax element ref_pic_resampling_enabled_flag is equal to 0, the value of scaling_window_flag shall be equal to 0. The scaling window offset parameters scaling_win_left_offset, scaling_win_right_offset, scaling_win_top_offset, and scaling_win_bottom_offset specify the offsets in units of luma samples, which are applied to the picture size for the scaling ratio calculation. When the syntax element scaling_window_flag is equal to 0, the values of these four scaling window offset parameters are inferred to be 0. The sum of the left and right scaling window offsets of the picture shall be less than the picture width pic_width_in_luma_samples, and the sum of the top and bottom scaling window offsets shall be less than the picture height pic_height_in_luma_samples. The variable scaling window width PicOutputWidthL and scaling window height PicOutputHeightL are derived as follows: PicOutputWidthL = pic_width_in_luma_samples – (scaling_win_right_offset + scaling_win_left_offset), and PicOutputHeightL = pic_height_in_luma_samples – (scaling_win_bottom_offset + scaling_win_top_offset). Suppose ppsA and ppsB are any two PPSs that refer to the same SPS. When ppsA and ppsB have the same values of pic_width_in_luma_samples and pic_height_in_luma_samples respectively, bitstream compliance requires that ppsA and ppsB shall have the same values of scaling_win_left_offset, scaling_win_right_offset, scaling_win_top_offset, and scaling_win_bottom_offset respectively.

[0042] Same Scaling Window When Certain Tools are Enabled In some embodiments, the same scaling window constraint is applied when at least one of certain sets of assigned tools is enabled. In one example, sub-picture segmentation is one of the certain sets of assigned tools, such that the same scaling window constraint is applied when sub-picture segmentation is enabled. Thus, when the sub-picture presence flag subpics_present_flag is equal to 1, all pictures that refer to the same SPS will each have the same derived values for the scaling window width PicOutputWidthL and the scaling window height PicOutputHeightL. In one embodiment, when sub-picture segmentation is enabled, the same scaling constraint limits all pictures that refer to the same SPS to have the same scaling window offset parameters scaling_win_left_offset, scaling_win_right_offset, scaling_win_top_offset, and scaling_win_bottomg_offset. The proposed modification to the semantics related to the set of picture parameters RBSP is as follows. The syntax element scaling_window_flag equal to 1 specifies that the scaling window offset parameters are present in the PPS, and scaling_window_flag equal to 0 specifies that the scaling window offset parameters are not present in the PPS. When the reference picture resampling enable flag ref_pic_resampling_enabled_flag is equal to 0, the value of scaling_window_flag will be equal to 0. The scaling window offset parameters scaling_win_left_offset, scaling_win_right_offset, scaling_win_top_offset, and scaling_win_bottom_offset specify the scaling window offset in units of luma samples, which is applied to the picture size for the scaling ratio calculation. When the syntax element scaling_window_flag is equal to 0, the values of scaling_win_left_offset, scaling_win_right_offset, scaling_win_top_offset, and scaling_win_bottom_offset are inferred to be equal to 0. The sum of the left and right scaling window offsets of the picture will be less than the picture width pic_width_in_luma_samples, and the sum of the top and bottom scaling window offsets will be less than the picture height pic_height_in_luma_samples.The variable scaled window width PicOutputWidthL and the scaled window height PicOutputHeightL are derived as follows: PicOutputWidthL = pic_width_in_luma_samples – (scaling_win_right_offset + scaling_win_left_offset), and PicOutputHeightL = pic_height_in_luma_samples – (scaling_win_bottom_offset + scaling_win_top_offset). Assume that ppsA and ppsB are any two PPSs that refer to the same SPS. In this embodiment, when the subpicture present flag in the SPS is signaled as equal to 1, ppsA and ppsB need to have bitstream consistency for the same values of scaling_win_left_offset, scaling_win_right_offset, scaling_win_top_offset, and scaling_win_bottom_offset, respectively.

[0043] Disable reference picture resampling when certain tools are enabled. In some embodiments, when certain assigned groups of coding tools are enabled, the video encoder or decoder disables or restricts the use of RPR. In some other equivalent embodiments, when RPR is enabled, the video encoder or decoder disables or restricts the use of certain assigned groups of tools. In a particular embodiment, certain assigned groups of tools include sub-picture segmentation, and when sub-picture segmentation is enabled, the video encoder or decoder encodes or decodes video data according to bitstream consistency requirements (which disable RPR). For example, when the sub-picture segmentation present flag subpics_present_flag is equal to 1, the syntax element related to the RPR enable flag ref_pic_resampling_enabled_flag signaled in the SPS will be equal to 0. Thus, since RPR is disabled when sub-picture segmentation is enabled, all pictures referring to the same SPS will respectively have the same derived values of the scaled window width PicOutputWidthL and the scaled window height PicOutputHeightL. In an alternative embodiment, when RPR is enabled, the video encoder or decoder encodes or decodes video data according to bitstream consistency requirements (which disable sub-picture segmentation). For example, when the RPR enable flag ref_pic_resampling_enabled_flag is equal to 1, the video encoder or decoder infers the value of the sub-picture segmentation present flag subpics_present_flag to be equal to 0. Table 4 shows an example of a modified syntax table of the syntax elements signaled in the SPS. In this example, the sub-picture segmentation present flag subpics_present_flag is signaled in the SPS only when the RPR enable flag ref_pic_resampling_enabled_flag is equal to 0, i.e., sub-picture segmentation is allowed only when RPR is disabled. When this flag is not signaled in the SPS, the value of the sub-picture segmentation present flag subpics_present_flag is inferred to be 0. The sub-picture segmentation present flag equal to 1 specifies that sub-picture related parameters are present in the SPS, and the sub-picture segmentation present flag equal to 0 specifies that sub-picture related parameters are not present in the SPS.

[0044] Table 4

[0045]

[0046] Table 5 shows the syntax table of the SPS of another embodiment that disables sub-picture segmentation when RPR is enabled. In this embodiment, the signaling resolution change flag sps_res_change_in_clvs_allowed_flag in the SPS specifies that the picture spatial resolution is variable within the coded layer video sequence (CLVS) of the reference SPS. When this resolution change flag sps_res_change_in_clvs_allowed_flag is equal to 0, the picture spatial resolution does not change within any CLVS of the reference SPS, and when the resolution change flag sps_res_change_in_clvs_allowed_flag is equal to 1, the picture resolution can change within the CLVS of the reference SPS. The sub-picture segmentation present flag sps_subpic_info_present_flag being equal to 1 specifies that sub-picture information exists for the CLVS and there may be one or more sub-pictures in each picture of the CLVS, and sps_subpic_info_present_flag being equal to 0 specifies that sub-picture information does not exist for the CLVS and there is only one sub-picture in each picture of the CLVS. In this embodiment, when the resolution change flag sps_res_change_in_clvs_allowed_flag is equal to 1, the value of the sub-picture segmentation present flag sps_subpic_info_present_flag will be equal to 0. In other words, when the picture spatial resolution is variable within the CLVS of the reference SPS, sub-picture segmentation is restricted such that sub-picture information cannot exist for the CLVS in the SPS. In one embodiment, when RPR is enabled, the sub-picture segmentation present flag is allowed to be set to 1 only when the resolution change flag sps_res_change_in_clvs_allowed_flag is equal to 0. Similarly, when sub-picture segmentation is enabled, the RPR enable flag is allowed to be set to 1 only when the resolution change flag sps_res_change_in_clvs_allowed_flag is equal to 0.

[0047] Table 5

[0048]

[0049]

[0050] The syntax element sps_subpic_id_len_minus1 plus 1 specifies the number of bits used to represent the syntax element sps_subpic_id[i], the syntax element pps_subpic_id[i], and the syntax element sh_subpic_id. The value of sps_subpic_id_len_minus1 shall be in the range including 0 to 15. The value of 1<<(sps_subpic_id_len_minus1+1) shall be greater than or equal to sps_num_subpics_minus1+1. The syntax element sps_subpic_id_mapping_explicitly_signalled_flag being equal to 1 specifies that the sub-picture identifier (ID) mapping is signalled explicitly in the SPS referenced by the coded pictures of the CLVS or in the PPS. The syntax element being equal to 0 specifies that the sub-picture ID mapping is not signalled explicitly for the CLVS, and the value of this syntax element shall be inferred as 0 when not present. The syntax element sps_subpic_id_mapping_present_flag being equal to 1 specifies that the sub-picture ID mapping is signalled in the SPS when sps_subpic_id_mapping_explicilty_signalled_flag is equal to 1. When sps_subpic_id_mapping_explicitly_signalled_flag is equal to 1, this syntax element sps_subpic_id_mapping_present_flag being equal to 0 specifies that the sub-picture ID mapping is signalled in the PPS of the coded pictures of the CLVS. The syntax element sps_subpic_id[i] specifies the sub-picture ID of the i-th sub-picture. The length of the syntax element sps_subpic_id[i] is sps_subpic_id_len_minus1+1 bits.

[0051] Inferring Sub - picture ID Flags When Sub - picture Segmentation is Disabled In some embodiments of the present invention, a video encoder or decoder encodes or decodes a first syntax flag in a high - level syntax set that indicates whether sub - picture segmentation is enabled or disabled. For example, the first syntax element is the sub - pictures_present_flag signaled in the SPS. When the first syntax flag equals 1, the relevant sub - picture layout is signaled; otherwise, the relevant sub - picture information is not signaled. The video encoder or decoder may further encode or decode a second syntax flag that indicates whether information related to the sub - picture ID will be further signaled. In some embodiments of the present invention, when the sub - pictures_present_flag equals 0, the second syntax flag is not encoded / decoded and is inferred to be 0. An example of the second syntax flag is the sps_subpic_id_present_flag signaled in the SPS. The sps_subpic_id_present_flag equal to 1 specifies that the sub - picture ID mapping exists in the SPS, and the sps_subpic_id_present_flag equal to 0 specifies that the sub - picture ID mapping does not exist in the SPS. The exemplary syntax table shown in Table 6 represents an embodiment of inferring sub - picture ID flags when sub - picture segmentation is disabled. In this embodiment, the sps_subpic_id_present_flag is signaled in the SPS only when the sub - pictures_present_flag equals 1, which implies that the sps_subpic_id_present_flag does not exist in the SPS when the sub - pictures_present_flag equals 0. When it does not exist, the sps_subpic_id_present_flag is inferred to be equal to 0.

[0052] Table 6

[0053]

[0054]

[0055] Figure 4 and Figure 5 The exemplary flowchart of shows an exemplary flowchart of the mutually exclusive use of RPR and sub - picture segmentation. As Figure 4As shown, in this embodiment, the video encoding or decoding system disables sub-image segmentation when enabling RPR. In step S402, the video encoding or decoding system receives input video data related to the video image of the reference SPS, and in step S404, determines whether RPR is enabled in the video image of the reference SPS. For example, the video encoding system signals a first syntax element in the SPS to specify whether reference image resampling is enabled or disabled, and by parsing the first syntax element signaled in the SPS, the video decoding system determines whether RPR is enabled or disabled. In step 406, when RPR is determined to be enabled for the video image in step S404, sub-image segmentation is disabled for the video image of the reference SPS. For example, when the first syntax element indicates that RPR is enabled, the video encoding system will not signal a second syntax element to specify whether sub-image segmentation is enabled or disabled, or the video decoding system infers that the second syntax element is 0, indicating that sub-image segmentation is disabled. In step S408, the video encoding or decoding system encodes or decodes the video data of the video image by allowing reference to images with different resolutions in inter-frame prediction and by processing each image without splitting it into sub-images. In step S410, when it is determined in step S404 that RPR is prohibited from being used to encode or decode the video image, the video encoding or decoding further determines whether to allow sub-image segmentation for the video image of the reference SPS. For example, the video decoding system signals a second syntax element to specify whether sub-image segmentation is enabled or disabled, or the video decoding system determines whether sub-image segmentation is enabled or disabled by parsing the second syntax element from the SPS. In the case where it is determined in step S410 that sub-image segmentation is enabled, in step S412, the video encoding or decoding system encodes or decodes the video data in the video image by allowing only images with the same resolution to be referenced in inter-frame prediction and by splitting each image into one or more sub-images. Otherwise, when both RPR and sub-image segmentation are disabled, in step S414, the video encoding or decoding system encodes or decodes the video data in the video image by allowing only images with the same resolution to be referenced in inter-frame prediction and by processing each image without splitting it into sub-images.

[0056] Figure 5Exemplary flowchart of a video encoding or decoding system showing an embodiment of disabling RPR when sub-image segmentation is enabled. In step S502, the video encoding or decoding system receives input video data of a video image of a reference SPS, and in step S504, determines whether sub-image segmentation is enabled in the video image of the reference SPS. For example, by parsing the second syntax element signaled in the SPS, the video decoding system determines whether sub-image segmentation is enabled or disabled. When sub-image segmentation is enabled, in step S506, RPR is disabled for the video image of the reference SPS. For example, when sub-image segmentation is enabled, the video encoding system will not signal in the SPS the first syntax element specifying whether RPR is enabled or disabled, or when the second syntax element specifies that sub-image segmentation is enabled, the video decoding system infers that this first syntax element is 0, specifying that RPR is disabled. In step S508, the video encoding or decoding system encodes or decodes the video data in the video image of the reference SPS by only allowing references to images with the same resolution in inter-frame prediction and by splitting each image into one or more sub-images. In step S510, the video encoding or decoding system further determines whether RPR is enabled in the video image of the reference SPS. For example, the video encoding system signals in the SPS the first syntax element specifying whether RPR is enabled or disabled, or the video decoding system determines whether RPR is enabled or disabled by parsing the first syntax element from the SPS. When RPR is determined to be enabled, in step 512, the video data in the video image of the reference SPS is encoded or decoded by allowing references to images with different resolutions in inter-frame prediction and by processing each image without splitting it into sub-images. Since both RPR and sub-image segmentation are disabled, in step S514, the video data in the video image is encoded or decoded by only allowing references to images with the same resolution in inter-frame prediction and by processing each image without splitting it into sub-images.

[0057] Video encoder and decoder implementations The previously proposed video processing methods related to advanced information signaling can be implemented in a video encoder or decoder. Figure 6Exemplary system block diagram of a video encoder 600 implementing one or more of the various embodiments of the present invention is shown. According to some embodiments, the requirement for bitstream consistency is that RPR and sub-picture segmentation cannot be enabled simultaneously. For example, when RPR is enabled, the video encoder signals a first syntax element specifying that RPR is enabled in the SPS, and skips signaling a second syntax element specifying whether sub-picture segmentation is disabled or constrained. When the second syntax element does not exist in the SPS, it is inferred that sub-picture segmentation is disabled. In another example, when sub-picture segmentation is enabled, the video encoder signals a second syntax element specifying that sub-picture segmentation is enabled in the SPS, and skips signaling a first syntax element specifying whether RPR is disabled or constrained. When the first syntax element does not exist in the SPS, it is inferred that RPR is disabled. In some other embodiments, the requirement for bitstream consistency is that sub-picture segmentation is constrained when RPR is enabled, or RPR is constrained when sub-picture segmentation is enabled. The block structure segmentation 610 receives input data and performs sub-picture block segmentation. The intra prediction module 612 provides an intra prediction signal based on the reconstructed video data of the current picture. The inter prediction module 614 performs motion estimation (ME) and motion compensation (MC) based on the video data from one or more other pictures to provide an inter prediction signal. The intra prediction module 612 or the inter prediction 614 provides the selected prediction signal to the adder 618 to form a prediction error, also known as a prediction residual. The prediction residual of the current block is further processed by a transform module (T) 620, followed by a quantization module (Q) 622. Then, the transformed and quantized residual signal is encoded by the entropy encoder 634 to form a video bitstream. The video bitstream is then packed together with auxiliary information. Then, the transformed and quantized residual signal of the current block is processed by an inverse quantization module (IQ) 624 and an inverse transform module (IT) 626 to recover the prediction residual. As Figure 6 shown, the prediction residual is recovered by adding back the selected prediction signal at the reconstruction module (REC) 628 to produce reconstructed video data. The reconstructed video data can be stored in a reference picture buffer (Ref.Pict.Buffer) 632 and used for prediction of other pictures. The reconstructed video data recovered from the REC module 628 may be subject to various impairments due to the encoding process. Therefore, an in-loop processing filter 630 is applied to the reconstructed video data before it is stored in the reference picture buffer 632 to further enhance the image quality.

[0058] Figure 7 shown for decoding from Figure 6The corresponding video decoder 700 for the video bitstream generated by the video encoder 600 shown. In some examples of the present invention, a first syntax element is signaled in a high-level syntax set to indicate whether RPR is disabled or constrained, and a second syntax element is signaled in the high-level syntax set to indicate whether sub-image segmentation is disabled or constrained. For example, the first syntax element is an RPR enable flag, and the second syntax element is a sub-image segmentation presence flag signaled in the SPS. In one embodiment, when the sub-image segmentation presence flag is equal to 1, the RPR enable flag is inferred to be 0, while in another embodiment, when the RPR enable flag is equal to 1, the sub-image segmentation presence flag is inferred to be 0. In yet another embodiment, the first syntax element is a resolution transformation flag and the second syntax element is a sub-image segmentation presence flag, where the sub-image segmentation presence flag is constrained such that it is only allowed to be set to 1 when the resolution transformation flag is equal to 0. The video bitstream is an input to the video decoder 700 and is decoded by an entropy decoder 710 to parse and recover the transformed and quantized residual signals and other system information. The decoding process of the decoder 700 is similar to the reconstruction loop at the encoder 600, except that the decoder 700 only requires motion compensation prediction in the inter-frame prediction 716. The block structure segmentation 712 receives the input data and performs sub-image block segmentation. Each block is decoded by an intra-frame prediction module 714 or an inter-frame prediction module 716. A switch 718 selects an intra-frame prediction sub from the intra-frame prediction module 714 or an inter-frame prediction sub from the inter-frame prediction module 716 according to the decoded mode information. The transformed and quantized residual signal associated with each block is recovered by an inverse quantization module (IQ) 722 and an inverse transform module (IT) 724. The recovered residual signal is reconstructed by adding the prediction sub in the REC module 720 to produce a reconstructed video. The reconstructed video is further processed by an in-loop processing filter (filter) 726 to generate a final decoded video. If the currently decoded image is a reference image for subsequent images in the decoding order, the reconstructed video of the currently decoded image is also stored in the reference image buffer 728.

[0059] Figure 6 and Figure 7The various components of the mid - video encoder 600 and the video decoder 700 can be implemented by hardware components, one or more processors configured to execute program instructions stored in a memory, or a combination of hardware and processors. For example, the processor executes program instructions to control the reception of input data associated with video images that reference a high - level syntax set. The processor is equipped with a single or multiple processing cores. In some instances, the processor executes program instructions to perform functions in some of the components of the encoder 600 and the decoder 700, and the memory electrically coupled to the processor is used to store program instructions, information corresponding to the reconstructed images of blocks, and / or intermediate data during the encoding or decoding process. In some embodiments, the memory includes a non - transitory computer - readable medium, such as semiconductor or solid - state memory, random access memory (RAM), read - only memory (ROM), hard disk, optical disk, or other suitable storage medium. The memory can also be a combination of two or more of the non - transitory computer - readable media listed above. As Figure 6 and Figure 7 shown, the encoder 600 and the decoder 700 can be implemented in the same electronic device. Thus, if implemented in the same electronic device, the various functional components of the encoder 600 and the decoder 700 can be shared or reused.

[0060] Embodiments of the video processing method for encoding or decoding can be implemented in a circuit integrated into a video compression chip or in code integrated into video compression software to perform the above - mentioned processing. For example, encoding or decoding that utilizes sub - image segmentation or reference - image resampling can be implemented in code executed on a computer processor, a digital signal processor (DSP), a microprocessor, or a field - programmable gate array (FPGA). These processors can be configured to perform specific tasks according to the present invention by executing machine - readable software code or firmware code that defines the specific methods embodied by the present invention.

[0061] References in this specification to "one embodiment", "some embodiments", or similar language mean that a particular feature, structure, or characteristic described in connection with the embodiment can be included in at least one embodiment of the present invention. Thus, the phrases "in one embodiment" or "in some embodiments" that appear throughout this specification do not necessarily all refer to the same embodiment, and these embodiments can be implemented individually or in combination with one or more other embodiments. Additionally, the described features, structures, or characteristics can be combined in any suitable manner in one or more embodiments. However, those of ordinary skill in the relevant art will recognize that the present invention can be practiced without one or more of the specific details, or with other methods, components, etc. In other cases, well - known structures or operations are not shown or described in detail to avoid obscuring aspects of the present invention.

[0062] Without departing from its spirit or essential characteristics, the present invention may be embodied in other specific forms. The described embodiments are to be considered in all respects only as illustrative and not restrictive. Thus, the scope of the present invention is indicated by the appended claims rather than by the foregoing description. All changes which come within the meaning and range of equivalency of the claims are to be embraced within their scope.

Claims

1. A video processing method in a video encoding or decoding system, comprising: receiving input video data of a video image that refers to a high-level syntax set; determining a first syntax element that is signaled or to be signaled in the high-level syntax set, where the first syntax element is a resolution change flag that specifies whether the image spatial resolution can vary within the encoded layer video sequence; determining a second syntax element that indicates whether sub-image segmentation is disabled or restricted for the video image related to the high-level syntax set, where the first syntax element and the second syntax element are restricted to disable or restrict the sub-image segmentation when the resolution change is enabled, and encoding or decoding the video image by allowing reference to images with different resolutions in inter-frame prediction when reference image resampling is enabled, or allowing reference to only images with the same resolution in inter-frame prediction when the reference image resampling is disabled, and processing each image by splitting each image into one or more sub-images when the sub-image segmentation is enabled, or not splitting each image into sub-images when the sub-image segmentation is disabled.

2. The video processing method in a video encoding or decoding system according to claim 1, wherein, the high-level syntax set is a sequence parameter set.

3. The video processing method in a video encoding or decoding system according to claim 1, wherein, the second syntax element is conditionally signaled in the high-level syntax set and the second syntax element is only signaled when the reference image resampling is disabled.

4. The video processing method in a video encoding or decoding system according to claim 1, wherein, when the sub-image segmentation is enabled, all video images referring to the high-level syntax set have the same derived value for a scaled window width and the same derived value for a scaled window height.

5. The video processing method in a video encoding or decoding system according to claim 4, wherein, the scaled window width of the image is derived from an image width, a left scaled window offset, and a right scaled window offset signaled in an image parameter set referred to by the image, and the scaled window height of the image is derived from an image height, an upper scaled window offset, and a bottom scaled window offset.

6. The video processing method in a video encoding or decoding system according to claim 1, wherein, when reference image resampling is enabled, the second syntax element is restricted to disable or restrict the sub-image segmentation, or when the sub-image segmentation is enabled, a fourth syntax element is restricted to disable or restrict the reference image resampling.

7. The video processing method in a video encoding or decoding system according to claim 6, wherein, the first syntax element is conditionally signaled in the high-level syntax set and the first syntax element is only signaled when the sub-image segmentation is disabled.

8. The video processing method in a video encoding or decoding system according to claim 6, wherein, The fourth syntax element is a reference picture resampling enable flag that specifies whether the reference picture resampling is enabled, and the second syntax element is a sub-picture segmentation presence flag that specifies whether sub-picture parameters exist in the high-level syntax set.

9. The video processing method in the video encoding or decoding system according to claim 8, wherein, when the sub-picture segmentation presence flag is equal to 1, the reference picture resampling enable flag is inferred to be equal to 0 or not signaled, and when the sub-picture segmentation presence flag is equal to 0, the reference picture resampling enable flag is only signaled in the high-level syntax set, wherein when the reference picture resampling enable flag is not signaled, the reference picture resampling enable flag is inferred to be 0.

10. The video processing method in the video encoding or decoding system according to claim 8, wherein, when the reference picture resampling enable flag is equal to 1, the sub-picture segmentation presence flag is inferred to be equal to 0 or not signaled, and when the reference picture resampling enable flag is equal to 0, the sub-picture segmentation presence flag is only signaled in the high-level syntax set, wherein when the sub-picture segmentation presence flag is not signaled, the sub-picture segmentation presence flag is inferred to be 0.

11. The video processing method in the video encoding or decoding system according to claim 1, wherein, and the second syntax element is a sub-picture segmentation presence flag that specifies whether sub-picture information exists for the coded layer video sequence in the high-level syntax set.

12. The video processing method in the video encoding or decoding system according to claim 11, wherein, by inferring the sub-picture segmentation presence flag to be 0 when the resolution change flag is equal to 1, when the picture spatial resolution is variable in the coded layer video sequence referring to the high-level syntax set, the sub-picture information does not exist for the coded layer video sequence in the high-level syntax set.

13. The video processing method in the video encoding or decoding system according to claim 1, wherein, further comprising determining a third syntax element according to the second syntax element, wherein the second syntax element is a sub-picture segmentation presence flag that specifies whether sub-picture information exists for a coded layer video sequence in the high-level syntax set, and the third syntax element is a sub-picture identifier flag that specifies whether a sub-picture identifier mapping exists in the high-level syntax element.

14. The video processing method in the video encoding or decoding system according to claim 13, wherein, further comprising determining a related sub-picture layout when the sub-picture segmentation presence flag is equal to 1.

15. The video processing method in the video encoding or decoding system according to claim 13, wherein, when the sub-picture segmentation presence flag is equal to 0, the sub-picture identifier flag is not coded and is inferred to be 0, and when the sub-picture segmentation presence flag is equal to 1, the sub-picture identifier flag is only signaled in the high-level syntax set.

16. An apparatus for processing video data in a video encoding or decoding system, the apparatus comprising one or more electronic circuits for: receiving input video data of a video image with reference to a high-level syntax set; determining a first syntax element signaled or to be signaled in the high-level syntax set, the first syntax element being a resolution change flag that specifies whether the image spatial resolution is variable within an encoded layer video sequence; determining a second syntax element that indicates whether sub-image segmentation is disabled or constrained for the video image associated with the high-level syntax set, wherein the first syntax element and the second syntax element are restricted to disabling or constraining the sub-image segmentation when the resolution change is enabled, and encoding or decoding the video image by allowing reference to images with different resolutions in inter-frame prediction when reference image resampling is enabled, or allowing reference to only images with the same resolution in inter-frame prediction when the reference image resampling is disabled, and processing each image by splitting each image into one or more sub-images when the sub-image segmentation is enabled, or processing each image without splitting each image into sub-images when the sub-image segmentation is disabled.

17. A non-transitory computer-readable medium storing program instructions that cause a processing circuit of a device to perform a video processing method of video data, and the method comprises: receiving input video data of a video image with reference to a high-level syntax set; determining a first syntax element signaled or to be signaled in the high-level syntax set, the first syntax element being a resolution change flag that specifies whether the image spatial resolution is variable within an encoded layer video sequence; determining a second syntax element that indicates whether sub-image segmentation is disabled or constrained for the video image associated with the high-level syntax set, wherein the first syntax element and the second syntax element are restricted to disabling or constraining the sub-image segmentation when the resolution change is enabled, and encoding or decoding the video image by allowing reference to images with different resolutions in inter-frame prediction when reference image resampling is enabled, or allowing reference to only images with the same resolution in inter-frame prediction when the reference image resampling is disabled, and by splitting each image into one or more sub-images when the sub-image segmentation is enabled, or processing each image without splitting each image into sub-images when the sub-image segmentation is disabled to encode or decode the video image.