Scaling window in video coding

CN115699756BActive Publication Date: 2026-09-15DOUYIN VISION CO LTD +1
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202180036537.1
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Priority Date
2020-05-21
Filing Date
2021-05-21
Publication Date
2026-09-15
Estimated Expiration
2041-05-21

Smart Images

  • Figure CN115699756B_ABST
    Figure CN115699756B_ABST
Patent Text Reader

Abstract

Techniques of video processing, including video coding, video decoding, and video transcoding, are described. One example method includes performing a conversion between a video comprising video pictures and a bitstream of the video according to a rule, wherein the rule specifies that a syntax element indicates a first width and a first height of a scaling window of a video picture, and wherein the rule specifies that a range of allowed values of the syntax element includes values that are greater than or equal to twice a second width of the video picture and twice a second height of the video picture.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] Cross-references to related applications

[0002] This application is based on International Patent Application No. PCT / CN2021 / 095012, filed on May 21, 2021, which claims priority and interest in International Patent Application No. PCT / CN2020 / 091533, filed on May 21, 2020. All of the aforementioned patent applications are incorporated herein by reference in their entirety. Technical Field

[0003] The patent document relates to image and video encoding and decoding. Background Technology

[0004] Digital video accounts for the largest share of bandwidth usage on the Internet and other digital communication networks. As the number of connected user devices capable of receiving and displaying video increases, the bandwidth demand for digital video is expected to continue to grow. Summary of the Invention

[0005] This document discloses techniques that can be used by video encoders and decoders to process the codec representation of video using control information useful for decoding the codec representation.

[0006] In one example aspect, a video processing method is disclosed. The method includes performing a conversion between a video comprising video images and a bitstream of the video according to rules, wherein the rules specify that syntax elements indicate a first width and a first height of a scaling window for the video images, and wherein the rules specify that the range of allowed values ​​for the syntax elements includes values ​​greater than or equal to twice the second width and twice the second height of the video images.

[0007] In another example, a video processing method is disclosed. The method includes a conversion between a video image and a bitstream of a video, determining (1) whether or how a first width or first height of a first scaled window of a reference video image and (2) a second width or second height of a second scaled window of the current video image are constrained according to rules; and performing the conversion based on the determination.

[0008] In another example, a video processing method is disclosed. This method includes performing a conversion between a video and a video bitstream comprising video images containing one or more video stripes, wherein the bitstream conforms to a rule specifying the values ​​of a variable determining the width and height of a scaling window for the video images, the variable specifying whether resampling of the j-th reference video image in a list of i-th reference images is enabled for the video stripes of the video images, where i and j are integers.

[0009] In another example, a video processing method is disclosed. The method includes performing a conversion between a video and a video bitstream comprising one or more video images according to format rules, wherein the format rules specify that syntax elements are included in a sequence parameter set, and wherein the syntax elements indicate whether reference image resampling is enabled for a reference image and whether one or more slices of the current video image in the codec layer video sequence are allowed to reference a reference image in an active entry of a reference image list, and wherein the reference image has any one or more of the following six parameters that differ from the current video image: 1) video image width, 2) video image height, 3) left offset of the scaling window, 4) right offset of the scaling window, 5) top offset of the scaling window, and 6) bottom offset of the scaling window.

[0010] In another example, a video processing method is disclosed. The method includes performing a conversion between a video and a video bitstream comprising one or more video images according to format rules, wherein the format rules specify that syntax elements are included in a sequence parameter set, wherein the format rules specify that the value of the syntax element is based on (1) whether the current video layer of the reference sequence parameter set is not an independent video layer and (2) information about one or more reference video layers associated with the current video layer, wherein a video layer is an independent video layer when it does not depend on one or more other video layers, and wherein a video layer is not an independent video layer when it depends on one or more other video layers.

[0011] In another example, a video processing method is disclosed. This method includes performing a conversion between a video and a video bitstream comprising one or more video images according to format rules, wherein the format rules specify syntax elements in a set of sequence parameters referenced by the one or more video images; and wherein the syntax elements indicate whether reference image resampling is enabled for one or more reference images in the same layer as the one or more video images.

[0012] In another example, a video processing method is disclosed. The method includes performing a conversion between a video comprising one or more video images and a codec representation of the video, wherein the codec representation conforms to a format rule; wherein the format rule specifies that one or more syntax elements contained in the codec representation associated with a scaling window are allowed to have values ​​indicating that the height or width of the scaling window is greater than or equal to the height or width of the corresponding video image.

[0013] In another example, a different video processing method is disclosed. This method includes performing a conversion between a video and a video codec representation comprising one or more video images in one or more video layers, wherein the codec representation conforms to format rules; wherein the format rules specify that the allowed values ​​of one or more syntax elements included in the codec representation associated with a scaling window are subject to constraint rules, wherein the constraint rules depend on the relationship between a first layer of the current image and a second layer of a reference image of the current image.

[0014] In another example, a different video processing method is disclosed. This method includes performing a conversion between a video and a codec representation of a video comprising one or more video layers containing one or more video images, wherein the codec representation conforms to a format rule specifying that syntax elements are included in the codec representation in a parameter set, wherein the syntax elements indicate whether reference image resampling is enabled for non-independent video layers, and wherein the value of the syntax element is a function of the reference layer of the non-independent video layer.

[0015] In yet another example, a video encoder apparatus is disclosed. The video encoder includes a processor configured to implement the methods described above.

[0016] In yet another example, a video decoder apparatus is disclosed. The video decoder includes a processor configured to implement the methods described above.

[0017] In yet another example, a computer-readable medium on which code is stored is disclosed. This code embodies one of the methods described herein in the form of processor-executable code.

[0018] These and other features are described throughout this document. Attached Figure Description

[0019] Figure 1 An example of raster scan strip segmentation of an image is shown, where the image is divided into 12 slices and 3 raster scan strips.

[0020] Figure 2 The example shown is an image divided into rectangular strips, where the image is divided into 24 slices (6 slice columns and 4 slice rows) and 9 rectangular strips.

[0021] Figure 3 An example of an image divided into slices and rectangular strips is shown, where the image is divided into 4 slices (2 slice columns and 2 slice rows) and 4 rectangular strips.

[0022] Figure 4 The image is displayed as being divided into 15 slices, 24 strips, and 24 sub-images.

[0023] Figure 5 This demonstrates a typical sub-picture-based viewport-dependent 360° video encoding and decoding scheme.

[0024] Figure 6 An improved viewport-dependent 360° video encoding and decoding scheme based on sub-pictures and spatial scalability is demonstrated.

[0025] Figure 7 This is a block diagram of an example video processing system.

[0026] Figure 8 This is a block diagram of a video processing device.

[0027] Figure 9 This is a flowchart of an example method for video processing.

[0028] Figure 10 This is a block diagram illustrating a video encoding and decoding system according to some embodiments of the present invention.

[0029] Figure 11 This is a block diagram illustrating an encoder according to some embodiments of the present invention.

[0030] Figure 12 This is a block diagram illustrating a decoder according to some embodiments of the present invention.

[0031] Figures 13 to 18 This is a flowchart of an example method for video processing. Detailed Implementation

[0032] The use of chapter headings in this document is for ease of understanding and does not limit the applicability of the techniques and embodiments disclosed in each chapter to that chapter only. Furthermore, the use of H.266 terminology in some descriptions is merely for ease of understanding and not to limit the scope of the disclosed techniques. Therefore, the techniques described herein are also applicable to other video codec protocols and designs.

[0033] 1. Introduction

[0034] This patent document relates to video codec technology. Specifically, it concerns improved support for reference picture resampling (RPR) in video codecs, including specifying a range of scaling window offset values ​​and signaling notifications controlling RPR and resolution variations within a layer's codec video sequence. These ideas can be applied individually or in various combinations to any video codec standard or non-standard video codec that supports multi-layer video codecs (e.g., the Multi-Functional Video Codec (VVC) under development).

[0035] 2. Abbreviation

[0036] APS (Adaptation Parameter Set)

[0037] AU (Access Unit)

[0038] AUD (Access Unit Delimiter)

[0039] AVC (Advanced Video Coding)

[0040] CLVS (Coded Layer Video Sequence)

[0041] CPB (Coded Picture Buffer) is a buffer for encoding and decoding pictures.

[0042] CRA (Clean Random Access)

[0043] CTU (Coding Tree Unit)

[0044] CVS (Coded Video Sequence) is a video sequence encoding and decoding mechanism.

[0045] DCI (Decoding Capability Information)

[0046] DPB (Decoded Picture Buffer)

[0047] EOB (End Of Bitstream) - End of Bitstream

[0048] EOS (End Of Sequence) - End of Sequence

[0049] GDR (Gradual Decoding Refresh) Gradual Decoding Refresh

[0050] HEVC (High Efficiency Video Coding)

[0051] HRD (Hypothetical Reference Decoder)

[0052] IDR (Instantaneous Decoding Refresh)

[0053] Inter-Layer Prediction (ILP)

[0054] ILRP (Inter-Layer Reference Picture)

[0055] JEM (Joint Exploration Model)

[0056] LTRP (Long-Term Reference Picture)

[0057] MCTS (Motion-Constrained Tile Sets)

[0058] NAL (Network Abstraction Layer)

[0059] OLS (Output Layer Set)

[0060] PH (Picture Header)

[0061] PPS (Picture Parameter Set)

[0062] PTL (Profile, Tier, and Level)

[0063] PU (Picture Unit)

[0064] RAP (Random Access Point)

[0065] RBSP (Raw Byte Sequence Payload)

[0066] SE (Syntax Element)

[0067] SEI (Supplemental Enhancement Information)

[0068] SPS (Sequence Parameter Set)

[0069] STRP (Short-Term Reference Picture)

[0070] SVC (Scalable Video Coding)

[0071] VCL (Video Coding Layer)

[0072] VPS (Video Parameter Set)

[0073] VTM (VVC Test Model)

[0074] VUI (Video Usability Information)

[0075] VVC (Versatile Video Coding) is a multi-functional video codec.

[0076] 3. Preliminary Discussion

[0077] Video codec standards have primarily evolved through the development of well-known ITU-T and ISO / IEC standards. ITU-T developed the H.261 and H.263 standards, while ISO / IEC developed the MPEG-1 and MPEG-4 Visual standards. The two organizations jointly developed the H.262 / MPEG-2 video standard, the H.264 / MPEG-4 Advanced Video Codec (AVC) standard, and the H.265 / HEVC standard. Starting with H.262, video codec standards are based on a hybrid video codec architecture, utilizing temporal prediction plus transform coding. To explore future video codec technologies beyond HEVC, the Joint Video Exploration Team (JVET) was jointly established by VCEG and MPEG in 2015. Since then, JVET has adopted many new methods and incorporated them into reference software called the Joint Exploration Model (JEM). JVET meetings are held quarterly, and the goal of new codec standards is to reduce the bitrate by 50% compared to HEVC. The new video codec standard was officially named Multifunctional Video Coding (VVC) at the JVET meeting in April 2018, and the first version of the VVC Test Model (VTM) was also released at that time. With ongoing efforts to standardize VVC, new codec technologies have been adopted into the VVC standard at each JVET meeting. The VVC working draft and test model VTM are updated after each meeting. The current goal of the VVC project is to achieve Technical Completion (FDIS) at the meeting in July 2020.

[0078] 3.1. Image Segmentation Schemes in HEVC

[0079] HEVC includes four different image segmentation schemes: regular striping, non-independent striping, slice, and wavefront parallel processing (WPP). These can be applied to maximum transfer unit (MTU) size matching, parallel processing, and reduction of end-to-end latency.

[0080] Regular slices are similar to those in H.264 / AVC. Each regular slice is encapsulated in its own NAL unit, and intra-frame prediction (intra-sample prediction, motion information prediction, and encoding / decoding mode prediction) and entropy encoding / decoding dependencies across slice boundaries are disabled. Therefore, regular slices can be reconstructed independently of other regular slices within the same image (although interdependencies may still exist due to loop filtering operations).

[0081] Regular striping is the only tool available for parallelization, and it is available in almost the same form in H.264 / AVC. Parallelization based on regular striping requires minimal inter-processor or inter-core communication (except for inter-processor or inter-core data sharing for motion compensation during decoding prediction of encode-decode images, which is typically much heavier than inter-processor or inter-core data sharing due to intra-frame image prediction). However, for the same reason, using regular striping can incur significant encoding / decoding overhead due to the bit cost of the stripe header and the lack of prediction across stripe boundaries. Furthermore, due to the intra-image independence of regular striping and the fact that each regular stripe is encapsulated in its own NAL, regular striping (compared to other tools mentioned below) can also serve as a key mechanism for bitstream segmentation to match MTU size requirements. In many cases, the goals of parallelization and MTU size matching are contradictory in terms of stripe layout requirements within the image. The implementation of such situations led to the development of the parallelization tools mentioned below.

[0082] Non-independent striping has a short stripe header and allows the bitstream to be split at tree block boundaries without disrupting any intra-picture predictions. Essentially, non-independent striping divides a regular stripe into multiple NAL units, reducing end-to-end latency by allowing a portion of the regular stripe to be sent before the entire regular stripe's encoding is complete.

[0083] In WPP, images are segmented into single-row coding tree blocks (CTBs). Entropy decoding and prediction are allowed to utilize data from CTBs in other segments. Parallel processing is possible through parallel decoding of CTB rows, where decoding of a CTB row begins with a delay of two CTBs to ensure that data associated with CTBs above and to the right of the main CTB is available before the main CTB being decoded. Using this staggered start (which looks like a wavefront when represented graphically), as many processors / cores as images containing CTB rows can be parallelized. Because intra-image prediction is allowed between adjacent tree block rows within an image, the inter-processor / inter-core communication required to achieve intra-image prediction can be substantial. WPP segmentation does not result in additional NAL units compared to segmentation without WPP application, therefore WPP is not a tool for MTU size matching. However, if MTU size matching is required, regular striping can be used with WPP, but with some encoding / decoding overhead.

[0084] A slice is defined by the horizontal and vertical boundaries that divide an image into slice columns and slice rows. Slice columns extend from the top to the bottom of the image. Similarly, slice rows extend from the left to the right of the image. The number of slices in an image can be simply derived by multiplying the number of slice columns by the number of slice rows.

[0085] Before decoding the top left CTB of the next slice in the order of the slice raster scans of the image, the scan order of the CTBs is changed to the local scan order within the slice (according to the order of the slice's CTB raster scans). Similar to regular stripes, slices break the intra-image prediction dependency and the entropy decoding dependency. However, they do not need to be contained in separate NAL units (the same as WPP in this respect); therefore, slices cannot be used for MTU size matching. Each slice can be processed by one processor / core, and in the case of a stripe spanning multiple slices, the inter-processor / inter-core communication required for decoding intra-image predictions between processing units of adjacent slices is limited to transmitting the shared stripe header and loop filtering associated with the sharing of reconstructed samples and metadata. When a stripe contains more than one slice or WPP segment, the entry point byte offset of each slice or WPP segment in the stripe, except for the first one, is signaled in the stripe header.

[0086] For simplicity, HEVC specifies restrictions on the application of four different image segmentation schemes. For most levels specified in HEVC, a given codec video sequence cannot simultaneously contain slices and wavefronts. For each strip and slice, one or both of the following conditions must be met: 1) All codec tree blocks in a strip belong to the same slice; 2) All codec tree blocks in a slice belong to the same strip. Finally, a wavefront segment contains exactly one CTB line, and when using WPP, if a strip begins within a CTB line, then that strip must end within the same CTB line.

[0087] The most recent revision of HEVC is specified in the JCT-VC output document JCTVC-AC1005, by J. Boyce, A. Ramasubramonian, R. Skupin, G. J. Sullivan, A. Tourapis, and Y.-K. Wang (editors). "HEVC Additional Supplemental Enhancement Information (Draft 4)," published on October 24, 2017, is available here: http: / / phenix.intevry.fr / jct / doc_end_user / documents- / 29_Macau / wg11 / JCTVC-AC1005-v2.zip. Including this revision, HEVC specifies three SEI messages related to MCT: the i.e., the domain MCTS SEI message, the MCTS extracted information set SEI message, and the MCTS extracted information nested SEI message.

[0088] The temporal MCTS SEI message indicates the presence of an MCTS in the bitstream and signals this to the MCTS. For each MCTS, motion vectors are restricted to pointing to full-sample locations within the MCTS and fractional-sample locations that require interpolation only from full-sample locations within the MCTS. Motion vector candidates derived from temporal motion vector predictions from blocks outside the MCTS are not allowed. This allows each MCTS to be decoded independently even if there are no slices not included in the MCTS.

[0089] The MCTS Extraction Information Set (SEI) message provides supplementary information that can be used for MCTS sub-bitstream extraction (defined as part of the semantics of the SEI message) to generate a bitstream conforming to the MCTS group. This information consists of multiple extraction information sets, each defining multiple MCTS groups and containing RBSP bytes for replacing the VPS, SPS, and PPS to be used during the MCTS sub-bitstream extraction process. When extracting the sub-bitstream according to the MCTS sub-bitstream extraction process, the parameter sets (VPS, SPS, and PPS) need to be rewritten or replaced, and the slice header needs to be slightly updated because one or all syntax elements related to the slice address (including first_slice_segment_in_pic_flag and slice_segment_address) typically need to have different values.

[0090] 3.2. Image Segmentation in VVC

[0091] In VVC, an image is divided into one or more slice rows and one or more slice columns. A slice is a sequence of CTUs covering a rectangular area of ​​the image. The CTUs within a slice are scanned in raster scan order within that slice.

[0092] A strip consists of an integer number of complete slices or an integer number of consecutive complete CTU lines within a slice of an image.

[0093] Two stripe modes are supported: raster scan stripe mode and rectangular stripe mode. In raster scan stripe mode, a stripe contains a complete stripe sequence within a sheet raster scan of the image. In rectangular stripe mode, a stripe contains multiple complete sheets that together form a rectangular area of ​​the image, or multiple consecutive complete CTU rows of a sheet that together forms a rectangular area of ​​the image. Sheets within a rectangular stripe are scanned in sheet raster scan order within the rectangular area corresponding to that stripe.

[0094] A sub-image contains one or more stripes that collectively cover a rectangular area of ​​the image.

[0095] Figure 1 An example of raster scan strip segmentation of an image is shown, where the image is divided into 12 slices and 3 raster scan strips.

[0096] Figure 2 An example of rectangular strip segmentation of an image is shown, where the image is divided into 24 slices (6 slice columns and 4 slice rows) and 9 rectangular slices.

[0097] Figure 3 An example of an image divided into slices and rectangular strips is shown, where the image is divided into 4 slices (2 slice columns and 2 slice rows) and 4 rectangular strips.

[0098] Figure 4 An example of sub-image segmentation of an image is shown, where the image is segmented into 18 slices: 12 slices on the left (each covering a strip with 4x4 CTUs) and 6 slices on the right (each covering two vertically stacked strips with 2x2 CTUs), resulting in a total of 24 strips and 24 sub-images of different sizes (each strip being a sub-image).

[0099] 3.3. Image resolution variations within a sequence

[0100] In AVC and HEVC, the spatial resolution of an image cannot be changed unless a new sequence with a new SPS begins with an IRAP image. VVC allows changing the image resolution within a sequence at locations where IRAP images are not encoded; IRAP images are always intra-frame encoded and decoded. This feature is sometimes called Reference Image Resampling (RPR) because it requires resampling the reference image used for inter-frame prediction when the reference image has a different resolution than the current image being decoded.

[0101] The scaling ratio is limited to greater than or equal to 1 / 2 (2x downsampling from the reference image to the current image) and less than or equal to 8 (8x upsampling). Three sets of resampling filters with different frequency cutoffs are specified to handle various scaling ratios between the reference and current images. The three sets of resampling filters are applied to scaling ratios ranging from 1 / 2 to 1 / 1.75, from 1 / 1.75 to 1 / 1.25, and from 1 / 1.25 to 8, respectively. Each set of resampling filters has 16 phases for luminance and 32 phases for chrominance, which is the same as in motion-compensated interpolation filters. In fact, the normal MC interpolation process is a special case of the resampling process where the scaling ratio ranges from 1 / 1.25 to 8. The horizontal and vertical scaling ratios are derived based on the image width and height, as well as the left, right, top, and bottom scaling offsets specified for the reference and current images.

[0102] Other aspects of VVC designs that support this feature, unlike HEVC, include: i) signaling the image resolution and corresponding consistency window in the PPS instead of the SPS, while in the SPS the signaling indicates the maximum image resolution. ii) for a single-layer bitstream, each image storage (a slot storing one decoded image in the DPB) occupies the buffer size required to store the decoded image with the maximum image resolution.

[0103] 3.4. Scalable Video Codec (SVC) in General Purpose and VVC

[0104] Scalable video codec (SVC, sometimes also called scalability in video codec) refers to video codec using a base layer (BL) (sometimes called a reference layer (RL)) and one or more scalable enhancement layers (EL). In SVC, the base layer can carry video data with a base quality level. One or more enhancement layers can carry additional video data to support, for example, higher spatial, temporal, and / or signal-to-noise (SNR) levels. Enhancement layers can be defined relative to previously encoded layers. For example, the bottom layer can be used as a BL, while the top layer can be used as an EL. Intermediate layers can be used as ELs or RLs, or both. For example, an intermediate layer (e.g., a layer that is neither the lowest nor the highest layer) can be an EL of a layer below the intermediate layer (e.g., a base layer or any intermediate enhancement layer) and simultaneously serve as an RL of one or more enhancement layers above the intermediate layer. Similarly, in the HEVC standard's multiview or 3D extension, there may be multiple views, and information from one view can be used to encode or decode (e.g., encode or decode) information from another view (e.g., motion estimation, motion vector prediction, and / or other redundancy).

[0105] In SVC, the parameters used by the encoder or decoder are grouped into parameter sets based on the codec level in which they can be used (e.g., video level, sequence level, picture level, stripe level, etc.). For example, parameters that can be used by one or more codec video sequences at different layers in a bitstream can be included in the Video Parameter Set (VPS), and parameters that can be used by one or more pictures in a codec video sequence can be included in the Sequence Parameter Set (SPS). Similarly, parameters used by one or more stripes in a picture can be included in the Picture Parameter Set (PPS), and other parameters specific to a single strip can be included in the stripe header. Likewise, indications of which parameter set(s) a particular layer uses at a given time can be provided at various codec levels.

[0106] Because of VVC's support for Reference Picture Resampling (RPR), support for multi-layered bitstreams can be designed without requiring any additional signaling notification processing level codec tools. For example, two layers in VVC with SD and HD resolutions can be supported because the upsampling required for spatial scalability can be achieved using only RPR upsampling filters. However, supporting scalability requires a higher level of syntax changes (compared to not supporting scalability at all). Scalability support was specified in VVC version 1. Unlike scalability support in any earlier video codec standards, including extensions to AVC and HEVC, VVC scalability was designed to be as friendly as possible to single-layer decoder designs. The decoding capability of multi-layered bitstreams is specified as if there were only a single layer in the bitstream. For example, decoding capabilities such as DPB size are specified in a way that is independent of the number of layers in the bitstream to be decoded. Essentially, decoders designed for single-layered bitstreams do not require many changes to decode multi-layered bitstreams. Compared to the multi-layered extensions of AVC and HEVC, the HLS aspect is significantly simplified at the expense of some flexibility. For example, IRAP AU requires a picture of every layer present in CVS.

[0107] 3.5. Viewport-dependent 360° video stream based on sub-images

[0108] In 360° video (also known as omnidirectional video) streaming, at any given moment, only a subset of the entire omnidirectional video sphere (e.g., the current viewport) is presented to the user, who can at any time rotate their head to change their viewing orientation, thus changing the current viewport. While it's desirable to have at least some lower-quality representations of areas not covered by the current viewport on the client side, ready to be presented to the user in case they suddenly change their viewing orientation to any location on the sphere, the high-quality representation of the omnidirectional video is only needed for the current viewport being presented to the user. This optimization can be achieved by dividing the high-quality representation of the entire omnidirectional video into sub-pictures with appropriate granularity. Using VVC, these two representations can be encoded as two independent layers.

[0109] Figure 5 This demonstrates a typical sub-picture-based viewport-dependent 360° video encoding and decoding scheme.

[0110] A typical sub-image-based viewport-dependent 360° video transmission scheme is as follows: Figure 5 As shown, a higher resolution representation of the complete video consists of sub-pictures, while a lower resolution representation of the complete video does not use sub-pictures and can be encoded and decoded using lower-frequency random access points in the higher resolution representation. The client receives the lower resolution complete video, while for the higher resolution video, it only receives and decodes the sub-pictures covering the current viewport.

[0111] The latest VVC draft specification also supports improved 360° video codec schemes, such as... Figure 6 As shown. With Figure 5 The only difference between the methods shown is that inter-layer prediction (ILP) is applied. Figure 6 The method shown.

[0112] Figure 6 An improved viewport-dependent 360° video encoding and decoding scheme based on sub-pictures and spatial scalability is demonstrated.

[0113] 3.6. Parameter Set

[0114] AVC, HEVC, and VVC specify parameter sets. Parameter set types include SPS, PPS, APS, and VPS. All AVC, HEVC, and VVC versions support SPS and PPS. VPS was introduced with HEVC and is included in both HEVC and VVC. APS is not included in AVC or HEVC, but is included in the latest VVC draft text.

[0115] SPS is designed to carry sequence-level header information, and PPS is designed to carry infrequently changing image-level header information. Using SPS and PPS eliminates the need to repeat infrequently changing information for each sequence or image, thus avoiding redundant signaling notifications. Furthermore, using SPS and PPS enables out-of-band transmission of critical header information, thereby not only avoiding redundant transmission but also improving error recovery capabilities.

[0116] The purpose of introducing a VPS is to carry sequence-level header information common to all layers in a multi-layer bitstream.

[0117] The purpose of APS is to carry such image-level or stripe-level information, which requires a considerable number of bits for encoding and decoding, can be shared by multiple images, and can have many different variations in the sequence.

[0118] 4. The technical problem solved by the disclosed technical solution

[0119] The existing design in the latest VVC documentation (JVET-R2001-vA / v10) has the following issues:

[0120] 1) In the latest VVC text, the derivation of variables CurrPicScalWinWidthL and CurrPicScalWinHeightL is as follows:

[0121] CurrPicScalWinWidthL=pps_pic_width_in_luma_samples- (79)

[0122] SubWidthC*(pps_scaling_win_right_offset+pps_scaling_win_left_offset)

[0123] CurrPicScalWinHeightL=pps_pic_height_in_luma_samples- (80)

[0124] The syntax `SubHeightC*(pps_scaling_win_bottom_offset+pps_scaling_win_top_offset)` specifies the range of values ​​for the scaling window offset syntax element as follows:

[0125] The value of SubWidthC*(Abs(pps_scaling_win_left_offset)+Abs(pps_scaling_win_right_offset)) should be less than pps_pic_width_in_luma_samples, and the value of SubHeightC*(Abs(pps_scaling_win_top_offset)+Abs(pps_scaling_win_bottom_offset)) should be less than pps_pic_height_in_luma_samples.

[0126] Therefore, the value of scaling window width CurrPicScalWinWidthL is always less than twice the image width, i.e., pps_pic_width_in_luma_samples*2, and the value of scaling window height CurrPicScalWinHeightL is always less than twice the image height, i.e., pps_pic_height_in_luma_samples*2.

[0127] However, the range of values ​​for scaling the window width and height is too limited.

[0128] For example, for such Figure 6 The improved 360° video encoding / decoding scheme shown assumes that each EL image (covering the entire sphere of 360° × 180°) is divided into 12 × 8 = 96 sub-images, and the user's viewport is 120° × 90°. This means that approximately 4 × 4 sub-images of each EL image are received, decoded, and presented to the user by the decoder. During sub-bitstream extraction to generate the bitstream sent to the decoder (i.e., as...), Figure 6As shown in the lower part), the PPS (including the scaling window offset parameter) of the BL image remains unchanged, while the PPS of the EL image will be rewritten. Assuming that in the original bitstream (before sub-bitstream extraction, such as...)... Figure 6 As shown in the upper part), if the scaling window offset of the EL image is equal to 0, then the values ​​of CurrPicScalWinWidthL and CurrPicScalWinHeightL of the EL image in the original bitstream will be the same as the image width and height, respectively. However, in the bitstream received by the decoder (after sub-bitstream extraction, similar to...), Figure 6 (The lower part), the values ​​of CurrPicScalWinWidthL and CurrPicScalWinHeightL of the EL image will need to be equal to three times the image width (pps_pic_width_in_luma_samples*3) and twice the image height (pps_pic_height_in_luma_samples*2), respectively.

[0129] However, this would violate the existing value range of the scaling window offset syntax element.

[0130] 2) In the latest VVC text, the variable RprConstraintsActive[i][j] specifies whether RPR is enabled for the j-th reference image in the reference image list i of the current strip. The derivation is as follows:

[0131] RprConstraintsActive[i][j]=(pps_pic_width_in_luma_samples!=

[0132] refPicWidth||pps_pic_height_in_luma_samples! =refPicHeight||

[0133] pps_scaling_win_left_offset! =refScalingWinLeftOffset||

[0134] pps_scaling_win_right_offset! =refScalingWinRightOffset||

[0135] pps_scaling_win_top_offset! =refScalingWinTopOffset||

[0136] pps_scaling_win_bottom_offset! =refScalingWinBottomOffset)

[0137] Basically, as long as one of the six parameters of the current image is different from the parameter of the reference image, the value of this variable is equal to 1, and RPR is enabled: 1) image width, 2) image height, 3) zoom window left offset, 4) zoom window right offset, 5) zoom window top offset, and 6) zoom window bottom offset.

[0138] However, even when in the original bitstream (such as...) Figure 6 Even when the spatial resolution of the EL image and the BL image (at the top) is the same, this does not allow for... Figure 6 The EL images in the improved 360° video encoding and decoding scheme shown use some encoding and decoding tools such as DMVR, BDOF, and PRO.

[0139] 3) The semantics of sps_ref_pic_resampling_enabled_flag are as follows:

[0140] A value of 1 for sps_ref_pic_resampling_enabled_flag indicates that reference image resampling is enabled, and one or more image stripes in CLVS can reference reference images with different spatial resolutions in the active entry of the reference image list.

[0141] A value of 0 for sps_ref_pic_resampling_enabled_flag indicates that reference image resampling is disabled, and that there are no reference images with different spatial resolutions in the active entry of the reference image list in CLVS.

[0142] However, whether RPR is enabled depends on whether the current image and the following six parameters of the current image are all the same, not just the image width and height: 1) image width, 2) image height,

[0143] 3) Scale window left offset, 4) Scale window right offset, 5) Scale window top offset, and 6) Scale window bottom offset.

[0144] Therefore, the semantics are inconsistent with how to determine whether RPR is on or off.

[0145] 4) The value of sps_ref_pic_resampling_enabled_flag may need to be specified based on whether a certain layer associated with SPS is not an independent layer and information about the reference layer based on that layer.

[0146] 5) Lack of signaling notification control for enabling or disabling the RPL, which comes from a reference picture in the same layer as the current picture.

[0147] 5. List of Implementation Examples and Solutions

[0148] To address the aforementioned and other problems, methods summarized below are disclosed. This invention should be considered as an example of interpreting general concepts and should not be interpreted narrowly. Furthermore, these inventions can be applied individually or in any combination.

[0149] 1) To solve problem 1, specify the range of values ​​for the SE indicating the scaling window (e.g., pps_scaling_win_left_offset, pps_scaling_win_right_offset, pps_scaling_win_top_offset, and pps_scaling_win_bottom_offset) in such a way that the width and height of the scaling window can be greater than or equal to twice the width of the image and twice the height of the image, respectively.

[0150] a. In one example, the value of SubWidthC*(pps_scaling_win_left_offset+pps_scaling_win_right_offset) is specified to be greater than or equal to -pps_pic_width_in_luma_samples*M and less than pps_pic_width_in_luma_samples, and the value of SubHeightC*(pps_scaling_win_top_offset+pps_scaling_win_bottom_offset) is specified to be greater than or equal to -pps_pic_height_in_luma_samples*N and less than pps_pic_height_in_luma_samples, where M and N are positive integer values ​​equal to or greater than 2.

[0151] i. In some examples, M and N are the same and equal to 16, 8, 32 or 64.

[0152] ii. In some other examples, M and N are specified as depending on the layout of the sub-images.

[0153] a) In one example, M equals Ceil(pps_pic_width_in_luma_samples÷minSubpicWidth), where minSubpicWidth is the minimum subpicture width among all subpictures in the reference PPS image, and N equals Ceil(pps_pic_height_in_luma_samples÷minSubpicHeight), where minSubpicHeight is the minimum subpicture height among all subpictures in the reference PPS image.

[0154] 2) Whether and / or how to constrain the width of the zoom window for the reference image and the zoom window for the current image.

[0155] The height can depend on whether the reference image and the current image are on the same layer (e.g., whether they have the same layer ID).

[0156] a. In one example, when the reference image and the current image are on different layers, there are no restrictions on the width / height of the scaling window for the reference image and the scaling window for the current image.

[0157] b. Assume that the requirement for bitstream consistency is that all of the following conditions must be met:

[0158] –CurrPicScalWinWidthL*M is greater than or equal to refPicScalWinWidthL.

[0159] –CurrPicScalWinHeightL*M is greater than or equal to refPicScalWinHeightL.

[0160] –CurrPicScalWinWidthL is less than or equal to refPicScalWinWidthL*N.

[0161] –CurrPicScalWinHeightL is less than or equal to refPicScalWinHeightL*N.

[0162] –CurrPicScalWinWidthL*sps_pic_width_max_in_luma_samples is greater than or equal to refPicScalWinWidthL*(pps_pic_width_in_luma_samples-Max(K0,MinCbSizeY)).

[0163] –CurrPicScalWinHeightL*sps_pic_height_max_in_luma_samples is greater than or equal to refPicScalWinHeightL*(pps_pic_height_in_luma_samples-Max(K1,MinCbSizeY)).

[0164] Where CurrPicScalWinWidthL and CurrPicScalWinHeightL represent the width and height of the scaling window for the current image, respectively, and refPicScalWinWidthL and refPicScalWinHeightL represent the width and height of the scaling window for the reference image. Assuming the reference image and the current image are on the same layer, M = M1 and N = N1. When the reference image and the current image are on different layers, M = M2 and N = N2.

[0165] i. In one example, M1 <= M2.

[0166] ii. In one example, N1 >= N2.

[0167] iii. In one example, K0 / K1 are integer values.

[0168] iv. In one example, K0 / K1 is the minimum allowed image width / height, for example, 8.

[0169] 3) To solve problem 2, the derivation of the variable RprConstraintsActive[i][j] can be specified to depend only on the width and height of the scaling window. The variable RprConstraintsActive[i][j] specifies whether RPR is enabled for the reference image of the j-th entry in the reference image list i of the current strip.

[0170] a. In one example, the derivation of the variable RprConstraintsActive[i][j] is as follows:

[0171] RprConstraintsActive[i][j]=(CurrPicScalWinWidthL!=fRefWidth||CurrPicScalWinHeightL!=fRefHeight)

[0172] Where fRefWidth and fRefHeight are the values ​​of CurrPicScalWinWidthL and CurrPicScalWinHeightL of the reference image, respectively.

[0173] b. In one example, the derivation of the variable RprConstraintsActive[i][j] can depend on the layer information of the reference image and the current image.

[0174] i. In one example, the derivation of the variable RprConstraintsActive[i][j] can be specified differently depending on whether the reference image and the current image are in the same layer (e.g., whether they have the same layer ID). The variable RprConstraintsActive[i][j] specifies whether RPR is enabled for the reference image of the j-th entry in the reference image list i of the current strip.

[0175] 4) To solve problem 3, the semantics of `sps_ref_pic_resampling_enabled_flag` can be changed as follows:

[0176] `sps_ref_pic_resampling_enabled_flag` equal to 1 indicates that reference image resampling is enabled, and that one or more stripes of an image in CLVS can reference a reference image in the active entry of the reference image list, which has one or more of the following six parameters that differ from the current image: 1) image width, 2) image height, 3) scale window left offset, 4) scale window right offset, 5) scale window top offset, and 6) scale window bottom offset. `sps_ref_pic_resampling_enabled_flag` equal to 0 indicates that reference image resampling is disabled, and that stripes of an image in CLVS that do not have an image reference a reference image in the active entry of the reference image list, which has one or more of the above six parameters that differ from the current image.

[0177] 5) To address issue 4, the value of sps_ref_pic_resampling_enabled_flag is specified based on whether a layer associated with SPS is not an independent layer and information about the reference layer based on that layer.

[0178] a. In one example, when the layer whose nuh_layer_id is equal to the nuh_layer_id of the SPS is not an independent layer and at least one reference layer of the current layer has a different spatial resolution than the current layer, the constraint sps_ref_pic_resampling_enabled_flag should be equal to 1.

[0179] b. In one example, when at least one layer of the reference SPS is not an independent layer and at least one reference layer of that layer has a different spatial resolution than the current layer, the value of the constraint sps_ref_pic_resampling_enabled_flag should be equal to 1.

[0180] c. In one example, when referencing at least one stripe of at least one image currPic of SPS references a reference image in an ILRP entry of the reference image list, the value of the constraint sps_ref_pic_resampling_enabled_flag should be equal to 1, and the reference image has one or more of the following six parameters that are different from the parameters of the image currPic: 1) image width, 2) image height, 3) scale window left offset, 4) scale window right offset, 5) scale window top offset, and 6) scale window bottom offset.

[0181] 6) To resolve issue 5, the flag sps_res_change_in_clvs_allowed_flag has been renamed to sps_intra_layer_rpr_enabled_flag, and its semantics have been changed as follows:

[0182] `sps_intra_layer_rpr_enabled_flag` equal to 1 enables reference image resampling in the CLVS for inter-frame prediction from reference images in the same layer as the current image. Specifically, stripes in the CLVS can reference reference images in the active entry of the reference image list. This reference image has the same `nuh_layer_id` value as the current image and has one or more of the following six parameters that differ from the current image: 1) image width, 2) image height, 3) scale window left offset, 4) scale window right offset, 5) scale window top offset, and 6) scale window bottom offset. `sps_intra_layer_rpr_enabled_flag` equal to 0 disables reference image resampling from reference images in the same layer as the current image. It also specifies that stripes in the CLVS without images reference reference images in the active entry of the reference image list. This reference image has the same `nuh_layer_id` value as the current image and has one or more of the above six parameters that differ from the current image.

[0183] a. Alternatively, sps_intra_layer_rpr_enabled_flag equal to 1 specifies that reference image resampling is enabled for inter-frame prediction of images from reference images in the same layer as the current image in CLVS, and the reference image is associated with RprConstraintsActive[i][j] equal to 1, where the j-th entry in the reference image list i is used for the current strip and it has the same nul_layer_id value as the current image.

[0184] b. In one example, sps_intra_layer_rpr_enabled_flag and sps_res_change_in_clvs_allowed_flag (with semantics in JVET-R2001) can be signaled sequentially.

[0185] i.sps_res_change_in_clvs_allowed_flag can be signaled after sps_intra_layer_rpr_enabled_flag.

[0186] ii. If sps_intra_layer_rpr_enabled_flag is 0, then sps_res_change_in_clvs_allowed_flag is not signaled and is inferred to be 0.

[0187] 6. Example

[0188] The following are some example embodiments of aspects of the invention summarized in Section 5 of the previous article, which can be applied to the VVC specification. The modified text is based on the latest VVC text in JVET-R2001-vA / v10. Most of the relevant sections that have been added or modified are... The text is indicated, and some deleted sections are marked with double brackets (e.g., [[]]), with the deleted text enclosed within the brackets. There may also be other editable changes that are not highlighted.

[0189] 6.1. First Embodiment

[0190] This embodiment applies to projects 1, 1.a, 1.ai and 4.

[0191] 7.4.3.3 Sequence Parameter Set (RBSP) Semantics ...

[0193] A `sps_ref_pic_resampling_enabled_flag` value of 1 indicates that reference image resampling is enabled, and specifies that one or more stripes of an image in CLVS can reference a reference image in the active entry of the reference image list, which has a different spatial resolution. A `sps_ref_pic_resampling_enabled_flag` value of 0 indicates that reference image resampling is disabled, and specifies the reference image in the active entry of the strip reference reference image list that has no image in CLVS, and that the reference image has a different spatial resolution.

[0194] Note 2 — When sps_ref_pic_resampling_enabled_flag equals 1, for the current image, [[having different spatial resolutions]] can belong to the same layer or be a different layer from the layer containing the current image.

[0195] `sps_res_change_in_clvs_allowed_flag` equal to 1 indicates that the image spatial resolution can be changed within the CLVS of the reference SPS. `sps_res_change_in_clvs_allowed_flag` equal to 0 indicates that the image spatial resolution will not change within the CLVS of any reference SPS. When it does not exist, the value of `sps_res_change_in_clvs_allowed_flag` is inferred to be 0. ...

[0197] 7.4.3.4 Image Parameter Set RBSP Semantics ...

[0199] A value of 1 for `pps_scaling_window_explicit_signalling_flag` indicates that a scaling window offset parameter exists in PPS. A value of 0 for `pps_scaling_window_explicit_signalling_flag` indicates that a scaling window offset parameter does not exist in PPS. When `sps_ref_pic_resampling_enabled_flag` is 0, the value of `pps_scaling_window_explicit_signalling_flag` should be 0.

[0200] pps_scaling_win_left_offset, pps_scaling_win_right_offset, pps_scaling_win_top_offset, and pps_scaling_win_bottom_offset specify the offsets applied to the image size to calculate the scaling ratio. When not present, the values ​​of pps_scaling_win_left_offset, pps_scaling_win_right_offset, pps_scaling_win_top_offset, and pps_scaling_win_bottom_offset are inferred to be equal to pps_conf_win_left_offset, pps_conf_win_right_offset, pps_conf_win_top_offset, and pps_conf_win_bottom_offset, respectively.

[0201] The value of SubWidthC*([[Abs(]]pps_scaling_win_left_offset[[)]]+[[Abs(]]pps_scaling_win_right_offset[[)]]) should The value of SubHeightC*([[Abs(]]pps_scaling_win_top_offset[[)]]+[[Abs(]]pps_scaling_win_bottom_offset[[)]]) should be less than pps_pic_width_in_luma_samples, and the value of SubHeightC*([[Abs(]]pps_scaling_win_bottom_offset[[)]]) should be less than pps_pic_width_in_luma_samples. Less than pps_pic_height_in_luma_samples.

[0202] The derivation of variables CurrPicScalWinWidthL and CurrPicScalWinHeightL is as follows:

[0203] CurrPicScalWinWidthL=pps_pic_width_in_luma_samples- (79)

[0204] SubWidthC*(pps_scaling_win_right_offset+pps_scaling_win_left_offset)

[0205] CurrPicScalWinHeightL=pps_pic_height_in_luma_samples- (80)

[0206] SubHeightC*(pps_scaling_win_bottom_offset+pps_scaling_win_top_offset)

[0207] Assume that refPicScalWinWidthL and refPicScalWinHeightL are CurrPicScalWinWidthL and CurrPicScalWinHeightL of the reference image to the current image of this PPS, respectively. Bitstream consistency requires that all of the following conditions be met:

[0208] –CurrPicScalWinWidthL*2 is greater than or equal to refPicScalWinWidthL.

[0209] –CurrPicScalWinHeightL*2 is greater than or equal to refPicScalWinHeightL.

[0210] –CurrPicScalWinWidthL is less than or equal to refPicScalWinWidthL*8.

[0211] –CurrPicScalWinHeightL is less than or equal to refPicScalWinHeightL*8.

[0212] –CurrPicScalWinWidthL*sps_pic_width_max_in_luma_samples is greater than or equal to refPicScalWinWidthL*(pps_pic_width_in_luma_samples-Max(8,MinCbSizeY)).

[0213] –CurrPicScalWinHeightL*sps_pic_height_max_in_luma_samples is greater than or equal to refPicScalWinHeightL*(pps_pic_height_in_luma_samples-Max(8,MinCbSizeY)). ...

[0215] Figure 7This is a block diagram of an example video processing system 1900 that can implement the various techniques disclosed herein. Various implementations may include some or all of the components in system 1900. System 1900 may include an input 1902 for receiving video content. The video content may be received in a raw or uncompressed format (e.g., 8 or 10-bit multi-component pixel values), or in a compressed or encoded format. Input 1902 may represent a network interface, a peripheral bus interface, or a storage interface. Examples of network interfaces include wired interfaces (such as Ethernet, Passive Optical Networking (PON), etc.) and wireless interfaces (such as Wi-Fi or cellular interfaces).

[0216] System 1900 may include a codec component 1904 capable of implementing the various codec or encoding methods described in this document. Codec component 1904 can reduce the average bit rate of the video from input 1902 to the output of codec component 1904 to produce a codec representation of the video. Therefore, codec techniques are sometimes referred to as video compression or video transcoding techniques. The output of codec component 1904 can be stored or transmitted via connected communication, as represented by component 1906. The stored or communicated bitstream (or codec) representation of the video received at input 1902 can be used by component 1908 to generate pixel values ​​or displayable video that are sent to display interface 1910. The process of generating user-visible video from the bitstream representation is sometimes referred to as video decompression. Furthermore, although some video processing operations are referred to as "codec" operations or tools, it should be understood that the codec tool or operation is used at the encoder, and the corresponding decoding tool or operation will be inverted by the decoder to retrieve the result of the codec.

[0217] Examples of peripheral bus interfaces or display interfaces may include Universal Serial Bus (USB), High Definition Multimedia Interface (HDMI), or DisplayPort. Examples of storage interfaces include SATA (Serial Advanced Technology Accessory), PCI, IDE, etc. The technologies described in this document can be implemented in a variety of electronic devices, such as mobile phones, laptops, smartphones, or other devices capable of digital data processing and / or video display.

[0218] Figure 8This is a block diagram of a video processing apparatus 3600. Apparatus 3600 can be used to implement one or more of the methods described herein. Apparatus 3600 can be implemented in smartphones, tablets, computers, Internet of Things (IoT) receivers, etc. Apparatus 3600 may include one or more processors 3602, one or more memories 3604, and video processing hardware 3606. The processors(multiple) 3602 can be configured to implement one or more methods described herein. The memories(multiple) 3604 can be used to store data and code used to implement the methods and techniques described herein. The video processing hardware 3606 can be used to implement some of the techniques described herein in hardware circuitry.

[0219] Figure 10 This is a block diagram illustrating an example video codec system 100 that can utilize the techniques disclosed herein.

[0220] like Figure 10 As shown, the video encoding / decoding system 100 may include a source device 110 and a target device 120. The source device 110 generates encoded video data and may be referred to as a video encoding device. The target device 120 can decode the encoded video data generated by the source device 110 and may be referred to as a video decoding device.

[0221] The source device 110 may include a video source 112, a video encoder 114, and an input / output (I / O) interface 116.

[0222] Video source 112 may include sources such as video capture devices, interfaces for receiving video data from video content providers, and / or computer graphics systems that generate video data, or combinations of these sources. Video data may include one or more images. Video encoder 114 encodes the video data from video source 112 to generate a bitstream. The bitstream may include a sequence of bits forming a codec representation of the video data. The bitstream may include codec images and associated data. A codec image is a codec representation of an image. Associated data may include sequence parameter sets, image parameter sets, and other syntax elements. I / O interface 116 includes a modulator / demodulator (modem) and / or a transmitter. Encoded video data may be transmitted directly to target device 120 via network 130a through I / O interface 116. Encoded video data may also be stored on storage medium / server 130b for access by target device 120.

[0223] The target device 120 may include an I / O interface 126, a video decoder 124, and a display device 122.

[0224] I / O interface 126 may include a receiver and / or a modem. I / O interface 126 may acquire encoded video data from source device 110 or storage medium / server 130b. Video decoder 124 may decode the encoded video data. Display device 122 may display the decoded video data to a user. Display device 122 may be integrated with target device 120 or may be external to target device 120 configured to connect to an external display device.

[0225] The video encoder 114 and the video decoder 124 can operate according to video compression standards such as High Efficiency Video Codec (HEVC), Multi-Functional Video Codec (VVC), and other current and / or other standards.

[0226] Figure 11 This is a block diagram illustrating an example of a video encoder 200, which may be... Figure 10 The video encoder 114 in the system 100 shown in the figure.

[0227] The video encoder 200 can be configured to perform any or all of the techniques disclosed herein. Figure 11 In the example, the video encoder 200 includes multiple functional components. The techniques described in this disclosure can be shared among the various components of the video encoder 200. In some examples, the processor can be configured to perform any or all of the techniques described in this disclosure.

[0228] The functional components of the video encoder 200 may include a segmentation unit 201, a prediction unit 202 (which may include a mode selection unit 203, a motion estimation unit 204, a motion compensation unit 205, and an intra-frame prediction unit 206), a residual generation unit 207, a transform unit 208, a quantization unit 209, an inverse quantization unit 210, an inverse transform unit 211, a reconstruction unit 212, a buffer 213, and an entropy coding unit 214.

[0229] In other examples, the video encoder 200 may include more, fewer, or different functional components. In one example, the prediction unit 202 may include an intra-block copy (IBC) unit. The IBC unit may perform prediction in IBC mode, where at least one reference picture is the picture in which the current video block is located.

[0230] Furthermore, some components, such as the motion estimation unit 204 and the motion compensation unit 205, can be highly integrated, but for interpretive purposes... Figure 11 The examples are shown separately.

[0231] The segmentation unit 201 can segment an image into one or more video blocks. The video encoder 200 and the video decoder 300 can support various video block sizes.

[0232] The mode selection unit 203 can, for example, select one of the intra-frame or inter-frame encoding / decoding modes based on the error result, and provide the obtained intra-frame or inter-frame encoded / decoded blocks to the residual generation unit 207 to generate residual block data and to the reconstruction unit 212 to reconstruct the encoded / decoded blocks for use as reference images. In some examples, the mode selection unit 203 can select a combination of intra-frame and inter-frame prediction (CIIP) mode, where the prediction is based on inter-frame prediction signals and intra-frame prediction signals. The mode selection unit 203 can also select the resolution of the motion vector (e.g., sub-pixel or integer pixel precision) for blocks in the inter-frame prediction case.

[0233] To perform inter-frame prediction for the current video block, motion estimation unit 204 can generate motion information for the current video block by comparing one or more reference frames from buffer 213 with the current video block. Motion compensation unit 205 can determine the predicted video block for the current video block based on the motion information of the image from buffer 213 (rather than the image associated with the current video block) and decoded samples.

[0234] The motion estimation unit 204 and the motion compensation unit 205 can perform different operations on the current video block, for example, the different operations performed depend on whether the current video block is in an I-strip, a P-strip, or a B-strip.

[0235] In some examples, motion estimation unit 204 can perform unidirectional prediction of the current video block, and can search for a reference video block for the current video block in the reference images of list 0 or list 1. Motion estimation unit 204 can then generate a reference index indicating that the reference image in list 0 or list 1 contains the reference video block, and a motion vector indicating the spatial displacement between the current video block and the reference video block. Motion estimation unit 204 can output the reference index, prediction direction indicator, and motion vector as motion information for the current video block. Motion compensation unit 205 can generate a predicted video block for the current block based on the reference video block indicated by the motion information of the current video block.

[0236] In other examples, motion estimation unit 204 can perform bidirectional prediction of the current video block. Motion estimation unit 204 can search for a reference video block for the current video block in the reference images of list 0 and can also search for another reference video block for the current video block in the reference images of list 1. Motion estimation unit 204 can then generate a reference index indicating that the reference images in list 0 or list 1 contain the reference video block, and a motion vector indicating the spatial displacement between the reference video block and the current video block. Motion estimation unit 204 can output the reference index and the motion vector of the current video block as the motion information of the current video block. Motion compensation unit 205 can generate a predicted video block for the current video block based on the reference video block indicated by the motion information of the current video block.

[0237] In some examples, the motion estimation unit 204 can output the complete set of motion information for the decoder's decoding process.

[0238] In some examples, motion estimation unit 204 may not output the complete set of motion information for the current video. Instead, motion estimation unit 204 may signal the motion information of the current video block by referencing the motion information of another video block. For example, motion estimation unit 204 may determine that the motion information of the current video block is sufficiently similar to the motion information of neighboring video blocks.

[0239] In one example, the motion estimation unit 204 may indicate in the syntax structure associated with the current video block that the current video block has the same motion information value as another video block.

[0240] In another example, motion estimation unit 204 can identify another video block and motion vector difference (MVD) in the syntax structure associated with the current video block. The motion vector difference indicates the difference between the motion vector of the current video block and the motion vector indicating the current video block. Video decoder 300 can use the motion vector indicating the current video block and the motion vector difference to determine the motion vector of the current video block.

[0241] As discussed above, the video encoder 200 can predictively signal motion vectors. Two examples of predictive signaling notification techniques that can be implemented by the video encoder 200 include advanced motion vector prediction (AMVP) and merge pattern signaling notification.

[0242] Intra-prediction unit 206 can perform intra-prediction on the current video block. When intra-prediction unit 206 performs intra-prediction on the current video block, it can generate prediction data for the current video block based on decoded samples from other video blocks in the same frame. The prediction data for the current video block can include the predicted video block and various syntax elements.

[0243] The residual generation unit 207 can generate residual data for the current video block by subtracting (e.g., indicated by a minus sign) multiple predicted video blocks from the current video block. The residual data for the current video block can include residual video blocks corresponding to different sample components of the samples in the current video block.

[0244] In other examples, such as in skip mode, residual data for the current video block may not exist, and the residual generation unit 207 may not perform a subtraction operation.

[0245] The transform processing unit 208 can generate one or more transform coefficient video blocks of the current video block by applying one or more transforms to the residual video block associated with the current video block.

[0246] After the transform processing unit 208 generates a transform coefficient video block associated with the current video block, the quantization unit 209 can quantize the transform coefficient video block associated with the current video block based on one or more quantization parameter (QP) values ​​associated with the current video block.

[0247] The inverse quantization unit 210 and the inverse transform unit 211 can apply inverse quantization and inverse transform to the transform coefficient video block respectively to reconstruct the residual video block from the transform coefficient video block. The reconstruction unit 212 can add the reconstructed residual video block to the corresponding samples of one or more predicted video blocks generated by the prediction unit 202 to produce a reconstructed video block associated with the current block for storage in the buffer 213.

[0248] After the video block is reconstructed in reconstruction unit 212, a loop filtering operation can be performed to reduce video block artifacts in the video block.

[0249] Entropy encoding unit 214 can receive data from other functional components of video encoder 200. When entropy encoding unit 214 receives data, it can perform one or more entropy encoding operations to generate entropy encoded data and output a bitstream including the entropy encoded data.

[0250] Figure 12 This is a block diagram illustrating an example of a video decoder 300, which may be... Figure 10 The video decoder 114 in the system 100 shown in the figure.

[0251] The video decoder 300 can be configured to perform any or all of the techniques disclosed herein. Figure 12 In the example, the video decoder 300 includes multiple functional components. The techniques described in this disclosure can be shared among the various components of the video decoder 300. In some examples, the processor can be configured to perform any or all of the techniques described in this disclosure.

[0252] exist Figure 12 In the example, the video decoder 300 includes an entropy decoding unit 301, a motion compensation unit 302, an intra-frame prediction unit 303, an inverse quantization unit 304, an inverse transform unit 305, a reconstruction unit 306, and a buffer 307. In some examples, the video decoder 300 can perform operations related to the video encoder 200 ( Figure 11 The decoding process is the overall inversion of the encoding process described.

[0253] Entropy decoding unit 301 can retrieve the encoded bitstream. The encoded bitstream may include entropy-encoded video data (e.g., encoded blocks of video data). Entropy decoding unit 301 can decode the entropy-encoded video, and based on the entropy-encoded video data, motion compensation unit 302 can determine motion information including motion vectors, motion vector precision, reference image list index, and other motion information. Motion compensation unit 302 can determine such information, for example, by performing AMVP and merge modes.

[0254] The motion compensation unit 302 can generate motion compensation blocks, possibly based on interpolation filters. The identifier of the interpolation filter to be used at sub-pixel precision can be included in the syntax element.

[0255] The motion compensation unit 302 can use the interpolation filter used by the video encoder 200 during the encoding of the video block to calculate the interpolation values ​​of a sub-integer number of pixels of the reference block. The motion compensation unit 302 can determine the interpolation filter used by the video encoder 200 based on the received syntax information and use the interpolation filter to generate the prediction block.

[0256] The motion compensation unit 302 can use some syntax information to determine: the size of the blocks used to encode (multiple) frames and / or (multiple) stripes of the encoded video sequence, segmentation information describing how each macroblock of the image of the encoded video sequence is segmented, the mode indicating how each segment is encoded, one or more reference frames (and a list of reference frames) for each inter-frame codec block, and other information for decoding the encoded video sequence.

[0257] Intra-prediction unit 303 can use, for example, an intra-prediction mode received in the bitstream to form prediction blocks from spatially adjacent blocks. Inverse quantization unit 303 inverse quantizes (i.e., dequantizes) the quantized video block coefficients provided in the bitstream and decoded by entropy decoding unit 301. Inverse transform unit 303 applies an inverse transform.

[0258] The reconstruction unit 306 can sum the residual blocks using the corresponding prediction blocks generated by the motion compensation unit 202 or the intra-frame prediction unit 303 to form a decoded block. As desired, a deblocking filter can also be applied to filter the decoded block to remove block artifacts. The decoded video block is then stored in a buffer 307, which provides a reference block for subsequent motion compensation / intra-frame prediction and also produces the decoded video for presentation on the display device.

[0259] The following provides a list of preferred solutions for some embodiments.

[0260] The following solutions show example implementations of the techniques discussed in the previous chapter (e.g., Project 1).

[0261] 1. A video processing method (e.g., Figure 9 The method 900 described in the text includes: performing (902) a conversion between a video and a video codec representation comprising one or more video images, wherein the codec representation conforms to a format rule; wherein the format rule specifies that one or more syntax elements contained in the codec representation associated with a scaling window are allowed to have values ​​indicating that the height or width of the scaling window is greater than or equal to the height or width of the corresponding video image.

[0262] The following solutions show example implementations of the techniques discussed in the previous chapter (e.g., Project 2).

[0263] 2. A video processing method, comprising: performing a conversion between a video and a video codec representation comprising one or more video images in one or more video layers, wherein the codec representation conforms to a format rule; wherein the format rule specifies that the allowed values ​​of one or more syntax elements included in the codec representation associated with a scaling window are subject to a constraint rule, wherein the constraint rule depends on the relationship between a first layer of the current image and a second layer of a reference image of the current image.

[0264] 3. The method described in Solution 2, wherein the constraint rule stipulates that all values ​​are allowed if the relationship is the same at the first and second levels.

[0265] The following solutions show example implementations of the techniques discussed in the previous chapter (e.g., Project 3).

[0266] 4. A video processing method, comprising: performing a conversion between a video comprising one or more video images and a video codec representation, wherein the codec representation conforms to a format rule.

[0267] The format rules specify that a syntax field is included in the codec representation, which indicates whether reference image resampling is enabled for the j-th entry in the reference image list i of the video strip.

[0268] The following solutions show example implementations of the techniques discussed in the previous chapter (e.g., Project 5).

[0269] 5. A video processing method, comprising: performing a conversion between a video and a codec representation of a video comprising one or more video layers containing one or more video images, wherein the codec representation conforms to a format rule, wherein the format rule specifies that a syntax element is included in the codec representation in a parameter set, wherein the syntax element indicates whether reference image resampling is enabled for a non-independent video layer, wherein the value of the syntax element is a function of the reference layer of the non-independent video layer.

[0270] 6. The method of any one of solutions 1 to 5, wherein the conversion includes encoding the video into a codec representation.

[0271] 7. The method of any one of solutions 1 to 5, wherein the conversion includes decoding the codec representation to generate pixel values ​​for the video.

[0272] 8. A video decoding apparatus, comprising a processor configured to implement the method described in one or more of solutions 1 to 7.

[0273] 9. A video encoding apparatus, comprising a processor configured to implement the method described in one or more of solutions 1 to 7.

[0274] 10. A computer program product having computer code stored thereon, which, when executed by a processor, causes the processor to implement the method of any one of solutions 1 to 7.

[0275] 11. A method, apparatus or system described in this document.

[0276] In the solution described in this paper, the encoder conforms to the format rules by generating a codec representation based on those rules. In the solution described in this paper, the decoder can use the format rules to parse the syntax elements in the codec representation, determining the presence or absence of syntax elements according to the format rules to generate the decoded video.

[0277] Figure 13This is a flowchart of example method 1300 for video processing. Operation 1302 includes performing a conversion between a video including a video image and a video bitstream according to rules, wherein the rules specify that syntax elements indicate a first width and a first height of a scaling window for the video image, and wherein the rules specify that the range of allowed values ​​for the syntax elements includes values ​​greater than or equal to twice the second width and twice the second height of the video image.

[0278] In some embodiments of method 1300, the value of the syntax element specifies the offset applied to the size of the video image to calculate the scaling ratio. In some embodiments of method 1300, the syntax elements include pps_scaling_win_left_offset, pps_scaling_win_right_offset, pps_scaling_win_top_offset, and pps_scaling_win_bottom_offset. In some embodiments of method 1300, the rule specifies that the value of SubWidthC*(pps_scaling_win_left_offset+pps_scaling_win_right_offset) is greater than or equal to -pps_pic_width_in_luma_samples*M and less than pps_pic_width_in_luma_samples, where CurrPicScalWinWidthL is the first width of the scaling window, where pps_pic_width_in_luma_samples is the second width of the video picture, where SubWidthC is the width of the video block and is obtained from a table according to the chroma format of the video picture containing the video block, where CurrPicScalWinWidthL = pps_pic_width_in_luma_samples - SubWidthC*(pps_scaling_win_right_offset+pps_scaling_win_left_offset). In some embodiments of method 1300, M is equal to 16, 8, 32, or 64.

[0279] In some embodiments of method 1300, the rule specifies that the value of SubHeightC*(pps_scaling_win_top_offset+pps_scaling_win_bottom_offset) is greater than or equal to -pps_pic_height_in_luma_samples*N and less than pps_pic_height_in_luma_samples, where CurrPicScalWinHeightL is the first height of the scaling window, where pps_pic_height_in_luma_samples is the second height of the video picture, where SubHeightC is the height of the video block and is obtained from a table according to the chroma format of the video picture containing the video block, where CurrPicScalWinHeightL = pps_pic_height_in_luma_samples - SubHeightC*(pps_scaling_win_bottom_offset+pps_scaling_win_top_offset). In some embodiments of method 1300, N is equal to 16, 8, 32, or 64. In some embodiments of method 1300, M and N depend on the layout of the subpictures of the video picture. In some embodiments of method 1300, M is equal to Ceil(pps_pic_width_in_luma_samples ÷ minSubpicWidth), where minSubpicWidth is the minimum subpicture width of all subpictures in the picture of the reference picture parameter set, where pps_pic_width_in_luma_samples is the second width of the video picture, and where the Ceil() function rounds the number to the next largest integer.

[0280] In some embodiments of method 1300, N equals Ceil(pps_pic_height_in_luma_samples÷minSubpicHeight), where minSubpicHeight is the minimum subpic height of all subpics in the video picture of the reference picture parameter set, where pps_pic_height_in_luma_samples is the second height of the video picture, and where the Ceil() function rounds the number to the next largest integer.

[0281] Figure 14This is a flowchart of example method 1400 for video processing. Operation 1402 includes a conversion between a video picture and a bitstream of the video, determining (1) whether or how a first width or first height of a first scaling window of a reference video picture and (2) a second width or second height of a second scaling window of the current video picture are constrained according to rules. Operation 1404 includes performing the conversion based on the determination.

[0282] In some embodiments of method 1400, the rule specifies that, in response to the reference video image and the current video image being on different video layers, there are no constraints on the first width or first height of the first zoom window and the second width or second height of the second zoom window. In some embodiments of method 1400, the rules specify that: CurrPicScalWinWidthL*M is greater than or equal to refPicScalWinWidthL, CurrPicScalWinHeightL*M is greater than or equal to refPicScalWinHeightL, CurrPicScalWinWidthL is less than or equal to refPicScalWinWidthL*N, CurrPicScalWinHeightL is less than or equal to refPicScalWinHeightL*N, CurrPicScalWinWidthL*sps_pic_width_max_in_luma_samples is greater than or equal to refPicScalWinWidthL*(pps_pic_width_in_luma_samples-Max(K0,MinCbSizeY)), CurrPicScalWinHeightL*sps_ `pic_height_max_in_luma_samples` is greater than or equal to `refPicScalWinHeightL*(pps_pic_height_in_luma_samples-Max(K1,MinCbSizeY))`, ​​where `CurrPicScalWinWidthL` and `CurrPicScalWinHeightL` are the second width and second height of the second scaling window of the current video image, respectively; `refPicScalWinWidthL` and `refPicScalWinHeightL` are the first width and first height of the first scaling window of the reference video image, respectively; `sps_pic_width_max_in_luma_samples` is the maximum width of each video image in luminance samples of the reference sequence parameter set; and `pps_pic_width_in_luma_samples` is the width of each video image in the reference picture parameter set.

[0283] In some embodiments of method 1400, the rule specifies that, in response to the reference video image and the current video image being in the same video layer, M = M1 and N = N1; wherein, the rule specifies that, in response to the reference video image and the current video image being in different video layers, M = M2 and N = N2; wherein M1 is less than or equal to M2; and wherein N1 is greater than or equal to N2. In some embodiments of method 1400, K0 or K1 is an integer value. In some embodiments of method 1400, K0 or K1 is the minimum allowed video image width or the minimum allowed video image height, respectively. In some embodiments of method 1400, K0 or K1 is 8.

[0284] Figure 15 This is a flowchart of example method 1500 for video processing. Operation 1502 includes performing a conversion between a video and a video bitstream comprising video images containing one or more video stripes, wherein the bitstream conforms to a rule that specifies the values ​​of a variable determining the width and height of a scaling window for the video images, and the variable specifying whether resampling of the j-th reference video image in the i-th reference image list is enabled for the video stripe of the video image, where i and j are integers.

[0285] In some embodiments of method 1500, the variable is RprConstraintsActive[i][j], where the rule specifies that RprConstraintsActive[i][j] = (CurrPicScalWinWidthL != fRefWidth ||CurrPicScalWinHeightL != fRefHeight), where fRefWidth and fRefHeight are the values ​​of CurrPicScalWinWidthL and CurrPicScalWinHeightL of the j-th reference video image, respectively. In some embodiments of method 1500, the rule specifies that the value of the variable is based on the first video layer to which the j-th reference video image belongs and the second video layer to which the current video image belongs. In some embodiments of method 1500, the rule specifies that the value of the variable is based on whether the first video layer of the j-th reference video image and the second video layer of the current video image are the same video layer.

[0286] Figure 16This is a flowchart of example method 1600 for video processing. Operation 1602 includes performing a conversion between a video and a video bitstream comprising one or more video images according to format rules, wherein the format rules specify that syntax elements are included in a sequence parameter set, and wherein the syntax elements indicate whether reference image resampling is enabled for reference images and whether one or more stripes of the current video image in the codec layer video sequence are allowed to reference reference images in the active entry of the reference image list, and wherein the reference image has any one or more of the following six parameters that are different from those of the current video image: 1) video image width, 2) video image height, 3) scaling window left offset, 4) scaling window right offset, 5) scaling window top offset, and 6) scaling window bottom offset.

[0287] In some embodiments of method 1600, a syntax element with a value of 1 indicates that reference picture resampling is enabled, and indicates that reference pictures in one or more active entries of the slice reference reference picture list of the current video picture in the codec layer video sequence are allowed. In some embodiments of method 1600, a syntax element with a value of 0 indicates that reference picture resampling is disabled, and indicates that reference pictures in the active entries of the slice reference reference picture list of the current video picture in the codec layer video sequence are not allowed.

[0288] Figure 17 This is a flowchart of example method 1700 for video processing. Operation 1702 includes performing a conversion between a video and a video bitstream comprising one or more video images according to format rules, wherein the format rules specify that syntax elements are included in a sequence parameter set, wherein the format rules specify that the value of the syntax element is based on (1) whether the current video layer of the reference sequence parameter set is not an independent video layer and (2) information of one or more reference video layers associated with the current video layer, wherein the video layer is an independent video layer when it does not depend on one or more other video layers, and wherein the video layer is not an independent video layer when it depends on one or more other video layers.

[0289] In some embodiments of method 1700, the format rule specifies that the value of the syntax element is 1 in response to (1) the current video layer having the same nuh_layer_id as the sequence parameter set is not an independent video layer such that the current video layer depends on one or more other video layers, and in response to (2) at least one reference video layer of the current video layer having a different spatial resolution than the current video layer. In some embodiments of method 1700, the format rule specifies that the value of the syntax element is 1 in response to (1) at least one video layer of the reference sequence parameter set is not an independent video layer such that at least one video layer depends on one or more other video layers, and in response to (2) at least one reference video layer of at least one video layer having a different spatial resolution than the current video layer. In some embodiments of method 1700, at least one strip is included in at least one video picture, at least one video picture reference sequence parameter set, and wherein the format rules specify that, in response to a reference picture in an interlayer reference picture entry of at least one strip reference reference picture list, the value of the syntax element is 1, and the reference picture has any one or more of the following six parameters different from the at least one video picture: 1) video picture width, 2) video picture height, 3) zoom window left offset, 4) zoom window right offset, 5) zoom window top offset, and 6) zoom window bottom offset.

[0290] Figure 18 This is a flowchart of example method 1800 for video processing. Operation 1802 includes performing a conversion between a video and a video bitstream comprising one or more video images according to format rules, wherein the format rules specify that syntax elements are indicated in a sequence parameter set referenced by one or more video images; and wherein the syntax elements indicate whether reference image resampling is enabled for one or more reference images in the same layer as one or more video images.

[0291] In some embodiments of method 1800, the format rule specifies that the value of the syntax element is 1, which indicates: (a) enabling reference image resampling for one or more reference images on the same layer as one or more video images, and (b) a reference image in the active entry of a list of one or more stripes of one or more video images in the codec layer video sequence, wherein the reference image (1) has the same nuh_layer_id value as the current video image of one or more video images, and (2) has any one or more of the following six parameters that are different from the current video image: 1) video image width, 2) video image height, 3) left offset of the scaling window, 4) right offset of the scaling window, 5) top offset of the scaling window, and 6) bottom offset of the scaling window.

[0292] In some embodiments of method 1800, the format rules specify that the value of a syntax element is 0, which indicates that: (a) reference image resampling is disabled for one or more reference images on the same layer as one or more video images, and (b) the stripes of one or more video images in the codec layer video sequence do not reference reference images in the active entry of the reference image list, the reference images (1) having the same nuh_layer_id value as the current video image of one or more video images, and (2) having any one or more of the following six parameters that are different from the current video image: 1) video image width, 2) video image height, 3) scaling window left offset, 4) scaling window right offset, 5) scaling window top offset, and 6) scaling window bottom offset. In some embodiments of method 1800, the format rule specifies that the value of the syntax element is 1, which indicates that reference image resampling is enabled for a reference image in the same layer as one or more video images, and wherein the reference image is associated with RprConstraintsActive[i][j] equal to 1, and wherein the reference image is the j-th entry in the reference image list i of the current strip, and wherein the j-th entry in the reference image list i has the same nuh_layer_id value as the current video image of one or more video images.

[0293] In some embodiments of method 1800, the format rule specifies that a syntax element is included in the bitstream before a second syntax element, the second syntax element indicating whether the image spatial resolution changes within the codec layer video sequence of the reference sequence parameter set. In some embodiments of method 1800, the format rule specifies that the bitstream does not include the second syntax element, the second syntax element indicating whether the image spatial resolution changes within the codec layer video sequence of the reference sequence parameter set, wherein the format rule specifies that a second value of the second syntax element is inferred to be 0, and wherein the format rule specifies that the bitstream does not include the second syntax element in response to a value of 0.

[0294] In some embodiments of methods 1300-1800, performing the conversion includes encoding video into a bitstream. In some embodiments of methods 1300-1800, performing the conversion includes generating a bitstream from video, and the method further includes storing the bitstream in a non-transitory computer-readable recording medium. In some embodiments of methods 1300-1800, performing the conversion includes decoding video from the bitstream. In some embodiments, a video decoding apparatus includes a processor configured to implement the methods described in one or more embodiments of methods 1300-1800. In some embodiments, a video encoding / decoding apparatus includes a processor configured to implement the methods described in one or more embodiments of methods 1300-1800. In some embodiments, a computer program product having computer instructions stored thereon, which, when executed by a processor, cause the processor to implement the methods described in the embodiments of methods 1300-1800. In some embodiments, a non-transitory computer-readable storage medium stores a bitstream generated according to the methods in the embodiments of methods 1300-1800. In some embodiments, a non-transitory computer-readable storage medium stores instructions that cause a processor to perform the methods(s) of embodiments of methods 1300-1800. In some embodiments, a bitstream generation method includes: generating a bitstream of video according to the methods(s) of embodiments of methods 1300-1800, and storing the bitstream on a computer-readable program medium. In some embodiments, a method, apparatus, bitstream generated according to the disclosed methods or systems described in this document.

[0295] In this document, the term "video processing" can refer to video encoding, video decoding, video compression, or video decompression. For example, a video compression algorithm can be applied during the conversion from the pixel representation of a video to the corresponding bitstream representation, and vice versa. As defined in the syntax, the bitstream representation of the current video block can, for example, correspond to bits that are co-occurring or scattered at different positions within the bitstream. For example, a macroblock can be encoded based on the error residual values ​​of the transformation and encoding / decoding, and also using bits in the header and other fields in the bitstream. Furthermore, during the conversion, the decoder can, based on this determination, parse the bitstream knowing that some fields may or may not be present, as described in the solutions above. Similarly, the encoder can determine whether to include or exclude certain syntax fields and generate the codec representation accordingly by including or excluding syntax fields from the codec representation.

[0296] Some embodiments of the technology disclosed herein include determining or enabling a video processing tool or mode. In one example, when a video processing tool or mode is enabled, the encoder will use or implement the tool or mode in the processing of video blocks, but it is not necessary to modify the resulting bitstream based on the use of the tool or mode. In other words, when a video processing tool or mode is determined or enabled, the conversion from video blocks to a video bitstream will use the video processing tool or mode. In another example, when a video processing tool or mode is enabled, the decoder will utilize knowledge that the bitstream has already been modified based on the video processing tool or mode to process the bitstream. In other words, using a video processing tool or mode determined or enabled, the conversion from a video bitstream to video blocks will be performed.

[0297] Some embodiments of the technology disclosed herein include deciding or determining to disable video processing tools or modes. In one example, when video processing tools or modes are disabled, the encoder will not use the tools or modes in the conversion of video blocks to a bitstream representation of the video. In another example, when video processing tools or modes are disabled, the decoder will process the bitstream using knowledge that the bitstream has not yet been modified based on the video processing tools or modes that have been decided or determined to be disabled.

[0298] The disclosures and other schemes, examples, embodiments, modules, and functional operations described in this document can be implemented in digital electronic circuits or in computer software, firmware, or hardware, containing the structures disclosed in this document and their equivalents, or combinations thereof. The disclosed and other embodiments can be implemented as one or more computer program products encoded on a computer-readable medium, i.e., one or more computer program instruction modules for execution by a data processing apparatus or for controlling the operation of a data processing apparatus. The computer-readable medium can be a machine-readable storage device, a machine-readable storage substrate, a memory device, a complex influencing machine-readable propagating signals, or combinations thereof. The term "data processing apparatus" encompasses all means, devices, and machines for processing data, including, for example, a programmable processor, a computer, or multiple processors or computers. In addition to hardware, the apparatus may also include code that creates an execution environment for the computer program in question, such as code constituting processor firmware, a protocol stack, a database management system, an operating system, or combinations thereof. Propagating signals are artificially generated signals, such as machine-generated electrical, optical, or electromagnetic signals, which are generated to encode information for transmission to a suitable receiver device.

[0299] Computer programs (also known as programs, software, software applications, scripts, or code) can be written in any programming language, including compiled or interpreted languages, and can be deployed in any form, including standalone programs or modules, components, subroutines, or other units suitable for use in a computing environment. A computer program does not necessarily correspond to a file in a file system. A program can be stored in a portion of a file that holds other programs or data (e.g., one or more scripts stored in a markup language document), in a single file dedicated to the program in question, or in multiple co-located files (e.g., a file storing one or more modules, subroutines, or code portions). A computer program can be deployed to execute on one computer or on multiple computers located at a single site or distributed across multiple sites and interconnected by a communications network.

[0300] The processes and logic flows described in this document can be performed by one or more programmable processors executing one or more computer programs to perform the functions by manipulating input data and generating outputs. The processes and logic flows can also be performed by special-purpose logic circuitry (e.g., field-programmable gate arrays (FPGAs) or application-specific integrated circuits (ASICs), and the apparatus can be implemented as special-purpose logic circuitry (e.g., FPGAs or ASICs).

[0301] Processors suitable for executing computer programs include, for example, both general-purpose and special-purpose microprocessors, and any one or more processors in any type of digital computer. Typically, a processor receives instructions and data from read-only memory or random access memory, or both. The basic components of a computer are a processor for executing instructions and one or more memory devices for storing instructions and data. Typically, a computer will also include one or more mass storage devices (e.g., magneto-optical, magneto-optical, or optical disc) for storing data, or operatively coupled to receive data from or transfer data to a mass storage device (e.g., magneto-optical, magneto-optical, or optical disc), or both. However, a computer does not necessarily need to have such devices. Computer-readable media suitable for storing computer program instructions and data include all forms of non-volatile memory, media, and memory devices, including, for example, semiconductor memory devices (e.g., EPROM, EEPROM, and flash memory devices); magnetic disks (e.g., internal hard disks or removable disks); magneto-optical disks; and CD-ROM and DVD-ROM disks. Processors and memory may be supplemented by or incorporated into special-purpose logic circuitry.

[0302] While this patent document contains numerous details, these details should not be construed as limiting any subject matter or the scope of the claims, but rather as descriptions of features specific to particular embodiments of a particular art. In this patent document, certain features described in the context of individual embodiments may also be implemented in combination in a single embodiment. Conversely, various features described in the context of a single embodiment may also be implemented separately in multiple embodiments or in various suitable sub-combinations. Furthermore, although features may be described above as operating in certain combinations and even initially claimed in the same manner, in certain circumstances one or more features from the claimed combination may be removed from the combination, and the claimed combination may be for sub-combinations or variations thereof.

[0303] Similarly, although operations are depicted in a specific order in the accompanying drawings, this should not be construed as requiring such operations to be performed in the specific order or sequence shown, or to perform all the operations shown, in order to achieve the desired result. Furthermore, the separation of various system components in the embodiments described in this patent document should not be construed as requiring such separation in all embodiments.

[0304] Only a few implementations and examples are described, and other implementations, enhancements and variations can be made based on what is described and shown in this patent document.

Claims

1. A video processing method, comprising: Perform the conversion between the video, including video images, and the bitstream of the video according to the rules. The rule specifies that the first syntax element indicates the first width and first height of the zoom window for the video image, and The rule specifies that the allowed values ​​of the first width and first height, indicated by the first syntax element, are greater than or equal to twice the second width and twice the second height of the video image. The first syntax element includes pps_scaling_win_left_offset, pps_scaling_win_right_offset, pps_scaling_win_top_offset, and pps_scaling_win_bottom_offset. The rule states that SubWidthC The value of (pps_scaling_win_left_offset + pps_scaling_win_right_offset) is greater than or equal to pps_pic_width_in_luma_samples M, and less than pps_pic_width_in_luma_samples, Where CurrPicScalWinWidthL is the first width of the scaled window. Where pps_pic_width_in_luma_samples is the second width of the video image. Wherein, SubWidthC is the chroma width sampling variable of the video image, obtained from the table according to the chroma format of the video image. Wherein, CurrPicScalWinWidthL = pps_pic_width_in_luma_samples SubWidthC ( pps_scaling_win_right_offset + pps_scaling_win_left_offset ), and Where M is a positive integer value greater than or equal to 2.

2. The method according to claim 1, wherein, The value of the first syntax element specifies an offset, which is applied to the size of the video image to calculate the scaling ratio.

3. The method according to claim 1, wherein, The rule states that SubHeightC The value of (pps_scaling_win_top_offset + pps_scaling_win_bottom_offset) is greater than or equal to -pps_pic_height_in_luma_samples N is less than pps_pic_height_in_luma_samples Where CurrPicScalWinHeightL is the first height of the scaling window. Where pps_pic_height_in_luma_samples is the second height of the video image. Wherein, SubHeightC is the chroma height sampling variable of the video image, obtained from the table according to the chroma format of the video image. Wherein, CurrPicScalWinHeightL = pps_pic_height_in_luma_samples SubHeightC ( pps_scaling_win_bottom_offset + pps_scaling_win_top_offset ), and Where N is a positive integer value greater than or equal to 2.

4. The method according to claim 3, wherein, N equals M.

5. The method according to claim 1, wherein, The rule stipulates that the second syntax element is included in the sequence parameter set referenced by the video image. The second syntax element indicates whether reference image resampling is enabled and whether one or more stripes of the video image are allowed to reference reference images in the active entry of the reference image list. The reference image has any one or more of a plurality of parameters that are different from the video image. The plurality of parameters include the following six parameters specified by the first syntax element: 1) video image width, 2) video image height, 3) zoom window left offset, 4) zoom window right offset, 5) zoom window top offset, and 6) zoom window bottom offset.

6. The method according to claim 5, wherein, The second syntax element having a value of 1 indicates that the reference image resampling is enabled, and indicates that the video image is allowed to have one or more stripes referencing the reference image in the active entry of the reference image list.

7. The method according to claim 5, wherein, The second syntax element with a value of 0 indicates that the reference image resampling is disabled, and indicates that the stripe of the video image is not allowed to reference the reference image in the active entry of the reference image.

8. The method according to claim 1, wherein, Performing the conversion includes encoding the video into the bitstream.

9. The method according to claim 1, wherein, Performing the conversion includes decoding the video from the bitstream.

10. An apparatus for processing video data, comprising a processor and a non-transitory memory having instructions thereon, wherein, When the instruction is executed by the processor, the processor: Perform the conversion between the video, including video images, and the bitstream of the video according to the rules. The rule specifies that the first syntax element indicates the first width and first height of the zoom window for the video image, and The rule specifies that the allowed values ​​of the first width and first height, indicated by the first syntax element, are greater than or equal to twice the second width and twice the second height of the video image. The first syntax element includes pps_scaling_win_left_offset, pps_scaling_win_right_offset, pps_scaling_win_top_offset, and pps_scaling_win_bottom_offset. The rule states that SubWidthC The value of (pps_scaling_win_left_offset + pps_scaling_win_right_offset) is greater than or equal to pps_pic_width_in_luma_samples M, and less than pps_pic_width_in_luma_samples, Where CurrPicScalWinWidthL is the first width of the scaled window. Where pps_pic_width_in_luma_samples is the second width of the video image. Wherein, SubWidthC is the chroma width sampling variable of the video image, obtained from the table according to the chroma format of the video image. Wherein, CurrPicScalWinWidthL = pps_pic_width_in_luma_samples SubWidthC ( pps_scaling_win_right_offset + pps_scaling_win_left_offset ), and Where M is a positive integer value greater than or equal to 2.

11. The apparatus according to claim 10, wherein, The value of the first syntax element specifies an offset, which is applied to the size of the video image to calculate the scaling ratio.

12. The apparatus according to claim 10, wherein, The rule states that SubHeightC The value of (pps_scaling_win_top_offset + pps_scaling_win_bottom_offset) is greater than or equal to pps_pic_height_in_luma_samples N is less than pps_pic_height_in_luma_samples Where CurrPicScalWinHeightL is the first height of the scaling window. Where pps_pic_height_in_luma_samples is the second height of the video image. Wherein, SubHeightC is the chroma height sampling variable of the video image, obtained from the table according to the chroma format of the video image. Wherein, CurrPicScalWinHeightL = pps_pic_height_in_luma_samples SubHeightC ( pps_scaling_win_bottom_offset + pps_scaling_win_top_offset ), and Where N is a positive integer value greater than or equal to 2.

13. The apparatus according to claim 12, wherein, N equals M.

14. The apparatus according to claim 10, wherein, The rule stipulates that the second syntax element is included in the sequence parameter set referenced by the video image. The second syntax element indicates whether reference image resampling is enabled and whether one or more stripes of the video image are allowed to reference reference images in the active entry of the reference image list. The reference image has one or more parameters that differ from the video image. These parameters include the following six parameters specified by the first syntax element: 1) video image width, 2) video image height, 3) left offset of the zoom window, 4) right offset of the zoom window, 5) top offset of the zoom window, and 6) bottom offset of the zoom window. The second syntax element having a value of 1 indicates that the reference image resampling is enabled, and indicates that the video image is allowed to have one or more stripes referencing the reference image, which is in the active entry of the reference image list. The second syntax element having a value of 0 indicates that the reference image resampling is disabled, and indicates that the strip of the video image is not allowed to reference the reference image in the active entry of the reference image.

15. A non-transitory computer-readable storage medium for storing instructions, said instructions causing a processor to: Perform the conversion between the video, including video images, and the bitstream of the video according to the rules. in, The rule specifies that the first syntax element indicates the first width and first height of the zoom window for the video image, and The rule specifies that the allowed values ​​of the first width and first height, indicated by the first syntax element, are greater than or equal to twice the second width and twice the second height of the video image. The first syntax element includes pps_scaling_win_left_offset, pps_scaling_win_right_offset, pps_scaling_win_top_offset, and pps_scaling_win_bottom_offset. The rule states that SubWidthC The value of (pps_scaling_win_left_offset + pps_scaling_win_right_offset) is greater than or equal to pps_pic_width_in_luma_samples M, and less than pps_pic_width_in_luma_samples, Where CurrPicScalWinWidthL is the first width of the scaled window. Where pps_pic_width_in_luma_samples is the second width of the video image. Wherein, SubWidthC is the chroma width sampling variable of the video image, obtained from the table according to the chroma format of the video image. wherein, CurrPicScalWinWidthL = pps_pic_width_in_luma_samples SubWidthC ( pps_scaling_win_right_offset + pps_scaling_win_left_offset ), and Where M is a positive integer value greater than or equal to 2.

16. The non-transitory computer-readable storage medium according to claim 15, wherein, The value of the first syntax element specifies an offset, which is applied to the dimensions of the video image to calculate the scaling ratio. The rule states that SubHeightC The value of (pps_scaling_win_top_offset + pps_scaling_win_bottom_offset) is greater than or equal to pps_pic_height_in_luma_samples N is less than pps_pic_height_in_luma_samples Where CurrPicScalWinHeightL is the first height of the scaling window. Where pps_pic_height_in_luma_samples is the second height of the video image. Wherein, SubHeightC is the chroma height sampling variable of the video image, obtained from the table according to the chroma format of the video image. wherein, CurrPicScalWinHeightL = pps_pic_height_in_luma_samples SubHeightC ( pps_scaling_win_bottom_offset + pps_scaling_win_top_offset ), and Where N is a positive integer greater than or equal to 2, and N equals M. The rule stipulates that the second syntax element is included in the sequence parameter set referenced by the video image. The second syntax element indicates whether reference image resampling is enabled and whether one or more stripes of the video image are allowed to reference reference images in the active entry of the reference image list. The reference image has one or more parameters that differ from the video image. These parameters include the following six parameters specified by the first syntax element: 1) video image width, 2) video image height, 3) left offset of the zoom window, 4) right offset of the zoom window, 5) top offset of the zoom window, and 6) bottom offset of the zoom window. The second syntax element having a value of 1 indicates that the reference image resampling is enabled, and indicates that the video image is allowed to have one or more stripes referencing the reference image, which is in the active entry of the reference image list. The second syntax element having a value of 0 indicates that the reference image resampling is disabled, and indicates that the strip of the video image is not allowed to reference the reference image in the active entry of the reference image.

17. A non-transitory computer-readable recording medium for storing a bitstream of video and a computer program, wherein, When the computer program is executed by a processor, it performs the following steps to generate the bitstream: The bitstream of the video, including video images, is generated according to the format rules. The rule specifies that the first syntax element indicates the first width and first height of the zoom window for the video image, and The rule specifies that the allowed values ​​of the first width and first height, indicated by the first syntax element, are greater than or equal to twice the second width and twice the second height of the video image. The first syntax element includes pps_scaling_win_left_offset, pps_scaling_win_right_offset, pps_scaling_win_top_offset, and pps_scaling_win_bottom_offset. The rule states that SubWidthC The value of (pps_scaling_win_left_offset + pps_scaling_win_right_offset) is greater than or equal to pps_pic_width_in_luma_samples M, and less than pps_pic_width_in_luma_samples, Where CurrPicScalWinWidthL is the first width of the scaled window. Where pps_pic_width_in_luma_samples is the second width of the video image. Wherein, SubWidthC is the chroma width sampling variable of the video image, obtained from the table according to the chroma format of the video image. Wherein, CurrPicScalWinWidthL = pps_pic_width_in_luma_samples SubWidthC ( pps_scaling_win_right_offset + pps_scaling_win_left_offset ), and Where M is a positive integer value greater than or equal to 2.

18. The non-transitory computer-readable recording medium according to claim 17, wherein, The value of the first syntax element specifies an offset, which is applied to the dimensions of the video image to calculate the scaling ratio. The rule states that SubHeightC The value of (pps_scaling_win_top_offset + pps_scaling_win_bottom_offset) is greater than or equal to pps_pic_height_in_luma_samples N is less than pps_pic_height_in_luma_samples Where CurrPicScalWinHeightL is the first height of the scaling window. Where pps_pic_height_in_luma_samples is the second height of the video image. Wherein, SubHeightC is the chroma height sampling variable of the video image, obtained from the table according to the chroma format of the video image. Wherein, CurrPicScalWinHeightL = pps_pic_height_in_luma_samples SubHeightC ( pps_scaling_win_bottom_offset + pps_scaling_win_top_offset ), and Where N is a positive integer greater than or equal to 2, and N equals M. The rule stipulates that the second syntax element is included in the sequence parameter set referenced by the video image. The second syntax element indicates whether reference image resampling is enabled and whether one or more stripes of the video image are allowed to reference reference images in the active entry of the reference image list. The reference image has one or more parameters that differ from the video image. These parameters include the following six parameters specified by the first syntax element: 1) video image width, 2) video image height, 3) left offset of the zoom window, 4) right offset of the zoom window, 5) top offset of the zoom window, and 6) bottom offset of the zoom window. The second syntax element having a value of 1 indicates that the reference image resampling is enabled, and indicates that the video image is allowed to have one or more stripes referencing the reference image, which is in the active entry of the reference image list. The second syntax element having a value of 0 indicates that the reference image resampling is disabled, and indicates that the strip of the video image is not allowed to reference the reference image in the active entry of the reference image.

19. A method for storing a bitstream of video, comprising: The bitstream of the video, including video images, is generated according to the format rules; as well as The bitstream is stored in a non-transitory computer-readable recording medium. The rule specifies that the first syntax element indicates the first width and first height of the zoom window for the video image, and The rule specifies that the allowed values ​​of the first width and first height, indicated by the first syntax element, are greater than or equal to twice the second width and twice the second height of the video image. The first syntax element includes pps_scaling_win_left_offset, pps_scaling_win_right_offset, pps_scaling_win_top_offset, and pps_scaling_win_bottom_offset. The rule states that SubWidthC The value of (pps_scaling_win_left_offset + pps_scaling_win_right_offset) is greater than or equal to pps_pic_width_in_luma_samples M, and less than pps_pic_width_in_luma_samples, Where CurrPicScalWinWidthL is the first width of the scaled window. Where pps_pic_width_in_luma_samples is the second width of the video image. Wherein, SubWidthC is the chroma width sampling variable of the video image, obtained from the table according to the chroma format of the video image. Wherein, CurrPicScalWinWidthL = pps_pic_width_in_luma_samples SubWidthC ( pps_scaling_win_right_offset + pps_scaling_win_left_offset ), and Where M is a positive integer value greater than or equal to 2.