Method for signaling virtual boundary and wraparound motion compensation - Patents.com

By disabling in-loop filtering across virtual boundaries and using wraparound motion compensation, the method addresses visible artifacts in 360-degree video and ultra-low-delay applications, enhancing compression performance and reducing latency.

JP7680453B2Active Publication Date: 2025-05-20HFI INNOVATION INC
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
JP2022537463
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Priority Date
2019-12-30
Filing Date
2020-12-22
Publication Date
2025-05-20
Estimated Expiration
2040-12-22

AI Technical Summary

Technical Problem

Existing video coding standards face issues with visible surface seam artifacts due to in-loop filtering across discontinuities in frame-packed pictures, particularly in 360-degree video and ultra-low-delay applications, where virtual boundaries are not effectively managed.

Method used

The method involves disabling in-loop filtering operations across virtual boundaries by signaling their positions in the bitstream, using sequence parameter sets (SPS) or picture headers (PH), and employing wraparound motion compensation to mitigate artifacts in 360-degree video and gradual decoding refresh (GDR) applications.

Benefits of technology

This approach reduces visible artifacts and improves compression performance by preventing seam artifacts and minimizing latency in 360-degree video and ultra-low-delay applications.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007680453000001
    Figure 0007680453000001
  • Figure 0007680453000002
    Figure 0007680453000002
  • Figure 0007680453000003
    Figure 0007680453000003
Patent Text Reader

Abstract

This disclosure provides a method for picture processing, which may include receiving a bitstream including a set of pictures, determining, according to the received bitstream, whether a virtual boundary is signaled at a sequence level for the set of pictures, determining, in response to the virtual boundary being signaled at the sequence level, a position of the virtual boundary for the set of pictures, the position being constrained by a range signaled in the received bitstream, and disabling in-loop filtering operations that cross the virtual boundary.
Need to check novelty before this filing date? Find Prior Art

Description

[Technical field]

[0001] CROSS-REFERENCE TO RELATED APPLICATIONS

[0001] This disclosure claims priority to U.S. Provisional Patent Application No. 62 / 954,828, filed December 30, 2019, the entirety of which is incorporated by reference herein.

[0002] Technical Field FIELD OF THE DISCLOSURE

[0002] This disclosure relates generally to video processing, and more particularly to methods for signaling virtual boundaries and wraparound motion compensation. [Background technology]

[0003] background

[0003] A video is a set of static pictures (or "frames") that capture visual information. To reduce storage memory and transmission bandwidth, a video can be compressed before storage or transmission and decompressed before display. The compression process is usually called encoding, and the decompression process is usually called decoding. There are various video coding formats that use standardized video coding techniques, most commonly based on prediction, transformation, quantization, entropy coding, and in-loop filtering. Standardization organizations have developed video coding standards, such as the High Efficiency Video Coding (HEVC / H.265) standard, the Versatile Video Coding (VVC / H.266) standard, and the AVS standard, that specify certain video coding formats. As more and more advanced video coding techniques are adopted into video standards, the coding efficiency of the new video coding standards becomes higher. Summary of the Invention [Means for solving the problem]

[0004]

[0004] This disclosure provides a method for picture processing, which may include receiving a bitstream including a set of pictures, determining whether a virtual boundary is signaled at a sequence level for the set of pictures according to the received bitstream, determining a position of a virtual boundary for the set of pictures in response to the virtual boundary being signaled at the sequence level, the position being constrained by a range signaled in the received bitstream, and disabling in-loop filtering operations that cross the virtual boundary.

[0005]

[0005] An embodiment of the present disclosure further provides an apparatus for picture processing, which may include a memory storing a set of instructions and one or more processors, configured to execute the set of instructions to cause the apparatus to receive a bitstream including a set of pictures, determine whether a virtual boundary is signaled at a sequence level for the set of pictures according to the received bitstream, determine a position of the virtual boundary for the set of pictures in response to the virtual boundary being signaled at the sequence level, the position being constrained by a range signaled in the received bitstream, and disable in-loop filtering operations that cross the virtual boundary.

[0006]

[0006] An embodiment of the present disclosure further provides a non-transitory computer-readable medium storing a set of instructions, the set of instructions executable by at least one processor of the computer to cause the computer to perform a picture processing method, the method including receiving a bitstream including a set of pictures, determining in accordance with the received bitstream whether a virtual boundary is signaled at a sequence level for the set of pictures, in response to the virtual boundary being signaled at the sequence level, determining a position of the virtual boundary for the set of pictures, the position being constrained by a range signaled in the received bitstream, and disabling in-loop filtering operations that cross the virtual boundary.

[0007] BRIEF DESCRIPTION OF THE DRAWINGS

[0007] Embodiments and various aspects of the present disclosure are illustrated in the following detailed description and the accompanying drawings, in which the various features illustrated are not drawn to scale. [Brief description of the drawings]

[0008] [Figure 1] 1 illustrates an exemplary video sequence structure according to some embodiments of the present disclosure. [Figure 2A]

[0009] 1 shows a schematic diagram of an example encoding process for a hybrid video encoding system in accordance with some embodiments of the present disclosure. [Figure 2B]

[0010] 4 shows a schematic diagram of another example encoding process of a hybrid video encoding system in accordance with some embodiments of the present disclosure. [Figure 3A]

[0011] 1 shows a schematic diagram of an example decoding process of a hybrid video coding system in accordance with some embodiments of the present disclosure. [Figure 3B]

[0012] 4 shows a schematic diagram of another example decoding process of a hybrid video coding system in accordance with some embodiments of the present disclosure. [Figure 4]

[0013] 1 shows a block diagram of an example device for encoding or decoding video in accordance with some embodiments of the present disclosure. [Diagram 5]

[0014] 1 illustrates an example sequence parameter set (SPS) syntax for signaling virtual boundaries in accordance with some embodiments of the present disclosure. [Figure 6]

[0015] 1 illustrates an example picture header (PH) syntax for signaling virtual boundaries according to some embodiments of the present disclosure. [Figure 7A]

[0016] 1 illustrates an example horizontal wraparound motion compensation for equirectangular projection (ERP) in accordance with some embodiments of the present disclosure. [Figure 7B]

[0017] 1 illustrates an example horizontal wraparound motion compensation for padded ERP (PERP) in accordance with some embodiments of the present disclosure. [Figure 8]

[0018] 1 illustrates an example syntax for wraparound motion compensation in accordance with some embodiments of this disclosure. [Figure 9]

[0019] 1 illustrates an example SPS syntax for signaling maximum picture width and height values ​​according to some embodiments of the present disclosure. [Figure 10]

[0020] 1 illustrates an example syntax for signaling wraparound motion compensation in accordance with some embodiments of the present disclosure. [Figure 11]

[0021] 1 illustrates an example SPS syntax for signaling virtual boundaries, according to some embodiments of the present disclosure. [Figure 12]

[0022] 1 illustrates another example SPS syntax for signaling virtual boundaries, according to some embodiments of the present disclosure. [Figure 13]

[0023] 1 illustrates an example SPS syntax for signaling wraparound motion compensation in accordance with some embodiments of the present disclosure. [Figure 14]

[0024] 1 illustrates another example SPS syntax for signaling wraparound motion compensation, according to some embodiments of the present disclosure. [Figure 15]

[0025] 1 illustrates an example method for picture processing according to some embodiments of the present disclosure. DETAILED DESCRIPTION OF THE PREFERRED EMBODIMENTS

[0009] Detailed Description

[0026] Reference will now be made in detail to the exemplary embodiments, examples of which are illustrated in the accompanying drawings. The following description refers to the accompanying drawings, in which the same reference numerals in different drawings represent the same or similar elements unless otherwise indicated. The implementations illustrated in the following description of the exemplary embodiments do not represent all implementations in accordance with the present disclosure. Rather, they are merely examples of apparatus and methods in accordance with aspects related to the present disclosure as recited in the appended claims. Certain aspects of the present disclosure are described in more detail below. In case of conflict with terms and / or definitions incorporated by reference, the terms and definitions provided herein shall control.

[0010]

[0027] As mentioned above, a video is a frame arranged in a time sequence to store visual information. A video capture device (e.g., a camera) can be used to capture and store those pictures in a time sequence, and a video playback device (e.g., a television, a computer, a smartphone, a tablet computer, a video player, or any end-user terminal with a display function) can be used to display such pictures in a time sequence. In some applications, the video capture device can also transmit the captured video in real time to a video playback device (e.g., a computer with a monitor) for supervision, conferencing, live broadcast, etc.

[0011]

[0028] To reduce the storage space and transmission bandwidth required by such applications, video may be compressed before storage and transmission, and decompressed before display. Compression and decompression may be performed by software executed by a processor (e.g., a processor of a general-purpose computer) or by specialized hardware. A module for compression is generally referred to as an "encoder," and a module for decompression is generally referred to as a "decoder." The encoders and decoders may be collectively referred to as a "codec." The encoders and decoders may be implemented as any of a variety of suitable hardware, software, or combinations thereof. For example, hardware implementations of the encoders and decoders may include circuitry such as one or more microprocessors, digital signal processors (DSPs), application specific integrated circuits (ASICs), field programmable gate arrays (FPGAs), discrete logic, or any combination thereof. Software implementations of the encoders and decoders may include program code, computer executable instructions, firmware, or any suitable computer-implemented algorithms or processes fixed in a computer-readable medium. In some applications, a codec can recover video from a first encoding standard and recompress the recovered video using a second encoding standard, in which case the codec may be referred to as a "transcoder."

[0012]

[0029] A video coding process can identify and keep useful information that can be used to reconstruct a picture and ignore information that is not important for reconstruction. If the ignored, unimportant information cannot be perfectly reconstructed, such a coding process may be called "lossy". Otherwise, it may be called "lossless". Most coding processes are lossy, which is a tradeoff to reduce the required storage space and transmission bandwidth.

[0013]

[0030] Useful information of the picture being coded (called the "current picture") includes changes with respect to a reference picture (e.g., a previously coded and reconstructed picture). Such changes may include changes in pixel position, brightness, or color, among which position changes are the most important. Changes in position of a group of pixels representing an object may reflect the motion of the object between the reference picture and the current picture.

[0014]

[0031] A picture that is coded without reference to another picture (i.e., it is its own reference picture) is called an "I-picture". A picture that is coded using a previous picture as a reference picture is called a "P-picture". A picture that is coded using both previous and future pictures as reference pictures (i.e., the references are "bidirectional") is called a "B-picture".

[0015]

[0032] To achieve the same subjective quality as HEVC / H.265 using half the bandwidth, JVET is developing techniques that go beyond HEVC using the Joint Search Model (JEM) reference software. As coding techniques are incorporated into JEM, JEM has achieved substantially higher coding performance than HEVC.

[0016]

[0033] The VVC standard continues to include more encoding techniques that result in better compression performance. VVC is based on the same hybrid video encoding system that has been used in modern video compression standards such as HEVC, H.264 / AVC, MPEG2, and H.263. In applications such as 360-degree video, the layout of a particular projection format usually has multiple faces. For example, MPEG-I Part 2: Omnidirectional Media Format (OMAF) standardizes a cube-map-based projection format named CMP that has six faces. For projection formats that include multiple faces, regardless of what kind of compact frame packing configuration is used, a discontinuity occurs between two or more adjacent faces in the frame-packed picture. If an in-loop filtering operation is performed across this discontinuity, surface seam artifacts may become visible in the rendered reconstructed image. To mitigate the surface seam artifacts, in-loop filtering operations across the discontinuity in the frame-packed picture must be disabled. To prevent artifacts by disabling in-loop filtering across the virtual boundary, the virtual boundary can be set as one of the encoding tools for 360-degree video.

[0017]

[0034] 1 illustrates the structure of an exemplary video sequence 100 according to some embodiments of the present disclosure. Video sequence 100 may be live video or captured and archived video. Video 100 may be real video, computer-generated video (e.g., computer game video), or a combination thereof (e.g., real video with augmented reality effects). Video sequence 100 may be input from a video capture device (e.g., a camera), a video archive containing previously captured video (e.g., video files stored in a storage device), or a video supply interface for receiving video from a video content provider (e.g., a video broadcast transceiver).

[0018]

[0035] As shown in FIG. 1, a video sequence 100 may include a series of pictures arranged temporally along a timeline, including pictures 102, 104, 106, and 108. Pictures 102-106 are consecutive, with additional pictures between pictures 106 and 108. In FIG. 1, picture 102 is an I-picture whose reference picture is picture 102 itself. Picture 104 is a P-picture whose reference picture is picture 102, as indicated by the arrow. Picture 106 is a B-picture whose reference pictures are pictures 104 and 108, as indicated by the arrow. In some embodiments, the reference picture of a picture (e.g., picture 104) may not be immediately preceding or following that picture. For example, the reference picture of picture 104 may be the picture before picture 102. It should be noted that the reference pictures of pictures 102-106 are merely examples, and this disclosure does not limit the embodiments of reference pictures to the examples shown in FIG.

[0019]

[0036] Typically, video codecs do not encode or decode an entire picture at once due to the computational complexity of such a task. Rather, they may divide a picture into elementary segments and encode or decode a picture segment by segment. Such elementary segments are referred to as basic processing units ("BPUs") in this disclosure. For example, structure 110 in FIG. 1 illustrates an example structure of a picture (e.g., any of pictures 102-108) of video sequence 100. In structure 110, a picture is divided into 4x4 basic processing units, the boundaries of which are shown as dashed lines. In some embodiments, the basic processing units may be referred to as "macroblocks" in some video coding standards (e.g., MPEG family, H.261, H.263, or H.264 / AVC) or as "coding tree units" ("CTUs") in some other video coding standards (e.g., H.265 / HEVC or H.266 / VVC). The basic processing units can have any arbitrary shape and size of variable size or pixels in a picture, such as 128x128, 64x64, 32x32, 16x16, 4x8, 16x32, etc. The size and shape of the basic processing unit can be selected based on a balance between coding efficiency and the level of detail to be maintained in the basic processing unit for a picture.

[0020]

[0037] A basic processing unit may be a logical unit that may include a group of different kinds of video data stored in a computer memory (e.g., in a video frame buffer). For example, a basic processing unit for a color picture may include a luma component (Y) representing colorless luminance information, one or more chroma components (e.g., Cb and Cr) representing color information, and related syntax elements, where the luma and chroma components may have the same size of a basic processing unit. The luma and chroma components may be referred to as "coding tree blocks" ("CTBs") in some video coding standards (e.g., H.265 / HEVC or H.266 / VVC). Any operation performed on a basic processing unit may be performed repeatedly on each of its luma and chroma components.

[0021]

[0038] Video coding has multiple operation stages, examples of which are shown in detail in Figures 2A-2B and 3A-3B. At each stage, the size of the basic processing unit may still be too large for processing, and therefore may be further divided into segments, referred to as "basic processing subunits" in this disclosure. In some embodiments, the basic processing subunits may be referred to as "blocks" in some video coding standards (e.g., MPEG family, H.261, H.263, or H.264 / AVC), or as "coding units" ("CUs") in some other video coding standards (e.g., H.265 / HEVC or H.266 / VVC). The basic processing subunits may have the same or smaller size than the basic processing units. Similar to the basic processing units, the basic processing subunits are also logical units that may include groups of different types of video data (e.g., Y, Cb, Cr, and related syntax elements) stored in computer memory (e.g., in a video frame buffer). Any operation performed on a basic processing sub-unit may be repeatedly performed on each of its luma and chroma components. Note that such division may be performed to further levels as required for processing. Also note that different stages may use different schemes to divide the basic processing units.

[0022]

[0039] For example, in the mode decision stage (an example of which is shown in detail in FIG. 2B), the encoder can decide what prediction mode (e.g., intra-picture prediction or inter-picture prediction) to use for the basic processing unit, but the basic processing unit may be too large to make such a decision. The encoder can divide the basic processing unit into multiple basic processing sub-units (e.g., CUs, as in the case of H.265 / HEVC or H.266 / VVC) and decide the type of prediction for each individual basic processing sub-unit.

[0023]

[0040] As another example, in the prediction stage (an example of which is shown in detail in FIGS. 2A-2B), the encoder can perform prediction operations at the level of the basic processing subunit (e.g., CU). However, in some cases, the basic processing subunit may still be too large to process. The encoder can further divide the basic processing subunit into smaller segments (e.g., referred to as "prediction blocks" or "PBs" in H.265 / HEVC or H.266 / VVC), at which the prediction operations can be performed.

[0024]

[0041] As another example, in the transform stage (an example of which is shown in detail in FIG. 2A-2B), the encoder can perform transform operations for residual basic processing subunits (e.g., CUs). However, in some cases, the basic processing subunits may still be too large to process. The encoder can further divide the basic processing subunits into smaller segments (e.g., referred to as "transform blocks" or "TBs" in H.265 / HEVC or H.266 / VVC), at which level the transform operations can be performed. Note that the division scheme of the same basic processing subunit may be different in the prediction stage and the transform stage. For example, in H.265 / HEVC or H.266 / VVC, the prediction blocks and transform blocks of the same CU may have different sizes and numbers.

[0025]

[0042] 1, the basic processing units 112 are further divided into 3×3 basic processing sub-units, the boundaries of which are shown as dotted lines. Different basic processing units of the same picture may be divided into basic processing sub-units in different ways.

[0026]

[0043] In some implementations, to provide parallel processing and error resilience capabilities to video encoding and decoding, a picture can be divided into regions for processing, such that the encoding or decoding process does not have to rely on information about a picture region from any other region of the picture. In other words, each region of a picture can be processed independently. By doing so, the codec can process different regions of a picture in parallel, thereby increasing the coding efficiency. Also, when data of a region is corrupted during processing or lost during network transmission, the codec can also correctly encode or decode other regions of the same picture without relying on the corrupted or lost data, thereby providing error resilience capabilities. In some video coding standards, a picture can be divided into different types of regions. For example, H.265 / HEVC and H.266 / VVC provide two types of regions: "slices" and "tiles". It should also be noted that different pictures of the video sequence 100 may have different partitioning schemes for dividing the picture into regions.

[0027]

[0044] For example, in Figure 1, structure 110 is divided into three regions 114, 116, and 118, whose boundaries are shown as solid lines inside structure 110. Region 114 includes four basic processing units. Regions 116 and 118 each include six basic processing units. It should be noted that the basic processing units, basic processing sub-units, and regions of structure 110 in Figure 1 are merely examples, and the present disclosure does not limit the embodiments thereof.

[0028]

[0045] FIG. 2A illustrates a schematic diagram of an exemplary encoding process 200A according to an embodiment of the present disclosure. For example, the encoding process 200A may be performed by an encoder. As shown in FIG. 2A, the encoder may encode a video sequence 202 into a video bitstream 228 according to the process 200A. Similar to the video sequence 100 in FIG. 1, the video sequence 202 may include a set of pictures (referred to as "original pictures") arranged in a temporal order. Similar to the structure 110 in FIG. 1, each original picture of the video sequence 202 may be divided into basic processing units, basic processing sub-units, or regions for processing by the encoder. In some embodiments, the encoder may perform the process 200A at the level of the basic processing units for each original picture of the video sequence 202. For example, the encoder may perform the process 200A in an iterative manner, in which case the encoder may encode a basic processing unit in one iteration of the process 200A. In some embodiments, the encoder may perform process 200A in parallel for regions of each original picture of video sequence 202 (eg, regions 114-118).

[0029]

[0046] 2A , the encoder may provide a fundamental processing unit (referred to as an “original BPU”) of an original picture of a video sequence 202 to a prediction stage 204 to generate prediction data 206 and a prediction BPU 208. The encoder may subtract the prediction BPU 208 from the original BPU to generate a residual BPU 210. The encoder may provide the residual BPU 210 to a transform stage 212 and a quantization stage 214 to generate quantized transform coefficients 216. The encoder may provide the prediction data 206 and the quantized transform coefficients 216 to a binary coding stage 226 to generate a video bitstream 228. The components 202, 204, 206, 208, 210, 212, 214, 216, 226, and 228 may be referred to as a “forward path.” During process 200A, after quantization stage 214, the encoder may provide quantized transform coefficients 216 to an inverse quantization stage 218 and an inverse transform stage 220 to generate a reconstructed residual BPU 222. The encoder may add the reconstructed residual BPU 222 to a prediction BPU 208 to generate a prediction reference 224, which is used in prediction stage 204 for the next iteration of process 200A. The components 218, 220, 222, and 224 of process 200A may be referred to as a "reconstruction path." The reconstruction path may be used to ensure that both the encoder and decoder use the same reference data for prediction.

[0030]

[0047] The encoder may iteratively perform process 200A to encode each original BPU of the original picture (in the forward path) and generate (in the reconstruction path) a prediction reference 224 for encoding the next original BPU of the original picture. After encoding all the original BPUs of the original picture, the encoder may proceed to encode the next picture in the video sequence 202.

[0031]

[0048] Referring to process 200A, an encoder may receive a video sequence 202 generated by a video capture device (e.g., a camera). As used herein, the term "receive" may refer to receiving, inputting, acquiring, getting, obtaining, reading, accessing, or any act by any method for inputting data.

[0032]

[0049] In the prediction step 204, in the current iteration, the encoder may receive the original BPU and a prediction reference 224, perform a prediction operation, and generate prediction data 206 and a predicted BPU 208. The prediction reference 224 may be generated from a reconstruction path of a previous iteration of the process 200A. The purpose of the prediction step 204 is to reduce information redundancy by extracting the prediction data 206, which can be used to reconstruct the original BPU as a predicted BPU 208 from the prediction data 206 and the prediction reference 224.

[0033]

[0050] Ideally, the predicted BPU 208 may be identical to the original BPU. However, due to non-ideal prediction and reconstruction operations, the predicted BPU 208 generally differs slightly from the original BPU. To record such differences, after generating the predicted BPU 208, the encoder may subtract it from the original BPU to generate the residual BPU 210. For example, the encoder may subtract the values ​​(e.g., grayscale or RGB values) of pixels of the predicted BPU 208 from the values ​​of corresponding pixels of the original BPU. Each pixel of the residual BPU 210 may have a residual value as a result of such a subtraction between the corresponding pixels of the original BPU and the predicted BPU 208. Compared to the original BPU, the predicted data 206 and the residual BPU 210 may have fewer bits, which may be used to reconstruct the original BPU without significant quality degradation. Therefore, the original BPU is compressed.

[0034]

[0051] To further compress the residual BPU 210, in the transform stage 212, the encoder can reduce spatial redundancy of the residual BPU 210 by decomposing it into a set of two-dimensional "basis patterns", where each basis pattern is associated with a "transform coefficient". The basis patterns can have the same size (e.g., the size of the residual BPU 210). Each basis pattern can represent a change frequency (e.g., frequency of luminance change) component of the residual BPU 210. None of the basis patterns can be reproduced from any combination (e.g., linear combination) of any other basis patterns. In other words, the decomposition can decompose the changes of the residual BPU 210 into the frequency domain. Such a decomposition is similar to a discrete Fourier transform of a function, where the basis patterns are similar to the basis functions (e.g., trigonometric functions) of the discrete Fourier transform, and the transform coefficients are similar to the coefficients associated with the basis functions.

[0035]

[0052] Different transform algorithms may use different basis patterns. For example, various transform algorithms may be used in transform stage 212, such as discrete cosine transform, discrete sine transform, or the like. The transform in transform stage 212 is invertible. That is, the encoder may recover residual BPU 210 by inverting the transform (referred to as an "inverse transform"). For example, to recover pixels of residual BPU 210, the inverse transform may multiply the values ​​of corresponding pixels of the basis pattern by their associated coefficients and add up the products to generate a weighted sum. For video coding standards, the encoder and decoder may both use the same transform algorithm (and therefore the same basis pattern). Therefore, the encoder may record only the transform coefficients, and the decoder may reconstruct residual BPU 210 from the transform coefficients without receiving the basis pattern from the encoder. Compared to residual BPU 210, the transform coefficients may have fewer bits, but they may be used to reconstruct residual BPU 210 without significant quality degradation. Therefore, the residual BPU 210 is further compressed.

[0036]

[0053] The encoder can further compress the transform coefficients in the quantization stage 214. In the transform process, different basis patterns can represent different change frequencies (e.g., luminance change frequencies). Since the human eye is generally better at recognizing low-frequency changes, the encoder can ignore the information of high-frequency changes without causing significant quality degradation in decoding. For example, in the quantization stage 214, the encoder can generate the quantized transform coefficients 216 by dividing each transform coefficient by an integer value (referred to as a "quantization parameter") and rounding the quotient to its nearest integer. After such an operation, some transform coefficients of the high-frequency basis patterns can be converted to zero, and the transform coefficients of the low-frequency basis patterns can be converted to smaller integers. The encoder can ignore the zero-valued quantized transform coefficients 216, which further compresses the transform coefficients. The quantization process can also be inverted, in which case the quantized transform coefficients 216 can be reconstructed into transform coefficients in the inverse operation of quantization (referred to as "dequantization").

[0037]

[0054] The quantization stage 214 may be lossy because the encoder ignores such division remainders in rounding operations. Typically, the quantization stage 214 may contribute the greatest information loss in the process 200A. The greater the information loss, the fewer bits the quantized transform coefficients 216 may require. To obtain different information loss levels, the encoder may use different values ​​of the quantization parameter or any other parameter of the quantization process.

[0038]

[0055] In the binary encoding stage 226, the encoder may encode the prediction data 206 and the quantized transform coefficients 216 using a binary encoding technique, such as, for example, entropy coding, variable length coding, arithmetic coding, Huffman coding, context-adaptive binary arithmetic coding, or any other lossless or lossy compression algorithm. In some embodiments, besides the prediction data 206 and the quantized transform coefficients 216, the encoder may encode other information in the binary encoding stage 226, such as, for example, a prediction mode used in the prediction stage 204, parameters of the prediction operation, a type of transformation in the transformation stage 212, parameters of the quantization process (e.g., quantization parameters), encoder control parameters (e.g., bitrate control parameters), or the like. The encoder may generate a video bitstream 228 using the output data of the binary encoding stage 226. In some embodiments, the video bitstream 228 may be further packetized for network transmission.

[0039]

[0056] Referring to the reconstruction path of process 200A, in an inverse quantization stage 218, the encoder may perform inverse quantization on the quantized transform coefficients 216 to generate reconstructed transform coefficients. In an inverse transform stage 220, the encoder may generate a reconstructed residual BPU 222 based on the reconstructed transform coefficients. The encoder may add the reconstructed residual BPU 222 to the prediction BPU 208 to generate a prediction reference 224 to be used in the next iteration of process 200A.

[0040]

[0057] It should be noted that other variations of the process 200A may be used to encode the video sequence 202. In some embodiments, the stages of the process 200A may be performed in a different order by the encoder. In some embodiments, one or more stages of the process 200A may be combined into a single stage. In some embodiments, a single stage of the process 200A may be split into multiple stages. For example, the transform stage 212 and the quantization stage 214 may be combined into a single stage. In some embodiments, the process 200A may include additional stages. In some embodiments, the process 200A may omit one or more stages in FIG. 2A.

[0041]

[0058] 2B shows a schematic diagram of another exemplary encoding process 200B according to an embodiment of the present disclosure. The process 200B may be modified from the process 200A. For example, the process 200B may be used by an encoder compliant with a hybrid video coding standard (e.g., H.26x series). Compared to the process 200A, the forward path of the process 200B additionally includes a mode decision stage 230 and divides the prediction stage 204 into a spatial prediction stage 2042 and a temporal prediction stage 2044. The reconstruction path of the process 200B additionally includes a loop filter stage 232 and a buffer 234.

[0042]

[0059] Generally, prediction techniques can be categorized into two types: spatial prediction and temporal prediction. Spatial prediction (e.g., intra-picture prediction or "intra prediction") can use pixels from one or more already coded neighboring BPUs in the same picture to predict the current BPU. That is, the prediction reference 224 in spatial prediction can include neighboring BPUs. Spatial prediction can reduce the inherent spatial redundancy of a picture. Temporal prediction (e.g., inter-picture prediction or "inter prediction") can use regions from one or more already coded pictures to predict the current BPU. That is, the prediction reference 224 in temporal prediction can include coded pictures. Temporal prediction can reduce the inherent temporal redundancy of a picture.

[0043]

[0060] Referring to process 200B, in the forward path, the encoder performs prediction operations in a spatial prediction stage 2042 and a temporal prediction stage 2044. For example, in the spatial prediction stage 2042, the encoder may perform intra prediction. For an original BPU of a picture being encoded, the prediction reference 224 may include one or more neighboring BPUs in the same picture that are coded (in the forward path) and reconstructed (in the reconstruction path). The encoder may generate the predicted BPU 208 by extrapolating the neighboring BPUs. The extrapolation technique may include, for example, linear extrapolation or interpolation, polynomial extrapolation or interpolation, or the like. In some embodiments, the encoder may perform the extrapolation at a pixel level, such as by extrapolating, for each pixel of the predicted BPU 208, the value of the corresponding pixel. The neighboring BPUs used for extrapolation can be located relative to the original BPU from various directions, such as vertically (e.g., above the original BPU), horizontally (e.g., to the left of the original BPU), diagonally (e.g., bottom-left, bottom-right, top-left, or top-right of the original BPU), or any direction defined in the video coding standard used. For intra prediction, the prediction data 206 can include, for example, the locations (e.g., coordinates) of the neighboring BPUs used, the size of the neighboring BPUs used, parameters of the extrapolation, the orientation of the neighboring BPUs used relative to the original BPU, or the like.

[0044]

[0061] As another example, in the temporal prediction stage 2044, the encoder may perform inter prediction. For an original BPU of the current picture, the prediction reference 224 may include one or more pictures (referred to as "reference pictures") that have been coded (in the forward path) and reconstructed (in the reconstruction path). In some embodiments, the reference pictures may be coded and reconstructed for each BPU. For example, the encoder may add the reconstructed residual BPU 222 to the prediction BPU 208 to generate a reconstructed BPU. When all reconstructed BPUs of the same picture have been generated, the encoder may generate the reconstructed picture as the reference picture. The encoder may perform a "motion estimation" operation to search for a matching region within a range (referred to as a "search window") of the reference picture. The location of the search window in the reference picture may be determined based on the location of the original BPU of the current picture. For example, the search window may be centered at a location in the reference picture that has the same coordinates as the original BPU in the current picture, and may extend outward for a predetermined distance. When the encoder identifies a region similar to the original BPU within the search window (e.g., by using a pixel-recursive algorithm, a block matching algorithm, or the like), the encoder may determine such a region as a matching region. The matching region may have different dimensions (e.g., smaller than, equal to, larger than, or of a different shape) than the original BPU. Because the reference picture and the current picture are temporally separated in a timeline (e.g., as shown in FIG. 1), the matching region may be considered to "move" to the location of the original BPU as time progresses. The encoder may record the direction and distance of such movement as a "motion vector." When multiple reference pictures are used (e.g., as picture 106 in FIG. 1), the encoder may search for the matching region for each reference picture and determine its associated motion vector. In some embodiments, the encoder may assign weights to the pixel values ​​of the matching region in each matching reference picture.

[0045]

[0062] Motion estimation may be used to identify various types of motion, such as, for example, translation, rotation, zooming, or the like. For inter prediction, the prediction data 206 may include, for example, the location (e.g., coordinates) of the matching region, a motion vector associated with the matching region, a number of reference pictures, weights associated with the reference pictures, or the like.

[0046]

[0063] To generate the predicted BPU 208, the encoder may perform a "motion compensation" operation. Motion compensation may be used to reconstruct the predicted BPU 208 based on the prediction data 206 (e.g., motion vectors) and the prediction reference 224. For example, the encoder may shift the matching region of the reference picture according to the motion vector, in which case the encoder may predict the original BPU of the current picture. When multiple reference pictures are used (e.g., like picture 106 in FIG. 1), the encoder may shift the matching region of the reference picture according to the respective motion vectors and average the pixel values ​​of the matching region. In some embodiments, if the encoder weights the pixel values ​​of the matching region of each matching reference picture, the encoder may add a weighted sum of pixel values ​​to the shifted matching region.

[0047]

[0064] In some embodiments, inter prediction can be unidirectional or bidirectional. Unidirectional inter prediction can use one or more reference pictures in the same temporal direction relative to the current picture. For example, picture 104 in FIG. 1 is a unidirectional inter predicted picture whose reference picture (i.e., picture 102) precedes picture 104. Bidirectional inter prediction can use one or more reference pictures in both temporal directions relative to the current picture. For example, picture 106 in FIG. 1 is a bidirectional inter predicted picture whose reference pictures (i.e., pictures 104 and 108) are in both temporal directions relative to picture 104.

[0048]

[0065] Still referring to the forward path of the process 200B, after the spatial prediction 2042 and temporal prediction stages 2044, in a mode decision stage 230, the encoder may select a prediction mode (e.g., one of intra prediction or inter prediction) for the current iteration of the process 200B. For example, the encoder may perform a rate-distortion optimization technique. In this technique, the encoder may select a prediction mode to minimize the value of a cost function that depends on the bitrate of the candidate prediction mode and the distortion of the reconstructed reference picture under the candidate prediction mode. Depending on the selected prediction mode, the encoder may generate a corresponding prediction BPU 208 and prediction data 206.

[0049]

[0066] Within the reconstruction path of the process 200B, if an intra prediction mode is selected within the forward path, after generating the prediction reference 224 (e.g., the current BPU encoded and reconstructed in the current picture), the encoder can directly provide the prediction reference 224 to the spatial prediction stage 2042 for later use (e.g., for extrapolation of the next BPU of the current picture). If an inter prediction mode is selected within the forward path, after generating the prediction reference 224 (e.g., the current picture in which all BPUs are encoded and reconstructed), the encoder can provide the prediction reference 224 to the loop filter stage 232, where the encoder can apply a loop filter to the prediction reference 224 to reduce or eliminate distortions (e.g., blocking artifacts) introduced by the inter prediction. The encoder can apply various loop filter techniques in the loop filter stage 232, such as, for example, deblocking, sample adaptive offset, adaptive loop filter, or the like. The loop filtered reference picture may be stored in a buffer 234 (or a "decoded picture buffer") for later use (e.g., to be used as an inter-prediction reference picture for future pictures of the video sequence 202). The encoder may store one or more reference pictures in the buffer 234 for use in the temporal prediction stage 2044. In some embodiments, the encoder may encode parameters of the loop filter (e.g., loop filter strength) along with the quantized transform coefficients 216, the prediction data 206, and other information in the binary encoding stage 226.

[0050]

[0067] FIG. 3A shows a schematic diagram of an example decoding process 300A according to an embodiment of the present disclosure. Process 300A may be a decompression process corresponding to compression process 200A in FIG. 2A. In some embodiments, process 300A may be similar to the reconstruction path of process 200A. A decoder may decode video bitstream 228 into video stream 304 according to process 300A. Video stream 304 may be similar to video sequence 202. However, due to information loss in the compression and decompression process (e.g., quantization stage 214 in FIGS. 2A-2B), video stream 304 is generally not identical to video sequence 202. Similar to processes 200A and 200B in FIGS. 2A-2B, a decoder may perform process 300A at the level of a basic processing unit (BPU) for each picture encoded in video bitstream 228. For example, the decoder may perform process 300A in an iterative manner, where the decoder may decode a basic processing unit in one iteration of process 300A. In some embodiments, the decoder may perform process 300A in parallel for regions (e.g., regions 114-118) of each picture encoded in video bitstream 228.

[0051]

[0068] In FIG. 3A, a decoder may provide a portion of a video bitstream 228 associated with a basic processing unit (referred to as an “encoded BPU”) of a coded picture to a binary decoding stage 302. In the binary decoding stage 302, the decoder may decode the portion into prediction data 206 and quantized transform coefficients 216. The decoder may provide the quantized transform coefficients 216 to an inverse quantization stage 218 and an inverse transform stage 220 to generate a reconstructed residual BPU 222. The decoder may provide the prediction data 206 to a prediction stage 204 to generate a prediction BPU 208. The decoder may add the reconstructed residual BPU 222 to the prediction BPU 208 to generate a prediction reference 224. In some embodiments, the prediction reference 224 may be stored in a buffer (e.g., a decoded picture buffer in a computer memory). The decoder may provide the prediction reference 224 to a prediction stage 204 for performing a prediction operation in a next iteration of the process 300A.

[0052]

[0069] The decoder may iteratively perform the process 300A to decode each coded BPU of the coded picture and generate a prediction reference 224 for coding the next coded BPU of the coded picture. After decoding all coded BPUs of the coded picture, the decoder may output the picture to the video stream 304 for display and proceed to decode the next coded picture in the video bitstream 228.

[0053]

[0070] In the binary decoding stage 302, the decoder may perform an inverse operation of the binary encoding technique used by the encoder (e.g., entropy encoding, variable length encoding, arithmetic encoding, Huffman encoding, context-adaptive binary arithmetic encoding, or any other lossless compression algorithm). In some embodiments, in addition to the prediction data 206 and the quantized transform coefficients 216, the decoder may decode other information in the binary decoding stage 302, such as, for example, a prediction mode, parameters of the prediction operation, a type of transform, parameters of the quantization process (e.g., quantization parameters), encoder control parameters (e.g., bitrate control parameters), or the like. In some embodiments, if the video bitstream 228 is transmitted in the form of packets over a network, the decoder may depacketize the video bitstream 228 before providing it to the binary decoding stage 302.

[0054]

[0071] 3B shows a schematic diagram of another exemplary decoding process 300B according to an embodiment of the present disclosure. The process 300B may be modified from the process 300A. For example, the process 300B may be used by a decoder compliant with a hybrid video coding standard (e.g., H.26x series). Compared to the process 300A, the process 300B additionally divides the prediction stage 204 into a spatial prediction stage 2042 and a temporal prediction stage 2044, and additionally includes a loop filter stage 232 and a buffer 234.

[0055]

[0072] In the process 300B, for a coding basic processing unit (referred to as a "current BPU") of a coding picture being decoded (referred to as a "current picture"), the prediction data 206 decoded by the decoder from the binary decoding stage 302 may include various kinds of data depending on what prediction mode was used by the encoder to code the current BPU. For example, if intra prediction was used by the encoder to code the current BPU, the prediction data 206 may include a prediction mode indicator (e.g., a flag value) indicating intra prediction, parameters of the intra prediction operation, or the like. The parameters of the intra prediction operation may include, for example, the location (e.g., coordinates) of one or more neighboring BPUs used as references, the size of the neighboring BPUs, parameters of extrapolation, the orientation of the neighboring BPUs relative to the original BPU, or the like. As another example, if inter prediction was used by the encoder to code the current BPU, the prediction data 206 may include a prediction mode indicator (e.g., a flag value) indicating inter prediction, parameters of the inter prediction operation, or the like. Parameters for the inter prediction operation may include, for example, the number of reference pictures associated with the current BPU, weights respectively associated with the reference pictures, locations (e.g., coordinates) of one or more matching regions within each reference picture, one or more motion vectors respectively associated with the matching regions, or the like.

[0056]

[0073] Based on the prediction mode indicator, the decoder may determine whether to perform spatial prediction (e.g., intra prediction) in the spatial prediction stage 2042 or perform temporal prediction (e.g., inter prediction) in the temporal prediction stage 2044. Details of performing such spatial or temporal prediction are described in FIG. 2B and will not be repeated below. After performing such spatial or temporal prediction, the decoder may generate a prediction BPU 208. The decoder may add the prediction BPU 208 and the reconstructed residual BPU 222 to generate a prediction reference 224, as described in FIG. 3A.

[0057]

[0074] In the process 300B, the decoder can provide the prediction reference 224 to the spatial prediction stage 2042 or the temporal prediction stage 2044 to perform the prediction operation in the next iteration of the process 300B. For example, if the current BPU is decoded using intra prediction in the spatial prediction stage 2042, after generating the prediction reference 224 (e.g., the decoded current BPU), the decoder can directly provide the prediction reference 224 to the spatial prediction stage 2042 for later use (e.g., for extrapolation of the next BPU of the current picture). If the current BPU is decoded using inter prediction in the temporal prediction stage 2044, after generating the prediction reference 224 (e.g., the reference picture decoded by all the BPUs), the encoder can provide the prediction reference 224 to the loop filter stage 232 to reduce or eliminate distortion (e.g., blocking artifacts). The decoder can apply a loop filter to the prediction reference 224 in the manner as described in FIG. 2B. The loop filtered reference picture may be stored in a buffer 234 (e.g., a decoded picture buffer in a computer memory) for later use (e.g., to be used as an inter-prediction reference picture for future encoded pictures of the video bitstream 228). The decoder may store one or more reference pictures in the buffer 234 for use in the temporal prediction stage 2044. In some embodiments, when the prediction mode indicator of the prediction data 206 indicates that inter prediction was used to encode the current BPU, the prediction data may further include parameters of a loop filter (e.g., loop filter strength).

[0058]

[0075] FIG. 4 is a block diagram of an exemplary device 400 for encoding or decoding video according to an embodiment of the present disclosure. As shown in FIG. 4, the device 400 may include a processor 402. When the processor 402 executes instructions described herein, the device 400 may become a specialized machine for video encoding or decoding. The processor 402 may be any type of circuitry capable of manipulating or processing information. For example, the processor 402 may include any number and combination of a central processing unit (or "CPU"), a graphics processing unit (or "GPU"), a neural processing unit ("NPU"), a microcontroller unit ("MCU"), an optical processor, a programmable logic controller, a microcontroller, a microprocessor, a digital signal processor, an intellectual property (IP) core, a programmable logic array (PLA), a programmable array logic (PAL), a generic array logic (GAL), a complex programmable logic device (CPLD), a field programmable gate array (FPGA), a system on a chip (SoC), an application specific integrated circuit (ASIC), or the like. In some embodiments, processor 402 may be a set of processors grouped together as a single logical entity. For example, as shown in FIG. 4, processor 402 may include multiple processors, including processor 402a, processor 402b, and processor 402n.

[0059]

[0076] The device 400 may also include a memory 404 configured to store data (e.g., a set of instructions, computer code, intermediate data, or the like). For example, as shown in FIG. 4, the stored data may include program instructions (e.g., program instructions for performing steps in processes 200A, 200B, 300A, or 300B) and data for processing (e.g., video sequence 202, video bitstream 228, or video stream 304). The processor 402 may access (e.g., via bus 410) the program instructions and data for processing, execute the program instructions, and perform operations or manipulations on the data for processing. The memory 404 may include a high-speed random access storage device, or a non-volatile storage device. In some embodiments, the memory 404 may include any number and combination of random access memory (RAM), read-only memory (ROM), optical disks, magnetic disks, hard drives, solid-state drives, flash drives, security digital (SD) cards, memory sticks, compact flash (CF) cards, or the like. The memory 404 may also be a group of memories (not shown in FIG. 4) grouped as a single logical entity.

[0060]

[0077] Bus 410 may be a communication device that transfers data between components internal to device 400, such as an internal bus (e.g., a CPU-memory bus), an external bus (e.g., a Universal Serial Bus port, a Peripheral Component Interconnect Express port), or the like.

[0061]

[0078] For ease of explanation and without creating ambiguity, the processor 402 and other data processing circuitry are collectively referred to in this disclosure as "data processing circuitry." The data processing circuitry may be implemented entirely as hardware, or as a combination of software, hardware, or firmware. In addition, the data processing circuitry may be a single independent module, or may be fully or partially combined with any other components of the device 400.

[0062]

[0079] The device 400 may further include a network interface 406 for providing wired or wireless communication with a network (e.g., the Internet, an intranet, a local area network, a mobile communication network, or the like). In some embodiments, the network interface 406 may include any number and combination of a network interface controller (NIC), a radio frequency (RF) module, a transponder, a transceiver, a modem, a router, a gateway, a wired network adapter, a wireless network adapter, a Bluetooth® adapter, an infrared adapter, a near field communication ("NFC") adapter, a cellular network chip, or the like.

[0063]

[0080] In some embodiments, optionally, apparatus 400 may further include a peripheral interface 408 for providing a connection to one or more peripheral devices. As shown in FIG. 4A, the peripheral devices may include, but are not limited to, a cursor control device (e.g., a mouse, a touchpad, or a touch screen), a keyboard, a display (e.g., a cathode ray tube display, a liquid crystal display, or a light emitting diode display), a video input device (e.g., a camera, or an input interface coupled to a video archive), or the like.

[0064]

[0081] It should be noted that a video codec (e.g., a codec performing processes 200A, 200B, 300A, or 300B) may be implemented as any combination of any software or hardware modules within device 400. For example, some or all of the stages of processes 200A, 200B, 300A, or 300B may be implemented as one or more software modules of device 400, such as program instructions that may be loaded into memory 404. As another example, some or all of the stages of processes 200A, 200B, 300A, or 300B may be implemented as one or more hardware modules of device 400, such as specialized data processing circuitry (e.g., FPGA, ASIC, NPU, or the like).

[0065]

[0082] According to the disclosed embodiment, several coding tools may be used to code 360-degree video or progressive decoding refresh (GDR). Virtual boundary is one of these coding tools. In applications such as 360-degree video, the layout of a particular projection format usually has multiple faces. For example, MPEG-I Part 2: Omnidirectional Media Format (OMAF) standardizes a cube-map-based projection format named CMP with six faces. For a projection format that includes multiple faces, no matter what kind of compact frame packing configuration is used, a discontinuity occurs between two or more adjacent faces in a frame-packed picture. If an in-loop filtering operation is performed across this discontinuity, surface seam artifacts may become visible in the rendered reconstructed image. To mitigate the surface seam artifacts, in-loop filtering operations across the discontinuity in the frame-packed picture must be disabled. Therefore, according to the disclosed embodiment, a concept called virtual boundary, in which loop filtering operations are disabled to cross, may be used. An encoder may set the discontinuity boundary as a virtual boundary, and may not apply loop filters on the discontinuity boundary. The location of the virtual border is signaled in the bitstream, and the encoder can change the location of the virtual border according to the current projection format.

[0066]

[0083] Besides 360-degree video, the virtual boundary can also be used for gradual decoding refresh (GDR), which is mainly used in ultra-low-delay applications. In ultra-low-delay applications, inserting an intra-coded picture as a random access point picture may cause unacceptable transmission latency due to the large size of the intra-coded picture. To reduce the latency, GDR is adopted, in which a picture is gradually refreshed by inserting an intra-coded region in a B / P picture. To prevent error propagation, pixels in a refreshed region in a picture cannot refer to those in an unrefreshed region of the current picture or reference picture. Therefore, loop filtering cannot be applied across the boundary between the refreshed region and the unrefreshed region. With the above virtual boundary scheme, the encoder can set the boundary between the refreshed region and the unrefreshed region as a virtual boundary, and in that case, loop filtering operations cannot be applied across the boundary.

[0067]

[0084] According to some embodiments, virtual boundaries can be signaled in the sequence parameter set (SPS) or in the picture header (PH). The picture header carries information about a particular picture and contains information common to all slices belonging to the same picture. The PH may contain information related to virtual boundaries. One PH is set per picture. The sequence parameter set contains syntax elements related to the coded layer video sequence (CLVS). The SPS contains sequence level information shared by all pictures in the entire coded layer video sequence (CLVS) and can give a complete picture of what the bitstream contains and how the information in the bitstream can be used. In the SPS, the virtual boundary present flag "sps_virtual_boundaries_present_flag" is signaled first. If the flag is true, the number of virtual boundaries and the position of each virtual boundary are signaled for the picture that references the SPS. If "sps_virtual_boundaries_present_flag" is false, another virtual boundary present flag "ph_virtual_boundaries_present_flag" may be signaled in the PH. Similarly, if "ph_virtual_boundaries_present_flag" is true, the number of virtual boundaries and the location of each virtual boundary may be signaled for the picture associated with the PH.

[0068]

[0085] The SPS syntax for virtual boundaries is shown in Table 1 of Figure 5. The semantics of the SPS syntax in Figure 5 are as follows:

[0069]

[0086] "sps_virtual_boundaries_present_flag" equal to 1 specifies that virtual boundary information is signaled in the SPS. "sps_virtual_boundaries_present_flag" equal to 0 specifies that virtual boundary information is not signaled in the SPS. If one or more virtual boundaries are signaled in the SPS, in-loop filtering operations across the virtual boundaries in pictures that reference the SPS are disabled. In-loop filtering operations include deblocking filters, sample adaptive offset filters, and adaptive loop filter operations.

[0070]

[0087] "sps_num_ver_virtual_boundaries" specifies the number of "sps_virtual_boundaries_pos_x[i]" syntax elements that are in the SPS. If "sps_num_ver_virtual_boundaries" is absent, its value is inferred to be equal to 0.

[0071]

[0088] "sps_virtual_boundaries_pos_x[i]" specifies the position of the i-th vertical virtual boundary in luma samples divided by 8. The value of "sps_virtual_boundaries_pos_x[i]" is in the range of 1 to Ceil(pic_width_in_luma_samples÷8)-1.

[0072]

[0089] "sps_num_hor_virtual_boundaries" specifies the number of "sps_virtual_boundaries_pos_y[i]" syntax elements in the SPS. If "sps_num_hor_virtual_boundaries" is absent, its value is inferred to be equal to 0.

[0073]

[0090] "sps_virtual_boundaries_pos_y[i]" specifies the position of the i-th horizontal virtual boundary in luma samples divided by 8. The value of "sps_virtual_boundaries_pos_y[i]" is in the range of 1 to Ceil(pic_height_in_luma_samples÷8)-1.

[0074]

[0091] The PH syntax for virtual boundaries is shown in Table 2 of Figure 6. The semantics of the PH syntax in Figure 6 are as follows:

[0075]

[0092] "ph_num_ver_virtual_boundaries" specifies the number of "ph_virtual_boundaries_pos_x[i]" syntax elements in PH. If "ph_num_ver_virtual_boundaries" is absent, its value is inferred to be equal to 0.

[0076]

[0093] The parameter VirtualBoundariesNumVer is derived as follows: VirtualBoundariesNumVer=sps_virtual_boundaries_present_flag? sps_num_ver_virtual_boundaries:ph_num_ver_virtual_boundaries

[0077]

[0094] "ph_virtual_boundaries_pos_x[i]" specifies the position of the i-th vertical virtual boundary in luma samples divided by 8. The value of "ph_virtual_boundaries_pos_x[i]" is in the range of 1 to Ceil(pic_width_in_luma_samples÷8)-1.

[0078]

[0095] The position of the vertical virtual boundary “VirtualBoundariesPosX[i]” in luma samples is derived as follows: VirtualBoundariesPosX[i]=(sps_virtual_boundaries_present_flag? sps_virtual_boundaries_pos_x[i]:ph_virtual_boundaries_pos_x[i])*8

[0079]

[0096] The distance between any two vertical virtual borders can be greater than or equal to CtbSizeY luma samples.

[0080]

[0097] "ph_num_hor_virtual_boundaries" specifies the number of "ph_virtual_boundaries_pos_y[i]" syntax elements in PH. If "ph_num_hor_virtual_boundaries" is absent, its value is inferred to be equal to 0.

[0081]

[0098] The parameter VirtualBoundariesNumHor is derived as follows: VirtualBoundariesNumHor=sps_virtual_boundaries_present_flag? sps_num_hor_virtual_boundaries:ph_num_hor_virtual_boundaries

[0082]

[0099] "ph_virtual_boundaries_pos_y[i]" specifies the position of the i-th horizontal virtual boundary in luma samples divided by 8. The value of "ph_virtual_boundaries_pos_y[i]" is in the range of 1 to Ceil(pic_height_in_luma_samples÷8)-1.

[0083]

[0100] The position of the horizontal virtual boundary “VirtualBoundariesPosY[i]” in luma samples is derived as follows: VirtualBoundariesPosY[i]=(sps_virtual_boundaries_present_flag? sps_virtual_boundaries_pos_y[i]:ph_virtual_boundaries_pos_y[i])*8

[0084]

[0101] The distance between any two horizontal virtual boundaries can be greater than or equal to CtbSizeY luma samples.

[0085]

[0102] Wrap-around motion compensation is another 360-degree video coding tool. In traditional motion compensation, when a motion vector references a sample beyond the picture boundary of a reference picture, repetition padding is applied to derive the value of the out-of-boundary sample by copying from the nearest neighbor on the corresponding picture boundary. In 360-degree video, this method of repetition padding is not suitable and may cause visual artifacts called "seam artifacts" in the reconstructed viewport video. Since 360-degree video is captured on a sphere and does not have inherent "boundaries", reference samples outside the boundary of a reference picture in the projection region can always be obtained from neighboring samples in the spherical region. It may be difficult to derive the corresponding neighboring samples in the spherical region in common projection formats, since it involves 2D-to-3D and 3D-to-2D coordinate transformations and sample interpolation for fractional sample positions. The problem is very simple for the left and right boundaries of Equirectangular Projection (ERP) or Padded ERP (PERP) formats, since the spherical neighborhood outside the left picture boundary can be obtained from samples inside the right picture boundary, and vice versa. Given the widespread use of, and relative ease of implementation of, the ERP or PERP projection formats, horizontal wraparound motion compensation can be used to improve the visual quality of 360-degree video encoded in the ERP projection format.

[0086]

[0103] FIG. 7A shows the horizontal wraparound motion compensation process. If a part of a reference block is outside the left (or right) boundary of a reference picture in the projection region, instead of repeated padding, the "out-of-bounds" part is taken from the corresponding sphere neighborhood located in the reference picture toward the right (or left) boundary in the projection region. Repeated padding is used only for the top and bottom picture boundaries. As shown in FIG. 7B, horizontal wraparound motion compensation can be combined with the non-normative padding method often used in 360-degree video coding. This can be achieved by signaling a high-level syntax element to indicate the wraparound offset to be set to the width of the ERP picture before padding. This syntax is used to adjust the position of the horizontal wraparound accordingly. This syntax is not affected by the specific padding amount on the left and right picture boundaries, and therefore naturally supports asymmetric padding of ERP pictures (with different left and right padding). Horizontal wraparound motion compensation provides more meaningful information for motion compensation when the reference samples are outside the left and right boundaries of the reference picture. This tool improves compression performance not only in terms of rate-distortion performance, but also in terms of reducing seam artifacts and improving subjective quality of the reconstructed 360-degree video. Horizontal wraparound motion compensation can also be used for other single-plane projection formats with constant sampling density in the horizontal direction, such as adjusted equal-area projection.

[0087]

[0104] According to some embodiments, wrap-around motion compensation is signaled in the SPS. First, an enable flag is signaled. If the enable flag is true, the wrap-around offset is signaled. The SPS syntax is shown in Figure 8 and the corresponding semantics are given below.

[0088]

[0105] "sps_ref_wraparound_enabled_flag" equal to 1 specifies that horizontal wraparound motion compensation is applied in inter prediction. "sps_ref_wraparound_enabled_flag" equal to 0 specifies that horizontal wraparound motion compensation is not applied. If the value of (CtbSizeY / MinCbSizeY+1) is greater than (pic_width_in_luma_samples / MinCbSizeY-1) and "pic_width_in_luma_samples" is the value of "pic_width_in_luma_samples" in any picture parameter set (PPS) that references an SPS, the value of "sps_ref_wraparound_enabled_flag" is equal to 0.

[0089]

[0106] "sps_ref_wraparound_offset_minus1" plus 1 specifies the offset used to calculate the horizontal wraparound position in MinCbSizeY luma samples. The value of "ref_wraparound_offset_minus1" is in the range (CtbSizeY / MinCbSizeY)+1 to (pic_width_in_luma_samples / MinCbSizeY)-1, and "pic_width_in_luma_samples" is the value of "pic_width_in_luma_samples" in any PPS that references the SPS.

[0090]

[0107] CtbSizeY is the luma size of the coding tree block (CTB), MinCbSizeY is the minimum size of a luma coding block, and "pic_width_in_luma_samples" is the picture width in luma samples.

[0091]

[0108] According to some embodiments, the maximum width and height of all pictures in a picture sequence are signaled in an SPS, and then the picture width and height for the current picture are signaled in each PPS. The syntax for signaling the maximum picture width and height is shown in Table 4 of Figure 9, and the syntax for signaling the picture width and height is shown in Table 5 of Figure 10. The semantics corresponding to Figures 9 and 10 are given below.

[0092]

[0109] "ref_pic_resampling_enabled_flag" equal to 1 specifies that reference picture resampling can be applied when decoding a coded picture in a coding layer video sequence (CLVS) that references an SPS. ref_pic_resampling_enabled_flag equal to 0 specifies that reference picture resampling is not applied when decoding a picture in a CLVS that references an SPS. For example, the decoding program may decode each of the frames. If the decoding program determines that the resolution of the current frame is different from the resolution of the reference picture, the decoding program may perform appropriate resampling on the reference picture and use the generated resampled reference picture as a reference picture for the current frame. That is, resampling of the reference picture is necessary if the spatial resolution of the picture is allowed to be changed in the video sequence. The appropriate resampling of the reference picture may be downsampling or upsampling of the reference picture.

[0093]

[0110] "pic_width_max_in_luma_samples" specifies the maximum width in luma samples of each decoded picture that references an SPS. "pic_width_max_in_luma_samples" does not have to be equal to 0 and may be an integer multiple of Max(8,MinCbSizeY).

[0094]

[0111] "pic_height_max_in_luma_samples" specifies the maximum height in luma samples of each decoded picture that references an SPS. "pic_height_max_in_luma_samples" does not have to be equal to 0 and may be an integer multiple of Max(8,MinCbSizeY).

[0095]

[0112] "pic_width_in_luma_samples" specifies the width in luma samples of each decoded picture that references the PPS. "pic_width_in_luma_samples" may not be equal to 0. Rather, "pic_width_in_luma_samples" may be an integer multiple of Max(8,MinCbSizeY) and may be less than or equal to "pic_width_max_in_luma_samples".

[0096]

[0113] If "subpics_present_flag" is equal to 1 or "ref_pic_resampling_enabled_flag" is equal to 0, then the value of "pic_width_in_luma_samples" is equal to "pic_width_max_in_luma_samples".

[0097]

[0114] "pic_height_in_luma_samples" specifies the height in luma samples of each decoded picture that references the PPS. "pic_height_in_luma_samples" may not be equal to 0. Rather, "pic_height_in_luma_samples" may be an integer multiple of Max(8,MinCbSizeY) and may be less than or equal to "pic_height_max_in_luma_samples".

[0098]

[0115] If "subpics_present_flag" is equal to 1 or "ref_pic_resampling_enabled_flag" is equal to 0, then the value of "pic_height_in_luma_samples" is equal to "pic_height_max_in_luma_samples".

[0099]

[0116] The SPS signaling of virtual boundaries mentioned above may cause some ambiguity. Specifically, the ranges of "sps_virtual_boundaries_pos_x[i]" and "sps_virtual_boundaries_pos_y[i]" signaled in the SPS are 0 to Ceil(pic_width_in_luma_samples÷8)-1 and 0 to Ceil(pic_height_in_luma_samples÷8)-1, respectively. However, as explained above, "pic_width_in_luma_samples" and "pic_height_in_luma_samples" are signaled in the PPS, and they may vary from PPS to PPS. Since there may be multiple PPSs referencing the same SPS, it is not clear whether "pic_width_in_luma_samples" and "pic_height_in_luma_samples" should be set as upper bounds for "sps_virtual_boundaries_pos_x[i]" and "sps_virtual_boundaries_pos_y[i]".

[0100]

[0117] Furthermore, the SPS signaling of wraparound motion compensation described above may have some problems. Specifically, "sps_ref_wraparound_enabled_flag" and "sps_ref_wraparound_offset_minus1" are syntax elements signaled in the SPS, but there is a conformance constraint of "sps_ref_wraparound_enabled_flag" that depends on "pic_width_in_luma_samples" signaled on the PPS. The range of "sps_ref_wraparound_offset_minus1" also depends on "pic_width_in_luma_samples" signaled on the PPS. These dependencies cause some problems. First, it is not an efficient way to restrict the value of an SPS syntax element by a syntax element in all of the related PPSs, because the SPS is a higher level than the PPS. Furthermore, it is usually understood that a high-level syntax should not refer to a low-level syntax. Second, according to the current design, "sps_ref_wraparound_enabled_flag" can only be true if the widths of all pictures in the sequence that reference the SPS meet the constraint. Therefore, even if only one frame in the entire sequence does not meet the constraint, wraparound motion compensation cannot be used. Therefore, the benefit of wraparound motion compensation for the entire sequence is lost because of only one frame.

[0101]

[0118] The present disclosure provides methods for solving the above problems associated with signaling virtual boundaries or wraparound motion compensation. Several exemplary embodiments according to the disclosed methods are described in detail below.

[0102]

[0119] In some example embodiments, to solve the above problems related to signaling virtual boundaries, the upper bounds of the values ​​of "sps_virtual_boundaries_pos_x[i]" and "sps_virtual_boundaries_pos_y[i]" are changed to the minimum of the width and height of a picture in the sequence. Thus, for each picture, the positions of the virtual boundaries ("sps_virtual_boundaries_pos_x[i]" and "sps_virtual_boundaries_pos_y[i]") signaled in the SPS do not cross picture boundaries.

[0103]

[0120] The semantics according to these embodiments are described below.

[0104]

[0121] "sps_virtual_boundaries_pos_x[i]" specifies the position of the i-th vertical virtual boundary in luma samples divided by 8. The value of "sps_virtual_boundaries_pos_x[i]" is in the range 1 to Ceil(pic_width_in_luma_samples÷8)-1, and "pic_width_in_luma_samples" is the value of "pic_width_in_luma_samples" in any PPS that references the SPS.

[0105]

[0122] "sps_virtual_boundaries_pos_y[i]" specifies the position of the i-th horizontal virtual boundary in luma samples divided by 8. The value of "sps_virtual_boundaries_pos_y[i]" is in the range 1 to Ceil(pic_height_in_luma_samples÷8)-1, and "pic_height_in_luma_samples" is the value of "pic_width_in_luma_samples" in any PPS that references the SPS.

[0106]

[0123] In some example embodiments, to solve the above problems related to signaling virtual borders, the values ​​of "sps_virtual_boundaries_pos_x[i]" and "sps_virtual_boundaries_pos_y[i]" are upper bounded to the maximum values ​​of the width and height of a picture in the sequence, "pic_width_max_in_luma_samples" and "pic_height_max_in_luma_samples", respectively. For each picture, if the position of the virtual border signaled in the SPS ("sps_virtual_boundaries_pos_x[i]" and "sps_virtual_boundaries_pos_y[i]") exceeds the picture border, the virtual border is clipped or discarded within the border.

[0107]

[0124] The semantics according to these embodiments are described below.

[0108]

[0125] "sps_virtual_boundaries_pos_x[i]" specifies the position of the i-th vertical virtual boundary in luma samples divided by 8. The value of "sps_virtual_boundaries_pos_x[i]" is in the range of 1 to Ceil(pic_width_max_in_luma_samples÷8)-1.

[0109]

[0126] "sps_virtual_boundaries_pos_y[i]" specifies the position of the i-th horizontal virtual boundary in luma samples divided by 8. The value of "sps_virtual_boundaries_pos_y[i]" is in the range of 1 to Ceil(pic_height_max_in_luma_samples÷8)-1.

[0110]

[0127] As an example, for each picture, the positions of the virtual boundaries signaled in the SPS are clipped within the current picture boundary. In this example, the derived positions of the virtual boundaries VirtualBoundariesPosX[i] and VirtualBoundariesPosY[i] and the numbers of virtual boundaries VirtualBoundariesNumVer and VirtualBoundariesNumHor are derived as follows: VirtualBoundariesNumVer=sps_virtual_boundaries_present_flag? sps_num_ver_virtual_boundaries:ph_num_ver_virtual_boundaries VirtualBoundariesPosX[i]=(sps_virtual_boundaries_present_flag? min(Ceil(pic_width_in_luma_samples÷8)-1,sps_virtual_boundaries_pos_x[i]):ph_virtual_boundaries_pos_x[i])*8 VirtualBoundariesNumHor=sps_virtual_boundaries_present_flag? sps_num_hor_virtual_boundaries:ph_num_hor_virtual_boundaries VirtualBoundariesPosY[i]=(sps_virtual_boundaries_present_flag? min(Ceil(pic_width_in_luma_samples÷8)-1,sps_virtual_boundaries_pos_y[i]):ph_virtual_boundaries_pos_y[i])*8

[0111]

[0128] The distance between any two vertical virtual borders can be 0 or greater than or equal to CtbSizeY luma samples.

[0112]

[0129] The distance between any two horizontal virtual boundaries can be 0 or greater than or equal to CtbSizeY luma samples.

[0113]

[0130] As another example, for each picture, if the position of the virtual border signaled in the SPS exceeds the current picture boundary, the virtual border is not used in the current picture. The derived virtual border positions VirtualBoundariesPosX[i] and VirtualBoundariesPosY[i] and the numbers of virtual borders VirtualBoundariesNumVer and VirtualBoundariesNumHor are derived as follows: VirtualBoundariesNumVer=sps_virtual_boundaries_present_flag? sps_num_ver_virtual_boundaries:ph_num_ver_virtual_boundaries VirtualBoundariesPosXInPic[i]=(sps_virtual_boundaries_present_flag? sps_virtual_boundaries_pos_x[i]:ph_virtual_boundaries_pos_x[i])*8,(i=0..VirtualBoundariesNumVer) VirtualBoundariesNumHor=sps_virtual_boundaries_present_flag? sps_num_hor_virtual_boundaries:ph_num_hor_virtual_boundaries VirtualBoundariesPosYInPic[i]=(sps_virtual_boundaries_present_flag? sps_virtual_boundaries_pos_y[i]:ph_virtual_boundaries_pos_y[i])*8,(i=0..VirtualBoundariesNumHor) for(i=0,j=0;i <VirtualBoundariesNumVer;i++){ if(VirtualBoundariesPosXInPic[i]<=Ceil(pic_width_in_luma_samples÷8)-1){ VirtualBoundariesPosX[j++]=VirtualBoundariesPosXInPic[i] } } VirtualBoundariesNumVer=j for(i=0,j=0;i <VirtualBoundariesNumHor;i++){ if(VirtualBoundariesPosYInPic[i]<=Ceil(pic_height_in_luma_samples÷8)-1){ VirtualBoundariesPosY[j++]=VirtualBoundariesPosYInPic[i] } } VirtualBoundariesNumHor=j

[0114]

[0131] Instead, the positions of the derived virtual boundaries "VirtualBoundariesPosX[i]", "VirtualBoundariesPosY[i]" and the numbers of virtual boundaries VirtualBoundariesNumVer and VirtualBoundariesNumHor are derived as follows. for(i=0,j=0;i <sps_num_ver_virtual_boundaries;i++){ if(sps_virtual_boundaries_pos_x[i]<=Ceil(pic_width_in_luma_samples÷8)-1){ VirtualBoundariesPosX[j++]=sps_virtual_boundaries_pos_x[i] } } VirtualBoundariesNumVer=j for(i=0,j=0;i<sps_num_hor_virtual_boundaries;i++){ if(sps_virtual_boundaries_pos_y[i]<=Ceil(pic_height_in_luma_samples÷8)-1){ VirtualBoundariesPosY[j++]=sps_virtual_boundaries_pos_y[i] } } VirtualBoundariesNumHor=j VirtualBoundariesNumVer=sps_virtual_boundaries_present_flag? VirtualBoundariesNumVer:ph_num_ver_virtual_boundaries VirtualBoundariesPosX[i]=(sps_virtual_boundaries_present_flag? VirtualBoundariesPosX[i]:ph_virtual_boundaries_pos_x[i])*8,(i=0..VirtualBoundariesNumVer) VirtualBoundariesNumHor=sps_virtual_boundaries_present_flag? VirtualBoundariesNumHor:ph_num_hor_virtual_boundaries VirtualBoundariesPosY[i]=(sps_virtual_boundaries_present_flag? VirtualBoundariesPosY[i]:ph_virtual_boundaries_pos_y[i])*8,(i=0..VirtualBoundariesNumHor)

[0115]

[0132] In some example embodiments, to solve the above problems related to signaling of virtual boundaries, the values ​​of "sps_virtual_boundaries_pos_x[i]" and "sps_virtual_boundaries_pos_y[i]" are upper bounded to the maximum values ​​of the width and height of a picture in the sequence, "pic_width_max_in_luma_samples" and "pic_height_max_in_luma_samples", respectively. For each picture, the positions of the virtual boundaries signaled in the SPS ("sps_virtual_boundaries_pos_x[i]" and "sps_virtual_boundaries_pos_y[i]") are scaled according to the ratio between the maximum picture width and height signaled in the SPS and the width and height of the current picture signaled in the PPS.

[0116]

[0133] The semantics according to these embodiments are described below.

[0117]

[0134] "sps_virtual_boundaries_pos_x[i]" specifies the position of the i-th vertical virtual boundary in luma samples divided by 8. The value of "sps_virtual_boundaries_pos_x[i]" is in the range of 1 to Ceil(pic_width_max_in_luma_samples÷8)-1.

[0118]

[0135] "sps_virtual_boundaries_pos_y[i]" specifies the position of the i-th horizontal virtual boundary in luma samples divided by 8. The value of "sps_virtual_boundaries_pos_y[i]" is in the range of 1 to Ceil(pic_height_max_in_luma_samples÷8)-1.

[0119]

[0136] To derive the location of the virtual border for each picture, a scaling ratio is first calculated, and then the location of the virtual border signaled in the SPS is scaled as follows: VBScaleX=((pic_width_max_in_luma_samples<<14)+(pic_width_in_luma_samples>>1)) / pic_width_in_luma_samples VBScaleY=((pic_height_max_in_luma_samples<<14)+(pic_height_in_luma_samples>>1)) / pic_height_in_luma_samples SPSVirtualBoundariesPosX[i]=(sps_virtual_boundaries_pos_x[i]×VBScaleX+(1<<13))>>14 SPSVirtualBoundariesPosY[i]=(sps_virtual_boundaries_pos_y[i]×VBScaleY+(1<<13))>>14

[0120]

[0137] For example, "SPSVirtualBoundariesPosX[i]" and "SPSVirtualBoundariesPosY[i]" can be further rounded to an 8-pixel grid as follows: SPSVirtualBoundariesPosX[i]=((SPSVirtualBoundariesPosX[i]+4)>>3)<<3 SPSVirtualBoundariesPosY[i]=((SPSVirtualBoundariesPosY[i]+4)>>3)<<3

[0121]

[0138] Finally, the positions of the derived virtual boundaries "VirtualBoundariesPosX[i]" and "VirtualBoundariesPosY[i]" and the numbers of virtual boundaries "VirtualBoundariesNumVer" and "VirtualBoundariesNumHor" are derived as follows. VirtualBoundariesNumVer=sps_virtual_boundaries_present_flag? sps_num_ver_virtual_boundaries:ph_num_ver_virtual_boundaries VirtualBoundariesPosX[i]=(sps_virtual_boundaries_present_flag? SPSVirtualBoundariesPosX[i]:ph_virtual_boundaries_pos_x[i])*8,(i=0..VirtualBoundariesNumVer) VirtualBoundariesNumHor=sps_virtual_boundaries_present_flag? sps_num_hor_virtual_boundaries:ph_num_hor_virtual_boundaries VirtualBoundariesPosY[i]=(sps_virtual_boundaries_present_flag? SPSVirtualBoundariesPosY[i]:ph_virtual_boundaries_pos_y[i])*8,(i=0..VirtualBoundariesNumHor)

[0122]

[0139] In some exemplary embodiments, to solve the above problems related to signaling virtual borders, signaling of virtual borders at the sequence level and changing picture width and height are used mutually exclusive. For example, if virtual borders are signaled in the SPS, picture width and height cannot be changed in the sequence. If picture width and height are changed in the sequence, virtual borders cannot be signaled in the SPS.

[0123]

[0140] In an embodiment of the present invention, the upper limit of the values ​​of "sps_virtual_boundaries_pos_x[i]" and "sps_virtual_boundaries_pos_y[i]" is changed to "pic_width_max_in_luma_samples" and "pic_width_max_in_luma_samples", which are the maximum values ​​of the width and height of a picture in the sequence, respectively.

[0124]

[0141] The semantics of "sps_virtual_boundaries_pos_x[i]" and "sps_virtual_boundaries_pos_y[i]" according to an embodiment of the present invention are as follows:

[0125]

[0142] "sps_virtual_boundaries_pos_x[i]" specifies the position of the i-th vertical virtual boundary in luma samples divided by 8. The value of "sps_virtual_boundaries_pos_x[i]" is in the range of 1 to Ceil(pic_width_max_in_luma_samples÷8)-1.

[0126]

[0143] "sps_virtual_boundaries_pos_y[i]" specifies the position of the i-th horizontal virtual boundary in luma samples divided by 8. The value of "sps_virtual_boundaries_pos_y[i]" is in the range of 1 to Ceil(pic_height_max_in_luma_samples÷8)-1.

[0127]

[0144] As an example, bitstream conformance requirements for "pic_width_in_luma_samples" and "pic_height_in_luma_samples" can be imposed as follows:

[0128]

[0145] The value of "pic_width_in_luma_samples" is equal to "pic_width_max_in_luma_samples" if (1) "subpics_present_flag" is equal to 1, or (2) "ref_pic_resampling_enabled_flag" is equal to 0, or (3) "sps_virtual_boundaries_present_flag" is equal to 1. Subject to the third condition of this constraint (i.e., "sps_virtual_boundaries_present_flag" is equal to 1), if the virtual boundaries are within the SPS, then each picture in the sequence has the same width equal to the maximum width of any picture in the sequence.

[0129]

[0146] The value of "pic_height_in_luma_samples" is equal to "pic_height_max_in_luma_samples" if (1) "subpics_present_flag" is equal to 1, or (2) "ref_pic_resampling_enabled_flag" is equal to 0, or (3) "sps_virtual_boundaries_present_flag" is equal to 1. Subject to the third condition of this constraint (i.e., "sps_virtual_boundaries_present_flag" is equal to 1), if the virtual boundaries are within the SPS, then each picture in the sequence has the same height equal to the maximum height of any picture in the sequence.

[0130]

[0147] As another example, bitstream conformance requirements for "pic_width_in_luma_samples" and "pic_height_in_luma_samples" can be imposed as follows:

[0131]

[0148] The value of "pic_width_in_luma_samples" is equal to "pic_width_max_in_luma_samples" if (1) "subpics_present_flag" is equal to 1, or (2) "ref_pic_resampling_enabled_flag" is equal to 0, or (3) "sps_num_ver_vritual_boundaries" is not equal to 0. Subject to the third condition of this constraint (i.e., "sps_num_ver_vritual_boundaries" is not equal to 0), if the number of vertical virtual borders is greater than 0 (i.e., there is at least one vertical virtual border), then each picture in the sequence has the same width, equal to the maximum width of any picture in the sequence.

[0132]

[0149] The value of "pic_height_in_luma_samples" is equal to "pic_height_max_in_luma_samples" if (1) "subpics_present_flag" is equal to 1, or (2) "ref_pic_resampling_enabled_flag" is equal to 0, or (3) "sps_num_hor_vritual_boundaries" is not equal to 0. Subject to the third condition of this constraint (i.e., "sps_num_hor_vritual_boundaries" is not equal to 0), if the number of vertical virtual borders is greater than 0 (i.e., there is at least one vertical virtual border), then each picture in the sequence has the same width, equal to the maximum width of any picture in the sequence.

[0133]

[0150] As another example, a bitstream conformance requirement for "sps_virtual_boundaries_present_flag" can be imposed as follows: It is a bitstream conformance requirement that "sps_virtual_boundaries_present_flag" is 0 if "ref_pic_resampling_enabled_flag" is 1. According to this constraint, if reference picture resampling is enabled, the virtual boundary should not be within the SPS. The principle of this constraint is as follows: Reference picture resampling is used when the current picture has a different resolution than the reference picture. If the resolution of a picture is allowed to change within a sequence, the location of the virtual boundary can also change from picture to picture, since different pictures may have different resolutions. Therefore, the location of the virtual boundary can be signaled at the picture level (e.g., PPS). Signaling of the virtual boundary at the sequence level (i.e., signaled within the SPS) is not appropriate.

[0134]

[0151] As another example, "sps_virtual_boundaries_present_flag" is conditionally signaled based on "ref_pic_resampling_enabled_flag". For example, "ref_pic_resampling_enabled_flag" equal to 0 specifies that reference picture resampling is not applied when decoding pictures in the CLVS that reference an SPS, and "sps_virtual_boundaries_present_flag" is signaled based on "!ref_pic_resampling_enabled_flag" having a value of 1. This syntax is shown in Figure 11, with changes to the syntax in Table 1 (Figure 5) in italics. The associated semantics are described as follows:

[0135]

[0152] "sps_virtual_boundaries_present_flag" equal to 1 specifies that virtual boundary information is signaled in the SPS. "sps_virtual_boundaries_present_flag" equal to 0 specifies that virtual boundary information is not signaled in the SPS. If one or more virtual boundaries are signaled in the SPS, in-loop filtering operations that cross the virtual boundaries in pictures that reference the SPS are disabled. In the absence of "sps_virtual_boundaries_present_flag", its value is inferred to be 0. In-loop filtering operations include a deblocking filter, a sample adaptive offset filter, and an adaptive loop filter operation.

[0136]

[0153] As another example, "ref_pic_resampling_enabled_flag" is conditionally signaled based on "sps_virtual_boundaries_present_flag". This syntax is shown in Figure 12, where changes to the syntax in Table 1 (Figure 5) are shown using italics and strikethrough. The associated semantics are described below.

[0137]

[0154] "ref_pic_resampling_enabled_flag" equal to 1 specifies that reference picture resampling can be applied when decoding a coded picture in the CLVS that references an SPS. "ref_pic_resampling_enabled_flag" equal to 0 specifies that reference picture resampling is not applied when decoding a picture in the CLVS that references an SPS. If "ref_pic_resampling_enabled_flag" is not present, its value is inferred to be 0.

[0138]

[0155] Furthermore, to solve the above problems related to signaling of wraparound motion compensation, the "sps_ref_wraparound_enabled_flag" can be true only if the widths of all pictures in the sequence that reference the SPS satisfy the constraint, and the following embodiment is provided by this disclosure.

[0139]

[0156] In some embodiments, the constraints on the wraparound motion compensation enable flag and the wraparound offset are modified to depend on the maximum picture width signaled in the SPS. At the picture level, a picture width check is introduced. Wraparound can only be applied to pictures with a width that satisfies the condition. For pictures whose width does not satisfy the condition, wraparound is turned off even if "sps_ref_wraparound_enabled_flag" is true. By doing so, there is no need to restrict "sps_ref_wraparound_enabled_flag" and "sps_ref_wraparound_offset_minus1" signaled in the SPS with the picture width signaled in the PPS. The syntax remains unchanged and the semantics of "sps_ref_wraparound_enabled_flag" and "sps_ref_wraparound_offset_minus1" are as follows:

[0140]

[0157] "sps_ref_wraparound_enabled_flag" equal to 1 specifies that horizontal wraparound motion compensation may be applied in inter prediction. sps_ref_wraparound_enabled_flag equal to 0 specifies that horizontal wraparound motion compensation is not applied. If the value of (CtbSizeY / MinCbSizeY+1) is greater than (pic_width_max_in_luma_samples / MinCbSizeY-1), the value of sps_ref_wraparound_enabled_flag shall be equal to 0.

[0141]

[0158] "sps_ref_wraparound_offset_minus1" plus 1 specifies the maximum value of the offset used to calculate the horizontal wraparound position in MinCbSizeY luma samples. The value of sps_ref_wraparound_offset_minus1 shall be in the range of (CtbSizeY / MinCbSizeY)+1 to (pic_width_max_in_luma_samples / MinCbSizeY)-1.

[0142]

[0159] As per VVC draft 7, pic_width_max_in_luma_samples is the maximum width in luma samples of each decoded picture that references an SPS. CtbSizeY and MinCbSizeY are as specified in VVC draft 7.

[0143]

[0160] For each picture in the sequence, the variable "PicRefWraparoundEnableFlag" is defined as follows:

[0144]

[0161] PicRefWraparoundEnableFlag=sps_ref_wraparound_enabled_flag&&sps_ref_wraparound_offset_minus1<=(pic_width_in_luma_samples / MinCbSizeY-1))

[0145]

[0162] where "pic_width_in_luma_samples" is the picture width referring to the PPS at which "pic_width_in_luma_samples" is signaled, as per VVC draft 7. CtbSizeY and MinCbSizeY are as defined in VVC draft 7.

[0146]

[0163] The variable "PicRefWraparoundEnableFlag" is used to determine whether wraparound MC can be enabled for the current picture.

[0147]

[0164] In an alternative method, for each picture in the sequence, two variables "PicRefWraparoundEnableFlag" and "PicRefWraparoundOffset" are defined as follows:

[0148]

[0165] PicRefWraparoundEnableFlag=sps_ref_wraparound_enabled_flag&&((ctbSizeY / MinCbSizeY+1)<=(pic_width_in_luma_samples / MinCbSizeY-1))

[0149]

[0166] PicRefWraparoundOffset=min(sps_ref_wraparound_ofset_minus+1,(pic_width_in_luma_samples / MinCbSizeY))

[0150]

[0167] where "pic_width_in_luma_samples" is the picture width referring to the PPS at which "pic_width_in_luma_samples" is signaled, as per VVC draft 7. CtbSizeY and MinCbSizeY are as defined in VVC draft 7.

[0151]

[0168] The variable "PicRefWraparoundEnableFlag" is used to determine whether wraparound MC can be enabled for the current picture. If wraparound MC can be enabled, the offset "PicRefWraparoundOffset" can be used in the motion compensation process.

[0152]

[0169] Some embodiments provide for mutually exclusive use of wrap-around motion compensation and sequence-level virtual boundary signaling and picture width change: if wrap-around motion compensation is enabled, picture width cannot be changed within a sequence, if picture width is changed within a sequence, wrap-around motion compensation is disabled.

[0153]

[0170] As an example, the semantics of "sps_ref_wraparound_enabled_flag" and "sps_ref_wraparound_offset_minus1" according to an embodiment of the present invention are described as follows:

[0154]

[0171] "sps_ref_wraparound_enabled_flag" equal to 1 specifies that horizontal wraparound motion compensation is applied in inter prediction. "sps_ref_wraparound_enabled_flag" equal to 0 specifies that horizontal wraparound motion compensation is not applied. If the value of (CtbSizeY / MinCbSizeY+1) is greater than (pic_width_max_in_luma_samples / MinCbSizeY-1), the value of sps_ref_wraparound_enabled_flag may be equal to 0.

[0155]

[0172] "sps_ref_wraparound_offset_minus1" plus 1 specifies the offset used to calculate the horizontal wraparound position in MinCbSizeY luma samples. The value of "ref_wraparound_offset_minus1" can be in the range of (CtbSizeY / MinCbSizeY)+1 to (pic_width_max_in_luma_samples / MinCbSizeY)-1.

[0156]

[0173] Bitstream conformance requirements for "pic_width_in_luma_samples" and "pic_height_in_luma_samples" can be imposed as follows:

[0157]

[0174] If "subpics_present_flag" is equal to 1, or "ref_pic_resampling_enabled_flag" is equal to 0, or "sps_ref_wraparound_enabled_flag" is equal to 1, then the value of "pic_width_in_luma_samples" is equal to "pic_width_max_in_luma_samples".

[0158]

[0175] As another example, a bitstream conformance requirement for "sps_ref_wraparound_enabled_flag" may be imposed. The semantics of "sps_ref_wraparound_enabled_flag" and "sps_ref_wraparound_offset_minus1" according to an embodiment of the present invention are described as follows:

[0159]

[0176] "sps_ref_wraparound_enabled_flag" equal to 1 specifies that horizontal wraparound motion compensation is applied in inter prediction. "sps_ref_wraparound_enabled_flag" equal to 0 specifies that horizontal wraparound motion compensation is not applied. If the value of (CtbSizeY / MinCbSizeY+1) is greater than (pic_width_max_in_luma_samples / MinCbSizeY-1), the value of sps_ref_wraparound_enabled_flag is equal to 0.

[0160]

[0177] If ref_pic_resampling_enabled_flag is 1, the value of "sps_ref_wraparound_enabled_flag" may be 0. ref_pic_resampling_enabled_flag specifies whether reference resampling is enabled. Reference resampling is used to resample a reference picture when the resolution of the reference picture is different from the resolution of the current picture. Therefore, at the picture level (e.g., PPS, PH), wraparound motion compensation is not used to predict the current picture when the resolution of the reference picture is different from the current picture.

[0161]

[0178] "sps_ref_wraparound_offset_minus1" plus 1 specifies the offset used to calculate the horizontal wraparound position in MinCbSizeY luma samples. The value of "ref_wraparound_offset_minus1" is in the range of (CtbSizeY / MinCbSizeY)+1 to (pic_width_max_in_luma_samples / MinCbSizeY)-1.

[0162]

[0179] As another example, "sps_ref_wraparound_enabled_flag" is conditionally signaled based on "ref_pic_resampling_enabled_flag". This syntax is shown in Figure 13, with the changes to the syntax in Table 3 (Figure 8) in italics. The associated semantics are described as follows:

[0163]

[0180] "sps_ref_wraparound_enabled_flag" equal to 1 specifies that horizontal wraparound motion compensation is applied in inter prediction. "sps_ref_wraparound_enabled_flag" equal to 0 specifies that horizontal wraparound motion compensation is not applied. If the value of (CtbSizeY / MinCbSizeY+1) is greater than (pic_width_max_in_luma_samples / MinCbSizeY-1), the value of "sps_ref_wraparound_enabled_flag" is equal to 0. If "sps_ref_wraparound_enabled_flag" is not present, its value is inferred to be 0.

[0164]

[0181] "sps_ref_wraparound_offset_minus1" plus 1 specifies the offset used to calculate the horizontal wraparound position in MinCbSizeY luma samples. The value of ref_wraparound_offset_minus1 is in the range of (CtbSizeY / MinCbSizeY)+1 to (pic_width_max_in_luma_samples / MinCbSizeY)-1.

[0165]

[0182] As an example, "ref_pic_resampling_enabled_flag" is conditionally signaled based on "sps_ref_wraparound_enabled_flag". This syntax is shown as Table 9 in Figure 14, with changes to the syntax in Table 3 (Figure 8) shown using italics and strikethrough. The associated semantics are as follows:

[0166]

[0183] "ref_pic_resampling_enabled_flag" equal to 1 specifies that reference picture resampling can be applied when decoding a coded picture in the CLVS that references an SPS. ref_pic_resampling_enabled_flag equal to 0 specifies that reference picture resampling is not applied when decoding a picture in the CLVS that references an SPS. If ref_pic_resampling_enabled_flag is not present, its value is inferred to be 0.

[0167]

[0184] 15 illustrates a flowchart of an example method 1500 according to some embodiments of the present disclosure. In some embodiments, the method 1500 may be performed by one or more software or hardware components of an encoder, an apparatus (e.g., apparatus 400 of FIG. 4). For example, a processor (e.g., processor 402 of FIG. 4) may perform the method 1500. In some embodiments, the method 1500 may be implemented by a computer program product embodied in a computer-readable medium that includes computer-executable instructions, such as program code, executed by a computer (e.g., apparatus 400 of FIG. 4).

[0168]

[0185] A virtual boundary can be set as one of the coding tools for 360-degree video or gradual decoding refresh (GDR) coding so that in-loop filtering can be disabled to prevent errors and artifacts. This is illustrated in the manner shown in FIG.

[0169]

[0186] In step 1501, a bitstream including a sequence of pictures is received. The sequence of pictures may be a set of pictures. As described, a basic processing unit of a color picture may include a luma component (Y) representing colorless luminance information, one or more chroma components (e.g., Cb and Cr) representing color information, and related syntax elements, where the luma and chroma components may have the same size of the basic processing unit. The luma and chroma components may be referred to as "coding tree blocks" ("CTBs") in some video coding standards (e.g., H.265 / HEVC or H.266 / VVC). Any operation performed on a basic processing unit may be repeatedly performed on each of its luma and chroma components.

[0170]

[0187] In step 1503, it is determined whether virtual borders are signaled at the sequence level (i.e., in the SPS) according to the received stream. If virtual borders are signaled at the sequence level, the number of virtual borders and the position of each virtual border are signaled for the picture that references the SPS.

[0171]

[0188] Virtual boundaries can be signaled in the sequence parameter set (SPS) or in the picture header (PH). In the SPS, the virtual boundary present flag "sps_virtual_boundaries_present_flag" is signaled first. If the flag is true, the number of virtual borders and the position of each virtual border are signaled for the picture that references the SPS.

[0172]

[0189] In step 1505, optionally, in response to the virtual boundaries not being signaled at the sequence level for the set of pictures, it is determined whether the virtual boundaries are signaled at the picture level (e.g., PPS or PH) for the pictures in the set of pictures according to the received stream. For example, if "sps_virtual_boundaries_present_flag" is false, another virtual boundary presence flag "ph_virtual_boundaries_present_flag" may be signaled in the PH. Similarly, if "ph_virtual_boundaries_present_flag" is true, the number of virtual boundaries and the position of each virtual boundary may be signaled for the picture associated with the PH. Step 1505 is optional, and in some embodiments, the method may include step 1505 in response to the virtual boundaries not being signaled at the sequence level. In some embodiments, in response to the virtual boundaries not being signaled at the sequence level, the method may end without performing step 1505. In response to the virtual border being signaled at the picture level, the method proceeds to step 1507 for determining the location of the virtual border.

[0173]

[0190] In some embodiments, it may be determined whether any condition is met that indicates that the virtual boundary is not signaled at the sequence level. As described above in the example, the value of "sps_virtual_boundaries_present_flag" in Table 1 of FIG. 5 may be determined, with a value of 1 indicating that the virtual boundary is signaled at the sequence level. However, there are other conditions that indicate that the virtual boundary is not applied at the sequence level. The first condition that indicates that there is no signaling of the virtual boundary at the sequence level is that reference resampling is enabled. The second condition is that a change in resolution of the pictures in the set of pictures is allowed. If either of these two conditions is met, the virtual boundary is not signaled at the sequence level. If neither of these two conditions is met, and if the value of "sps_virtual_boundaries_present_flag" is determined to be 1, as in this embodiment, it is determined that the virtual boundary is signaled at the sequence level for the set of pictures. The method proceeds to step 1507 for determining the location of the virtual boundary.

[0174]

[0191] In some embodiments, it is determined whether reference resampling is enabled for the sequence of pictures according to the received bitstream, and in response to reference resampling being enabled for the sequence of pictures, no virtual boundary is signaled at the sequence level.

[0175]

[0192] The decoding program may decode each of the frames. If the decoding program determines that the resolution of the current frame is different from the resolution of the reference picture, the decoding program may perform appropriate resampling on the reference picture and may use the generated resampled reference block as a reference block for decoding the block in the current frame. Reference picture resampling is necessary if the spatial resolution of the picture is allowed to be changed in the video sequence. If resampling is enabled, resolution change may or may not be allowed. If the resolution of the picture is allowed to be changed, resampling is necessary, so resampling is enabled. The appropriate resampling of the reference picture may be downsampling or upsampling of the reference picture. For example, as shown in FIG. 9, "ref_pic_resampling_enabled_flag" equal to 1 specifies that reference picture resampling can be applied when decoding a coded picture in the CLVS that references the SPS. ref_pic_resampling_enabled_flag equal to 0 specifies that reference picture resampling is not applied when decoding a picture in the CLVS that references the SPS.

[0176]

[0193] For example, "sps_virtual_boundaries_present_flag" in Table 1 of FIG. 5 may be conditionally signaled based on "ref_pic_resampling_enabled_flag". The value of "ref_pic_resampling_enabled_flag" may be determined. "ref_pic_resampling_enabled_flag" equal to 0 specifies that reference picture resampling is not applied when decoding pictures in the CLVS that reference an SPS. If "ref_pic_resampling_enabled_flag" is 1, then "sps_virtual_boundaries_present_flag" is 0.

[0177]

[0194] In some embodiments, determining whether a change in resolution of a picture in a sequence of pictures is allowed, and in response to the change in resolution being allowed, determining that the virtual boundary is not signaled at a sequence level.

[0178]

[0195] Reference picture resampling is used when the current picture has a different resolution than the reference picture. If the resolution of a picture is allowed to be changed in a sequence, the position of the virtual border can also be changed for each picture, since different pictures may have different resolutions. Therefore, the position of the virtual border can be signaled at the picture level (e.g., PPS and PH), while the signaling of the virtual border at the sequence level (i.e., signaled in the SPS) is not appropriate. Therefore, it is determined that the virtual border is not signaled at the sequence level when the resolution of a picture is allowed to be changed. At the same time, if resampling is allowed but the resolution of a picture is not allowed to be changed, no constraint is imposed on the signaling of the virtual border.

[0179]

[0196] In some embodiments, in response to reference resampling being enabled for the sequence of pictures, it is determined that wraparound motion compensation is disabled for the sequence of pictures.

[0180]

[0197] For example, it is determined that the value of "ref_pic_resampling_enabled_flag" in FIG. 13 is 1, "sps_ref_wraparound_enabled_flag" is 0, and horizontal wraparound motion compensation is not applied. "ref_pic_resampling_enabled_flag" specifies whether reference resampling is enabled. Reference resampling is used to resample a reference picture when the resolution of the reference picture is different from that of the current picture. At the picture level (e.g., PPS), in this example, when the resolution of the reference picture is different from that of the current picture, wraparound motion compensation is not used to predict the current picture. If it is determined that the value of "ref_pic_resampling_enabled_flag" is 0, "sps_ref_wraparound_enabled_flag" is 1, and horizontal wraparound motion compensation is applied.

[0181]

[0198] In some embodiments, in response to allowing a change in resolution of a picture within the sequence of pictures, it is determined that wraparound motion compensation is disabled for the sequence of pictures.

[0182]

[0199] In step 1507, in response to the virtual borders being signaled at the sequence level, determine positions of virtual borders for the sequence. The positions are bounded by ranges based on the widths and heights of the pictures in the sequence. The virtual borders may include vertical and horizontal borders. The positions of the virtual borders may include vertical border position points and horizontal border position points. For example, as shown in Table 1 of FIG. 5, "sps_virtual_boundaries_pos_x[i]" specifies the position of the i-th vertical virtual border in luma sample units divided by 8, while "sps_virtual_boundaries_pos_y[i]" specifies the position of the i-th horizontal virtual border in luma sample units divided by 8.

[0183]

[0200] In some embodiments, the vertical extent of this position is less than or equal to the maximum width allowed for each picture in the sequence, and the horizontal extent is less than or equal to the maximum height allowed for each picture in the set. The maximum width and height are signaled in the received stream. An example of the semantics is as follows:

[0184]

[0201] "sps_virtual_boundaries_pos_x[i]" specifies the position of the i-th vertical virtual boundary in luma samples divided by 8. The value of "sps_virtual_boundaries_pos_x[i]" is in the range of 1 to Ceil(pic_width_max_in_luma_samples÷8)-1.

[0185]

[0202] "sps_virtual_boundaries_pos_y[i]" specifies the position of the i-th horizontal virtual boundary in luma samples divided by 8. The value of "sps_virtual_boundaries_pos_y[i]" is in the range of 1 to Ceil(pic_height_max_in_luma_samples÷8)-1.

[0186]

[0203] If virtual boundaries are in the SPS, i.e. "sps_virtual_boundaries_present_flag" is equal to 1, then each picture in the sequence has the same width equal to the maximum width of any picture in the sequence.

[0187]

[0204] In some embodiments, the location of the virtual boundary for a picture in a set of pictures may be determined in response to a virtual boundary being signaled at the picture level for that picture. One or more boundaries may be signaled for one or more pictures in the set. For example, as shown in the semantics of the PH syntax in Figure 6, "ph_num_ver_virtual_boundaries" specifies the number of "ph_virtual_boundaries_pos_x[i]" syntax elements in the PH. If "ph_num_ver_virtual_boundaries" is not present, its value is inferred to be equal to 0.

[0188]

[0205] "ph_virtual_boundaries_pos_x[i]" specifies the position of the i-th vertical virtual boundary in luma samples divided by 8. The value of "ph_virtual_boundaries_pos_x[i]" is in the range of 1 to Ceil(pic_width_in_luma_samples÷8)-1.

[0189]

[0206] "ph_virtual_boundaries_pos_y[i]" specifies the position of the i-th horizontal virtual boundary in luma samples divided by 8. The value of "ph_virtual_boundaries_pos_y[i]" is in the range of 1 to Ceil(pic_height_in_luma_samples÷8)-1.

[0190]

[0207] As mentioned, in applications such as 360-degree video, the layout of a particular projection format usually has multiple faces. For example, MPEG-I Part 2: Omnidirectional Media Format (OMAF) standardizes a cube-map-based projection format named CMP with six faces. For a projection format that includes multiple faces, regardless of what kind of compact frame packing configuration is used, a discontinuity occurs between two or more adjacent faces in the frame-packed picture. If an in-loop filtering operation is performed across this discontinuity, surface seam artifacts may become visible in the rendered reconstructed image. To mitigate the surface seam artifacts, in-loop filtering operations across the discontinuity in the frame-packed picture must be disabled. A virtual boundary can be used, which the loop filtering operation across is disabled. The encoder can set the discontinuity boundary as a virtual boundary, thereby not applying loop filters on the discontinuity boundary. Besides 360-degree video, the virtual boundary can also be used for gradual decoding refresh (GDR), which is mainly used in ultra-low latency applications. In ultra-low delay applications, inserting an intra-coded picture as a random access point picture may cause unacceptable transmission latency due to the large size of the intra-coded picture. To reduce the latency, GDR is adopted, in which pictures are gradually refreshed by inserting intra-coded regions in B / P pictures. To prevent error propagation, pixels in the refreshed region in a picture cannot refer to those in the unrefreshed region of the current picture or reference picture. Therefore, loop filtering cannot be applied across the boundary between the refreshed region and the unrefreshed region. With the above virtual boundary scheme, the encoder can set the boundary between the refreshed region and the unrefreshed region as a virtual boundary, and therefore cannot apply loop filtering operations across the boundary.

[0191]

[0208] Step 1509 disables in-loop filtering operations (e.g., look filter stage 232 in FIG. 3A) that cross a virtual boundary of a picture. If one or more virtual boundaries are signaled in the SPS, in-loop filtering operations that cross a virtual boundary in a picture that references the SPS may be disabled. In-loop filtering operations include a deblocking filter, a sample adaptive offset filter, and an adaptive loop filter operation.

[0192]

[0209] According to some embodiments of the present disclosure, another exemplary method 1600 is provided. The method may be performed by one or more software or hardware components of an encoder, an apparatus (e.g., apparatus 400 of FIG. 4). For example, a processor (e.g., processor 402 of FIG. 4) may perform the method 1600. In some embodiments, the method 1600 may be implemented by a computer program product embodied in a computer-readable medium that includes computer-executable instructions, such as program code, executed by a computer (e.g., apparatus 400 of FIG. 4). The method may include the following steps:

[0193]

[0210] In step 1601, a bitstream including a sequence of pictures is received. The sequence of pictures may be a set of pictures. As described, a basic processing unit of a color picture may include a luma component (Y) representing colorless luminance information, one or more chroma components (e.g., Cb and Cr) representing color information, and related syntax elements, where the luma and chroma components may have the same size of the basic processing unit. The luma and chroma components may be referred to as "coding tree blocks" ("CTBs") in some video coding standards (e.g., H.265 / HEVC or H.266 / VVC). Any operation performed on a basic processing unit may be repeatedly performed on each of its luma and chroma components.

[0194]

[0211] Step 1603 determines whether reference resampling is enabled for the sequence of pictures according to the received bitstream.

[0195]

[0212] Step 1605 determines that, in response to reference resampling being enabled for the sequence of pictures, wraparound motion compensation is disabled for the sequence of pictures.

[0196]

[0213] For example, it is determined that the value of "ref_pic_resampling_enabled_flag" in FIG. 13 is 1, "sps_ref_wraparound_enabled_flag" is 0, and horizontal wraparound motion compensation is not applied. "ref_pic_resampling_enabled_flag" specifies whether reference resampling is enabled. Reference resampling is used to resample a reference picture when the resolution of the reference picture is different from that of the current picture. At the picture level (e.g., PPS), in this example, when the resolution of the reference picture is different from that of the current picture, wraparound motion compensation is not used to predict the current picture. If it is determined that the value of "ref_pic_resampling_enabled_flag" is 0, "sps_ref_wraparound_enabled_flag" is 1, and horizontal wraparound motion compensation is applied.

[0197]

[0214] Therefore, step 1703, which is an alternative to step 1603, consists in determining whether the resolution of a reference picture of the sequence of pictures differs from the resolution of the current picture of the sequence of pictures.

[0198]

[0215] Step 1705, which is an alternative to step 1605, determines that wraparound motion compensation is disabled for the current picture in response to a resolution of a reference picture of the sequence of pictures being different from a resolution of a current picture of the sequence of pictures.

[0199]

[0216] In some embodiments, a non-transitory computer-readable storage medium is also provided that includes instructions that can be executed by a device (such as the encoder and decoder of the present disclosure) to perform the above-described methods. Common forms of non-transitory media include, for example, floppy disks, flexible disks, hard disks, solid state drives, magnetic tape or any other magnetic data storage medium, CD-ROMs, any other optical data storage medium, any physical medium with a pattern of holes, RAM, PROMs and EPROMs, FLASH-EPROMs or any other flash memory, NVRAM, caches, registers, any other memory chips or cartridges, and networked versions thereof. A device may include one or more processors (CPUs), input / output interfaces, network interfaces, and / or memory.

[0200]

[0217] The disclosed embodiments may be further described using the following clauses: 1. Receiving a set of pictures; determining the width and height of the pictures in the set; and determining a position of a virtual border for the set of pictures based on said width and height; The method includes: 2. Determining the minimum of the determined width and height of the picture determining a position of a virtual border for the set of pictures based on the width and height further comprises: determining a position of a virtual boundary for the set of pictures based on said minimum value; 2. The method of claim 1, further comprising: 3. Determining the maximum of the determined width and height of the picture determining a position of a virtual border for the set of pictures based on the width and height further comprises: determining a position of a virtual boundary for the set of pictures based on said maximum value; 2. The method of claim 1, further comprising: 4. Determining whether the width of the first picture in the set satisfies a given condition; and Disabling horizontal wraparound motion compensation for the first picture in response to determining that the width of the first picture satisfies the given condition. 2. The method of claim 1, further comprising: 5. Disabling in-loop filtering operations that cross a virtual boundary of a set of pictures 2. The method of claim 1, further comprising: 6. Determining whether resampling has been performed on the first picture in the set; Disabling at least one of virtual boundary or wraparound motion compensation signaling for the first picture in response to determining that resampling for the first picture has not occurred. 2. The method of claim 1, further comprising: 7. The method according to clause 1, wherein the location of the virtual boundary is signaled in at least one of a sequence parameter set (SPS) or a picture header (PH). 8. A memory for storing a set of instructions; and one or more processors, the one or more processors comprising: receiving a set of pictures; determining the width and height of the pictures in the set; and determining a position of a virtual border for the set of pictures based on said width and height; The apparatus is configured to execute a set of instructions to cause the apparatus to perform the 9. One or more processors may: Determining the minimum of the determined width and height of the picture determining a position of a virtual border for a set of pictures based on the width and height; determining a position of a virtual boundary for the set of pictures based on said minimum value; 13. The device according to clause 8, further comprising: 10. One or more processors may: Determining the maximum of the determined width and height of the picture determining a position of a virtual border for a set of pictures based on the width and height; determining a position of a virtual boundary for the set of pictures based on said maximum value; 13. The device according to clause 8, further comprising: 11. One or more processors may: determining whether the width of a first picture in the set satisfies a given condition; and Disabling horizontal wraparound motion compensation for the first picture in response to determining that the width of the first picture satisfies the given condition. 9. The apparatus of claim 8, configured to execute a set of instructions to cause the apparatus to further: 12. One or more processors may: Disabling in-loop filtering operations across virtual boundaries of a set of pictures 9. The apparatus of claim 8, configured to execute a set of instructions to cause the apparatus to further: 13. One or more processors may: determining whether resampling has been performed on a first picture in the set; Disabling at least one of virtual boundary or wraparound motion compensation signaling for the first picture in response to determining that resampling for the first picture has not occurred. 9. The apparatus of claim 8, configured to execute a set of instructions to cause the apparatus to further: 14. The device of clause 8, wherein the location of the virtual boundary is signaled in at least one of a sequence parameter set (SPS) or a picture header (PH). 15. A non-transitory computer-readable medium storing a set of instructions, the set of instructions being executable by at least one processor of a computer to cause the computer to perform a video processing method, the method comprising: receiving a set of pictures; determining the width and height of the pictures in the set; and determining a position of a virtual border for the set of pictures based on said width and height; A non-transitory computer readable medium comprising: 16. A set of instructions is Determining the minimum of the determined width and height of the picture determining a position of a virtual boundary for the set of pictures based on the width and height; determining a position of a virtual boundary for the set of pictures based on said minimum value; 16. The non-transitory computer-readable medium of clause 15, further comprising: 17. A set of instructions is Determining the maximum of the determined width and height of the picture determining a position of a virtual boundary for the set of pictures based on the width and height; determining a position of a virtual boundary for the set of pictures based on said maximum value; 16. The non-transitory computer-readable medium of clause 15, further comprising: 18. A set of instructions is determining whether the width of a first picture in the set satisfies a given condition; and Disabling horizontal wraparound motion compensation for the first picture in response to determining that the width of the first picture satisfies the given condition. 16. The non-transitory computer-readable medium of claim 15, executable by a computer to further cause the computer to perform the steps of: 19. A set of instructions is Disabling in-loop filtering operations across virtual boundaries of a set of pictures 16. The non-transitory computer-readable medium of claim 15, executable by a computer to further cause the computer to perform the steps of: 20. A set of instructions is determining whether resampling has been performed on a first picture in the set; Disabling at least one of virtual boundary or wraparound motion compensation signaling for the first picture in response to determining that resampling for the first picture has not occurred. 16. The non-transitory computer-readable medium of claim 15, executable by a computer to further cause the computer to perform the steps of: 21. The non-transitory computer-readable medium of clause 15, wherein the location of the virtual boundary is signaled in at least one of a sequence parameter set (SPS) or a picture header (PH).

[0201]

[0218] It should be noted that relational terms herein, such as "first" and "second," are used merely to distinguish one entity or operation from another, and do not require or imply any actual relationship or order between those entities or operations. Furthermore, the words "comprise," "have," "contain," and "include," and other similar forms, are intended to be equivalent in meaning and that the element or elements following any of these words are not meant to be a definitive listing of such elements or elements, or to be limited to only the listed element or elements.

[0202]

[0219] As used herein, unless specifically stated otherwise, the term "or" includes all possible combinations unless impracticable. For example, if it is stated that a database can include A or B, then the database can include A or B, or A and B, unless specifically stated otherwise or impracticable. As a second example, if it is stated that a database can include A, B, or C, then the database can include A, B, or C, or A and B, A and C, or B and C, or A and B and C, unless specifically stated otherwise or impracticable.

[0203]

[0220] It is understood that the above-described embodiments can be implemented by hardware, or software (program code), or a combination of hardware and software. If implemented by software, it can be stored in the above-described computer-readable medium. The software, when executed by a processor, can perform the methods of the present disclosure. The computational units and other functional units described in the present disclosure can be implemented by hardware, or software, or a combination of hardware and software. Those skilled in the art will also understand that multiple ones of the above-described modules / units can be combined into one module / unit, and each of the above-described modules / units can be further divided into multiple sub-modules / sub-units.

[0204]

[0221] In the above specification, the embodiments have been described with reference to many specific details that may vary from implementation to implementation. Certain adaptations and modifications of the above-described embodiments may be made. Other embodiments may become apparent to those skilled in the art from consideration of the specification and practice of the invention disclosed herein. It is intended that the specification and examples be considered as exemplary only, with the true scope and spirit of the invention being indicated by the appended claims. It is also intended that the sequences of steps depicted in the figures are for illustrative purposes only, and are not intended to be limited to any particular sequence of steps. Thus, one skilled in the art can appreciate that these steps may be performed in different orders while performing the same method.

[0205]

[0222]

[0033] In the drawings and specification, illustrative embodiments have been disclosed. However, many variations and modifications to these embodiments may be made. Thus, although specific terminology is employed, it is used in a generic and descriptive sense only and not for purposes of limitation.

Claims

1. 1. A video decoding method, comprising: receiving a bitstream associated with a set of pictures; determining, according to the received bitstream, whether a resolution of a first picture in the set of pictures differs from a resolution of a reference picture associated with the first picture; determining, in response to the resolution of the first picture differing from the resolution of the reference picture associated with the first picture, that wraparound motion compensation is disabled with respect to the first picture; determining whether a virtual boundary is signaled at a sequence level for the set of pictures according to the received bitstream; and Controlling an in-loop filtering operation based on whether the virtual boundary is signaled at the sequence level. Including, The controlling of the in-loop filtering operation comprises: determining a location of the virtual boundary with respect to the set of pictures in response to the virtual boundary being signaled at the sequence level, the location being constrained by a range signaled in the received bitstream; and Disabling in-loop filtering operations that cross the virtual boundary. Including, the range to which the position is limited includes at least one of a vertical range or a horizontal range; the vertical extent is less than or equal to a first value associated with a maximum width of each picture of the set of pictures, the maximum width being signaled in the received bitstream; and the horizontal extent is less than or equal to a second value related to a maximum height of each picture of the set of pictures, the maximum height being signaled in the received bitstream. method.

2. the first value is equal to Ceil(pic_width_max_in_luma_samples÷8)−1, the second value is equal to Ceil(pic_height_max_in_luma_samples÷8)−1, The method of claim 1 , wherein pic_width_max_in_luma_samples represents the maximum width in luma samples units of each picture in the set of pictures, and pic_height_max_in_luma_samples represents the maximum height in luma samples units of each picture in the set of pictures.

3. Determining whether the virtual boundary is signaled at the sequence level according to the received bitstream comprises: determining whether changing the resolution of the first picture is allowed according to the received bitstream; and determining, in response to the resolution change of the first picture being permitted, that the virtual boundary is not signaled at the sequence level; The method of claim 1 , comprising:

4. determining, in response to the resolution change of the first picture being permitted, that wraparound motion compensation is disabled for the set of pictures; The method of claim 3 further comprising:

5. A method for storing a bitstream of a set of pictures, the method comprising: setting a virtual boundary for said set of pictures; disabling in-loop filtering operations that cross the virtual boundary; determining whether a resolution of a first picture in the set of pictures differs from a resolution of a reference picture associated with the first picture; disabling wraparound motion compensation with respect to the first picture in response to the resolution of the first picture being different from the resolution of the reference picture associated with the first picture; generating a bitstream including coded information signaling a maximum value indicative of a range within which the location of the virtual boundary is constrained; and storing said bitstream on a non-transitory computer readable medium; Including, the range comprises at least one of a vertical range or a horizontal range, and the maximum value comprises at least one of a maximum width of each picture of the set of pictures and a maximum height of each picture of the set of pictures; the vertical extent is less than or equal to a first value related to the maximum width; the horizontal extent is less than or equal to a second value related to the maximum height; method.

6. the first value is equal to Ceil(pic_width_max_in_luma_samples÷8)−1, the second value is equal to Ceil(pic_height_max_in_luma_samples÷8)−1, The method of claim 5 , wherein pic_width_max_in_luma_samples represents the maximum width in luma samples units of each picture in the set of pictures, and pic_height_max_in_luma_samples represents the maximum height in luma samples units of each picture in the set of pictures.

7. The method comprises: determining a flag indicating whether changing the resolution of the first picture is permitted; If the change in resolution of the first picture is allowed, the coded information is without signaling the virtual boundary at a sequence level of the set of pictures; If changing the resolution of the first picture is not allowed, the coded information is: The method of claim 5 , further comprising signaling the virtual boundary at the sequence level.

8. The method comprises: disabling the wraparound motion compensation for the set of pictures in response to the resolution change of the first picture being permitted. The method of claim 7 further comprising:

9. The method of claim 7, wherein if a change in the resolution of the first picture is allowed, the encoded information signals the virtual boundary at a picture level for one or more pictures in the set.

10. 1. A video encoding method, comprising: setting a virtual boundary for the set of pictures; Disabling in-loop filtering operations that cross the virtual boundary. determining whether a resolution of a first picture in the set of pictures differs from a resolution of a reference picture associated with the first picture; disabling wraparound motion compensation with respect to the first picture in response to the resolution of the first picture being different from the resolution of the reference picture associated with the first picture; signalling in the bitstream a maximum value indicating a range limiting the location of said virtual boundary; Including, the range comprises at least one of a vertical range or a horizontal range, and the maximum value comprises at least one of a maximum width of each picture of the set of pictures or a maximum height of each picture of the set of pictures; the vertical extent is less than or equal to a first value related to the maximum width; the horizontal extent is less than or equal to a second value related to the maximum height; method.

11. the first value is equal to Ceil(pic_width_max_in_luma_samples÷8)−1, the second value is equal to Ceil(pic_height_max_in_luma_samples÷8)−1, The method of claim 10 , wherein pic_width_max_in_luma_samples represents the maximum width in luma samples units of each picture in the set of pictures, and pic_height_max_in_luma_samples represents the maximum height in luma samples units of each picture in the set of pictures.

12. determining whether changing the resolution of the first picture is permitted; signaling the virtual boundary at a sequence level of the set of pictures in response to the resolution change of the first picture being not permitted. skipping signaling the virtual boundary at the sequence level in response to the resolution change of the first picture being permitted. The method of claim 10 further comprising:

13. disabling wraparound motion compensation for the set of pictures in response to the resolution change of the first picture being permitted. The method of claim 12 further comprising:

Citation Information

Patent Citations

  • Method for encoding 360-degree panoramic video, encoding device, and computer program

    JP2018534827A

  • 360-degree video encoding using geometry projection

    JP2019525563A

  • Video encoding method and apparatus with in-loop filtering process not applied to reconstructed blocks located at image content discontinuity edge and associated video decoding method and apparatus

    US20180054613A1

  • Handling face discontinuities in 360-degree video coding

    WO2019060443A1