How to perform wrap-around motion compensation
Patent Information
- Application Number
- JP2025108943
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Priority Date
- 2019-12-17
- Filing Date
- 2025-06-27
- Publication Date
- 2026-08-21
- Estimated Expiration
- 2040-12-17
Smart Images

Figure 0007909665000002 
Figure 0007909665000003 
Figure 0007909665000004
Abstract
Description
[Technical Field]
[0001] Cross-reference of related applications
[0001] This disclosure claims priority and the benefit of priority to U.S. Provisional Patent Application No. 62 / 949,396, filed on 17 December 2019. The Provisional Application is incorporated herein by reference in its entirety.
[0002] Technical field
[0002] This disclosure relates, in general terms, to image processing, and more particularly to methods and systems for performing wrap-around motion compensation. [Background technology]
[0003] background
[0003] Video is a set of static pictures (or "frames") that capture visual information. To reduce memory and transmission bandwidth, video can be compressed before storage or transmission and restored before display. The compression process is usually called encoding, and the restoration process is usually called decoding. Most commonly, there are various video encoding formats that use standardized video encoding techniques based on prediction, transformation, quantization, entropy coding, and in-loop filtering. Video encoding standards, such as High Efficiency Video Coding (e.g., HEVC / H.265), Versatile Video Coding (e.g., VVC / H.266), and the standard AVS, which specify a particular video encoding format, are developed by standardization organizations. As more advanced video encoding techniques are adopted into video standards, the encoding efficiency of new video encoding standards increases. [Overview of the project] [Means for solving the problem]
[0004] Summary of Disclosure
[0004] Embodiments of the present disclosure provide a method for performing motion compensation. The method includes receiving a first wrap-around motion compensation flag, which is associated with a picture; determining whether the first wrap-around motion compensation flag is valid; receiving a wrap-around motion compensation offset, which is associated with a picture, in response to the determination that the first wrap-around motion compensation flag is valid; and performing wrap-around motion compensation on the picture in accordance with the first wrap-around motion compensation flag and the wrap-around motion compensation offset.
[0005]
[0005] Embodiments of the present disclosure further provide a system for performing motion compensation. The system comprises a memory for storing a set of instructions and a processor, the processor configured to execute a set of instructions such that it causes the system to: receive a first wrap-around motion compensation flag, associate the first wrap-around motion compensation flag with a picture; determine whether the first wrap-around motion compensation flag is valid; receive a wrap-around motion compensation offset, in response to the determination that the first wrap-around motion compensation flag is valid, associate the wrap-around motion compensation offset with a picture; and perform wrap-around motion compensation on the picture in accordance with the first wrap-around motion compensation flag and the wrap-around motion compensation offset.
[0006]
[0006] Embodiments of the present disclosure are non - transitory computer - readable media storing a set of instructions, the set of instructions being executable by one or more processors of a device to initiate a method for performing motion compensation on the device, the method comprising receiving a first wraparound motion compensation flag, wherein the first wraparound motion compensation flag is associated with a picture; determining whether the first wraparound motion compensation flag is valid; in response to determining that the first wraparound motion compensation flag is valid, receiving a wraparound motion compensation offset, wherein the wraparound motion compensation offset is associated with the picture; and performing wraparound motion compensation on the picture according to the first wraparound motion compensation flag and the wraparound motion compensation offset. The present disclosure further provides a non - transitory computer - readable media including the above.
[0007] Brief Description of the Drawings
[0007] Embodiments and various aspects of the present disclosure are illustrated in the following detailed description and the accompanying drawings. The various features shown in the drawings are not drawn to scale.
Brief Description of the Drawings
[0008] [Figure 1]
[0008] Shown is the structure of an exemplary video sequence according to some embodiments of the present disclosure. [Figure 2A]
[0009] Shown is a schematic diagram of an exemplary encoding process according to some embodiments of the present disclosure. [Figure 2B]
[0010] Shown is a schematic diagram of another exemplary encoding process according to some embodiments of the present disclosure. [Figure 3A]
[0011] Shown is a schematic diagram of an exemplary decoding process according to some embodiments of the present disclosure. [Figure 3B]
[0012] Shown is a schematic diagram of another exemplary decoding process according to some embodiments of the present disclosure. [Figure 4]
[0013] A block diagram of an exemplary device for encoding or decoding video, according to some embodiments of the present disclosure, is shown. [Figure 5A]
[0014] A schematic diagram of an exemplary blend operation for generating a reconstructed equirectangular projection, according to some embodiments of this disclosure, is shown. [Figure 5B]
[0015] A schematic diagram of exemplary cropping operations for generating a reconstructed equirectangular projection, according to some embodiments of the present disclosure, is shown. [Figure 6A]
[0016] A schematic diagram of an exemplary horizontal wrap-around motion compensation process for equirectangular projection, according to some embodiments of this disclosure, is shown. [Figure 6B]
[0017] A schematic diagram of an exemplary horizontal wraparound motion compensation process for padded equirectangular projection, according to some embodiments of the present disclosure, is shown. [Figure 7]
[0018] The syntax of an exemplary sequence parameter set for wrap-around motion compensation according to some embodiments of this disclosure is shown. [Figure 8]
[0019] The semantics of exemplary sequence parameter sets for wrap-around motion compensation according to some embodiments of this disclosure are shown below. [Figure 9]
[0020] The syntax of exemplary sequence parameter sets for improved wrap-around motion compensation, according to some embodiments of this disclosure, is shown below. [Figure 10]
[0021] The meaning of exemplary sequence parameter sets for improved wrap-around motion compensation, according to some embodiments of this disclosure, is shown. [Figure 11]
[0022] The meaning of exemplary sequence parameter sets for improved wrap-around motion compensation using maximum picture width, according to some embodiments of this disclosure, is shown. [Figure 12]
[0023] Examples of derivation of the variables "PicRefWraparoundEnableFlag" and "PicRefWraparoundOffset" according to some embodiments of this disclosure are shown below. [Figure 13]
[0024] Examples of deriving sample positions used for motion compensation according to some embodiments of this disclosure are shown. [Figure 14]
[0025] The syntax of exemplary sequence and picture parameter sets for wrap-around motion compensation using a wrap-around motion compensation offset in a picture parameter set, according to some embodiments of this disclosure, is shown. [Figure 15]
[0026] The meaning of exemplary sequence and picture parameter sets for wrap-around motion compensation using wrap-around motion compensation offsets in picture parameter sets, according to some embodiments of this disclosure, is shown. [Figure 16]
[0027] The syntax of exemplary sequence parameter sets for improved wrap-around motion compensation without wrap-around motion compensation offset, according to some embodiments of this disclosure, is shown below. [Figure 17]
[0028] The syntax of an exemplary picture parameter set for improved wrap-around motion compensation using a wrap-around motion compensation offset, according to some embodiments of this disclosure, is shown. [Figure 18]
[0029] The meaning of exemplary sequence and picture parameter sets for improved wrap-around motion compensation using a wrap-around motion compensation offset in a picture parameter set, according to some embodiments of this disclosure, is shown. [Figure 19]
[0030] Examples of derivation of the variables "PicRefWraparoundEnableFlag" and "PicRefWraparoundOffset" according to some embodiments of this disclosure are shown below. [Figure 20]
[0031] The syntax of an exemplary sequence parameter set for improved wrap-around motion compensation without a wrap-around motion compensation offset in the sequence parameter set, according to some embodiments of the present disclosure, is shown. [Figure 21]
[0032] The syntax of an exemplary picture parameter set for improved wrap-around motion compensation using a wrap-around motion control flag, according to some embodiments of this disclosure, is shown. [Figure 22]
[0033] The meaning of exemplary sequence and picture parameter sets for improved wrap-around motion compensation using wrap-around control flags in picture parameter sets, according to some embodiments of this disclosure, is shown. [Figure 23]
[0034] The meaning of exemplary sequence and picture parameter sets for improved wrap-around motion compensation using wrap-around control flags in picture parameter sets, according to some embodiments of this disclosure, is shown. [Figure 24]
[0035] The following are examples of the derivation of the variable "PicRefWraparoundOffset" according to some embodiments of this disclosure. [Figure 25]
[0036] The meaning of exemplary sequence parameter sets and picture parameter sets for improved wrap-around motion compensation using limitations on picture size, according to some embodiments of this disclosure, is shown. [Figure 26]
[0037] The meaning of exemplary sequence parameter sets for improved wrap-around motion compensation, using constraints imposed on the variables "pic_width_max_in_luma_samples", "CtbSizeY", and "MinCbSizeY", according to some embodiments of this disclosure, is shown. [Figure 27]
[0038] The meaning of exemplary picture parameter sets for improved wrap-around motion compensation using the constraints imposed on the variable "pic_width_in_luma_samples" in some embodiments of this disclosure is shown. [Figure 28]
[0039] A flowchart illustrating an exemplary method for performing motion compensation according to some embodiments of this disclosure is shown. [Figure 29]
[0040] A flowchart illustrates an exemplary method for performing motion compensation using a limited range for sequence wrap-around motion compensation offsets, according to some embodiments of the present disclosure. [Figure 30]
[0041] A flowchart illustrates an exemplary method for performing motion compensation using a picture associated with a sequence wrap-around motion compensation offset, according to some embodiments of the present disclosure. [Figure 31]
[0042] A flowchart illustrating an exemplary method for performing motion compensation using a limited maximum picture size, according to some embodiments of this disclosure, is shown. [Modes for carrying out the invention]
[0009] Detailed explanation
[0043] Next, we will refer in detail to exemplary embodiments illustrated in the accompanying drawings. The following description refers to the accompanying drawings, where the same reference numerals in different drawings represent the same or similar elements unless otherwise indicated. The implementations shown in the following description of exemplary embodiments do not represent all implementations according to this disclosure. Rather, they are merely examples of apparatus and methods according to the aspects relating to this disclosure as enumerated in the accompanying claims. Specific aspects of this disclosure are described in more detail below. In the event of any conflict between terms and / or definitions incorporated by reference and those provided herein, the terms and definitions provided herein shall prevail.
[0010]
[0044] The Joint Video Experts Team (JVET) of the ITU-T Video Coding Expert Group (ITU-T VCEG) and the ISO / IEC Moving Picture Expert Group (ISO / IEC MPEG) is currently developing the Multipurpose Video Coding (VVC / H.266) standard. The VVC standard aims to double the compression efficiency of its predecessor, the High Efficiency Video Coding (HEVC / H.265) standard. In other words, the goal of VVC is to achieve the same subjective quality as HEVC / H.265 using half the bandwidth.
[0011]
[0045] To achieve the same subjective quality as HEVC / H.265 using half the bandwidth, the Joint Video Experts Team (JVET) is developing a technology that surpasses HEVC using the Joint Exploration Model (JEM) reference software. Because the encoding technology has been incorporated into JEM, JEM has achieved substantially higher encoding performance than HEVC. VCEG and MPEG have also officially begun development of next-generation video compression standards that surpass HEVC.
[0012]
[0046] The VVC standard is a relatively recent development and continues to incorporate more encoding techniques to deliver superior compression performance. VVC is based on the same hybrid video encoding system used in modern video compression standards such as HEVC, H.264 / AVC, MPEG2, and H.263.
[0013]
[0047] Video is a set of static pictures (or "frames") arranged in chronological order to store visual information. A video capture device (e.g., a camera) can be used to capture and store these pictures in chronological order, and a video playback device (e.g., a television, computer, smartphone, tablet computer, video player, or any end-user terminal with display capabilities) can be used to display such pictures in chronological order. Depending on the application, the video capture device can also transmit the captured video in real time to a video playback device (e.g., a computer with a monitor) for purposes such as supervision, conference management, or live broadcasting.
[0014]
[0048] To reduce the memory space and transmission bandwidth required for such applications, video can be compressed. For example, video can be compressed before storage and transmission, and decompressed before display. Compression and decompression can be performed by software executed by a processor (e.g., a general-purpose computer processor) or by specialized hardware. The module or circuit configuration for compression is generally called an “encoder,” and the module or circuit configuration for decompression is generally called a “decoder.” Encoders and decoders can be collectively called a “codec.” Encoders and decoders can be implemented as any of a variety of suitable hardware, software, or combinations thereof. For example, hardware implementations of encoders and decoders can include circuit mechanisms such as one or more microprocessors, digital signal processors ("DSPs"), application-specific integrated circuits ("ASICs"), field-programmable gate arrays ("FPGAs"), discrete logic, or any combination thereof. Software implementations of encoders and decoders may include program code, computer-executable instructions, firmware, or any suitable computer-implementable algorithm or process fixed within a computer-readable medium. Video compression and decompression can be performed using various algorithms or standards, such as MPEG-1, MPEG-2, MPEG-4, H.26x series, or similar. Depending on the application, a codec may decompress video from a first encoding standard and recompress the decompressed video using a second encoding standard. In this case, the codec may be referred to as a "transcoder."
[0015]
[0049] A video encoding process can identify and maintain useful information that can be used to reconstruct a picture. If information ignored in the video encoding process cannot be fully reconstructed, the encoding process may be called "lossy." Otherwise, it may be called "lossy." Most encoding processes are lossy, which is a trade-off to reduce the required memory space and transmission bandwidth.
[0016]
[0050] In many cases, useful information about the picture being encoded (referred to as the "current picture") includes changes relative to the reference picture (e.g., a previously encoded or reconstructed picture). Such changes may include changes in pixel position, brightness, or color. Changes in the position of a group of pixels representing an object can reflect the movement of the object between the reference picture and the current picture.
[0017]
[0051] A picture encoded without referencing another picture (i.e., it is its own reference picture) is called an "I picture". A picture encoded using a previous picture as its reference picture is called a "P picture". A picture encoded using both a previous picture and a future picture as reference pictures (i.e., the reference is "bidirectional") is called a "B picture".
[0018]
[0052] Figure 1 shows the structure of an exemplary video sequence 100 according to some embodiments of the present disclosure. As shown in Figure 1, the video sequence 100 can be live video or captured and archived video. The video 100 can be real video, computer-generated video (e.g., computer game video), or a combination thereof (e.g., real video with augmented reality effects). The video sequence 100 may be input from a video capture device (e.g., a camera), a video archive containing previously captured video (e.g., video files stored in a storage device), or a video supply interface for receiving video from a video content provider (e.g., a video broadcast transceiver).
[0019]
[0053] As shown in Figure 1, the video sequence 100 may include a series of pictures arranged in time along a timeline, including pictures 102, 104, 106, and 108. Pictures 102-106 are consecutive, with further pictures between pictures 106 and 108. In Figure 1, picture 102 is an I picture, and its reference picture is picture 102 itself. Picture 104 is a P picture, and its reference picture is picture 102, as indicated by the arrows. Picture 106 is a B picture, and its reference pictures are pictures 104 and 108, as indicated by the arrows. Depending on the embodiment, the reference picture of a picture (e.g., picture 104) may not be immediately before or after that picture. For example, the reference picture of picture 104 may be the picture before picture 102. Please note that the reference pictures 102-106 are merely examples, and this disclosure does not limit the embodiments of the reference pictures to the examples shown in Figure 1.
[0020]
[0054] Typically, video codecs do not encode or decode an entire picture at once due to the computational complexity of such tasks. Rather, they can divide the picture into basic segments and encode or decode the picture segment by segment. Such basic segments are referred to in this disclosure as basic processing units ("BPUs"). For example, structure 110 in Figure 1 shows an exemplary structure of a picture (e.g., any of pictures 102-108) in video sequence 100. In structure 110, the picture is divided into 4x4 basic processing units, their boundaries shown as dashed lines. Depending on the embodiment, basic processing units may be referred to as "macroblocks" in some video encoding standards (e.g., the MPEG family, H.261, H.263, or H.264 / AVC) or as "coding tree units" ("CTUs") in some other video encoding standards (e.g., H.265 / HEVC or H.266 / VVC). The basic processing unit can have a variable size in the picture, or any arbitrary shape and size of pixels, such as 128×128, 64×64, 32×32, 16×16, 4×8, or 16×32. The size and shape of the basic processing unit can be selected for the picture based on a balance between encoding efficiency and the level of detail that should be maintained in the basic processing unit.
[0021]
[0055] A basic processing unit can be a logical unit that can contain groups of different types of video data stored in computer memory (for example, in a video frame buffer). For example, a basic processing unit for a color picture may include a luminance component (Y) representing colorless luminance information, one or more chroma components (e.g., Cb and Cr) representing color information, and associated syntactic elements, where the luminance and chroma components may have the same size as the basic processing unit. The luminance and chroma components may be referred to as a "coding tree block" (CTB) in some video encoding standards (e.g., H.265 / HEVC or H.266 / VVC). Any operation performed on a basic processing unit may be repeated on each of its luminance and chroma components.
[0022]
[0056] Video encoding has multiple computational stages, examples of which are shown in Figures 2A-2B and 3A-3B. At each stage, the size of the basic processing unit may still be too large for processing and can therefore be further divided into segments referred to in this disclosure as “basic processing subunits.” Depending on the embodiment, a basic processing subunit may be referred to as a “block” in some video encoding standards (e.g., the MPEG family, H.261, H.263, or H.264 / AVC) or as an “encoding unit” (“CU”) in some other video encoding standards (e.g., H.265 / HEVC or H.266 / VVC). A basic processing subunit may be the same size as or smaller than a basic processing unit. Like a basic processing unit, a basic processing subunit is also a logical unit that can contain groups of different types of video data (e.g., Y, Cb, Cr, and associated syntactic elements) stored in computer memory (e.g., in a video frame buffer). Any operation performed on a basic processing subunit can be repeated on each of its luma and chroma components. Note that such divisions can be carried out to further levels as needed for processing. Also note that different stages can divide the basic processing unit using different methods.
[0023]
[0057] For example, during the mode determination phase (an example of which is shown in Figure 2B), the encoder can determine which prediction mode (e.g., intrapicture prediction or interpicture prediction) to use for the basic processing unit, but the basic processing unit may be too large to make such a decision. The encoder can divide the basic processing unit into multiple basic processing subunits (e.g., CUs, as in the case of H.265 / HEVC or H.266 / VVC) and determine the type of prediction for each basic processing subunit.
[0024]
[0058] As another example, in the prediction phase (an example of which is shown in Figures 2A and 2B), the encoder can perform prediction calculations at the level of the basic processing subunit (e.g., CU). However, in some cases, the basic processing subunit may still be too large to process. The encoder can further divide the basic processing subunit into smaller segments (e.g., referred to as "prediction blocks" or "PBs" in H.265 / HEVC or H.266 / VVC), at which level prediction calculations can be performed.
[0025]
[0059] As another example, in the transformation stage (an example of which is shown in Figures 2A and 2B), the encoder may perform transformation operations for residual basic processing subunits (e.g., CUs). However, in some cases, the basic processing subunit may still be too large to process. The encoder may further divide the basic processing subunit into smaller segments (e.g., referred to as "transform blocks" or "TBs" in H.265 / HEVC or H.266 / VVC), and transformation operations may be performed at that level. Note that the division method for the same basic processing subunit may differ in the prediction and transformation stages. For example, in H.265 / HEVC or H.266 / VVC, the prediction blocks and transform blocks of the same CU may have different sizes and numbers.
[0026]
[0060] In the structure 110 of Figure 1, the basic processing unit 112 is further divided into 3x3 basic processing subunits, and their boundaries are shown as dotted lines. Different basic processing units of the same picture may be divided into basic processing subunits in different ways.
[0027]
[0061] Depending on the implementation, a picture can be divided into processing regions to bring parallel processing and error tolerance to video encoding and decoding. This means that the encoding or decoding process does not have to rely on information from any other regions of the picture with respect to a given region. In other words, each region of a picture can be processed independently. By doing so, the codec can process different regions of the picture in parallel, thus increasing encoding efficiency. Furthermore, if the data in a region is corrupted during processing or lost during network transmission, the codec can correctly encode or decode other regions of the same picture without relying on the corrupted or lost data, thus providing error tolerance. Some video encoding standards allow a picture to be divided into different types of regions. For example, H.265 / HEVC and H.266 / VVC offer two types of regions: "slices" and "tiles". It should also be noted that different pictures in video sequence 100 may have different division schemes for dividing the picture into regions.
[0028]
[0062] For example, in Figure 1, structure 110 is divided into three regions 114, 116, and 118, their boundaries shown as solid lines within structure 110. Region 114 contains four basic processing units. Regions 116 and 118 each contain six basic processing units. Note that the basic processing units, basic processing subunits, and regions of structure 110 in Figure 1 are merely examples, and this disclosure does not limit its embodiments.
[0029]
[0063] Figure 2A shows a schematic diagram of an exemplary encoding process 200A according to some embodiments of the present disclosure. For example, the encoding process 200A shown in Figure 2A may be performed by an encoder. As shown in Figure 2A, the encoder can encode a video sequence 202 into a video bitstream 228 according to process 200A. Similar to video sequence 100 in Figure 1, video sequence 202 may include a set of pictures arranged in chronological order (referred to as “original pictures”). Similar to structure 110 in Figure 1, each original picture in video sequence 202 may be divided by the encoder into a basic processing unit, basic processing subunit, or region for processing. In some embodiments, the encoder may perform process 200A at the level of a basic processing unit for each original picture in video sequence 202. For example, the encoder may perform process 200A in an iterative manner, in which case the encoder may encode a basic processing unit in a single iteration of process 200A. Depending on the embodiment, the encoder can perform process 200A in parallel for each region of the original picture in the video sequence 202 (e.g., regions 114-118).
[0030]
[0064] In Figure 2A, the encoder can supply the basic processing unit (referred to as the "original BPU") of the original picture of the video sequence 202 to the prediction stage 204, generating prediction data 206 and prediction BPU 208. The encoder can subtract the prediction BPU 208 from the original BPU to generate residual BPU 210. The encoder can supply the residual BPU 210 to the conversion stage 212 and the quantization stage 214, generating quantization conversion coefficients 216. The encoder can supply the prediction data 206 and quantization conversion coefficients 216 to the binary coding stage 226, generating video bitstream 228. Components 202, 204, 206, 208, 210, 212, 214, 216, 226, and 228 may be referred to as the "forward path". During process 200A, after the quantization stage 214, the encoder may supply the quantization conversion coefficients 216 to the inverse quantization stage 218 and the inverse conversion stage 220 to generate the reconstructed residual BPU 222. The encoder may add the reconstructed residual BPU 222 to the prediction BPU 208 to generate a prediction criterion 224, which will be used in the prediction stage 204 for the next iteration of process 200A. Components 218, 220, 222, and 224 of process 200A may be referred to as the “reconstruction path”. The reconstruction path may be used to ensure that both the encoder and the decoder use the same reference data for prediction.
[0031]
[0065] The encoder can iteratively perform process 200A to encode each original BPU of the original picture (in the forward path) and to generate a prediction criterion 224 (in the reconstruction path) for encoding the next original BPU of the original picture. After encoding all the original BPUs of the original picture, the encoder can proceed to encode the next picture in the video sequence 202.
[0032]
[0066] Referring to process 200A, the encoder can receive a video sequence 202 generated by a video acquisition device (e.g., a camera). As used herein, the term “receive” can mean receiving, inputting, acquiring, obtaining, getting, reading, accessing, or any act in any way for inputting data.
[0033]
[0067] In prediction stage 204, in the current iteration, the encoder receives the original BPU and prediction criterion 224, performs prediction calculations, and can generate prediction data 206 and prediction BPU 208. The prediction criterion 224 may be generated from the reconstruction path of a previous iteration of process 200A. The objective of prediction stage 204 is to reduce information redundancy by extracting prediction data 206, which can be used to reconstruct the original BPU as prediction BPU 208 from the prediction data 206 and prediction criterion 224.
[0034]
[0068] Ideally, the predicted BPU 208 can be identical to the original BPU. However, due to non-ideal prediction and reconstruction operations, the predicted BPU 208 is generally slightly different from the original BPU. To record such differences, after generating the predicted BPU 208, the encoder can subtract it from the original BPU to generate the residual BPU 210. For example, the encoder can subtract the pixel values (e.g., grayscale values or RGB values) of the predicted BPU 208 from the corresponding pixel values of the original BPU. Each pixel of the residual BPU 210 may have a residual value resulting from such a subtraction between the original BPU and the corresponding pixel of the predicted BPU 208. Compared to the original BPU, the predicted data 206 and residual BPU 210 may have fewer bits, but they can be used to reconstruct the original BPU without significant quality degradation. Therefore, the original BPU is compressed.
[0035]
[0069] To further compress the residual BPU210, in the transformation step 212, the encoder can reduce the spatial redundancy of the residual BPU210 by decomposing it into a set of two-dimensional "basis patterns," each basis pattern associated with a "transformation coefficient." The basis patterns can have the same size (e.g., the size of the residual BPU210). Each basis pattern can represent the changing frequency components of the residual BPU210 (e.g., the frequency of brightness changes). No basis pattern can be reconstructed from any combination (e.g., a linear combination) of any other basis pattern. In other words, the decomposition can decompose the changes in the residual BPU210 into the frequency domain. Such a decomposition is analogous to the discrete Fourier transform of a function, in which case the basis patterns are analogous to the basis functions of the discrete Fourier transform (e.g., trigonometric functions), and the transformation coefficients are analogous to the coefficients associated with the basis functions.
[0036]
[0070] Different transformation algorithms can use different basis patterns. For example, various transformation algorithms such as discrete cosine transform, discrete sine transform, or similar can be used in transformation stage 212. The transformation in transformation stage 212 is inversely operable. That is, the encoder can recover the residual BPU 210 by the inverse operation of the transformation (referred to as "inverse transform"). For example, to recover the pixels of the residual BPU 210, the inverse transform can be generated by multiplying the values of the corresponding pixels in the basis pattern by their respective associated coefficients, adding the products, and so on. For video coding standards, both the encoder and decoder can use the same transformation algorithm (and therefore the same basis pattern). Therefore, the encoder can record only the transformation coefficients, and the decoder can reconstruct the residual BPU 210 from the transformation coefficients without receiving the basis pattern from the encoder. Compared to the residual BPU 210, the transformation coefficients may have fewer bits, but they can be used to reconstruct the residual BPU 210 without significant quality degradation. Therefore, the residual BPU210 is further compressed.
[0037]
[0071] The encoder can further compress the conversion coefficients in the quantization stage 214. In the conversion process, different basis patterns can represent different change frequencies (e.g., brightness change frequencies). Since the human eye is generally better at recognizing low-frequency changes, the encoder can ignore information about high-frequency changes without causing significant quality degradation in decoding. For example, in the quantization stage 214, the encoder can generate quantization conversion coefficients 216 by dividing each conversion coefficient by an integer value (referred to as a "quantization parameter") and rounding the quotient to its nearest integer. After such an operation, some conversion coefficients of high-frequency basis patterns may be converted to 0, and conversion coefficients of low-frequency basis patterns may be converted to smaller integers. The encoder can ignore quantization conversion coefficients 216 that are 0, thereby further compressing the conversion coefficients. The quantization process is also inversely operable, in which case the quantization conversion coefficients 216 can be reconstructed into conversion coefficients in the inverse operation of quantization (referred to as "inverse quantization").
[0038]
[0072] Because the encoder rounds off the remainder of such divisions, the quantization stage 214 can be irreversible. Typically, the quantization stage 214 can contribute the greatest information loss in process 200A. The greater the information loss, the fewer bits the quantization conversion coefficient 216 may require. To obtain different levels of information loss, the encoder can use different values for the quantization parameters or any other parameters of the quantization process.
[0039]
[0073] In the binary coding stage 226, the encoder may encode the predicted data 206 and the quantization conversion coefficients 216 using a binary coding technique such as entropy coding, variable-length coding, arithmetic coding, Huffman coding, context-adaptive binary arithmetic coding, or any other lossless or lossy compression algorithm. In some embodiments, in addition to the predicted data 206 and the quantization conversion coefficients 216, the encoder may encode other information in the binary coding stage 226, such as the prediction mode used in the prediction stage 204, the parameters of the prediction operation, the type of conversion in the conversion stage 212, the parameters of the quantization process (e.g., quantization parameters), the encoder control parameters (e.g., bitrate control parameters), or the like. The encoder may generate a video bitstream 228 using the output data from the binary coding stage 226. In some embodiments, the video bitstream 228 may be further packetized for network transmission.
[0040]
[0074] Referring to the reconstruction path of process 200A, in the inverse quantization step 218, the encoder can perform inverse quantization on the quantization transformation coefficients 216 to generate reconstruction transformation coefficients. In the inverse transformation step 220, the encoder can generate reconstruction residual BPU 222 based on the reconstruction transformation coefficients. The encoder can add the reconstruction residual BPU 222 to the prediction BPU 208 to generate a prediction criterion 224, which will be used in the next iteration of process 200A.
[0041]
[0075] It should be noted that other variations of process 200A may also be used to encode the video sequence 202. In some embodiments, the steps of process 200A may be performed in a different order by the encoder. In some embodiments, one or more steps of process 200A may be combined into a single step. In some embodiments, a single step of process 200A may be divided into multiple steps. For example, the conversion step 212 and the quantization step 214 may be combined into a single step. In some embodiments, process 200A may include additional steps. In some embodiments, process 200A may omit one or more steps in Figure 2A.
[0042]
[0076] Figure 2B shows a schematic diagram of another exemplary encoding process 200B according to several embodiments of the present disclosure. As shown in Figure 2B, process 200B may be modified from process 200A. For example, process 200B may be used with an encoder compliant with a hybrid video encoding standard (e.g., the H.26x series). Compared to process 200A, the forward path of process 200B additionally includes a mode determination stage 230 and divides the prediction stage 204 into a spatial prediction stage 2042 and a temporal prediction stage 2044. The reconstruction path of process 200B additionally includes a loop filter stage 232 and a buffer 234.
[0043]
[0077] Generally, prediction techniques can be classified into two types: spatial prediction and temporal prediction. Spatial prediction (e.g., intra-picture prediction or "intra-prediction") can use pixels from one or more already encoded neighboring BPUs within the same picture to predict the current BPU. That is, the prediction criterion 224 in spatial prediction can include neighboring BPUs. Spatial prediction can reduce the inherent spatial redundancy of a picture. Temporal prediction (e.g., inter-picture prediction or "inter-prediction") can use regions from one or more already encoded pictures to predict the current BPU. That is, the prediction criterion 224 in temporal prediction can include encoded pictures. Temporal prediction can reduce the inherent temporal redundancy of a picture.
[0044]
[0078] Referring to process 200B, within the forward path, the encoder performs prediction calculations in spatial prediction stage 2042 and temporal prediction stage 2044. For example, in spatial prediction stage 2042, the encoder may perform intra-prediction. For the original BPU of the picture being encoded, the prediction criterion 224 may include one or more neighboring BPUs encoded (within the forward path) and reconstructed (within the reconstruction path) within the same picture. The encoder may generate a prediction BPU 208 by extrapolating neighboring BPUs. The extrapolation technique may include, for example, linear extrapolation or interpolation, polynomial extrapolation or interpolation, or similar. In some embodiments, the encoder may perform extrapolation at the pixel level, for example, by extrapolating the values of the corresponding pixels for each pixel of the prediction BPU 208. The adjacent BPUs used for extrapolation can be positioned relative to the original BPU from various directions, such as vertically (e.g., above the original BPU), horizontally (e.g., to the left of the original BPU), diagonally (e.g., to the lower left, lower right, upper left, or upper right of the original BPU), or any direction defined in the video encoding standard used. For intra-prediction, the prediction data 206 may include, for example, the location (e.g., coordinates) of the adjacent BPUs used, the size of the adjacent BPUs used, the extrapolation parameters, the orientation of the adjacent BPUs used relative to the original BPU, or similar.
[0045]
[0079] As another example, in the temporal prediction stage 2044, the encoder can perform interpretation. For the original BPU of the current picture, the prediction criterion 224 may include one or more pictures (referred to as "reference pictures") that have been encoded (in the forward path) and reconstructed (in the reconstruction path). In some embodiments, the reference pictures may be encoded and reconstructed for each BPU. For example, the encoder may add the reconstructed residual BPU 222 to the prediction BPU 208 to generate a reconstructed BPU. When all reconstructed BPUs for the same picture have been generated, the encoder may generate the reconstructed picture as a reference picture. The encoder may perform a "motion estimation" operation to search for matching regions within the range of the reference picture (referred to as the "search window"). The location of the search window in the reference picture may be determined based on the location of the original BPU of the current picture. For example, the search window may have its center in the reference picture at a location with the same coordinates as the original BPU in the current picture and may extend outward over a predetermined distance. When the encoder identifies a region in the search window that is similar to the original BPU (for example, by using a pixel-recursive (pel-recursive) algorithm, a block-matching algorithm, or similar), the encoder can determine such a region to be a matching region. The matching region may have different dimensions from the original BPU (for example, smaller than, equal to, larger than, or of a different shape than the original BPU). Because the reference picture and the current picture are temporally separated in the timeline (for example, as shown in Figure 1), the matching region can be considered to "move" to the location of the original BPU over time. The encoder can record the direction and distance of such movement as a "motion vector". When multiple reference pictures are used (for example, as picture 106 in Figure 1), the encoder can search for a matching region for each reference picture and determine its associated motion vector. In some embodiments, the encoder can assign weights to the pixel values of the matching region for each matching reference picture.
[0046]
[0080] Motion estimation can be used to identify various types of motion, such as translation, rotation, zooming, or similar. For interpretation, the prediction data 206 may include, for example, the location of the matching region (e.g., coordinates), the motion vector associated with the matching region, the number of reference pictures, the weights associated with the reference pictures, or similar.
[0047]
[0081] To generate a predicted BPU 208, the encoder can perform a “motion compensation” operation. Motion compensation can be used to reconstruct the predicted BPU 208 based on prediction data 206 (e.g., motion vectors) and prediction criteria 224. For example, the encoder can move the matching region of a reference picture according to the motion vector, in which case the encoder can predict the original BPU of the current picture. When multiple reference pictures are used (e.g., as picture 106 in Figure 1), the encoder can move the matching region of each reference picture according to its respective motion vector and average the pixel values of the matching region. In some embodiments, if the encoder weights the pixel values of the matching region of each matching reference picture, the encoder can add the weighted sum of the pixel values to the moved matching region.
[0048]
[0082] Depending on the embodiment, interpretation can be unidirectional or bidirectional. Unidirectional interpretation can use one or more reference pictures in the same time direction relative to the current picture. For example, picture 104 in Figure 1 is a unidirectional interpretation picture in which the reference picture (i.e., picture 102) precedes picture 104. Bidirectional interpretation can use one or more reference pictures in both time directions relative to the current picture. For example, picture 106 in Figure 1 is a bidirectional interpretation picture in which the reference pictures (i.e., pictures 104 and 108) are in both time directions relative to picture 104.
[0049]
[0083] Still referring to the forward path of process 200B, after the spatial prediction stage 2042 and the temporal prediction stage 2044, in the mode determination stage 230, the encoder may select a prediction mode (e.g., one of intra-prediction or inter-prediction) for the current iteration of process 200B. For example, the encoder may perform a rate-distortion optimization technique. In this technique, the encoder may select a prediction mode that minimizes the value of a cost function that depends on the bit rate of the candidate prediction mode and the distortion of the reconstructed reference picture under the candidate prediction mode. Depending on the selected prediction mode, the encoder may generate the corresponding prediction BPU 208 and prediction data 206.
[0050]
[0084] If the intra-prediction mode is selected in the forward path within the reconstruction path of process 200B, after generating the prediction criterion 224 (e.g., the current BPU encoded and reconstructed in the current picture), the encoder can directly supply the prediction criterion 224 to the spatial prediction stage 2042 for later use (e.g., for extrapolation of the next BPU of the current picture). If the inter-prediction mode is selected in the forward path, after generating the prediction criterion 224 (e.g., the current picture with all BPUs encoded and reconstructed), the encoder can supply the prediction criterion 224 to the loop filter stage 232, where the encoder can apply a loop filter to the prediction criterion 224 to reduce or eliminate distortions (e.g., blocking artifacts) introduced by inter-prediction. The encoder can apply various loop filtering techniques in the loop filter stage 232, such as deblocking, sample-adaptive offset, adaptive loop filtering, or similar. Loop-filtered reference pictures may be stored in buffer 234 (or “Decoded Picture Buffer”) for later use (e.g., to be used as interprediction reference pictures for future pictures in video sequence 202). The encoder may store one or more reference pictures in buffer 234 for use in the temporal prediction stage 2044. In some embodiments, the encoder may encode loop filter parameters (e.g., loop filter strength) together with quantization transformation coefficients 216, prediction data 206, and other information in the binary coding stage 226.
[0051]
[0085] Figure 3A shows a schematic diagram of an exemplary decoding process 300A according to several embodiments of the present disclosure. As shown in Figure 3A, process 300A can be a decompression process corresponding to the compression process 200A in Figure 2A. In some embodiments, process 300A can be similar to the reconstruction path of process 200A. The decoder can decode the video bitstream 228 into a video stream 304 according to process 300A. The video stream 304 can be very similar to the video sequence 202. However, due to information loss in the compression and decompression processes (e.g., the quantization stage 214 in Figures 2A-2B), the video stream 304 is generally not identical to the video sequence 202. Similar to processes 200A and 200B in Figures 2A-2B, the decoder can perform process 300A at the level of a basic processing unit (BPU) for each picture encoded within the video bitstream 228. For example, the decoder can perform process 300A in an iterative manner, in which case the decoder can decode the basic processing unit in a single iteration of process 300A. Depending on the embodiment, the decoder can perform process 300A in parallel for each region of picture encoded in the video bitstream 228 (e.g., regions 114-118).
[0052]
[0086] In Figure 3A, the decoder can supply a portion of the video bitstream 228 associated with the basic processing unit of the encoded picture (referred to as the "encoded BPU") to the binary decoding stage 302. In the binary decoding stage 302, the decoder can decode this portion into prediction data 206 and quantization conversion coefficients 216. The decoder can supply the quantization conversion coefficients 216 to the inverse quantization stage 218 and the inverse conversion stage 220 to generate the reconstructed residual BPU 222. The decoder can supply the prediction data 206 to the prediction stage 204 to generate the prediction BPU 208. The decoder can add the reconstructed residual BPU 222 to the prediction BPU 208 to generate the prediction criterion 224. Depending on the embodiment, the prediction criterion 224 can be stored in a buffer (e.g., a decoded picture buffer in computer memory). The decoder can supply the prediction criterion 224 to the prediction stage 204 to perform the prediction calculation in the next iteration of process 300A.
[0053]
[0087] The decoder can iteratively perform process 300A to decode each encoding BPU of the encoded picture and generate a prediction criterion 224 for encoding the next encoding BPU of the encoded picture. After decoding all encoding BPUs of the encoded picture, the decoder can output the picture to the video stream 304 for display and proceed to decode the next encoded picture in the video bitstream 228.
[0054]
[0088] In the binary decoding stage 302, the decoder can perform the inverse operation of the binary coding technique used by the encoder (e.g., entropy coding, variable-length coding, arithmetic coding, Huffman coding, context-adaptive binary arithmetic coding, or any other lossless compression algorithm). Depending on the embodiment, in addition to the prediction data 206 and quantization conversion coefficients 216, the decoder can decode other information in the binary decoding stage 302, such as, for example, the prediction mode, the parameters of the prediction operation, the type of conversion, the parameters of the quantization process (e.g., quantization parameters), the encoder control parameters (e.g., bitrate control parameters), or the like. Depending on the embodiment, if the video bitstream 228 is transmitted in the form of packets over the network, the decoder can depacket the video bitstream 228 before supplying it to the binary decoding stage 302.
[0055]
[0089] Figure 3B shows a schematic diagram of another exemplary decoding process 300B according to several embodiments of the present disclosure. As shown in Figure 3B, process 300B may be modified from process 300A. For example, process 300B may be used with a decoder compliant with a hybrid video coding standard (e.g., the H.26x series). Compared to process 300A, process 300B further divides the prediction stage 204 into a spatial prediction stage 2042 and a temporal prediction stage 2044, and additionally includes a loop filter stage 232 and a buffer 234.
[0056]
[0090] In process 300B, the prediction data 206 decoded by the decoder from binary decoding stage 302 for the encoding base processing unit ("current BPU") of the encoded picture being decoded ("current picture") can include various types of data, depending on which prediction mode was used by the encoder to encode the current BPU. For example, if intra-prediction is used by the encoder to encode the current BPU, the prediction data 206 can include intra-prediction, parameters of the intra-prediction operation, or prediction mode indicators (e.g., flag values) indicating the same. Parameters of the intra-prediction operation can include, for example, the locations (e.g., coordinates) of one or more adjacent BPUs used as reference, the size of the adjacent BPUs, extrapolation parameters, the orientation of the adjacent BPUs relative to the original BPU, or the same. As another example, if inter-prediction is used by the encoder to encode the current BPU, the prediction data 206 can include inter-prediction, parameters of the inter-prediction operation, or prediction mode indicators (e.g., flag values) indicating the same. The parameters for the interpretation calculation may include, for example, the number of reference pictures associated with the current BPU, the weights associated with each reference picture, the locations (e.g., coordinates) of one or more matching regions within each reference picture, one or more motion vectors associated with each matching region, or similar.
[0057]
[0091] Based on the prediction mode indicator, the decoder can determine whether to perform a spatial prediction (e.g., intra-prediction) in the spatial prediction stage 2042 or a temporal prediction (e.g., inter-prediction) in the temporal prediction stage 2044. Details of performing such spatial or temporal predictions are shown in Figure 2B and will not be repeated below. After performing such spatial or temporal predictions, the decoder can generate a prediction BPU 208. The decoder can then add the prediction BPU 208 and the reconstructed residual BPU 222 to generate a prediction criterion 224, as described in Figure 3A.
[0058]
[0092] In process 300B, the decoder may supply the prediction criterion 224 to the spatial prediction stage 2042 or the temporal prediction stage 2044 to perform the prediction calculation in the next iteration of process 300B. For example, if the current BPU is decoded using intra-prediction in the spatial prediction stage 2042, after generating the prediction criterion 224 (e.g., the decoded current BPU), the decoder may supply the prediction criterion 224 directly to the spatial prediction stage 2042 for later use (e.g., for extrapolation of the next BPU of the current picture). If the current BPU is decoded using inter-prediction in the temporal prediction stage 2044, after generating the prediction criterion 224 (e.g., the reference picture with all BPUs decoded), the encoder may supply the prediction criterion 224 to the loop filter stage 232 to reduce or eliminate distortion (e.g., blocking artifacts). The decoder may apply the loop filter to the prediction criterion 224 in the manner described in Figure 2B. Loop-filtered reference pictures may be stored in buffer 234 (e.g., a decoded picture buffer in computer memory) for later use (e.g., to be used as an inter-prediction reference picture for future encoded pictures of the video bitstream 228). The decoder may store one or more reference pictures in buffer 234 for use in the temporal prediction stage 2044. In some embodiments, the prediction data 206 may further include loop filter parameters (e.g., loop filter strength) when a prediction mode indicator indicates that inter-prediction was used to encode the current BPU.
[0059]
[0093] Figure 4 is a block diagram of an exemplary apparatus 400 for encoding or decoding video, according to some embodiments of the present disclosure. As shown in Figure 4, the apparatus 400 may include a processor 402. When the processor 402 executes instructions as described herein, the apparatus 400 can become a specialized machine for video encoding or decoding. The processor 402 may be any kind of circuit mechanism having the ability to manipulate or process information. For example, the processor 402 may include any number and any combination of the following: a central processing unit (or "CPU"), a graphics processing unit (or "GPU"), a neural processing unit ("NPU"), a microcontroller unit ("MCU"), an optical processor, a programmable logic controller, a microcontroller, a microprocessor, a digital signal processor, an intellectual property (IP) core, a programmable logic array (PLA), a programmable array logic (PAL), a generic array logic (GAL), a complex programmable logic device (CPLD), a field-programmable gate array (FPGA), a system on a chip (SoC), an application-specific integrated circuit (ASIC), or similar components. Depending on the embodiment, the processor 402 may also be a set of processors grouped as a single logical component. For example, as shown in Figure 4, the processor 402 may include multiple processors, including processor 402a, processor 402b, and processor 402n.
[0060]
[0094] The device 400 may also include a memory 404 configured to store data (e.g., a set of instructions, computer code, intermediate data, or the like). For example, as shown in Figure 4, the stored data may include program instructions (e.g., program instructions for carrying out steps in processes 200A, 200B, 300A, or 300B) and data for processing (e.g., video sequence 202, video bitstream 228, or video stream 304). The processor 402 can access the program instructions and data for processing (e.g., via bus 410), execute the program instructions, and perform arithmetic or operations on the data for processing. The memory 404 may include a high-speed random-access storage device or a non-volatile storage device. Depending on the embodiment, memory 404 may include any number and any combination of random-access memory (RAM), read-only memory (ROM), optical disks, magnetic disks, hard drives, solid-state drives, flash drives, security digital (SD) cards, memory sticks, CompactFlash® (CF) cards, or similar. Memory 404 may also be a group of memories grouped as a single logical component (not shown in Figure 4).
[0061]
[0095] Bus 410 can be a communication device that transfers data between internal components of the device 400, such as an internal bus (e.g., a CPU-memory bus), an external bus (e.g., a universal serial bus port, a peripheral component interconnect express port), or similar.
[0062]
[0096] To facilitate explanation without creating ambiguity, the processor 402 and other data processing circuits are collectively referred to as “data processing circuits” in this disclosure. The data processing circuits may be implemented entirely as hardware, or as a combination of software, hardware, or firmware. In addition, the data processing circuits may be a single, standalone module, or may be fully or partially combined with any other component of the device 400.
[0063]
[0097] The device 400 may further include a network interface 406 for providing wired or wireless communication to a network (e.g., the Internet, an intranet, a local area network, a mobile communication network, or the same). Depending on the embodiment, the network interface 406 may include any number or any combination of a network interface controller (NIC), a radio frequency (RF) module, a transponder, a transceiver, a modem, a router, a gateway, a wired network adapter, a wireless network adapter, a Bluetooth® adapter, an infrared adapter, a near-field communication ("NFC") adapter, a cellular network chip, or the same.
[0064]
[0098] Depending on the embodiment, the apparatus 400 may further include a peripheral interface 408 for providing connectivity to one or more peripheral devices. As shown in Figure 4, the peripheral devices may include, but are not limited to, a cursor control device (e.g., a mouse, touchpad, or touchscreen), a keyboard, a display (e.g., a cathode ray tube display, a liquid crystal display, or a light-emitting diode display), a video input device (e.g., a camera, or an input interface communicatively coupled to a video archive), or the like.
[0065]
[0099] It should be noted that the video codec (for example, the codec that performs processes 200A, 200B, 300A, or 300B) may be implemented as any combination of any software or hardware modules within the device 400. For example, some or all stages of processes 200A, 200B, 300A, or 300B may be implemented as one or more software modules of the device 400, such as program instructions that can be loaded into memory 404. As another example, some or all stages of processes 200A, 200B, 300A, or 300B may be implemented as one or more hardware modules of the device 400, such as special data processing circuits (e.g., FPGA, ASIC, NPU, or similar).
[0066]
[0100] In quantization and dequantization function blocks (e.g., quantization 214 and dequantization 218 in Figure 2A or 2B, and dequantization 218 in Figure 3A or 3B), the quantization parameter (QP) is used to determine the amount of quantization (and dequantization) applied to the prediction residual. The initial QP value used for encoding a picture or slice can be signaled at a high level, for example, using the init_qp_minus26 syntax element in the Picture Parameter Set (PPS) and the slice_qp_delta syntax element in the slice header. Furthermore, the QP value can be adapted at a local level per CU using the delta QP value transmitted with the granularity of the quantization group.
[0067]
[0101] The equirectangular projection ("ERP") format is a common projection format used to represent 360-degree images and videos. The projection maps meridians to vertical lines at regular intervals and latitude circles to horizontal lines at regular intervals. Because the relationship between the position of image pixels on the map and their corresponding geographical location on the sphere is particularly simple, ERP is one of the most common projections used for 360-degree images and videos.
[0068]
[0102] The algorithm description of projection format conversion and the video quality criteria output by JVET provide the introduction and coordinate conversion between ERP and the sphere. For the coordinate conversion from 2D to 3D, assuming the sampling position is (m, n), (u, v) can be calculated based on the following formulae. u = (m + 0.5) / W, 0 ≤ m < W Equation (1) v = (n + 0.5) / H, 0 ≤ n < H Equation (2)
[0069]
[0103] Next, the longitude and latitude (φ, θ) of the sphere can be calculated from (u, v) based on the following formulae. φ = (u - 0.5)×(2×π) Equation (3) θ = (0.5 - v)×π Equation (4)
[0070]
[0104] The coordinates (X, Y, Z) can be calculated based on the following formulae. X = cos(θ)cos(φ) Equation (5) Y = sin(θ) Equation (6) Z = -cos(θ)sin(φ) Equation (7)
[0071]
[0105] For the coordinate conversion from 3D to 2D starting from (X, Y, Z), (φ, θ) can be calculated based on the following formulae. Next, (u, v) is calculated based on the formulae. Finally, (m, n) can be calculated based on the formulae. φ = tan -1 (-Z / X) Equation (8) θ = sin -1 (Y / (X 2 +Y 2 +Z 2 ) 1 / 2 ) Equation (9)
[0072]
[0106] To reduce seam artifacts in the reconstructed viewport that encompasses the left and right boundaries of the ERP picture, a new format called padded orthographic cylindrical projection (「PERP」) is provided by padding samples on each of the left and right sides of the ERP picture.
[0073]
[0107] When PERP is used to represent a 360-degree video, the PERP picture is encoded. After decoding, the reconstructed PERP is converted to the reconstructed ERP by blending the duplicated samples or cropping the padded area.
[0074]
[0108] FIG. 5A shows a schematic diagram of an exemplary blending operation for generating a reconstructed orthographic cylindrical projection according to some embodiments of the present disclosure. Unless otherwise specified, "recPERP" is used to indicate the reconstructed PERP before post-processing, and "recERP" is used to indicate the reconstructed ERP after post-processing. As shown in FIG. 5A, the duplicated samples of recPERP can be blended by applying a distance-based weighted average operation. For example, region A can be generated by blending region A1 with A2, and region B is generated by blending region B1 with B2.
[0075]
[0109] In the following description, the width and height of the unpadded recERP are denoted as "W" and "H", respectively. The left padding width and the right padding width are denoted as "P L " and "P R ", respectively. The total padding width is denoted as "P W ", and P W can be the sum of P L and P R . In some embodiments, recPERP can be converted to recERP by a blending operation. For example, for the samples recERP(j,i) in A, i = [0, P R-1 and j = [0, H - 1], recERP(j,i) can be determined according to the following formula. 及びj=[0,H-1]について、recERP(j,i)は、以下の式に従って決定され得る。 A = w × A1+(1 - w)×A2, where w is from P L / P w to 1 Equation (10) The recERP(j,i) within A = (recPERP(j,i + P L )×(i + P L)+recPERP(j,i+P L+W )×(P R-i )+(P W >>1)) / P W Formula (11)
[0076]
[0110] Depending on the embodiment, the sample recERP(j,i) in B, i=[WP L For [0,H-1] and j=[0,H-1], recERP(j,i) can be determined according to the following formula. B = k × B1 + (1 - k) × B2, where k ranges from 0 to P. L / P w Formula (12) recERP(j,i) = (recPERP(j,i+P L )×(P R-i +W)+recPERP(j,i+P L-w )×(i-W+P L )+(P w >>1)) / P W Formula (13)
[0077]
[0111] Figure 5B shows a schematic diagram of an exemplary cropping operation for generating a reconstructed equirectangular projection according to some embodiments of the present disclosure. As shown in Figure 5B, during the cropping process, padded samples in recPERP may be discarded directly to obtain recERP. For example, padded samples B1 and A2 may be discarded, padded area A being equal to A1, and padded area B being equal to B2.
[0078]
[0112] In some embodiments, horizontal wrap-around motion compensation may be used to improve the encoding performance of ERP. For example, horizontal wrap-around motion compensation may be used in the VVC standard as a 360-degree-specific encoding tool designed to improve the visual quality of 360-degree video reconstructed in ERP or PERP format. In conventional motion compensation, when a motion vector references a sample that exceeds the picture boundary of a reference picture, iterative padding is applied to derive the value of the sample that exceeds the boundary by copying from the nearest neighbor on the corresponding picture boundary. In the case of 360-degree video, this method of iterative padding is not suitable and can cause a visual artifact called a "seam artifact" in the reconstructed viewport video. Since 360-degree video is captured on a sphere and does not inherently have a "boundary," reference samples outside the boundary of the reference picture in the projection region can be obtained from neighboring samples within the spherical region. In the case of common projection formats, it can be difficult to derive the corresponding neighboring samples within the spherical region because it involves not only 2D-to-3D and 3D-to-2D coordinate transformations but also sample interpolation for fractional sample positions. This problem can be solved for the left and right boundaries of the ERP or PERP projection format, since the vicinity of the sphere outside the left picture boundary can be obtained from the sample inside the right picture boundary, and vice versa. Given the wide use of the ERP or PERP projection format and the relative ease of its embodiment, horizontal wrap-around motion compensation has been adapted to VVC to improve the visual quality of 360-degree images encoded in the ERP or PERP projection format.
[0079]
[0113] Figure 6A shows a schematic diagram of an exemplary horizontal wrap-around motion compensation process for equirectangular projection according to some embodiments of the present disclosure. As shown in Figure 6A, when a portion of the reference block lies outside the left (or right) boundary of the reference picture within the projection region, instead of repetitive padding, the “beyond the boundary” portion may be taken from the vicinity of the corresponding sphere located within the reference picture relative to the right (or left) boundary of the projection region. In some embodiments, repetitive padding is used only for the upper and lower picture boundaries.
[0080]
[0114] Figure 6B shows a schematic diagram of an exemplary horizontal wrap-around motion compensation process for padded equirectangular projection according to some embodiments of the present disclosure. As shown in Figure 6B, horizontal wrap-around motion compensation can be combined with non-standard padding methods often used in 360-degree video coding. In some embodiments, this is achieved by signaling a high-level syntactic element indicating a wrap-around motion compensation offset, which may be set to the ERP picture width before padding. This syntax may be used to adjust the position of the horizontal wrap-around accordingly. In some embodiments, this syntax is not affected by a specific amount of padding at the left or right picture boundary. As a result, this syntax can naturally support asymmetric padding of the ERP picture, where the left and right padding are different. In some embodiments, the wrap-around motion compensation may be determined according to the following formula:
number
[0081]
[0115] Horizontal wrap-around motion compensation can provide more meaningful information for motion compensation when the reference sample is outside the left and right boundaries of the reference picture. Under common test conditions for 360-degree video, this tool can improve compression performance not only in terms of rate distortion but also in terms of reduced seam artifacts and subjective quality of the reconstructed 360-degree video. Horizontal wrap-around motion compensation can also be used for other single-plane projection formats that have a constant sampling density in the horizontal direction, such as adjusted iso-area projection.
[0082]
[0116] Depending on the embodiment, limitations may be imposed on the wrap-around motion compensation offset. The offset value may be derived from the range of (CtbSizeY / MinCbSizeY+2) to (pic_width_in_luma_samples / MinCbSizeY), where the variable "CtbSizeY" refers to the luma size of the coded tree block ("CTB"), the variable "MinCbSizeY" refers to the minimum size of the luma coded block, and the variable "pic_width_in_luma_samples" refers to the picture width of the luma samples, thereby avoiding repeated wrap-arounds that are unnecessary for practical use but impose a burden on the hardware embodiment.
[0083]
[0117] Figure 7 shows the syntax of an exemplary sequence parameter set for wraparound motion compensation according to some embodiments of the present disclosure. As shown in Figure 5, in VVC (e.g., VVC Draft 7), the enable flag "sps_ref_wraparound_enabled_flag" and the offset "sps_ref_wraparound_offset_minus1" may be signaled in the sequence parameter set ("PPS") for wraparound motion compensation.
[0084]
[0118] Figure 8 shows the meaning of an exemplary sequence parameter set for wrap-around motion compensation in some embodiments of the present disclosure. It should be understood that the meaning shown in Figure 8 may correspond to the syntax shown in Figure 7. As shown in Figure 8, in some embodiments, “sps_ref_wraparound_enabled_flag” may indicate whether horizontal wrap-around motion compensation is applied to interprediction. For example, a value of 1 may indicate that horizontal wrap-around motion compensation is applied, and a value of 0 may indicate that horizontal wrap-around motion compensation is not applied. In some embodiments, when the value of (CtbSizeY / MinCbSizeY+1) is greater than (pic_width_in_luma_samples / MinCbSizeY-1), the value of sps_ref_wraparound_enabled_flag is equal to 0, in which case “pic_width_in_luma_samples” is the value of “pic_width_in_luma_samples” in any PPS that references the SPS.
[0085]
[0119] In some embodiments, as shown in Figure 8, “sps_ref_wraparound_offset_minus1”+1 may represent an offset used to calculate the horizontal wraparound position in “MinCbSizeY” luma sample units. In some embodiments, the value of ref_wraparound_offset_minus1 is within the range of (CtbSizeY / MinCbSizeY)+1 to (pic_width_in_luma_samples / MinCbSizeY)-1, where pic_width_in_luma_samples is the value of pic_width_in_luma_samples in any PPS that references the SPS.
[0086]
[0120] There are several problems with the syntax shown in Figure 7 and the meaning shown in Figure 8. In particular, "sps_ref_wraparound_enabled_flag" and "sps_ref_wraparound_offset_minus1" are syntactic elements signaled in the SPS, but there are compatibility constraints that depend on all of "pic_width_in_luma_samples" which are signaled in the PPS. Since the SPS is a higher-level syntax than the PPS, and higher-level syntax should not normally refer to lower-level syntax, restricting the values of SPS syntactic element values by syntactic elements in all related PPS can be problematic. Furthermore, in some embodiments, wraparound motion compensation is controlled at the sequence level, while changes in picture size are permitted in the VVC draft (e.g., VVC Draft 7). At the same time, "sps_ref_wraparound_enabled_flag" can only be true if the width of all pictures in the sequence referencing the SPS satisfies the constraint. Therefore, wraparound motion compensation cannot be used if only one frame does not meet the size requirement. This means that the benefits of wrap-around motion compensation for the entire sequence can be lost due to a single frame.
[0087]
[0121] In addition, depending on the embodiment, "sps_ref_wraparound_offset_minus1" ranges from (CtbSizeY / MinCbSizeY)+1 to (pic_width_in_luma_samples / MinCbSizeY)-1. Therefore, the minimum value signaled in the bitstream for sps_ref_wraparound_offset_minus1 is (CtbSizeY / MinCbSizeY)+1, which may not be zero. Generally, larger values require more bits for signaling than smaller values. As a result, signaling syntactic elements with a range of values that do not start from zero is inefficient.
[0088]
[0122] Embodiments of this disclosure provide an improved method for solving the problems described above. Figure 9 shows the syntax of an exemplary sequence parameter set for improved wrap-around motion compensation according to some embodiments of this disclosure. In some embodiments, the signaling overhead of the wrap-around motion compensation ("MC") offset can be preserved. In order to preserve the bits specified for the wrap-around motion compensation offset, (CtbSizeY / MinCbSizeY)+2 can be subtracted from the wrap-around motion compensation offset before the wrap-around motion compensation offset is signaled. As a result, the minimum value of this syntax element may be 0.
[0089]
[0123] Figure 10 illustrates the meaning of an exemplary sequence parameter set for improved wrap-around motion compensation according to some embodiments of this disclosure. Changes from the previous VVC are shown in italics, as shown in Figure 10. It should be understood that the meaning shown in Figure 10 may correspond to the syntax shown in Figure 9.
[0090]
[0124] In some embodiments, as shown in Figure 10, “sps_ref_wraparound_enabled_flag” may indicate whether horizontal wraparound motion compensation is applied to interpretation. For example, a value of 1 may indicate that horizontal wraparound motion compensation is applied, and a value of 0 may indicate that horizontal wraparound motion compensation is not applied. In some embodiments, when the value of (CtbSizeY / MinCbSizeY+1) is greater than (pic_width_in_luma_samples / MinCbSizeY-1), the value of “sps_ref_wraparound_enabled_flag” is equal to 0.
[0091]
[0125] Depending on the embodiment, as shown in Figure 10, “sps_ref_wraparound_offset” + (CtbSizeY / MinCbSizeY) + 2 may represent an offset used to calculate the horizontal wraparound position in “MinCbSizeY” luma sample units. The value of “sps_ref_wraparound_offset” can be in the range of 0 to (pic_width_in_luma_samples / MinCbSizeY) - (CtbSizeY / MinCbSizeY) - 2, where pic_width_in_luma_samples is the value of pic_width_in_luma_samples in any PPS that references the SPS.
[0092]
[0126] As mentioned above, another problem with conventional designs is that wraparound motion compensation is disabled for all pictures in a video sequence, even if one picture in the video sequence has dimensions that violate the compliance requirements. In some embodiments, constraints on syntactic element values are removed. The wraparound motion compensation control flag sps_ref_wraparound_enabled_flag is first signaled in SPS. In some embodiments, if sps_ref_wraparound_enabled_flag is true, the offset value sps_ref_wraparound_offset_minus1 is signaled.
[0093]
[0127] Figure 11 illustrates the meaning of an exemplary sequence parameter set for improved wrap-around motion compensation using maximum picture width, according to some embodiments of the present disclosure. As shown in Figure 11, changes from the previous VVC are indicated in italics, and proposed deleted meanings are further indicated in strikethrough.
[0094]
[0128] Depending on the embodiment, as shown in Figure 11, “sps_ref_wraparound_enabled_flag” may indicate whether horizontal wraparound motion compensation is applied to interprediction. For example, a value of 1 may indicate that horizontal wraparound motion compensation can be applied to interprediction, while a value of 0 may indicate that horizontal wraparound motion compensation is not applied.
[0095]
[0129] In some embodiments, as shown in Figure 11, “sps_ref_wraparound_offset_minus1”+1 may represent the maximum offset used to calculate the horizontal wraparound position in “MinCbSizeY” luma sample units. In some embodiments, the value of “sps_ref_wraparound_offset_minus1” is within the range of (CtbSizeY / MinCbSizeY)+1 to (pic_width_max_in_luma_samples / MinCbSizeY)-1.
[0096]
[0130] In some embodiments, "pic_width_max_in_luma_samples" is the maximum width in luma samples for each decoded picture referencing the SPS.
[0097]
[0131] Depending on the embodiment, two variables, "PicRefWraparoundEnableFlag" and "PicRefWraparoundOffset," may be defined for each picture in a sequence. Figure 12 shows examples of the derivation of the variables "PicRefWraparoundEnableFlag" and "PicRefWraparoundOffset" according to some embodiments of the present disclosure. As shown in Figure 12, "pic_width_in_luma_samples" may refer to the width of the picture that "pic_width_in_luma_samples" refers to, which signals the PPS.
[0098]
[0132] In some embodiments, as shown in Figure 12, the variable "PicRefWraparoundEnableFlag" may be used to determine whether wrap-around MC can be enabled for the current picture. For example, if the value of "PicRefWraparoundEnableFlag" indicates that wrap-around MC can be enabled for the current picture, the offset "PicRefWraparoundOffset" is used in the motion compensation process.
[0099]
[0133] Figure 13 shows an example of deriving a sample position used for motion compensation according to some embodiments of the present disclosure. As shown in Figure 13, the sample position (xInt i ,yInt i ) refers to the sample position before the wraparound, and the sample position (xInt i ,yInt i) can be determined. Depending on the embodiment, the variable "picW" is equal to the variable "pic_width_in_luma_samples". Depending on the embodiment, the functions "ClipH" and "Clip3" may be performed according to the formulas shown in Figure 13.
[0100]
[0134] In some embodiments, the wraparound motion compensation control flag may also be signaled in the SPS, while the wraparound motion compensation offset is signaled in the PPS rather than the SPS. Figure 14 shows exemplary sequence and picture parameter set syntax for wraparound motion compensation using a wraparound motion compensation offset in a picture parameter set, according to some embodiments of the present disclosure. As shown in Figure 14, changes from the previous VVC are indicated in italics, and proposed deleted syntax is further indicated in strikethrough. In some embodiments, as shown in Figure 14, “sps_ref_wraparound_enabled_flag” is signaled in the SPS, and “pps_ref_wraparound_offset_minus1” is signaled in the PPS.
[0101]
[0135] Figure 15 shows the meaning of exemplary sequence and picture parameter sets for wrap-around motion compensation using a wrap-around motion compensation offset in a picture parameter set, according to some embodiments of this disclosure. As shown in Figure 15, changes from the previous VVC are shown in italics, and proposed deleted meanings are further shown in strikethrough. It should be understood that the meanings shown in Figure 15 may correspond to the syntax shown in Figure 14.
[0102]
[0136] Depending on the embodiment, as shown in Figure 15, “sps_ref_wraparound_enabled_flag” may indicate whether horizontal wraparound motion compensation is applied to interpretation. For example, a value of 1 may indicate that horizontal wraparound motion compensation is applied, and a value of 0 may indicate that horizontal wraparound motion compensation is not applied.
[0103]
[0137] In some embodiments, as shown in Figure 15, "pps_ref_wraparound_offset_minus1" + 1 may represent an offset used to calculate the horizontal wraparound position in "MinCbSizeY" luma sample units. In some embodiments, "pps_ref_wraparound_offset_minus1" is equal to 0 when "sps_ref_wraparound_enabled_flag" is equal to 0 or when the value of (CtbSizeY / MinCbSizeY+1) is greater than (pic_width_in_luma_samples / MinCbSizeY-1). Otherwise, the value of "pps_ref_wraparound_offset_minus1" is within the range of (CtbSizeY / MinCbSizeY)+1 to (pic_width_in_luma_samples / MinCbSizeY)-1.
[0104]
[0138] In some embodiments, two variables, "PicRefWraparoundEnableFlag" and "PicRefWraparoundOffset," may be defined for each picture in a sequence. In some embodiments, "PicRefWraparoundEnableFlag" may be determined as shown in Figure 12. In some embodiments, "PicRefWraparoundOffset" may be determined as "pps_ref_wraparound_offset_minus1" + 1.
[0105]
[0139] Depending on the embodiment, "PicRefWraparoundEnableFlag" and "PicRefWraparoundOffset" may be used for wraparound motion compensation during the decoding process. For example, the sample position (xInt) used for motion compensation. i ,yInt i ) can be derived in the manner shown in Figure 13. As shown in Figure 13, in some embodiments, the variable "picW" may be equal to "pic_width_in_luma_samples".
[0106]
[0140] In some embodiments, the wraparound motion compensation control flag may also be signaled, but the wraparound motion compensation offset is signaled in the PPS rather than the SPS. Furthermore, “pps_ref_wraparound_offset” may also indicate the use of wraparound motion compensation for pictures referencing the PPS. Figure 16 shows the syntax of an exemplary sequence parameter set for improved wraparound motion compensation without a wraparound motion compensation offset, according to some embodiments of the present disclosure. As shown in Figure 16, changes from the previous VVC are shown in italics, and proposed deleted syntax is further shown in strikethrough. As shown in Figure 14, “sps_ref_wraparound_enabled_flag” may be signaled in the SPS.
[0107]
[0141] Figure 17 shows the syntax of an exemplary picture parameter set for improved wrap-around motion compensation using a wrap-around motion compensation offset, according to some embodiments of the present disclosure. Changes from the previous VVC are shown in italics, as shown in Figure 17. It should be understood that the PPS shown in Figure 17 may correspond to the SPS shown in Figure 16. As shown in Figure 17, "pps_ref_wraparound_offset" may be signaled in the PPS. In some embodiments, "pps_ref_wraparound_offset" signaled in the PPS may also indicate the use of wrap-around motion compensation for pictures referencing the PPS. In other words, the encoder may disable wrap-around motion compensation at the PPS level by setting pps_ref_wraparound_offset to a special value.
[0108]
[0142] Figure 18 shows the semantics of exemplary sequence and picture parameter sets for improved wrap-around motion compensation using a wrap-around motion compensation offset in a picture parameter set, according to some embodiments of this disclosure. As shown in Figure 18, changes from the previous VVC are shown in italics, and proposed deleted semantics are further shown in strikethrough. It should be understood that the semantics shown in Figure 18 may correspond to the syntax shown in Figures 16 and 17.
[0109]
[0143] Depending on the embodiment, as shown in Figure 18, “sps_ref_wraparound_enabled_flag” may indicate whether horizontal wraparound motion compensation is applied to interprediction. For example, a value of 1 may indicate that horizontal wraparound motion compensation can be applied to interprediction, while a value of 0 may indicate that horizontal wraparound motion compensation is not applied.
[0110]
[0144] Depending on the embodiment, as shown in Figure 18, "pps_ref_wraparound_offset" + 1 may represent the value of the offset used to calculate the horizontal wraparound position in MinCbSizeY luma samples. For example, when "pps_ref_wraparound_offset" is equal to 0, wraparound motion compensation is disabled. "pps_ref_wraparound_offset" is equal to 0 when "sps_ref_wraparound_enabled_flag" is equal to 0 or when the value of (CtbSizeY / MinCbSizeY+1) is greater than (pic_width_in_luma_samples / MinCbSizeY-1). Otherwise, the value of "pps_ref_wraparound_offset" is within the range of (CtbSizeY / MinCbSizeY)+1 to (pic_width_in_luma_samples / MinCbSizeY)-1.
[0111]
[0145] Depending on the embodiment, two variables, "PicRefWraparoundEnableFlag" and "PicRefWraparoundOffset," may be defined for each picture in the sequence. Figure 19 shows examples of the derivation of the variables "PicRefWraparoundEnableFlag" and "PicRefWraparoundOffset" according to some embodiments of this disclosure.
[0112]
[0146] Depending on the embodiment, "PicRefWraparoundEnableFlag" and "PicRefWraparoundOffset" may be used for wraparound motion compensation during the decoding process. For example, the sample position (xInt) used for motion compensation. i ,yInt i ) can be derived in the manner shown in Figure 11. As shown in Figure 11, in some embodiments, the variable "picW" may be equal to "pic_width_in_luma_samples".
[0113]
[0147] Depending on the embodiment, the syntax may be modified. The wrap-around motion compensation control flag may still be signaled in the SPS, but the wrap-around motion compensation offset may be signaled in the PPS instead of the SPS. Furthermore, the wrap-around motion compensation control flag at the PPS level may also be signaled. Figure 20 shows the syntax of an exemplary sequence parameter set for improved wrap-around motion compensation, without a wrap-around motion compensation offset in the sequence parameter set, according to some embodiments of the present disclosure. As shown in Figure 20, changes from the previous VVC are shown in italics, and proposed deleted syntax is further shown in strikethrough. As shown in Figure 20, "sps_ref_wraparound_enabled_flag" may be signaled in the SPS.
[0114]
[0148] Figure 21 shows the syntax of an exemplary picture parameter set for improved wraparound motion compensation using a wraparound control flag, according to some embodiments of the present disclosure. Changes from the previous VVC are shown in italics, as shown in Figure 21. As shown in Figure 21, “pps_ref_wraparound_enabled_flag” may be signaled in the PPS. In some embodiments, if “pps_ref_wraparound_enabled_flag” is true (e.g., its value is equal to 1), “pps_ref_wraparound_offset” may be signaled.
[0115]
[0149] Figure 22 illustrates the meaning of exemplary sequence and picture parameter sets for improved wrap-around motion compensation using a wrap-around control flag in the picture parameter set, according to some embodiments of this disclosure. As shown in Figure 22, changes from the previous VVC are shown in italics, and proposed deleted meanings are further shown in strikethrough. It should be understood that the meanings shown in Figure 22 may correspond to the syntax shown in Figures 20 and 21.
[0116]
[0150] Depending on the embodiment, as shown in Figure 22, “sps_ref_wraparound_enabled_flag” may indicate whether horizontal wraparound motion compensation is applied to interprediction. For example, a value of 1 may indicate that horizontal wraparound motion compensation can be applied to interprediction, while a value of 0 may indicate that horizontal wraparound motion compensation is not applied.
[0117]
[0151] In some embodiments, as shown in Figure 22, a "pps_ref_wraparound_enabled_flag" equal to 1 may indicate that horizontal wraparound motion compensation is applied in interpretation. A "pps_ref_wraparound_enabled_flag" equal to 0 may indicate that horizontal wraparound motion compensation is not applied. In some embodiments, "pps_ref_wraparound_enabled_flag" is equal to 0 when "sps_ref_wraparound_enabled_flag" is equal to 0 or when the value of (CtbSizeY / MinCbSizeY+1) is greater than (pic_width_in_luma_samples / MinCbSizeY-1).
[0118]
[0152] Depending on the embodiment, alternative meanings exist for the sequence parameter set and picture parameter set, as shown in Figure 22. Figure 23 shows exemplary sequence parameter set and picture parameter set meanings for improved wrap-around motion compensation using a wrap-around control flag in the picture parameter set, according to some embodiments of this disclosure. As shown in Figure 23, changes from the previous VVC are shown in italics, and proposed deleted meanings are further shown in strikethrough. It should be understood that the meanings shown in Figure 23 may correspond to the syntax shown in Figures 20 and 21.
[0119]
[0153] In some embodiments, as shown in Figure 23, a "pps_ref_wraparound_enabled_flag" equal to 1 may indicate that horizontal wraparound motion compensation is applied to interpretation. A "pps_ref_wraparound_enabled_flag" equal to 0 may indicate that horizontal wraparound motion compensation is not applied. In some embodiments, "pps_ref_wraparound_enabled_flag" is 0 when "sps_ref_wraparound_enabled_flag" is equal to 0 or when the value of (CtbSizeY / MinCbSizeY+1) is greater than (pic_width_in_luma_samples / MinCbSizeY-1). Otherwise, "pps_ref_wraparound_enabled_flag" is equal to 1.
[0120]
[0154] In some embodiments, as shown in Figure 23, "pps_ref_wraparound_offset" + (CtbSizeY / MinCbSizeY) + 2 may specify an offset value used to calculate the horizontal wraparound position in MinCbSizeY luma sample units. In some embodiments, the value of "pps_ref_wraparound_offset", if present, may be within the range of 0 to (pic_width_in_luma_samples / MinCbSizeY) - (CtbSizeY / MinCbSizeY) - 2.
[0121]
[0155] Depending on the embodiment, a variable "PicRefWraparoundOffset" may be defined for each picture in the sequence. For example, "PicRefWraparoundOffset" may be derived as pps_ref_wraparound_offset_minus1+1.
[0122]
[0156] In some embodiments, during the decoding process, the variables "pps_ref_wraparound_enabled_flag" and "PicRefWraparoundOffset" may be used for wraparound motion compensation. Figure 24 shows an example of the derivation of the variable "PicRefWraparoundOffset" according to some embodiments of this disclosure. As shown in Figure 24, the variable "PicRefWraparoundOffset" may be derived according to the variables "pps_ref_wraparound_offset", "CtbSizeY", and "MinCbSizeY". In some embodiments, the variable "PicRefWraparoundOffset" is a sample position (xInt) used for motion compensation, similar to the sample position shown in Figure 13. i ,yInt i It can also be used to determine ). As shown in Figure 13, the variable "picW" may be equal to "pic_width_in_luma_samples".
[0123]
[0157] Depending on the embodiment, "sps_ref_wraparound_enabled_flag" may be removed, while "pps_ref_wraparound_enabled_flag" and "pps_ref_wraparound_offset" may be retained.
[0124]
[0158] Depending on the embodiment, the restrictions on the value ranges of "sps_ref_wraparound_enabled_flag" and "sps_ref_wraparound_offset_minus1" may be removed, and restrictions on the value range of the picture size signaled in the SPS and PPS may be added. Furthermore, no syntax changes are required. Figure 25 shows the meaning of exemplary sequence parameter sets and picture parameter sets for improved wraparound motion compensation using restrictions on picture size according to some embodiments of the present disclosure. As shown in Figure 25, changes from the previous VVC are indicated in italics, and proposed deleted meanings are further indicated in strikethrough.
[0125]
[0159] In some embodiments, as shown in Figure 25, “pic_width_max_in_luma_samples” may represent the maximum width in luma samples for each decoded picture referencing the SPS. In some embodiments, “pic_width_max_in_luma_samples” does not have to be equal to 0 and may be an integer multiple of max(8,MinCbSizeY).
[0126]
[0160] In some embodiments, as shown in Figure 25, "pic_height_max_in_luma_samples" may represent the maximum height in luma samples for each decoded picture referencing the SPS. In some embodiments, "pic_height_max_in_luma_samples" does not have to be equal to 0 and may be an integer multiple of max(8,MinCbSizeY).
[0127]
[0161] Depending on the embodiment, as shown in Figure 25, a "sps_ref_wraparound_enabled_flag" equal to 1 may indicate that horizontal wraparound motion compensation is applied to interpretation, while a "sps_ref_wraparound_enabled_flag" equal to 0 may indicate that horizontal wraparound motion compensation is not applied.
[0128]
[0162] In some embodiments, "sps_ref_wraparound_offset_minus1" + 1 may represent an offset used to calculate the horizontal wraparound position in units of "MinCbSizeY" luma samples. In some embodiments, the value of "sps_ref_wraparound_offset_minus1" is greater than or equal to (CtbSizeY / MinCbSizeY) + 1.
[0129]
[0163] Depending on the embodiment, restrictions may be imposed on "pic_width_max_in_luma_samples", "CtbSizeY", and "MinCbSizeY". Figure 26 shows the meaning of an exemplary sequence parameter set for improved wrap-around motion compensation using restrictions imposed on the variables "pic_width_max_in_luma_samples", "CtbSizeY", and "MinCbSizeY" according to some embodiments of the present disclosure. As shown in Figure 26, changes from the previous VVC are shown in italics, and proposed deleted meanings are further shown in strikethrough.
[0130]
[0164] Depending on the embodiment, a restriction may be imposed on "pic_width_in_luma_samples," which is signaled in the PPS. Figure 27 shows the meaning of an exemplary picture parameter set for improved wrap-around motion compensation using a restriction imposed on the variable "pic_width_in_luma_samples" according to some embodiments of this disclosure. As shown in Figure 27, changes from the previous VVC are shown in italics, and proposed deleted meanings are further shown in strikethrough.
[0131]
[0165] Depending on the embodiment, the methods shown in Figures 9-11 may be combined with any of the methods shown in Figures 11-27. To reduce signaling costs, a special value is subtracted from the wrap-around motion compensation offset before it is signaled (e.g., the method shown in Figures 9-10), so when the methods are combined, the range limit of the wrap-around motion compensation offset signaled in the bitstream may also be modified. For example, the same special value may be subtracted from both the upper and lower limits. Furthermore, if the lower limit after subtraction is 0, it may be removed, as this ensures that the offset signaled in the bitstream is a non-negative value in the VVC standard (e.g., VVC Draft 7).
[0132]
[0166] Embodiments of this disclosure further provide methods for performing motion compensation. Figure 28 shows a flowchart of an exemplary method for performing motion compensation according to some embodiments of this disclosure. It should be understood that the method 28000 shown in Figure 28 may be performed in accordance with the syntax and semantics shown in Figures 9 and 10.
[0133]
[0167] In step S28010, a sequence of pictures is received. The sequence is associated with a sequence wrap-around motion compensation flag and a sequence wrap-around motion compensation offset. The minimum value for the sequence wrap-around motion compensation offset is 0. For example, as shown in Figure 9, (CtbSizeY / MinCbSizeY)+2 may be subtracted from the wrap-around motion compensation offset before it is signaled in order to save the bits specified for the wrap-around motion compensation offset. As a result, the minimum value for this syntactic element can be 0.
[0134]
[0168] In step S28020, it is determined whether the sequence wrap-around motion compensation flag is enabled.
[0135]
[0169] In step S28030, depending on whether the sequence wrap-around motion compensation flag is enabled, wrap-around motion compensation is performed on the pictures in the sequence of pictures according to the sequence wrap-around motion compensation offset. Depending on the embodiment, motion compensation is performed according to the VVC standard.
[0136]
[0170] Embodiments of the present disclosure further provide methods for performing motion compensation with a limited range of sequence wrap-around motion compensation offsets. Figure 29 shows a flowchart of an exemplary method for performing motion compensation with a limited range of sequence wrap-around motion compensation offsets according to some embodiments of the present disclosure. It should be understood that the method 29000 shown in Figure 29 may be performed in accordance with the meaning shown in Figure 11.
[0137]
[0171] In step S29010, a sequence of pictures is received. The sequence is associated with a sequence wrap-around motion compensation flag and a sequence wrap-around motion compensation offset. The range for the sequence wrap-around motion compensation offset is limited according to the maximum width of the pictures in the sequence of pictures. For example, as shown in Figure 11, "pic_width_max_in_luma_samples" may represent the maximum width in luma samples for each decoded picture referencing the SPS. The value of "sps_ref_wraparound_offset_minus1" can be within the range of (CtbSizeY / MinCbSizeY)+1 to (pic_width_max_in_luma_samples / MinCbSizeY)-1.
[0138]
[0172] In step S29020, it is determined whether the sequence wrap-around motion compensation flag is enabled.
[0139]
[0173] In step S29030, depending on whether the sequence wraparound motion compensation flag is enabled, wraparound motion compensation is performed on the pictures in the sequence of pictures according to the sequence wraparound motion compensation offset. In some embodiments, motion compensation is performed according to the VVC standard. In some embodiments, wraparound motion compensation may be performed on multiple pictures in the sequence of pictures, and the multiple pictures may have different sizes. In some embodiments, wraparound motion compensation for a picture is performed according to the sequence wraparound motion compensation offset depending on whether the picture wraparound enabled flag is enabled. The picture wraparound enabled flag may be determined according to the sequence wraparound motion compensation flag. For example, as shown in Figure 12, the picture wraparound enabled flag may be determined from an expression that includes the variable "sps_ref_wraparound_enabled_flag".
[0140]
[0174] Embodiments of this disclosure further provide methods for performing motion compensation using pictures associated with sequence wrap-around motion compensation offsets. Figure 30 shows a flowchart of an exemplary method for performing motion compensation using pictures associated with sequence wrap-around motion compensation offsets, according to some embodiments of this disclosure. It should be understood that the method 30000 shown in Figure 30 can be performed in accordance with the syntax and semantics shown in Figures 14 and 15.
[0141]
[0175] In step S30010, a sequence of pictures is received. The sequence is associated with a sequence wraparound motion compensation flag, and the pictures within the sequence are associated with a picture wraparound motion compensation offset. For example, as shown in Figure 14, a new variable "pps_ref_wraparound_offset" may be included in the picture parameter set.
[0142]
[0176] In step S30020, it is determined whether the sequence wrap-around motion compensation flag is enabled.
[0143]
[0177] In step S30030, depending on whether the sequence wrap-around motion compensation flag is enabled, wrap-around motion compensation is performed on the pictures in the sequence of pictures according to the sequence wrap-around motion compensation offset. In some embodiments, motion compensation is performed according to the VVC standard. In some embodiments, wrap-around motion compensation may be performed on multiple pictures in the sequence of pictures, and the multiple pictures may have different sizes.
[0144]
[0178] In some embodiments, wrap-around motion compensation for a picture is performed according to a sequence wrap-around motion compensation offset, depending on whether the picture wrap-around enabled flag is enabled. The picture wrap-around enabled flag may be determined according to the sequence wrap-around motion compensation flag. For example, as shown in Figure 12, the picture wrap-around enabled flag may be determined from an expression including the variable "sps_ref_wraparound_enabled_flag". In some embodiments, the minimum value of the picture wrap-around motion compensation offset is 0. For example, as shown in Figure 18, the minimum value for the variable "pps_ref_wraparound_offset" may be 0.
[0145]
[0179] In some embodiments, a picture is associated with a picture wraparound motion compensation flag. Depending on whether the picture wraparound motion compensation flag is enabled, wraparound motion compensation may be performed on the picture according to a picture wraparound motion compensation offset. For example, as shown in Figure 21, a new variable "pps_ref_wraparound_enabled_flag" may be added to the picture parameter set. As shown in Figure 22, the variable "pps_ref_wraparound_enabled_flag" may indicate whether horizontal wraparound motion compensation is applied at the picture level. In some embodiments, depending on whether the picture wraparound motion compensation flag is enabled, a picture wraparound motion compensation offset may be signaled.
[0146]
[0180] Embodiments of the present disclosure further provide methods for performing motion compensation with a limited maximum picture size. Figure 31 shows a flowchart of an exemplary method for performing motion compensation with a limited maximum picture size according to some embodiments of the present disclosure. It should be understood that the method 31000 shown in Figure 31 may be performed in accordance with the meaning shown in Figure 25.
[0147]
[0181] In step S31010, a sequence of pictures is received. The sequence is associated with a sequence wrap-around motion compensation flag, and the pictures within the sequence are associated with a picture wrap-around motion compensation offset.
[0148]
[0182] In step S31020, it is determined whether the sequence wrap-around motion compensation flag is enabled.
[0149]
[0183] In step S31030, depending on whether the sequence wraparound motion compensation flag is enabled, wraparound motion compensation is performed on the pictures in the sequence of pictures according to the sequence wraparound motion compensation offset. The maximum size of a picture is limited to a minimum value according to the sequence wraparound motion compensation offset. For example, as shown in Figure 26, the maximum picture width may be determined according to an expression including "sps_ref_wraparound_offset_minus1". In some embodiments, motion compensation is performed according to the VVC standard. In some embodiments, wraparound motion compensation may be performed on multiple pictures in the sequence of pictures, and the multiple pictures may have different sizes. In some embodiments, the size of a picture is limited to a minimum value according to the sequence wraparound motion compensation offset. For example, as shown in Figure 27, the picture width may be determined according to an expression including "sps_ref_wraparound_offset_minus1".
[0150]
[0184] Depending on the embodiment, a non-temporary computer-readable storage medium containing instructions may also be provided, which may be executed by a device (such as an encoder and decoder of the Disclosure) to carry out the methods described above. Common forms of non-temporary media include, for example, floppy disks, flexible disks, hard disks, solid-state drives, magnetic tapes, or any other magnetic data storage media, CD-ROMs, any other optical data storage media, any physical media having a pattern of holes, RAM, PROMs, and EPROMs, FLASH®-EPROMs, or any other flash memory, NVRAMs, caches, registers, any other memory chips or cartridges, and networked versions thereof. A device may include one or more processors (CPUs), input / output interfaces, network interfaces, and / or memory.
[0151]
[0185] It should be noted that relational terms in this specification, such as "first" and "second," are used merely to distinguish one entity or action from another, and do not imply or require any actual relationship or order between these entities or actions. Furthermore, the words "comprising," "having," "containing," and "including," as well as other similar forms, are intended to be open-ended in that they are equivalent in meaning, and the elements or groups of elements that follow any of these words are not meant to be an exhaustive enumeration of such elements or groups of elements, nor are they meant to be limited to only the enumerated elements or groups of elements.
[0152]
[0186] As used herein, unless otherwise specified, the term "or" encompasses all possible combinations, except in cases where it is not feasible. For example, if it is stated that a database may contain A or B, then unless otherwise specified or it is not feasible, the database may contain A, or B, or A and B. As a second example, if it is stated that a database may contain A, B, or C, then unless otherwise specified or it is not feasible, the database may contain A, B, or C, or A and B, A and C, or B and C, or A and B and C.
[0153]
[0187] It will be understood that the embodiments described above may be implemented by hardware, software (program code), or a combination of hardware and software. When implemented by software, it may be stored in the computer-readable medium described above. The software can perform the methods of this disclosure when executed by a processor. The computing units and other functional units described in this disclosure may be implemented by hardware, software, or a combination of hardware and software. Those skilled in the art will also understand that some of the above modules / units may be combined into a single module / unit, and each of the above modules / units may be further divided into some submodules / subunits.
[0154]
[0188] In the above-described specification, embodiments have been described with reference to numerous specific details that may differ depending on the implementation. Specific adaptations and modifications of the above-described embodiments can be made. Other embodiments may become apparent to those skilled in the art from the considerations herein and the implementations of the invention disclosed herein. The specification and examples are intended to be considered as examples only, and the true scope and spirit of the invention are indicated by the appended claims. Furthermore, the arrangement of steps shown in the figures is for illustrative purposes only and is not intended to limit the invention to any particular arrangement of steps. Therefore, those skilled in the art will understand that these steps may be performed in different orders while carrying out the same method.
[0155]
[0189] Embodiments can be further described using the following clauses. 1. A method for performing motion compensation, Receiving a first wrap-around motion compensation flag, wherein the first wrap-around motion compensation flag is associated with a picture, To determine whether the first wrap-around motion compensation flag is enabled, In response to the determination that the first wrap-around motion compensation flag is valid, the wrap-around motion compensation offset is received, and the wrap-around motion compensation offset is associated with the picture. Perform motion compensation on the picture according to the first wrap-around motion compensation flag and wrap-around motion compensation offset, Methods that include... 2. Receiving a second wrap-around motion compensation flag, wherein the first wrap-around motion compensation flag is associated with a set of pictures including the picture associated with the first wrap-around motion compensation flag, Determine whether the second wrap-around motion compensation flag is disabled, In accordance with the determination that the second wrap-around motion compensation flag is invalid, the first wrap-around motion compensation flag is also determined to be invalid. The method described in Clause 1, further including the method described in Clause 1. 3. To determine whether the first wrap-around motion compensation flag is enabled, Determining the picture width of the picture associated with the first wrap-around motion compensation flag, The first wrap-around motion compensation flag is determined based on the picture width, and The method described in Clause 2, further including the method described in Clause 2. 4. Determine whether the Luma coding tree block size + 1 for the smallest coding block unit is greater than the picture width - 1 for the smallest coding block unit, The first motion compensation flag is invalid when it is determined that the Luma coding tree block size + 1 for the smallest coding block unit is greater than the picture width - 1 for the smallest coding block unit. The method described in Clause 3, further including the method described in Clause 3. 5. Determine whether the second wrap-around motion compensation flag is enabled, In response to the determination that the second wrap-around motion compensation flag is enabled, it is determined that the picture width of the picture is greater than or equal to the Luma coding tree block size + offset, The method described in any one of the clauses 2 to 4, further including the method described in any one of the clauses 2 to 4. 6. The method according to Clause 5, further comprising determining, in response to the determination that the second wrap-around motion compensation flag is valid, that the Luma coding tree block size + 1 of the minimum coding block unit is less than or equal to the picture width - 1 of the minimum coding block unit picture. 7. The method according to any one of clauses 2 to 6, wherein a second wrap-around motion compensation flag is signaled in the sequence parameter set, and a first wrap-around motion compensation flag and a wrap-around motion compensation offset are signaled in the picture parameter set. 8. Motion compensation is performed in accordance with the Multipurpose Video Coding Standard, as described in any one of Clauses 1 to 7. 9. The method described in any one of the clauses 1 to 8, further comprising performing motion compensation for multiple pictures, wherein the multiple pictures have different sizes. 10. Performing motion compensation on a picture according to a wrap-around motion compensation offset is, The second wrap-around motion compensation offset is determined by adding an offset to the wrap-around motion compensation offset received from the bitstream, Perform motion compensation on the picture according to the second wrap-around motion compensation offset, The method described in any one of the clauses 1 to 9, further including the method described in any one of the clauses 1 to 9. 11. A system that performs motion compensation, Memory that stores a set of instructions, Equipped with a processor, the processor is Receiving a first wrap-around motion compensation, wherein the first wrap-around motion compensation offset is associated with picture i, To determine whether the first wrap-around motion compensation flag is enabled, In response to the determination that the first wrap-around motion compensation flag is valid, the wrap-around motion compensation offset is received, and the wrap-around motion compensation offset is associated with the picture. Perform motion compensation on the picture according to the first wrap-around motion compensation flag and wrap-around motion compensation offset, A system configured to execute a set of instructions, causing the system to perform a specific action. 12. The processor, Receiving a second wrap-around motion compensation flag, wherein the second wrap-around motion compensation flag is associated with a set of pictures including the picture associated with the first wrap-around motion compensation flag, Determine whether the second wrap-around motion compensation flag is disabled, In accordance with the determination that the second wrap-around motion compensation flag is invalid, the first wrap-around motion compensation flag is also determined to be invalid. The system described in Clause 11, further configured to execute a set of instructions, causing the system to perform the following actions. 13. The processor, Determining the picture width of the picture associated with the first wrap-around motion compensation flag, The first wrap-around motion compensation flag is determined based on the picture width, and The system described in Clause 12, further configured to execute a set of instructions, causing the system to perform the following actions. 14. The processor, Determine whether the Luma coding tree block size + 1 for the smallest coding block unit is greater than the picture width - 1 for the smallest coding block unit, The first motion compensation flag is invalid when it is determined that the Luma coding tree block size + 1 for the smallest coding block unit is greater than the picture width - 1 for the smallest coding block unit. The system described in Clause 13, further configured to execute a set of instructions, causing the system to perform the following actions. 15. The processor, Determine whether the second wrap-around motion compensation flag is enabled, In response to the determination that the second wrap-around motion compensation flag is enabled, it is determined that the picture width of the picture is greater than or equal to the Luma coding tree block size + offset, A system as described in any one of clauses 12 to 14, further configured to execute a set of instructions to cause the system to perform the following actions. 16. The processor, In response to the determination that the second wrap-around motion compensation flag is enabled, it is determined that the Luma coding tree block size + 1 for the smallest coding block unit is less than or equal to the picture width - 1 for the smallest coding block unit picture. The system described in Clause 15, further configured to execute a set of instructions, causing the system to perform the following actions. 17. The system described in any one of clauses 12 to 16, wherein a second wrap-around motion compensation flag is signaled in the sequence parameter set, and a first wrap-around motion compensation flag and a wrap-around motion compensation offset are signaled in the picture parameter set. 18. The processor, Perform motion compensation for multiple pictures. It is further configured to execute a set of instructions so that the system can perform the following actions: A system described in any one of clauses 11 to 17, in which multiple pictures have different sizes. 19. The processor, The second wrap-around motion compensation offset is determined by adding an offset to the wrap-around motion compensation offset received from the bitstream, Perform motion compensation on the picture according to the second wrap-around motion compensation offset, A system as described in any one of clauses 11 to 18, further configured to execute a set of instructions to cause the system to perform the following actions. 20. A non-temporary computer-readable medium for storing a set of instructions, wherein the set of instructions is executable by one or more processors of the device in order to cause the device to initiate a method for performing motion compensation, Receiving a first wrap-around motion compensation flag, wherein the first wrap-around motion compensation flag is associated with a picture in a set of pictures, To determine whether the first wrap-around motion compensation flag is enabled, In response to the determination that the first wrap-around motion compensation flag is valid, the wrap-around motion compensation offset is received, and the wrap-around motion compensation offset is associated with the picture. Perform motion compensation on the picture according to the first wrap-around motion compensation flag and wrap-around motion compensation offset, Non-temporary computer-readable media, including [specific examples of such media]. twenty one. A set of instructions Determining the picture width of the picture associated with the first wrap-around motion compensation flag, The first wrap-around motion compensation flag is determined based on the picture width, and A non-temporary computer-readable medium as described in Clause 20, which is executable by at least one processor of the computer system in order to further execute the computer system.
[0156]
[0190] Exemplary embodiments are disclosed in the drawings and specification. However, many variations and modifications can be made to these embodiments. Therefore, where certain terms are used, they are merely for general descriptive purposes and not for limiting purposes.
Claims
1. A method of performing motion compensation using a decoder, Receiving a first wrap-around motion compensation flag that is associated with one or more pictures and indicates whether horizontal wrap-around motion compensation is enabled for the one or more pictures, Depending on whether the horizontal wrap-around motion compensation is effective for one or more pictures, the parameter associated with the wrap-around motion compensation offset is received, wherein the parameter associated with the wrap-around motion compensation offset is associated with one or more pictures, the value of the parameter associated with the wrap-around motion compensation offset is less than or equal to the difference value -2, and the difference value is obtained by subtracting the quotient of the luma-encoded tree block size divided by the minimum luma-encoded block size from the quotient of the picture width of the luma sample divided by the minimum luma-encoded block size. A method comprising setting the first wrap-around motion compensation flag and the parameter associated with the wrap-around motion compensation offset in a picture parameter set.
2. The method according to claim 1, wherein the difference value is obtained by (pic_width_in_luma_samples / MinCbSizeY) - (CtbSizeY / MinCbSizeY), where pic_width_in_luma_samples is the picture width of the luma samples, MinCbSizeY is the minimum luma coding block size, and CtbSizeY is the luma coding tree block size.
3. Determining the picture width of the picture associated with the first wrap-around motion compensation flag, Determining whether the first wrap-around motion compensation flag is enabled based on the picture width, The method according to claim 1, further comprising:
4. Receiving a second wrap-around motion compensation flag, wherein the second wrap-around motion compensation flag is associated with a sequence of pictures including the one or more pictures. To determine whether the second wrap-around motion compensation flag is enabled, In accordance with the determination that the second wrap-around motion compensation flag is enabled, it is determined that the value obtained by adding 1 to the quotient of the Luma-encoded tree block size divided by the minimum Luma-encoded block size is less than or equal to the value obtained by subtracting 1 from the quotient of the picture width of the picture divided by the minimum Luma-encoded block size, The method according to claim 1, further comprising:
5. The method according to claim 4, wherein the second wrap-around motion compensation flag is set in the sequence parameter set.
6. The method according to claim 1, further comprising performing motion compensation on a plurality of pictures, wherein the plurality of pictures have different sizes.
7. Determining a second parameter associated with a second wrap-around motion compensation offset by adding an offset to the parameter associated with the wrap-around motion compensation offset received from the bitstream, Performing motion compensation on the picture according to the second parameter associated with the second wrap-around motion compensation offset, The method according to claim 1, further comprising:
8. A method of performing motion compensation using an encoder, Setting a first wrap-around motion compensation flag that is associated with one or more pictures and indicates whether horizontal wrap-around motion compensation is enabled for the one or more pictures, Depending on whether the horizontal wrap-around motion compensation is effective for one or more pictures, a parameter associated with the wrap-around motion compensation offset is set, wherein the parameter associated with the wrap-around motion compensation offset is associated with one or more pictures, the value of the parameter associated with the wrap-around motion compensation offset is less than or equal to the difference value -2, and the difference value is obtained by subtracting the quotient of the luma coding tree block size divided by the minimum luma coding block size from the quotient of the picture width of the luma sample divided by the minimum luma coding block size. A method comprising the first wrap-around motion compensation flag and the parameter being set in a picture parameter set.
9. The method according to claim 8, wherein the difference value is obtained by (pic_width_in_luma_samples / MinCbSizeY) - (CtbSizeY / MinCbSizeY), where pic_width_in_luma_samples is the picture width of the luma samples, MinCbSizeY is the minimum luma coding block size, and CtbSizeY is the luma coding tree block size.
10. Determining the picture width of the picture associated with the first wrap-around motion compensation flag, Determining whether the first wrap-around motion compensation flag is enabled based on the picture width, The method according to claim 8, further comprising:
11. Before setting the first wrap-around motion compensation flag associated with one or more pictures, the method may: Setting a second wrap-around motion compensation flag, wherein the second wrap-around motion compensation flag is associated with a sequence of pictures including the one or more pictures. To determine whether the second wrap-around motion compensation flag is enabled, In accordance with the determination that the second wrap-around motion compensation flag is enabled, it is determined that the value obtained by adding 1 to the quotient of the Luma-encoded tree block size divided by the minimum Luma-encoded block size is less than or equal to the value obtained by subtracting 1 from the quotient of the picture width of the picture divided by the minimum Luma-encoded block size, The method according to claim 8, further comprising:
12. The method according to claim 11, wherein the second wrap-around motion compensation flag is set in the sequence parameter set.
13. The method according to claim 8, further comprising performing motion compensation on a plurality of pictures, wherein the plurality of pictures have different sizes.
14. The second parameter associated with a second wrap-around motion compensation offset is determined by adding an offset to the parameter associated with the wrap-around motion compensation offset signaled in the bitstream, Performing motion compensation on the picture according to the second parameter associated with the second wrap-around motion compensation offset, The method according to claim 8, further comprising:
15. A method for storing a bitstream of a video sequence, Receiving a video sequence, Encoding one or more pictures of the aforementioned video sequence, The process of generating a bitstream based on the aforementioned encoding, The bitstream is stored in a non-temporary computer-readable storage medium, Includes, The above encoding is Setting a first wrap-around motion compensation flag that is associated with one or more pictures and indicates whether horizontal wrap-around motion compensation is enabled for the one or more pictures, In accordance with the validity of the horizontal wrap-around motion compensation for one or more pictures, the parameter associated with the wrap-around motion compensation offset is signaled such that the parameter associated with the wrap-around motion compensation offset is associated with one or more pictures, the value of the parameter associated with the wrap-around motion compensation offset is less than or equal to the difference value -2, and the difference value is obtained by subtracting the quotient of the luma-encoded tree block size divided by the minimum luma-encoded block size from the quotient of the picture width of the luma sample divided by the minimum luma-encoded block size. A method comprising the first wrap-around motion compensation flag and the parameter being set in a picture parameter set.
16. The method according to claim 15, wherein the difference value is obtained by (pic_width_in_luma_samples / MinCbSizeY) - (CtbSizeY / MinCbSizeY), where pic_width_in_luma_samples is the picture width of the luma samples, MinCbSizeY is the minimum luma coding block size, and CtbSizeY is the luma coding tree block size.
17. Before setting the first wrap-around motion compensation flag associated with one or more pictures, the encoding is performed Setting a second wrap-around motion compensation flag, wherein the second wrap-around motion compensation flag is associated with a sequence of pictures including the one or more pictures. To determine whether the second wrap-around motion compensation flag is enabled, In accordance with the determination that the second wrap-around motion compensation flag is enabled, it is determined that the value obtained by adding 1 to the quotient of the Luma-encoded tree block size divided by the minimum Luma-encoded block size is less than or equal to the value obtained by subtracting 1 from the quotient of the picture width of the picture divided by the minimum Luma-encoded block size. The method according to claim 15, further comprising:
18. The method according to claim 17, wherein the second wrap-around motion compensation flag is set in the sequence parameter set.
19. The method according to claim 15, further comprising performing motion compensation on a plurality of pictures, wherein the plurality of pictures have different sizes.
20. The above encoding is The second parameter associated with the second wrap-around motion compensation offset is determined by adding an offset to the parameter associated with the wrap-around motion compensation offset signaled in the bitstream, Performing motion compensation on the picture according to the second parameter associated with the second wrap-around motion compensation offset, The method according to claim 15, further comprising:
Citation Information
Patent Citations
360-degree panorama video encoding method, encoding device, and computer program
JP2018534827A