Method of signaling video encoded data
By receiving the surrounding motion compensation flag and image width difference value, performing motion compensation, and signaling specific variables in the image parameter set, the problem of inefficient signal notification of surrounding motion compensation and strip layout information in the prior art is solved, and more efficient video encoding is achieved.
Patent Information
- Application Number
- CN202510317570.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2020-03-26
- Filing Date
- 2021-03-26
- Publication Date
- 2025-05-30
- Estimated Expiration
- 2041-03-26
AI Technical Summary
Existing video encoding technologies have inefficient problems in surrounding motion compensation, strip layout and information signaling notification of strip addresses.
By receiving the surround motion compensation flag, it is determined whether surround motion compensation is enabled, and motion compensation is performed in response to the difference data between the received image width and the horizontal surround position offset when enabled. Furthermore, in the image parameter set, variables indicating the number of stripes in the video frame minus 2 or subtract the number of sub-images minus 1 are signaled to more efficiently notify the video coded data.
It improves the surround motion compensation efficiency during video encoding, optimizes the strip layout and signal notification of strip addresses, reduces the size of the bitstream, and improves the encoding performance.
Smart Images

Figure CN120075459A_ABST
Abstract
Description
Cross - reference to related applications
[0001] This disclosure claims the benefit and priority of U.S. Provisional Patent Application No. 63 / 000,443, filed on March 26, 2020. The provisional application is incorporated herein by reference in its entirety. Technical Field
[0002] This disclosure generally relates to video data processing, and more particularly, to methods and apparatuses for signaling information regarding circular motion compensation, stripe layout, and stripe addresses. Background Art
[0003] Video is a set of static images (or "frames") that capture visual information. To reduce storage memory and transmission bandwidth, video can be compressed before storage or transmission and then decompressed before display. The compression process is generally referred to as encoding, and the decompression process is generally referred to as decoding. There are various video coding formats using standardized video coding techniques, most commonly based on prediction, transformation, quantization, entropy coding, and in - loop filtering. Standardization organizations have developed video coding standards, such as the High Efficiency Video Coding (HEVC / H.265) standard, the Versatile Video Coding (VVC / H.266) standard, and the AVS standard, which specify specific video coding formats. As more and more advanced video coding techniques are adopted in video standards, the coding efficiency of new video coding standards is getting higher and higher. Summary of the Invention
[0004] Embodiments of the present disclosure provide a method for signaling video coding data, the method comprising: receiving a circular motion compensation flag; determining whether to enable circular motion compensation based on the circular motion compensation flag; in response to determining to enable the circular motion compensation, receiving data indicating a difference between the width of an image and an offset for determining a horizontal circular position; and performing motion compensation according to the circular motion compensation flag and the difference.
[0005] Embodiments of the present disclosure also provide a method for signaling video coding data, the method comprising: receiving an image for encoding, wherein the image includes one or more stripes; and signaling, in an image parameter set of the image, a variable indicating the number of stripes in a video frame minus 2.
[0006] Embodiments of the present disclosure also provide a method for signaling video coding data, the method comprising: receiving an image for encoding, wherein the image includes one or more stripes and one or more sub - images; and signaling, in an image parameter set of the image, a variable indicating the number of stripes in the image minus the number of sub - images in the image minus 1.
[0007] Embodiments of the present disclosure also provide a method for signaling video coding data, the method comprising: receiving an image to be encoded, wherein the image includes one or more slices; signaling a variable indicating whether an image header syntax structure of the image exists in the slice headers of the one or more slices; and signaling a slice address according to the variable.
[0008] Embodiments of the present disclosure also provide a system for performing video data processing, the system comprising: a memory storing an instruction set; and a processor configured to execute the instruction set to cause the system to perform: receiving a loop motion compensation flag; determining whether to enable loop motion compensation based on the loop motion compensation flag; in response to determining to enable the loop motion compensation, receiving data indicating a difference between a width of an image and an offset for determining a horizontal loop position; and performing motion compensation according to the loop motion compensation flag and the difference.
[0009] Embodiments of the present disclosure also provide a system for performing video data processing, the system comprising: a memory storing a set of instructions; and a processor configured to execute the set of instructions to cause the system to perform: receiving an image to be encoded, wherein the image includes one or more slices; and in an image parameter set of the image, signaling a variable indicating a number of slices in a video frame minus 2.
[0010] Embodiments of the present disclosure also provide a system for performing video data processing, the system comprising: a memory storing an instruction set; and a processor configured to execute the instruction set to cause the system to perform: receiving an image to be encoded, wherein the image includes one or more slices and one or more sub-images; and in an image parameter set of the image, signaling a variable indicating a number of slices in the image minus a number of sub-images in the image minus 1.
[0011] Embodiments of the present disclosure also provide a system for performing video data processing, the system comprising: a memory storing an instruction set; and a processor configured to execute the instruction set to cause the system to perform: receiving an image to be encoded, wherein the image includes one or more slices; signaling a variable indicating whether an image header syntax structure of the image exists in the slice headers of the one or more slices; and signaling a slice address according to the variable.
[0012] Embodiments of the present disclosure also provide a non-transitory computer-readable medium storing an instruction set, the instruction set being executable by one or more processors of a device to cause the device to initiate a method for performing video data processing, the method including: receiving a wrap-around motion compensation flag; determining whether to enable wrap-around motion compensation based on the wrap-around motion compensation flag; in response to determining to enable the wrap-around motion compensation, receiving data indicating a difference between a width of an image and an offset for determining a horizontal wrap-around position; and performing motion compensation according to the wrap-around motion compensation flag and the difference.
[0013] Embodiments of the present disclosure also provide a non-transitory computer-readable medium storing an instruction set, the instruction set being executable by one or more processors of a device to cause the device to initiate a method for performing video data processing, the method including: receiving an image for encoding, where the image includes one or more slices; and signaling, in an image parameter set of the image, a variable indicating a number of slices in a video frame minus 2.
[0014] Embodiments of the present disclosure also provide a non-transitory computer-readable medium storing an instruction set, the instruction set being executable by one or more processors of a device to cause the device to initiate a method for performing video data processing, the method including: receiving an image for encoding, where the image includes one or more slices and one or more sub-images; and signaling, in an image parameter set of the image, a variable indicating a number of slices in the image minus a number of sub-images in the image minus 1.
[0015] Embodiments of the present disclosure also provide a non-transitory computer-readable medium storing an instruction set, the instruction set being executable by one or more processors of a device to cause the device to initiate a method for performing video data processing, the method including: receiving an image for encoding, where the image includes one or more slices; signaling a variable indicating whether an image header syntax structure exists in a slice header of the one or more slices; and signaling a slice address according to the variable. BRIEF DESCRIPTION OF THE DRAWINGS
[0016] Embodiments and aspects of the present disclosure are shown in the following detailed description and the drawings. The various features shown in the drawings are not drawn to scale.
[0017] Figure 1 The structure of an example video sequence according to some embodiments of the present disclosure is shown.
[0018] Figure 2A A schematic diagram of an exemplary encoding process according to some embodiments of the present disclosure is shown.
[0019] Figure 2BShows a schematic diagram of another example encoding process according to some embodiments of the present disclosure.
[0020] Figure 3A Shows a schematic diagram of an exemplary decoding process according to some embodiments of the present disclosure.
[0021] Figure 3B Shows a schematic diagram of another example decoding process according to some embodiments of the present disclosure.
[0022] Figure 4 Shows a block diagram of an example apparatus for encoding or decoding video according to some embodiments of the present disclosure.
[0023] Figure 5A Shows a schematic diagram of an example hybrid operation for generating a reconstructed equirectangular projection according to some embodiments of the present disclosure.
[0024] Figure 5B Shows a schematic diagram of an example cropping operation for generating a reconstructed equirectangular projection according to some embodiments of the present disclosure.
[0025] Figure 6A Shows a schematic diagram of an exemplary horizontal wrap-around motion compensation process for equirectangular projection according to some embodiments of the present disclosure.
[0026] Figure 6B Shows a schematic diagram of an exemplary horizontal wrap-around motion compensation process for filling an equirectangular projection according to some embodiments of the present disclosure.
[0027] Figure 7 Shows the syntax of an example advanced wrap-around offset according to some embodiments of the present disclosure.
[0028] Figure 8 Shows the semantics of an example advanced wrap-around offset according to some embodiments of the present disclosure.
[0029] Figure 9 Shows a schematic diagram of an example strip and sub-image partitioning of an image according to some embodiments of the present disclosure.
[0030] Figure 10 Shows a schematic diagram of an example strip and sub-image partitioning of an image with different strips and sub-images according to some embodiments of the present disclosure.
[0031] Figure 11 Shows the syntax of an example picture parameter set for slice mapping and strip layout according to some embodiments of the present disclosure.
[0032] Figure 12 shows the semantics of an example picture parameter set for slice mapping and strip layout according to some embodiments of the present disclosure.
[0033] Figure 13 Shows the syntax of an example strip header according to some embodiments of the present disclosure.
[0034] Figure 14 Shows the semantics of an example strip header according to some embodiments of the present disclosure.
[0035] Figure 15 Shows the syntax of an example improved picture parameter set according to some embodiments of the present disclosure.
[0036] Figure 16 Shows the semantics of an example improved picture parameter set according to some embodiments of the present disclosure.
[0037] Figure 17 Shows the syntax of an example picture parameter set with variable wraparound_offset_type according to some embodiments of the present disclosure.
[0038] Figure 18 Shows the semantics of an example improved picture parameter set with variable wraparound_offset_type according to some embodiments of the present disclosure.
[0039] Figure 19 Shows the syntax of an example picture parameter set with variable num_slices_in_pic_minus2 according to some embodiments of the present disclosure.
[0040] Figure 20 shows the semantics of an example improved picture parameter set with variable num_slices_in_pic_minus2 according to some embodiments of the present disclosure.
[0041] Figure 21 Shows the syntax of an example picture parameter set with variable num_slices_in_pic_minus_subpic_num_minus1 according to some embodiments of the present disclosure.
[0042] Figure 22 shows the semantics of an example improved picture parameter set with variable num_slices_in_pic_minus_subpic_num_minus1 according to some embodiments of the present disclosure.
[0043] Figure 23 Shows the syntax of an example updated strip header according to some embodiments of the present disclosure.
[0044] Figure 24A flowchart of an example video coding method according to some embodiments of the present disclosure is shown, the method having a variable that signals a difference between a width of a video frame and an offset used to calculate a horizontal wrap-around position.
[0045] Figure 25 A flowchart of an example video coding method according to some embodiments of the present disclosure is shown, the method having a variable that signals a number of stripes in a video frame minus 2.
[0046] Figure 26 A flowchart of an example video coding method according to some embodiments of the present disclosure is shown, having a variable, wherein the variable signals a variable that indicates a number of stripes in a video frame minus a number of sub-images in the video frame minus 1.
[0047] Figure 27 A flowchart of an example video coding method according to some embodiments of the present disclosure is shown, having a variable that indicates whether an image header syntax structure exists in a stripe header of a video frame. Detailed Description
[0048] Reference will now be made in detail to the exemplary embodiments, examples of which are illustrated in the accompanying drawings. The following description refers to the accompanying drawings unless otherwise noted, where like numbers in different drawings represent the same or similar elements. The embodiments set forth in the following description of the exemplary embodiments do not represent all embodiments consistent with the present disclosure. Instead, they are merely examples of apparatus and methods consistent with aspects related to the present disclosure as set forth in the appended claims. Specific aspects of the present disclosure are described in more detail below. If there is a conflict with terms and / or definitions incorporated by reference, the terms and definitions provided herein shall prevail.
[0049] The Joint Video Exploration Team (JVET) of the ITU-T Video Coding Experts Group (ITU-T VCEG) and the ISO / IEC Moving Picture Experts Group (ISO / IEC MPEG) is currently developing the Versatile Video Coding (VVC / H.266) standard. The VVC standard aims to double the compression efficiency of its predecessor, the High Efficiency Video Coding (HEVC / H.265) standard. In other words, the goal of VVC is to achieve the same subjective quality as HEVC / H.265 using half the bandwidth.
[0050] To achieve the same subjective quality as HEVC / H.265 using half the bandwidth, JVET has been using the Joint Exploration Model (JEM) reference software to explore technologies beyond HEVC. As coding techniques are incorporated into JEM, JEM has achieved higher coding performance than HEVC.
[0051] The VVC standard has been recently developed and continues to include more encoding techniques that provide better compression performance. VVC is based on the hybrid video coding system that has been used in modern video compression standards such as HEVC, H.264 / AVC, MPEG2, H.263, etc.
[0052] Video is a set of static images (or "frames") arranged in chronological order to store visual information. These images can be acquired and stored in chronological order using a video acquisition device (e.g., a camera), and such images in the time series can be displayed using a video playback device (e.g., a television, a computer, a smartphone, a tablet computer, a video player, or any end-user terminal with a display function). In addition, in some applications, the video acquisition device can send the acquired video in real time to the video playback device (e.g., a computer with a monitor), such as for surveillance, conferencing, or live broadcasting.
[0053] To reduce the storage space and transmission bandwidth required for such applications, the video can be compressed before storage and transmission and decompressed before display. Compression and decompression can be implemented by software executed by a processor (e.g., the processor of a general-purpose computer) or dedicated hardware. The module for compression is usually referred to as an "encoder", and the module for decompression is usually referred to as a "decoder". Encoders and decoders can be collectively referred to as "codecs". Encoders and decoders can be implemented as any one of various suitable hardware, software, or a combination thereof. For example, the hardware implementation of an encoder and a decoder can include circuits such as one or more microprocessors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), discrete logic, or any combination thereof. The software implementation of an encoder and a decoder can include program code fixed in a computer-readable medium, computer-executable instructions, firmware, or any suitable computer-implemented algorithm or process. Video compression and decompression can be achieved through various algorithms or standards, such as MPEG-1, MPEG-2, MPEG-4, the H.26x series, etc. In some applications, a codec can decompress a video from a first coding standard and recompress the decompressed video using a second coding standard, in which case the codec can be referred to as a "transcoder".
[0054] The video encoding process can identify and retain useful information that can be used to reconstruct an image and ignore unimportant reconstruction information. If the unimportant information cannot be completely reconstructed after being ignored, such an encoding process can be called "lossy". Otherwise, it can be called "lossless". Most encoding processes are lossy, which is a trade-off to reduce the required storage space and transmission bandwidth.
[0055] In many cases, the useful information of the encoded image (referred to as the "current image") includes changes relative to a reference image (e.g., a previously encoded and reconstructed image). Such changes can include changes in the position of pixels, changes in brightness, or changes in color, with changes in position being of the most concern. The change in the position of a group of pixels representing an object can reflect the movement of the object between the reference image and the current image.
[0056] An image encoded without reference to another image (i.e., it is its own reference image) is referred to as an "I-image". An image encoded using a previous image as a reference image is referred to as a "P-image", and an image encoded using a previous image and a future image as reference images is referred to as a "B-image" (the reference is "bidirectional").
[0057] Figure 1 The structure of an example video sequence 100 according to some embodiments of the present disclosure is shown. The video sequence 100 can be a live video or a video that has been captured and archived. The video 100 can be a real-life video, a computer-generated video (e.g., a computer game video), or a combination of both (e.g., a real video with augmented reality effects). The video sequence 100 can be input from a video capture device (e.g., a camera), a video archive containing previously captured videos (e.g., a video file stored in a storage device), or a video feed interface (e.g., a video broadcast transceiver) that receives video from a video content provider.
[0058] As Figure 1 shown, the video sequence 100 can include a series of images arranged in time along a timeline, including images 102, 104, 106, and 108. Images 102 - 106 are consecutive, and there are more images between images 106 and 108. In Figure 1 this example, image 102 is an I-image, and its reference image is image 102 itself. Image 104 is a P-image, and its reference image is image 102, as indicated by the arrow. Image 106 is a B-image, and its reference images are image 104 and 108, as indicated by the arrows. In some embodiments, the reference image of an image (e.g., image 104) may not be immediately before or after the image. For example, the reference image of image 104 can be an image before image 102. It should be noted that the reference images of images 102 - 106 are merely examples, and the present disclosure does not limit the embodiments of the reference images as Figure 1 shown.
[0059] Typically, due to the computational complexity of the encoding and decoding tasks, video codecs do not encode or decode an entire image at once. Instead, they can divide the image into basic segments and encode or decode the image segments segment by segment. In the present disclosure, such a basic segment is referred to as a basic processing unit ("BPU"). For example, Figure 1The structure 110 therein shows an example structure of an image (e.g., any of the images 102 - 108) of the video sequence 100. In the structure 110, the image is divided into 4×4 basic processing units, the boundaries of which are shown as dashed lines. In some embodiments, the basic processing unit may be referred to as a "macroblock" in some video coding standards (e.g., the MPEG family, H.261, H.263, or H.264 / AVC), or as a "coding tree unit" ("CTU") in some other video coding standards (e.g., H.265 / HEVC or H.266 / VVC). The basic processing unit may have a variable size in the image, such as 128×128, 64×64, 32×32, 16×16, 4×8, 16×32, or pixels of any shape and size. The size and shape of the basic processing unit can be selected for the image based on a balance between coding efficiency and the level of detail to be maintained in the basic processing unit..
[0060] The basic processing unit may be a logical unit that may include a set of different types of video data stored in a computer memory (e.g., in a video frame buffer). For example, the basic processing unit of a color image may include a luminance component (Y) representing achromatic luminance information, one or more chrominance components (e.g., Cb and Cr) representing color information, and associated syntax elements, where the luminance and chrominance components may have the same size as the basic processing unit. In some video coding standards (e.g., H.265 / HEVC or H.266 / VVC), the luminance and chrominance components may be referred to as "coding tree blocks" ("CTB"). Any operation performed on the basic processing unit may be repeated for each of its luminance and chrominance components.
[0061] Video coding has multiple operation stages, examples of which are Figures 2A - 2B and Figures 3A - 3BAs shown. For each stage, the size of the basic processing unit may still be too large for processing, so it can be further divided into segments called "basic processing subunits" in the present disclosure. In some embodiments, the basic processing subunit may be called a "block" in some video coding standards (e.g., the MPEG family, H.261, H.263, or H.264 / AVC), or an "encoding unit" ("CU") in some other video coding standards (e.g., H.265 / HEVC or H.266 / VVC). The basic processing subunit may have the same size as the basic processing unit or a smaller size than the basic processing unit. Similar to the basic processing unit, the basic processing subunit is also a logical unit, which may include a set of different types of video data (e.g., Y, Cb, Cr, and associated syntax elements) stored in a computer memory (e.g., in a video frame buffer). Any operation performed on the basic processing subunit can be repeated for each of its luminance and chrominance components. It should be noted that this division can be performed to further levels according to processing needs. It should also be noted that different stages may use different schemes to divide the basic processing unit.
[0062] For example, in the mode decision stage (an example of which is shown in Figure 2B ), the encoder can decide what prediction mode (e.g., intra prediction or inter prediction) to use for the basic processing unit, which may be too large to make such a decision. The encoder can divide the basic processing unit into multiple basic processing subunits (e.g., CUs in H.265 / HEVC or H.266 / VVC), and decide the prediction type for each individual basic processing subunit.
[0063] For another example, in the prediction stage (an example of which is shown in Figures 2A - 2B ), the encoder can perform prediction operations at the level of the basic processing subunit (e.g., CU). However, in some cases, the basic processing subunit may still be too large to process. The encoder can further divide the basic processing subunit into smaller segments (e.g., called "prediction blocks" or "PBs" in H.265 / HEVC or H.266 / VVC), at which level the prediction operations can be performed.
[0064] For another example, in the transform stage (an example of which is shown in Figures 2A - 2BAs shown (in [description]), the encoder may perform a transformation operation on a residual basic processing subunit (e.g., a CU). However, in some cases, the basic processing subunit may still be too large to process. The encoder may further divide the basic processing subunit into smaller segments (e.g., called "transformation blocks" or "TBs" in H.265 / HEVC or H.266 / VVC), at which level the transformation operation can be performed. It should be noted that the partitioning scheme of the same basic processing subunit may be different in the prediction stage and the transformation stage. For example, in H.265 / HEVC or H.266 / VVC, the prediction blocks and transformation blocks of the same CU may have different sizes and numbers.
[0065] In Figure 1 the structure 110, the basic processing unit 112 is further divided into 3×3 basic processing subunits, the boundaries of which are shown as dashed lines. Different basic processing units of the same image may be divided into basic processing subunits in different schemes.
[0066] In some embodiments, to provide the ability for parallel processing and fault tolerance for video encoding and decoding, an image may be divided into regions for processing such that for a region of the image, the encoding or decoding process may not depend on information from any other region of the image. In other words, each region of the image can be processed independently. By doing so, the codec can process different regions of the image in parallel, thereby improving the encoding efficiency. Additionally, when the data of a region is corrupted during processing or lost during network transmission, the codec can correctly encode or decode other regions of the same image without depending on the corrupted or lost data, thereby providing fault tolerance. In certain video coding standards, an image may be divided into different types of regions. For example, H.265 / HEVC and H.266 / VVC provide two types of regions: "slices" and "tiles". It should also be noted that different images of the video sequence 100 may have different partitioning schemes for dividing the image into regions.
[0067] For example, in Figure 1 the structure 110 is divided into three regions 114, 116, and 118, the boundaries of which are shown as solid lines inside the structure 110. Region 114 includes four basic processing units. Regions 116 and 118 each include six basic processing units. It should be noted that Figure 1 the basic processing units, basic processing subunits, and structural regions in 110 are only examples, and the present disclosure does not limit its embodiments.
[0068] Figure 2A shows a schematic diagram of an exemplary encoding process 200A according to an embodiment of the present disclosure. For example, the encoding process 200A may be performed by an encoder. As Figure 2AAs shown, the encoder may encode video sequence 202 into video bitstream 228 according to process 200A. Similar to Figure 1 the video sequence 100 in Figure 1 , the video sequence 202 may include a set of images arranged in chronological order (referred to as "original images"). Similar to
[0069] the structure 110 in Figure 2A , each original image of the video sequence 202 may be divided by the encoder into basic processing units, basic processing subunits, or regions for processing. In some embodiments, the encoder may execute process 200A at the level of basic processing units for each original image of the video sequence 202. For example, the encoder may execute process 200A in an iterative manner, where the encoder may encode a basic processing unit in one iteration of process 200A. In some embodiments, the encoder may execute process 200A in parallel for regions (e.g., regions 114-118) of each original image of the video sequence 202.
[0070] Referring to , the encoder may feed the basic processing unit of the original image of the video sequence 202 (referred to as "original BPU") to prediction stage 204 to generate prediction data 206 and prediction BPU 208. The encoder may subtract the predicted BPU 208 from the original BPU to generate residual BPU 210. The encoder may feed the residual BPU 210 to transform stage 212 and quantization stage 214 to 216 generate quantized transform coefficients 216. The encoder may feed the prediction data 206 and the quantized transform coefficients 216 to binary coding stage 226 to generate video bitstream 228. Components 202, 204, 206, 208, 210, 212, 214, 216, 226, and 228 may be referred to as the "forward path". During process 200A, after quantization stage 214, the encoder may feed the quantized transform coefficients 216 to inverse quantization stage 218 and inverse transform stage 220 to generate reconstructed residual BPU 222. The encoder may add the reconstructed residual BPU 222 to the predicted BPU 208 to generate prediction reference 224, which is used in prediction stage 204 for the next iteration of process 200A. Components 218, 220, 222, and 224 of process 200A may be referred to as the "reconstruction path". The reconstruction path may be used to ensure that both the encoder and the decoder use the same reference data for prediction.
[0070] The encoder may iteratively execute process 200A to encode each original BPU (in the forward path) of the encoded original image and generate prediction reference 224 for the next original BPU (in the reconstruction path) of the encoded original image. After encoding all the original BPUs of the original image, the encoder may continue to encode the next image in the video sequence 202.
[0071] Referring to process 200A, the encoder may receive video sequence 202 generated by a video capture device (e.g., a camera). As used herein, the term "receive" may refer to any action of receiving, inputting, obtaining, retrieving, fetching, reading, accessing, or using for inputting data in any way.
[0072] In prediction stage 204, at the current iteration, the encoder may receive the original BPU and prediction reference 224 and perform a prediction operation to generate prediction data 206 and prediction BPU 208. The prediction reference 224 may be generated from the reconstruction path of a previous iteration of process 200A. The purpose of prediction stage 204 is to reduce information redundancy by extracting prediction data 206 from prediction data 206 and prediction reference 224 that can be used to reconstruct the original BPU into prediction BPU 208.
[0073] Ideally, the predicted BPU 208 may be the same as the original BPU. However, due to non-ideal prediction and reconstruction operations, the predicted BPU 208 is typically slightly different from the original BPU. To record these differences, when generating prediction BPU 208, the encoder may subtract it from the original BPU to generate residual BPU 210. For example, the encoder may subtract the value of the corresponding pixel of prediction BPU 208 (e.g., grayscale value or RGB value) from the value of the pixel of the original BPU. Each pixel of residual BPU 210 may have a residual value as the result of such subtraction between the corresponding pixels of the original BPU and prediction BPU 208. Compared with the original BPU, prediction data 206 and residual BPU 210 may have fewer bits, but they can be used to reconstruct the original BPU without significant quality degradation. Thus, the original BPU is compressed.
[0074] To further compress residual BPU 210, in transform stage 212, the encoder may reduce its spatial redundancy by decomposing residual BPU 210 into a set of two-dimensional "base patterns". Each base pattern is associated with a "transformation coefficient". The base patterns may have the same size (e.g., the size of residual BPU 210), and each base pattern may represent a component of the change frequency (e.g., the frequency of brightness change) of residual BPU 210. None of the base patterns can be reproduced from any combination (e.g., linear combination) of any other base patterns. In other words, the decomposition can decompose the change of residual BPU 210 into the frequency domain. This decomposition is similar to the discrete Fourier transform of a function, where the base image is similar to the basic function of the discrete Fourier transform (e.g., trigonometric function), and the transformation coefficient is similar to the coefficient associated with the basic function.
[0075] Different transformation algorithms can use different basic patterns. Various transformation algorithms can be used at the transformation stage 212, such as, for example, the discrete cosine transform, the discrete sine transform, etc. The transformation at the transformation stage 212 is reversible. That is, the encoder can recover the residual BPU 210 through the inverse operation of the transformation (referred to as "inverse transformation"). For example, in order to recover the pixels of the residual BPU 210, the inverse transformation can be to multiply the values of the corresponding pixels of the basic pattern by the corresponding correlation coefficients and sum the products to produce a weighted sum. For video coding standards, both the encoder and the decoder can use the same transformation algorithm (and thus have the same basic pattern). Therefore, the encoder can record only the transformation coefficients, and the decoder can reconstruct the residual BPU 210 from them without receiving the basic pattern from the encoder. Compared with the residual BPU 210, the transformation coefficients can have fewer bits, but they can be used to reconstruct the residual BPU 210 without significant quality degradation. Thus, the residual BPU 210 is further compressed.
[0076] The encoder can further compress the transformation coefficients at the quantization stage 214. During the transformation process, different basic patterns can represent different change frequencies (e.g., luminance change frequencies). Since the human eye is generally better at recognizing low-frequency changes, the encoder can ignore the information of high-frequency changes without causing significant quality degradation in decoding. For example, at the quantization stage 214, the encoder can generate the quantized transformation coefficients 216 by dividing each transformation coefficient by an integer value (referred to as "quantization parameter") and rounding the quotient to its nearest integer. After such an operation, some transformation coefficients of the high-frequency basic pattern can be converted to zero, and the transformation coefficients of the low-frequency basic pattern can be converted to smaller integers. The encoder can ignore the quantized transformation coefficients 216 with zero values, whereby the transformation coefficients are further compressed. This quantization process is also reversible, where the quantized transformation coefficients 216 can be reconstructed as transformation coefficients in the inverse operation of quantization (referred to as "inverse quantization").
[0077] Since the encoder ignores the remainder of this division in the rounding operation, the quantization stage 214 can be lossy. Generally, the quantization stage 214 can contribute the most information loss in the process 200A. The greater the information loss, the fewer bits required for the quantized transformation coefficients 216. To obtain different levels of information loss, the encoder can use different quantization parameter values or any other parameters of the quantization process.
[0078] In the binary coding stage 226, the encoder may use binary coding techniques to code the prediction data 206 and the quantized transform coefficients 216. The binary coding may be, for example, entropy coding, variable length coding, arithmetic coding, Huffman coding, context - adaptive binary arithmetic coding, or any other lossless or lossy compression algorithm. In some embodiments, in addition to the prediction data 206 and the quantized transform coefficients 216, the encoder may code other information in the binary coding stage 226. For example, the prediction mode used in the prediction stage 204, the parameters of the prediction operation, the type of transform at the transform stage 212, the parameters of the quantization process (e.g., quantization parameters), the encoder control parameters (e.g., bit - rate control parameters), etc. The encoder may use the output data of the binary coding stage 226 to generate the video bitstream 228. In some embodiments, the video bitstream 228 may be further packed for network transmission.
[0079] Referring to the reconstruction path of process 200A, in the inverse quantization stage 218, the encoder may perform inverse quantization on the quantized transform coefficients 216 to generate the reconstructed transform coefficients. In the inverse transform stage 220, the encoder may generate the reconstructed residual BPU 222 based on the reconstructed transform coefficients. The encoder may add the reconstructed residual BPU 222 to the prediction BPU 208 to generate the prediction reference 224 that will be used in the next iteration of process 200A.
[0080] It should be noted that other variants of process 200A may be used to code the video sequence 202. In some embodiments, the stages of process 200A may be executed by the encoder in a different order. In some embodiments, one or more stages of process 200A may be combined into a single stage. In some embodiments, a single stage of process 200A may be divided into multiple stages. For example, the transform stage 212 and the quantization stage 214 may be combined into a single stage. In some embodiments, process 200A may include additional stages. In some embodiments, process 200A may omit Figure 2A one or more of the
[0081] Figure 2B FIG. shows a schematic diagram of another exemplary coding process 200B according to an embodiment of the present disclosure. Process 200B may be modified from process 200A. For example, process 200B may be used by an encoder compliant with a hybrid video coding standard (e.g., the H.26x series). Compared with process 200A, the forward path of process 200B further includes a mode decision stage 230 and divides the prediction stage 204 into a spatial prediction stage 2042 and a temporal prediction stage 2044. The reconstruction path of process 200B further includes a loop filter stage 232 and a buffer 234.
[0082] Generally, prediction techniques can be divided into two types: spatial prediction and temporal prediction. Spatial prediction (e.g., intra-image prediction or "intra prediction") can use pixels from one or more already-encoded adjacent BPUs in the same image to predict the current BPU. That is, the prediction reference 224 in spatial prediction can include adjacent BPUs. Spatial prediction can reduce the spatial redundancy inherent in the image. Temporal prediction (e.g., inter-image prediction or "inter prediction") can use regions from one or more already-encoded images to predict the current BPU. That is, the prediction reference 224 in temporal prediction can include encoded images. Temporal prediction can reduce the temporal redundancy inherent in the image.
[0083] Referring to process 200B, in the forward path, the encoder performs prediction operations at the spatial prediction stage 2042 and the temporal prediction stage 2044. For example, at the spatial prediction stage 2042, the encoder can perform intra prediction. For the original BPU of the encoded image, the prediction reference 224 can include one or more adjacent BPUs that have been encoded (in the forward path) and reconstructed (in the reconstruction path) in the same image. The encoder can generate the predicted BPU 208 by interpolating the adjacent BPUs. The interpolation technique can include, for example, linear interpolation or interpolation, polynomial interpolation or interpolation, etc. In some embodiments, the encoder can perform interpolation at the pixel level, e.g., by interpolating the values of the corresponding pixels of each pixel of the predicted BPU 208. The adjacent BPUs used for interpolation can be located in various directions relative to the original BPU, e.g., in the vertical direction (e.g., at the top of the original BPU), the horizontal direction (e.g., to the left of the original BPU), the diagonal direction (e.g., bottom left, bottom right, top left, or top right of the original BPU), or any direction defined in the video coding standard being used. For intra prediction, the prediction data 206 can include, for example, the positions (e.g., coordinates) of the adjacent BPUs used, the sizes of the adjacent BPUs used, the parameters of the interpolation, the direction of the adjacent BPUs relative to the original BPU, etc.
[0084] For another example, during the temporal prediction stage 2044, the encoder may perform inter-frame prediction. For the original BPU of the current image, the prediction reference 224 may include one or more images (referred to as "reference images") that have been encoded (in the forward path) and reconstructed (in the reconstruction path). In some embodiments, the reference images may be encoded and reconstructed on a per-BPU basis. For example, the encoder may add the reconstructed residual BPU 222 to the prediction BPU 208 to generate a reconstructed BPU. When all the reconstructed BPUs of the same image have been generated, the encoder may generate a reconstructed image as a reference image. The encoder may perform an operation of "motion estimation" to search for a matching region within a range of the reference image (referred to as the "search window"). The position of the search window in the reference image may be determined based on the position of the original BPU in the current image. For example, the search window may be centered at a position in the reference image that has the same coordinates as the original BPU in the current image and may extend outward a predetermined distance. When the encoder identifies (e.g., by using a pel recursive algorithm, a block matching algorithm, etc.) a region in the search window that is similar to the original BPU, the encoder may determine such a region as the matching region. The matching region may have a different size (e.g., smaller than, equal to, larger than, or a different shape) from the original BPU. Since the reference image and the current image are temporally separated on the timeline (e.g., as Figure 1 shown), the matching region may be considered to "move" over time to the position of the original BPU. The encoder may record the direction and distance of this motion as a "motion vector". When multiple reference images are used (e.g., as in Figure 1 image 106), the encoder may search for the matching region and determine its associated motion vector for each reference image. In some embodiments, the encoder may assign weights to the pixel values of the matching regions of the respective matching reference images.
[0085] Motion estimation can be used to identify various types of motion, such as translation, rotation, scaling, etc. For inter-frame prediction, the prediction data 206 may include, for example, the position (e.g., coordinates) of the matching region, the motion vector associated with the matching region, the number of reference images, the weights associated with the reference images, etc.
[0086] To generate the predicted BPU 208, the encoder may perform an operation of "motion compensation". Motion compensation can be used to reconstruct the predicted BPU 208 based on the prediction data 206 (e.g., the motion vector) and the prediction reference 224. For example, the encoder may move the matching region of the reference image according to the motion vector, where the encoder may predict the original BPU of the current image. When multiple reference images are used (e.g., as in Figure 1In the image 106), the encoder can move the matching region of the reference image according to the respective motion vectors and average pixel values of the matching regions. In some embodiments, if the encoder has assigned weights to the pixel values of the matching regions of the respective matching reference images, the encoder can add the weighted sums of the pixel values of the moved matching regions.
[0087] In some embodiments, the inter-frame prediction can be unidirectional or bidirectional. Unidirectional inter-frame prediction can use one or more reference images in the same time direction relative to the current image. For example, Figure 1 the image 104 in is a unidirectional inter-frame prediction image, where the reference image (i.e., image 102) is before image 04. Bidirectional inter-frame prediction can use one or more reference images in two time directions relative to the current image. For example, Figure 1 the image 106 in is a bidirectional inter-frame prediction image, where the reference images (i.e., images 104 and 08) are in two time directions relative to image 104.
[0088] Still referring to the forward path of process 200B, after the spatial prediction 2042 and the temporal prediction stage 2044, at the mode decision stage 230, the encoder can select a prediction mode (e.g., one of intra-frame prediction or inter-frame prediction) for the current iteration of process 200B. For example, the encoder can perform rate-distortion optimization techniques, where the encoder can select a prediction mode to minimize the value of a cost function based on the bit rate of the candidate prediction mode and the distortion of the reconstructed reference image under the candidate prediction mode. According to the selected prediction mode, the encoder can generate the corresponding prediction BPU 208 and prediction data 206.
[0089] In the reconstruction path of process 200B, if an intra prediction mode has been selected in the forward path, after generating the prediction reference 224 (e.g., the current BPU that has been encoded and reconstructed in the current image), the encoder can directly feed the prediction reference 224 to the spatial prediction stage 2042 for later use (e.g., for interpolating the next BPU of the current image). If an inter prediction mode has been selected in the forward path, after generating the prediction reference 224 (e.g., the current image in which all BPUs have been encoded and reconstructed), the encoder can feed the prediction reference 224 to the loop filter stage 232. At this stage, the encoder can apply a loop filter to the prediction reference 224 to reduce or eliminate the distortion introduced by inter prediction (e.g., blocking artifacts). The encoder can apply various loop filter techniques at the loop filter stage 232, such as deblocking, sample adaptive compensation, adaptive loop filter, etc. The loop-filtered reference image can be stored in the buffer 234 (or "decoded image buffer") for later use (e.g., as an inter prediction reference image for future images of the video sequence 202). The encoder can store one or more reference images in the buffer 234 for use at the temporal prediction stage 2044. In some embodiments, the encoder can encode the parameters of the loop filter (e.g., loop filter strength) as well as the quantized transform coefficients 216, prediction data 206, and other information at the binary coding stage 226.
[0090] Figure 3A FIG. shows a schematic diagram of an exemplary decoding process 300A according to an embodiment of the present disclosure. The process 300A can be a decompression process corresponding to Figure 2A the compression process 200A in. In some embodiments, the process 300A can be similar to the reconstruction path of the process 200A. The decoder can decode the video bitstream 228 into a video stream 304 according to the process 300A. The video stream 304 can be very similar to the video sequence 202. However, due to information loss during the compression and decompression processes (e.g., Figures 2A - 2B the quantization stage 214 in), generally, the video stream 304 is different from the video sequence 202. Similar to Figures 2A - 2B the processes 200A and 200B in, the decoder can perform the process 300A on each image encoded in the video bitstream 228 at the basic processing unit (BPU) level. For example, the decoder can perform the process 300A in an iterative manner, where the decoder can decode the basic processing unit in one iteration of the process 300A. In some embodiments, the decoder can perform the process 300A in parallel for each region (e.g., regions 114-118) of each image encoded in the video bitstream 228.
[0091] As Figure 3AAs shown, the decoder can feed a portion of the video bitstream 228 associated with a basic processing unit of the encoded image (referred to as an "encoded BPU") into the binary decoding stage 302, where the decoder can decode the portion into prediction data 206 and quantized transform coefficients 216. The decoder can feed the quantized transform coefficients 216 into the inverse quantization stage 218 and the inverse transform stage 220 to generate a reconstructed residual BPU 222. The decoder can feed the prediction data 206 into the prediction stage 204 to generate a predicted BPU 208. The decoder can add the reconstructed residual BPU 222 to the predicted BPU 208 to generate a prediction reference 224. In some embodiments, the prediction reference 224 can be stored in a buffer (e.g., a decoded image buffer in a computer memory). The decoder can feed the prediction reference 224 into the prediction stage 204 for performing prediction operations in the next iteration of process 300A.
[0092] The decoder can iteratively execute process 300A to decode each encoded BPU of the encoded image and generate a prediction reference 224 for the next encoded BPU of the encoded image. After decoding all the encoded BPUs of the encoded image, the decoder can output the image to the video stream 304 for display and continue to decode the next encoded image in the video bitstream 228.
[0093] In the binary decoding stage 302, the decoder can perform the inverse operations of the binary coding techniques used by the encoder (e.g., entropy coding, variable length coding, arithmetic coding, Huffman coding, context - adaptive binary arithmetic coding, or any other lossless compression algorithm). In some embodiments, in addition to the prediction data 206 and the quantized transform coefficients 216, the decoder can decode other information in the binary decoding stage 302, such as prediction mode, parameters of the prediction operation, transform type, parameters of the quantization process (e.g., quantization parameter), encoder control parameters (e.g., bit - rate control parameter), etc. In some embodiments, if the video bitstream 228 is transmitted in packets over a network, the decoder can unpack the video bitstream 228 before feeding it into the binary decoding stage 302.
[0094] Figure 3B A schematic diagram of another example decoding process 300B according to an embodiment of the present disclosure is shown. Process 300B can be modified from process 300A. For example, process 300B can be used by a decoder compliant with a hybrid video coding standard (e.g., the H.26x series). Compared with process 300A, process 300B additionally divides the prediction stage 204 into a spatial prediction stage 2042 and a temporal prediction stage 2044, and additionally includes a loop filtering stage 232 and a buffer 234.
[0095] In process 300B, for an encoded basic processing unit (referred to as the "current BPU") of a decoded encoded image (referred to as the "current image"), the prediction data 206 decoded by the decoder from the binary decoding stage 302 can include various types of data, depending on what prediction mode the encoder uses to encode the current BPU. For example, if the encoder uses intra prediction to encode the current BPU, the prediction data 206 can include a prediction mode indicator (e.g., a flag value) indicating intra prediction, parameters of the intra prediction operation, etc. The parameters of the intra prediction operation can include, for example, the positions (e.g., coordinates) of one or more adjacent BPUs used as references, the sizes of the adjacent BPUs, interpolation parameters, the directions of the adjacent BPUs relative to the original BPU, etc. For another example, if the encoder uses inter prediction to encode the current BPU, the prediction data 206 can include a prediction mode indicator (e.g., a flag value) indicating inter prediction, parameters of the inter prediction operation, etc. The parameters of the inter prediction operation can include, for example, the number of reference images associated with the current BPU, the weights respectively associated with the reference images, the positions (e.g., coordinates) of one or more matching regions in the respective reference images, one or more motion vectors respectively associated with the matching regions, etc.
[0096] Based on the prediction mode indicator, the decoder can decide whether to perform spatial prediction (e.g., intra prediction) in the spatial prediction stage 2042 or temporal prediction (e.g., inter prediction) in the temporal prediction stage 2044. The details of performing such spatial prediction or temporal prediction are described in Figure 2B and will not be repeated hereinafter. After performing such spatial prediction or temporal prediction, the decoder can generate a predicted BPU 208. The decoder can add the predicted BPU 208 and the reconstructed residual BPU 222 to generate a prediction reference 224, as described in Figure 3A .
[0097] In process 300B, the decoder can feed the prediction reference 224 to the spatial prediction stage 2042 or the temporal prediction stage 2044 for performing prediction operations in the next iteration of process 300B. For example, if the current BPU is decoded using intra prediction in the spatial prediction stage 2042, after generating the prediction reference 224 (e.g., the decoded current BPU), the decoder can directly feed the prediction reference 224 to the spatial prediction stage 2042 for later use (e.g., for interpolating the next BPU of the current image). If the current BPU is decoded using inter prediction in the temporal prediction stage 2044, after generating the prediction reference 224 (e.g., the reference image in which all BPUs are decoded), the encoder can feed the prediction reference 224 to the loop filter stage 232 to reduce or eliminate distortion (e.g., blocking artifacts). The decoder can, as Figure 2BApply the loop filter to the prediction reference 224 in the manner shown. The reference image for loop filtering can be stored in buffer 234 (e.g., the decoded picture buffer in a computer memory) for later use (e.g., as an inter-prediction reference image for future coded pictures of the video bitstream 228). The decoder can store one or more reference images in buffer 234 for use at the temporal prediction stage 2044. In some embodiments, when the prediction mode indicator of the prediction data 206 indicates that inter-prediction is used for encoding the current BPU, the prediction data can further include the parameters of the loop filter (e.g., loop filter strength). The reconstructed picture from buffer 234 can also be sent to a display, such as a TV, PC, smartphone, or tablet, for viewing by an end user.
[0098] There can be four types of loop filters. For example, the loop filter can include a deblocking filter, a sample adaptive offset (“SAO”) filter, a luminance mapping with chroma scaling (“LMCS”) filter, and an adaptive loop filter (“ALF”). The order of applying the four types of loop filters can be the LMCS filter, the deblocking filter, the SAO filter, and the ALF. The LMCS filter can include two main components. The first component can be an in-loop mapping of the luminance component based on an adaptive piecewise linear model. The second component can be used for the chrominance component and can apply chrominance residual scaling related to luminance.
[0099] Figure 4 is a block diagram of an example apparatus 400 for encoding or decoding video according to an embodiment of the present disclosure. As Figure 4 shown, the apparatus 400 can include a processor 402. When the processor 402 executes the instructions described herein, the apparatus 400 can become a dedicated machine for video encoding or decoding. The processor 402 can be any type of circuit capable of manipulating or processing information. For example, the processor 402 can include any number of central processing units (or “CPUs”), graphics processing units (or “GPUs”), neural processing units (“NPUs”), microcontroller units (“MCUs”), optical processors, programmable logic controllers, microprocessors, digital signal processors, intellectual property (IP) cores, programmable logic arrays (PLAs), programmable array logics (PALs), generic array logics (GALs), complex programmable logic devices (CPLDs), a field programmable gate array (FPGA), a system on a chip (SoC), an application specific integrated circuit (ASIC), etc., in any combination. In some embodiments, the processor 402 can also be a group of processors grouped as a single logical component. For example, as Figure 4 shown, the processor 402 can include multiple processors, including processor 402a, processor 402b, and processor 402n.
[0100] The apparatus 400 may also include a memory 404 configured to store data (e.g., instruction sets, computer code, intermediate data, etc.). For example, as Figure 4 shown, the stored data may include program instructions (e.g., for implementing stages in processes 200A, 200B, 300A, or 300B) and data for processing (e.g., video sequence 202, video bitstream 228, or video stream 304). The processor 402 may access the program instructions and data for processing (e.g., via bus 410), and execute the program instructions to perform operations or manipulations on the data for processing. The memory 404 may include high-speed random access storage devices or non-volatile storage devices. In some embodiments, the memory 404 may include any combination of any number of random access memories (RAMs), read-only memories (ROMs), optical discs, magnetic disks, hard disk drives, solid state drives, flash drives, secure digital (SD) cards, memory sticks, compact flash (CF) cards, etc. The memory 404 may also be a group of memories grouped as a single logical component ( Figure 4 not shown in the figure).
[0101] The bus 410 may be a communication device for transferring data between components inside the apparatus 400, such as an internal bus (e.g., CPU-memory bus), an external bus (e.g., universal serial bus port, peripheral component interconnect express port), or the like.
[0102] For ease of explanation without causing ambiguity, in this disclosure, the processor 402 and other data processing circuits are collectively referred to as "data processing circuits". The data processing circuits may be implemented entirely in hardware, or as a combination of software, hardware, or firmware. In addition, the data processing circuits may be a single separate module, or may be fully or partially combined into any other component of the apparatus 400.
[0103] The apparatus 400 may also include a network interface 406 to provide wired or wireless communication with a network (e.g., the Internet, an intranet, a local area network, a mobile communication network, etc.). In some embodiments, the network interface 406 may include any combination of any number of network interface controllers (NICs), radio frequency (RF) modules, transponders, transceivers, modems, routers, gateways, wired network adapters, wireless network adapters, Bluetooth adapters, infrared adapters, near field communication ("NFC") adapters, cellular network chips, etc.
[0104] In some embodiments, optionally, the apparatus 400 may further include a peripheral interface 408 to provide connections to one or more peripheral devices. As Figure 4As shown, the peripheral device may include, but is not limited to, a cursor control device (e.g., a mouse, a touchpad, or a touch screen), a keyboard, a display (e.g., a cathode ray tube display, a liquid crystal display, or a light emitting diode display), a video input device (e.g., a camera or an input interface coupled to a video archive), etc.
[0105] It should be noted that a video codec (e.g., the codec that executes processes 200A, 200B, 300A, or 300B) may be implemented as any combination of any software or hardware modules in device 400. For example, some or all stages of processes 200A, 200B, 300A, or 300B may be implemented as one or more software modules of device 400, such as program instances that may be loaded into memory 404. For another example, some or all stages of processes 200A, 200B, 300A, or 300B may be implemented as one or more hardware modules of device 400, such as dedicated data processing circuits (e.g., FPGA, ASIC, NPU, etc.).
[0106] In the quantization and inverse quantization functional blocks (e.g., Figure 2A or Figure 2B the quantization 214 and inverse quantization 218 of Figure 3A or Figure 3B the inverse quantization 218 of ), a quantization parameter (QP) is used to determine the amount of quantization (and inverse quantization) applied to the prediction residual. The initial QP value used for encoding an image or a slice may be signaled at a higher level, e.g., using the init_qp_minus26 syntax element in the picture parameter set (PPS) and using the slice_qp_delta syntax element in the slice header. Additionally, incremental QP values sent in the granularity of quantization groups may be used to adapt the QP value at the local level of each CU.
[0107] Rectangular projection (“ERP”) formats such as etc. are common projection formats for representing 360-degree videos and images. The projection maps the meridians to vertical lines with a constant spacing and the latitude circles to horizontal lines with a constant spacing. Since the relationship between the position of an image pixel on the map and its corresponding geographical location on the sphere is particularly simple, ERP is one of the most common projections for 360-degree videos and images.
[0108] The algorithm descriptions of the projection format conversion and video quality metrics output by JVET give the coordinate conversion between ERP and sphere and an introduction. For the 2D to 3D coordinate conversion, given the sampling positions (m, n), (u, v) can be calculated based on the following formulas (1) and (2). u = (m + 0.5) / W, 0 ≤ m < W Formula (1) v = (n + 0.5) / H, 0 ≤ n < H Equation (2)
[0109] Then, the longitude and latitude (φ, θ) in the sphere can be calculated from (u, v) based on the following Equations (3) and (4). φ = (u - 0.5) × (2 × π) Equation (3) θ = (0.5 - v) × π Equation (4)
[0110] The 3D coordinates (X, Y, Z) can be calculated based on the following Equations (5) - (7). X = cos(θ)cos(φ) Equation (5) Y = sin(θ) Equation (6) Z = -cos(θ)sin(φ) Equation (7)
[0111] For 3D to 2D coordinate conversion starting from (X, Y, Z), (φ, θ) can be calculated based on the following Equations (8) and (9). Then, (u, v) are calculated based on Equations (3) and (4). Finally, the 2D coordinates (m, n) can be calculated according to Equations (1) and (2). φ = tan-1(-Z / X) Equation (8) θ = sin-1(Y / (X2 + Y2 + Z2)1 / 2) Equation (9)
[0112] To reduce seam artifacts in the reconstructed viewport that includes the left and right boundaries of the ERP image, a new format called padded equirectangular projection ("PERP") is provided by padding samples on each of the left and right sides of the ERP image.
[0113] When representing 360-degree video using PERP, the PERP image is encoded. After decoding, the reconstructed PERP is converted back to the reconstructed ERP by blending the replicated samples or cropping the padded regions.
[0114] Figure 5A A schematic diagram showing an example blending operation for generating a reconstructed equirectangular projection according to some embodiments of the present disclosure is shown. Unless otherwise specified, "recPERP" is used to represent the reconstructed PERP before post-processing, and "recERP" is used to represent the reconstructed ERP after post-processing. In Figure 5A where A1 and B2 are boundary regions in the ERP image, and B1 and A2 are padded regions, where A2 is padded from A1 and B1 is padded from B2. As Figure 5AAs shown, the replicated samples of recPERP can be blended by applying a distance-based weighted average operation. For example, region A can be generated by blending region A1 and A2, while region B can be generated by blending region B1 and b2.
[0115] In the following description, the width and height of the unfilled recERP are denoted as "W" and "H" respectively. The left and right padding widths are denoted as "P L " and "P R " respectively. The total padding width is denoted as "P w ", which can be the sum of P L and P R . In some embodiments, recPERP can be converted to recERP through a blending operation. For example, for the sample recERP(j, i) in A, where (j, i) are the coordinates in the ERP image, and i is in [0, P R - 1] and j is in [0, H - 1], recERP(j, i) can be determined according to the following formula. A = w × A1 + (1 – w) × A2, where w ranges from PL / Pw to 1 Formula (10) recERP(j,i) in A = (recPERP(j,i + PL) × (i + PL) + recPERP(j,i + PL + W) × (PR - i) + (PW >> 1)) / PW Formula (11) where, recPERP(y, x) is the sample on the reconstructed PERP image, where (y, x) are the coordinates of the sample in the PERP image.
[0116] In some embodiments, for the sample recERP(j, i) in B, where (j, i) are the coordinates in the ERP image, and i is in [W - P L , W - 1] and j is in [0, H - 1], recERP(j, i) can be generated according to the following equation: B = k × B1 + (1 – k) × B2, where k ranges from 0 to PL / Pw Formula (12) recERP(j,i) in B = (recPERP(j,i + PL) × (PR - i + W) + recPERP(j,i + PL - w) × (i – W + PL) + (Pw >> 1)) / PW Formula (13) where, recPERP(y, x) is the sample on the reconstructed PERP image, where (y, x) are the coordinates of the sample in the PERP image.
[0117] Figure 5B FIG. shows a schematic diagram of an example cropping operation for generating a reconstructed equirectangular projection according to some embodiments of the present disclosure. InFigure 5B where A1 and B2 are boundary regions within the ERP image, and B1 and A2 are filling regions, where A2 is filled from A1 and B1 is filled from B2. As Figure 5B shown, during the cropping process, the filled samples in recPERP can be directly discarded to obtain recERP. For example, the filled samples B1 and A2 can be discarded.
[0118] In some embodiments, horizontal wrap-around motion compensation can be used to improve the coding performance of ERP. For example, horizontal wrap-around motion compensation can be used as a 360-degree specific coding tool in the VVC standard, which is designed to improve the visual quality of the reconstructed 360-degree video in ERP format or PERP format. In conventional motion compensation, when the motion vector reference exceeds the samples of the image boundary of the reference image, repeated filling is applied by copying those nearest neighboring samples on the corresponding image boundary to derive the value of the out-of-bounds sample. For 360-degree video, this method of repeated filling is not suitable and may result in visual artifacts called "seam artifacts" in the reconstructed viewport video. Since 360-degree video is captured on a sphere and inherently has no "boundaries", the reference samples outside the boundary of the reference image in the projection domain can be obtained from neighboring samples in the spherical domain. For a general projection format, it may be difficult to derive the corresponding neighboring samples in the spherical domain because it involves 2D to 3D and 3D to 2D coordinate conversions, as well as sample interpolation at fractional sample positions. For the left and right boundaries of the ERP or PERP projection format, this problem can be solved because the spherical neighbor samples outside the left image boundary can be obtained from the samples within the right image boundary and vice versa. Given the wide use and relatively easy implementation of the ERP or PERP projection format, VVC adopts horizontal wrap-around motion compensation to improve the visual quality of 360-degree videos encoded in the ERP or PERP projection format.
[0119] Figure 6A FIG. shows a schematic diagram of an example horizontal wrap-around motion compensation process for equirectangular projection according to some embodiments of the present disclosure. As Figure 6A shown, when a part of the reference block is outside the left (or right) boundary of the reference image in the projection domain, instead of repeated filling, the "out-of-bounds" part can be obtained from the corresponding spherical neighbor part located within the reference image towards the right (or left) boundary in the projection domain. In some embodiments, repeated filling can be used for the top and bottom image boundaries.
[0120] Figure 6B FIG. shows a schematic diagram of an example horizontal wrap-around motion compensation process for filling equirectangular projection according to some embodiments of the present disclosure. As Figure 6BAs shown, horizontal wrap-around motion compensation can be combined with non-normative padding methods commonly used in 360-degree video coding. In some embodiments, this is achieved by signaling a high-level syntax element to indicate the wrap-around motion compensation offset, which can be set to the ERP image width before padding. This syntax can be used to adjust the position of the horizontal wrap-around accordingly. In some embodiments, this syntax is not affected by a specific amount of padding on the left or right image boundaries. As a result, this syntax can naturally support asymmetric padding of ERP images. In asymmetric padding of ERP images, the padding on the left and right can be different. In some embodiments, the wrap-around motion compensation can be determined according to the following formula: where the offset can be the wrap-around motion compensation offset signaled in the bitstream, picW can be the image width including the padding area before encoding, pos x can be the reference position determined by the current block position and the motion vector, and the output of the formula pos x _wrap can be the actual reference position of the reference block in the wrap-around motion compensation. To save the signaling overhead of the wrap-around motion compensation offset, it can be in units of the minimum luminance coded block, and thus, the offset can be replaced by offset w ×MinCbSizeY, where offset w is the wrap-around motion compensation offset in units of the minimum luminance coded block, which is signaled in the bitstream, and MinCbSizeY is the size of the minimum luminance coded block. In contrast, in traditional motion compensation, the actual reference position from which the reference block comes can be directly derived by clipping pos x within the range of 0 to picW - 1.
[0121] When the reference samples are outside the left and right boundaries of the reference image, horizontal wrap-around motion compensation can provide more meaningful information for motion compensation. Under common test conditions for 360-degree video, this tool can not only improve the compression performance in terms of rate-distortion, but also improve the compression performance in terms of reducing seam artifacts and the subjective quality of the reconstructed 360-degree video. Horizontal wrap-around motion compensation can also be used in other single-sided projection formats with a constant sampling density in the horizontal direction, such as the adjusted equal-area projection.
[0122] In VVC (e.g., VVC Draft 8), wrap-around motion compensation can be achieved by signaling the high-level syntax variable pps_ref_wraparound_offset to indicate the wrap-around offset. Sometimes, the wrap-around offset should be set to the ERP image width before padding. This syntax can be used to adjust the position of the horizontal wrap-around motion compensation accordingly. This syntax may not be affected by a specific padding amount on the left or right image boundary. As a result, this syntax can naturally support asymmetric padding of ERP images (e.g., when the left padding and the right padding are different). When the reference samples are outside the left and right boundaries of the reference image, the horizontal wrap-around motion compensation can provide more meaningful information for motion compensation.
[0123] Figure 7 The syntax of an example high-level wrap-around offset according to some embodiments of the present disclosure is shown. It should be understood that Figure 7 the shown syntax can be used in VVC (e.g., VVC Draft 8). As Figure 7 shown, when wrap-around motion compensation is enabled (e.g., pps_ref_wraparound_enabled_flag == 1), the wrap-around offset pps_ref_wraparound_offset can be directly signaled. As Figure 7 shown, the syntax elements or variables in the bitstream are shown in bold.
[0124] Figure 8 The semantics of an example high-level wrap-around offset according to some embodiments of the present disclosure are shown. It should be understood that Figure 8 the semantics shown in Figure 7 can correspond to the syntax shown in Figure 8 In some embodiments, as
[0125] shown, a value of 1 for pps_ref_wraparound_enabled_flag means that horizontal wrap-around motion compensation is applied in inter prediction. If the value of pps_ref_wraparound_enabled_flag is 0, horizontal wrap-around motion compensation is not applied. When the value of CtbSizeY / MinCbSizeY + 1 is greater than pic_width_in_luma_samples / MinCbSizeY - 1, the value of pps_ref_wraparound_enabled_flag should be equal to 0. When sps_ref_wraparound_enabled_flag is equal to 0, the value of pps_ref_wraparound_enabled_flag should be equal to 0. CtbSizeY is the size of the luma coding tree block.
[0125] In some embodiments, as Figure 8As shown, the value of pps_ref_wraparound_offset plus (CtbSizeY / MinCbSizeY)+2 can specify an offset for calculating the horizontal wraparound position in units of MinCbSizeY luma samples. The value of pps_ref_wraparound_offset can range from 0 to (pic_width_in_luma_samples / MinCbSizeY)-(CtbSizeY / MinCbSizeY)-2, inclusive. The variable PpsRefWraparoundOffset can be set to be equal to pps_ref_wraparound_offset+(CtbSizeY / MinCbSizeY)+2. In some embodiments, the variable PpsRefWraparoundOffset can be used to determine the luma position of a sub-block (e.g., in VVC Draft 8).
[0126] In VVC (e.g., VVC Draft 8), an image can be divided into one or more tile rows or one or more tile columns. A tile can be a sequence of coding tree units ("CTUs") that cover a rectangular region of the image. A slice can include an integer number of complete tiles or an integer number of consecutive complete CTU rows within a tile of the image.
[0127] Two modes of slices can be supported, namely raster scan slice mode and rectangular slice mode. In the raster scan slice mode, a slice can include a sequence of complete tiles in the raster scan of the tiles of the image. In the rectangular slice mode, a slice can include multiple complete tiles that together form a rectangular region of the image, or multiple consecutive complete CTU rows of a single tile that together form a rectangular region of the image. The tiles within a rectangular slice are scanned in tile raster scan order within the rectangular region corresponding to the slice.
[0128] A sub-image can include one or more slices that together cover a rectangular region of the image. Figure 9 A schematic diagram of an example slice and sub-image partitioning of an image according to some embodiments of the present disclosure is shown. As Figure 9 shown, the image is divided into 20 tiles, with 5 tile columns and 4 tile rows. There are 12 tiles on the left, each covering a slice of 4×4 CTUs. There are 8 tiles on the right, each covering a slice of 2 vertically stacked 2×2 CTUs. There are a total of 28 slices and 28 sub-images of different sizes (e.g., each slice can be a sub-image).
[0129] Figure 10 A schematic diagram of an example slice and sub-image partitioning with different slices and sub-images according to some embodiments of the present disclosure is shown. As Figure 10As shown, the image is divided into 20 tiles, having 5 tile columns and 4 tile rows. There are 12 tiles on the left, each tile covering a stripe of 4×4 CTUs. There are 8 tiles on the right, each tile covering a stripe of 2 vertically stacked 2×2 CTUs. There are a total of 28 stripes. For the 12 stripes on the left, each stripe is a sub-image. For the 16 stripes on the right, every 4 stripes form a sub-image. As a result, there are a total of 16 sub-images of the same size.
[0130] In VVC (e.g., VVC Draft 8), information about the stripe layout can be signaled in the Picture Parameter Set (“PPS”). In some embodiments, the Picture Parameter Set is a syntax structure that includes syntax elements or variables that apply to zero or more entire coded pictures determined by syntax elements found in each picture header. Figure 11 The syntax of an example Picture Parameter Set for tile mapping and stripe layout according to some embodiments of the present disclosure is shown. It should be understood that Figure 11 the shown syntax can be used in VVC (e.g., VVC Draft 8). As Figure 11 shown, syntax elements or variables in the bitstream are shown in bold. As Figure 11 shown, if the number of tiles in the current picture is greater than 1 and the rectangular stripe mode is used (e.g., rect_slice_flag == 1), then a flag called single_slice_per_subpic_flag can be signaled first to indicate that each sub-image includes only one stripe. In this case (e.g., single_slice_per_subpic_flag == 1), there is no need to further signal the stripe layout information, because it can be the same as the sub-image layout already signaled in the Sequence Parameter Set (“SPS”). In some embodiments, the SPS is a syntax structure that includes syntax elements that apply to zero or more entire coded layer video sequences (“CLVS”), the content of which is determined by syntax elements found in the Picture Parameter Set referenced by syntax elements found in each picture header. In some embodiments, the picture header is a syntax structure that includes syntax elements applicable to all stripes of the coded picture. In some embodiments, as Figure 11 shown, if the value of single_slice_per_subpic_flag is 0, then the number of stripes in the picture (e.g., num_slices_in_pic_minus1) can be signaled first, followed by the stripe position and size information of each stripe.
[0131] In some embodiments, to signal the number of slices, the number of slices minus 1 (e.g., num_slices_in_pic_minus1) may be signaled instead of directly signaling the number of slices, since there is at least 1 slice in the picture. Generally, signaling a smaller positive value can cost fewer bits and improve the overall efficiency of performing video processing.
[0132] FIG. 12 shows the semantics of an example picture parameter set for slice mapping and strip layout according to some embodiments of the present disclosure. It should be understood that the semantics 2 shown in FIG. 12 may correspond to Figure 11 the syntax shown. In some embodiments, as shown in FIG. 12, the semantics corresponds to VVC (e.g., VVC draft 8).
[0133] In some embodiments, as shown in FIG. 12, when the variable rect_slice_flag is equal to 0, the slices within each strip are in raster scan order and strip information is not signaled in the PPS. When the variable rect_slice_flag is equal to 1, the slices within each strip may cover a rectangular region of the picture and strip information may be signaled in the PPS. In some embodiments, when not present, the variable rect_slice_flag may be inferred to be equal to 1. In some embodiments, when the variable subpic_info_present_flag is equal to 1, the value of rect_slice_flag shall be equal to 1.
[0134] In some embodiments, as shown in FIG. 12, when the variable single_slice_per_subpic_flag is equal to 1, it means that each subpicture may include one and only one rectangular strip. When the variable single_slice_per_subpic_flag is equal to 0, each subpicture may include one or more rectangular strips. In some embodiments, when the variable single_slice_per_subpic_flag is equal to 1, the variable num_slices_in_pic_minus1 may be inferred to be equal to the variable sps_num_subpics_minus1. In some embodiments, when not present, the value of single_slice_per_subpic_flag may be inferred to be equal to 0.
[0135] In some embodiments, as shown in FIG. 12, the variable num_slices_in_pic_minus1 plus 1 is the number of rectangular stripes of the reference PPS in each picture. In some embodiments, the value of num_slices_in_pic_minus1 may be in the range of 0 to MaxSlicesPerPicture - 1, including the end values. In some embodiments, when the variable no_pic_partition_flag is equal to 1, it can be inferred that the value of num_slices_in_pic_minus1 is equal to 0.
[0136] In some embodiments, as shown in FIG. 12, if the variable tile_idx_delta_present_flag is equal to 0, the value of tile_idx_delta does not exist in the PPS, and all the rectangular stripes in the picture of the reference PPS are specified in raster order. In some embodiments, when the variable tile_idx_delta_present_flag is equal to 1, the value of tile_idx_delta may exist in the PPS, and all the rectangular stripes in the picture of the reference PPS are specified in the order indicated by the value of tile_idx_delta. In some embodiments, when it does not exist, it can be inferred that the value of tile_idx_delta_present_flag is equal to 0.
[0137] In some embodiments, as shown in FIG. 12, the variable slice_width_in_tiles_minus1[i] plus 1 specifies the width of the i-th rectangular stripe in terms of tiles. In some embodiments, the value of slice_width_in_tiles_minus1[i] should be in the range of 0 to NumTileColumns - 1, including the end values. In some embodiments, if slice_width_in_tiles_minus1[i] does not exist, the following can be applied: if NumTileColumns is equal to 1, it can be inferred that the value of slice_width_in_tiles_minus1[i] is equal to 0; otherwise, the value of slice_width_in_tiles_minus_1[i] can be inferred as the value specified in the clause in the VVC draft (e.g., VVC draft 8).
[0138] In some embodiments, as shown in FIG. 12, the variable slice_height_in_tiles_minus1 plus 1 specifies the height of the i-th rectangular strip in terms of tile rows. In some embodiments, the value of slice_height_in_tiles_minus1[i] should be in the range of 0 to NumTileRows - 1, inclusive. In some embodiments, when the variable slice_height_in_tiles_minus1[i] does not exist, the following applies: If NumTileRows is equal to 1, or the variable tile_idx_delta_present_flag is equal to 0 and tileIdx % NumTileColumns is greater than 0, then it is inferred that the value of slice_height_in_tiles_minus1[i] is equal to 0; otherwise (e.g., NumTileRows is not equal to 1, and tile_idx_delta_present_flag is equal to 1 or tileIdx % NumTileColumns is equal to 0), when tile_idx_delta_present_flag is equal to 1 or tileIdx % NumTileColumns is equal to 0, the value of slice_height_in_tiles_minus1[i] can be inferred to be equal to slice_height_in_tiles_minus1[i - 1].
[0139] In some embodiments, as shown in FIG. 12, the value of num_exp_slices_in_tile[i] specifies the number of strip heights explicitly provided in the current tile that includes multiple rectangular strips. In some embodiments, the value of num_exp_slices_in_tile[i] should be in the range of 0 to RowHeight[tileY] - 1, inclusive, where tileY is the tile row index that includes the i-th strip. In some embodiments, when it does not exist, it can be inferred that the value of num_exp_slices_in_tile[i] is equal to 0. In some embodiments, when num_exp_slices_in_tile[i] is equal to 0, the value of the variable NumSliceInTile[i] is derived to be equal to 1.
[0140] In some embodiments, as shown in FIG. 12, the value of exp_slice_height_in_ctus_minus1[j] plus 1 specifies the height of the j-th rectangular strip in the current slice in units of CTU rows. In some embodiments, the value of exp_slice_height_in_ctus_minus1[j] should be in the range of 0 to RowHeight[tileY] - 1, inclusive, where tileY is the slice row index of the current slice.
[0141] In some embodiments, as shown in FIG. 12, when the variable num_exp_slices_in_tile[i] is greater than 0, the variables NumSlicesInTile[i] and SliceHeightInCtusMinus1[i + k] where k is in the range of 0 to NumSlicesInTile[i] - 1 can be obtained. In some embodiments, as Figure 8 shown, when the value of num_exp_slices_in_tile[i] is greater than 0, the values of NumSlicesInTile[i] and SliceHeightInCtusMinus1[i + k] can be obtained.
[0142] In some embodiments, as shown in FIG. 12, the value of tile_idx_delta[i] specifies the difference between the slice index of the first slice in the i-th rectangular strip and the slice index of the first slice in the (i + 1)-th rectangular strip. The value of tile_idx_delta[i] should be in the range of -NumTilesInPic + 1 to NumTilesInPic - 1, inclusive. In some embodiments, when it does not exist, it can be inferred that the value of tile_idx_delta[i] is equal to 0. In some embodiments, when it exists, it can be inferred that the value of tile_idx_delta[i] is not equal to 0.
[0143] In VVC (e.g., VVC Draft 8), in order to locate each strip in the image, one or more strip addresses can be signaled in the strip header. Figure 13 The syntax of an example strip header according to some embodiments of the present disclosure is shown. It should be understood that Figure 13 the syntax 3 shown can be used in VVC (e.g., VVC Draft 8). As Figure 11 shown, the syntax elements or variables in the bitstream are shown in bold. As Figure 13 shown, if the variable picture_header_in_slice_header_flag is equal to 1, the picture header syntax structure exists in the strip header.
[0144] Figure 14 illustrates the semantics of an example strip header according to some embodiments of the present disclosure. It should be understood that Figure 14 the semantics shown in Figure 13 may correspond to the syntax shown in Figure 14 In some embodiments, the semantics shown in
[0145] In some embodiments, as Figure 14 shown, the requirement for bitstream consistency is that the value of picture_header_in_slice_header_flag should be the same in all coded strips in the CLVS.
[0146] In some embodiments, as Figure 14 shown, when the variable picture_header_in_slice_header_flag is equal to 1 for a coded strip, the requirement for bitstream consistency is that there should be no video coding layer ("VCL") network abstraction layer ("NAL") unit in the CLVS with a nal_unit_type equal to PH_NUT.
[0147] In some embodiments, as Figure 14 shown, when picture_header_in_slice_header_flag is equal to 0, all coded strips in the current picture should have a picture_header_in_slice_header_flag equal to 0, and the current PU should have a PH NAL unit.
[0148] In some embodiments, as Figure 14 shown, the variable slice_address can specify the strip address of a strip. In some embodiments, when it does not exist, the value of slice_address can be inferred to be equal to 0. When the variable rect_slice_flag is equal to 1 and NumSliceInSubpic[CurrSubpicIdx] is equal to 1, the value of slice_address is inferred to be equal to 0.
[0149] In some embodiments, as Figure 14 shown, if the variable rect_slice_flag is equal to 0, the following can be applied: the strip address can be the raster scan block slice index; the length of slice_address can be ceil(Log2(NumTilesInPic)) bits; and the value of slice_address should be in the range of 0 to NumTilesInPic - 1, inclusive.
[0150] In some embodiments, as Figure 14 shown, the bitstream consistency requirements apply the following constraints: If the variable rect_slice_flag is equal to 0 or the variable subpic_info_present_flag is equal to 0, the value of slice_address shall not be equal to the value of slice_address of any other coded slice NAL unit of the same coded picture; otherwise, the slice_subpic_id and slice_address value pair shall not be equal to the slice_subpic_id and slice_address value pair of any other coded slice NAL unit of the same coded picture.
[0151] In some embodiments, as Figure 14 shown, the shape of the slices of the picture shall be such that each CTU shall have, at decoding, its entire left boundary and its entire top boundary that includes the picture boundary or includes the boundary of a previously decoded CTU.
[0152] In some embodiments, as Figure 14 shown, the variable slice_address may be signaled only if one of the following two conditions is met: The rectangular slice mode is used and the number of slices in the current subpicture is greater than 1; or the rectangular slice mode is not used and the number of tiles in the current picture is greater than 1. In some embodiments, if neither of the above two conditions is met, there is only one slice in the current subpicture or the current picture. In this case, since the entire subpicture or the entire picture is a single slice, signaling the slice address is not required.
[0153] There are many problems in the current VVC design. First, the wrap-around offset pps_ref_wraparound_offset is signaled in the bitstream and shall be set to the ERP picture width before padding. To save bits, the minimum value of the wrap-around offset (e.g., (CtbSizeY / MinCbSizeY)+2) is subtracted from the wrap-around offset before signaling. However, the width of the padding area is much smaller than the width of the original ERP picture. This is especially true for coded ERP pictures where the width of the padding area can be 0. Assuming that the total width of the picture to be coded or decoded is known, signaling the width of the original ERP part may cost more bits than signaling the width of the padding area. As a result, the current signaling of the wrap-around offset that signals the original ERP width in units of MinCbSizeY is not efficient, the wrap-around offset.
[0154] In addition, the signaling of the stripe layout can be improved. For example, subtract 1 from the number of stripes before signaling the number of stripes, because the number of stripes in an image is always greater than or equal to 1, and signaling a smaller positive value requires fewer bits. However, in the current VVC (e.g., VVC Draft 8), a sub-image consists of an integer number of complete stripes. As a result, the number of stripes in an image is greater than or equal to the number of sub-images in the image. In the current VVC (e.g., VVC Draft 8), num_slices_in_pic_minus1 is signaled only when the variable single_slice_per_subpic_flag is 0, and the variable single_slice_per_subpic_flag being equal to 0 means that at least one sub-image contains multiple stripes. Therefore, in this case, the number of stripes must be greater than the number of sub-images with a minimum value of 1. As a result, the minimum value of num_slices_in_pic_minus1 is greater than zero. Signaling a non-negative value whose signaling range does not start from zero is inefficient.
[0155] In addition, there are other issues with stripe address signaling. When there is only one stripe in the current image, signaling the stripe address may not be necessary because the entire sub-image or the entire image is a single stripe. The stripe address can be inferred as 0. However, the two conditions for skipping stripe address signaling are not complete. For example, in the raster scan stripe mode, even if the number of tiles is greater than 1, there can be only one stripe that includes all the tiles in the image, and in this case, stripe address signaling can also be avoided.
[0156] Embodiments of the present disclosure provide a method for solving the above problems. In some embodiments, since the width of the padding region is usually smaller than the width of the original ERP image, and the width of the original ERP image can be the same as the wrap-around offset in the wrap-around motion compensation, it can be proposed to signal the difference between the width of the coded image in the bitstream and the width of the original ERP image. For example, it can be proposed to signal the difference between the width of the coded image and the wrap-around motion compensation offset in the bitstream, and perform a derivation after parsing the signaled difference to obtain the wrap-around offset on the decoder side. Since the difference between the width of the coded image and the wrap-around offset is usually smaller than the wrap-around offset itself, this method can save the bits for signaling.
[0157] Figure 15 The syntax of an exemplary improved picture parameter set according to some embodiments of the present disclosure is shown. As Figure 15 shown, the syntax elements or variables in the bitstream are shown in bold, and the changes for the previous VVC are shown in italic type (e.g., Figure 7 the syntax shown in Figure 15As shown, a new variable pps_pic_width_minus_wraparound_offset can be created. The variable the value of pps_pic_width_minus_wraparound_offset can be signaled according to the value of pps_ref_wraparound_enabled_flag. In some embodiments, the new variable pps_pic_width_minus_wraparound_offset can replace the variable pps_ref_wraparound_offset in the original VVC (e.g., VVC Draft 8).
[0158] Figure 16 Illustrated is the semantics of an exemplary improved picture parameter set according to some embodiments of the present disclosure. As Figure 16 shown, the changes to the previous VVC (e.g., Figure 8 the semantics shown therein) are shown in italic type, and the proposed deletion syntax is further shown in strikethrough. It should be understood that Figure 16 the semantics shown therein may correspond to Figure 15 the syntax shown. In some embodiments, Figure 16 the semantics shown correspond to VVC (e.g., VVC Draft 8).
[0159] As Figure 16 shown, the semantics of the new variable pps_pic_width_minus_wraparound_offset are different from the variable pps_ref_wraparound_offset (e.g., as Figure 8 shown). The variable pps_pic_width_minus_wraparound_offset can specify the difference between the picture width and the offset used to calculate the horizontal wraparound position in units of MinCbSizeY luma samples. In some embodiments, as Figure 16 shown, the value of pps_pic_width_minus_wraparound_offset should be less than or equal to (pic_width_in_luma_samples / MinCbSizeY)-(CtbSizeY / MinCbSizeY)-2. In some embodiments, the variable PpsRefWraparoundOffset can be set to be equal to pic_width_in_luma_samples / MinCbSizeY – pps_pic_width_minus_wraparound_offset.
[0160] In some embodiments, the flag wraparound_offset_type can be signaled to indicate whether the wraparound offset used for signaling is the original ERP image width or the difference between the coded image width and the original ERP image width. The encoder can select the smaller of these two values and signal it in the bitstream, thereby further reducing the signaling overhead. Figure 17 shows the syntax of an example picture parameter set with the variable wraparound_offset_type according to some embodiments of the present disclosure. As Figure 17 shown, the syntax elements or variables in the bitstream are shown in bold, and the changes to the previous VVC (e.g., Figure 7 the syntax shown in Figure 17 are shown in italic type. As
[0161] Figure 18 shown, the semantics of an example enhanced picture parameter set with the variable wraparound_offset_type according to some embodiments of the present disclosure. As Figure 18 shown, the changes to the previous VVC (e.g., Figure 8 the semantics shown in Figure 18 are shown in italic type, and the proposed deleted syntax is further shown with a strikethrough. It should be understood that Figure 17 the semantics shown in Figure 18 can correspond to the syntax shown in
[0162] In some embodiments, as Figure 18 shown, a new variable wraparound_offset_type can be added to specify the type of the variable pps_ref_wraparound_offset. The value of pps_ref_wraparound_offset should be in the range of 0 to ((pps_pic_width_in_luma_samples / MinCbSizeY) – (CtbSizeY / MinCbSizeY) – 2) / 2.
[0163] In some embodiments, as Figure 18As shown, the value of the variable PpsRefWraparoundOffset can be derived. For example, when the variable wraparound_offset_type is equal to 0, the signaled wrapound offset is the original ERP image width. As a result, the variable PpsRefWraparoundOffset is equal to pps_ref_wraparound_offset+(CrbSizeY / MinCbSizeY)+2. When the variable wraparound_offset_type is not equal to 0, the signaled wrapound offset is the difference between the coded image width and the original ERP image width. As a result, the variable PpsRefWraparoundOffset is equal to pps_pic_width_in_luma_samples / MinCbSizeY – pps_ref_wraparound_offset.
[0164] In a previous VVC (e.g., VVC Draft 8), when the variable single_slice_per_subpic_flag is 0, the number of slices is signaled in the PPS, which means that there is at least one subpicture that contains more than one slice. In this case, considering that each subpicture should contain one or more complete slices, the number of slices needs to be greater than the number of subpictures, which thus results in a minimum number of slices of 2. This is because there is at least one subpicture in an image.
[0165] Embodiments of the present disclosure provide a method for improving the signaling of the number of slices. Figure 19 The syntax of an example picture parameter set with the variable num_slices_in_pic_minus2 according to some embodiments of the present disclosure is shown. As Figure 19 shown, the syntax elements or variables in the bitstream are shown in bold, and the changes to the previous VVC (e.g., Figure 11 the syntax shown therein) are shown in italic type, where the proposed deleted syntax is further shown with a strikethrough. As Figure 19 shown, a new variable num_slices_in_pic_minus2 can be added to specify the value for determining PpsRefWraparoundOffset.
[0166] Figure 20 shows the semantics of an example enhanced picture parameter set with the variable num_slices_in_pic_minus2 according to some embodiments of the present disclosure. As shown in Figure 20, the changes to the previous VVC (e.g., the semantics shown in Figure 12) are shown in italic type, and the proposed deletion syntax is further shown with a strikethrough. It should be understood that the semantics shown in Figure 20 may correspond to Figure 19 the syntax shown. In some embodiments, the semantics shown in Figure 20 correspond to VVC (e.g., VVC Draft 8).
[0167] In some embodiments, as shown in Figure 20, adding 2 to the value of the variable num_slices_in_pic_minus2 may specify the number of rectangular stripes in each picture of the reference PPS. In some embodiments, as shown in Figure 20, the variable num_slices_in_pic_minus2 may replace the variable num_slices_in_pic_minus_1. In some embodiments, the value of num_slices_in_pic_minus2 plus 2 should be in the range of 0 to MaxSlicesPerPicture - 2, inclusive, where MaxSlicesPerPicture is specified in VVC (e.g., VVC Draft 8). When the variable no_pic_partition_flag is equal to 1, the value of num_slices_in_pic_minus2 is inferred to be equal to -1.
[0168] In some embodiments, a variable that is the number of stripes minus the number of sub - pictures and then minus 1 (e.g., num_slices_in_pic_minus_subpic_num_minus1) may be used to signal the number of stripes. Figure 21 Figure shows the syntax of an example picture parameter set with the variable num_slices_in_pic_minus_subpic_num_minus1 according to some embodiments of the present disclosure. As Figure 21 shown, the syntax elements or variables in the bitstream are shown in bold, and the changes to the previous VVC (e.g., the syntax shown in Figure 11 ) are shown in italic type, where the proposed deletion syntax is further shown with a strikethrough.
[0169] FIG. 22 illustrates the semantics of an exemplary modified picture parameter set with the variable num_slices_in_pic_minus_subpic_num_minus1 according to some embodiments of the present disclosure. As shown in FIG. 22, the changes to the previous VVC (e.g., the semantics shown in FIG. 12) are shown in italic type, and the proposed deletion syntax is further shown in strikethrough. It should be understood that the semantics shown in FIG. 22 may correspond to Figure 21 the syntax shown. In some embodiments, the semantics shown in FIG. 22 correspond to VVC (e.g., VVC Draft 8).
[0170] In VVC (e.g., VVC Draft 8), the number of stripes is signaled in the PPS only when the variable single_slice_per_subpic_flag is equal to 0, which may be equivalent to stating that there is at least one sub-picture containing more than one stripe. Therefore, the number of stripes should be greater than the number of sub-pictures, since a sub-picture should contain one or more complete stripes. As a result, the minimum number of stripes is equal to the number of sub-pictures plus 1. As Figure 21 shown in FIG. 22, the number of stripes signaled minus the number of sub-pictures and then minus 1 (e.g., num_slices_in_pic_minus_subpic_num_minus1) instead of the number of stripes minus 1 (e.g., num_slices_in_pic_minus1) to reduce the number of bits signaled.
[0171] In some embodiments, as shown in FIG. 22, the value of num_slices_in_pic_minus_subpic_num_minus1 plus the number of sub-images plus 1 may specify the number of rectangular stripes in each image of the reference PPS. In some embodiments, the value of num_slices_in_pic_minus_subpic_num_minus1 should be in the range of 0 to MaxSlicesPerPicture minus sps_num_subpics_minus1 minus 2, inclusive, where MaxSlicesPerPicture may be specified in VVC (e.g., VVC Draft 8). In some embodiments, when the variable no_pic_partition_flag is equal to 1, it can be inferred that the value of num_slices_in_pic_minus_subpic_num_minus1 is equal to (sps_num_subpics_minus1 + 1). In some embodiments, the variable SliceNumInPic can be derived as SliceNumInPic = num_slices_in_pic_minus_subpic_num_minus1 + sps_num_subpics_minusl + 2.
[0172] Embodiments of the present disclosure also provide a new way to signal the stripe address. Figure 23 The syntax of an exemplary updated strip header according to some embodiments of the present disclosure is shown. As Figure 23 shown, the syntax elements or variables in the bitstream are shown in bold, and the changes to the previous VVC (e.g., Figure 13 the syntax shown therein) are shown in italic type, and the proposed deleted syntax is further shown in strikethrough.
[0173] In some embodiments, as Figure 23 shown, the variable picture_header_in_slice_header_flag can be signaled in the strip header to indicate whether there is a PH syntax structure within the strip header. There are constraints in VVC (e.g., VVC Draft 8) on the presence or absence of the PH syntax structure in the strip header and the number of stripes in an image. When there is a PH syntax structure in the strip header, the image should have only one stripe. Therefore, it is also not necessary to signal the stripe address. Thus, the signaling of the stripe address can be conditional on picture_header_in_slice_header_flag. As Figure 23As shown, the value of picture_header_in_slice_header_flag can be used as another condition for determining whether to signal the variable slice_address. When picture_header_in_slice_header_flag is equal to 1, signaling of the variable slice_address is skipped. In some embodiments, when the variable slice_address is not signaled, it can be inferred to be 0.
[0174] Embodiments of the present disclosure also provide a method for performing video coding. Figure 24 A flowchart of an example video coding method according to some embodiments of the present disclosure is shown, the method having a variable that signals the difference between the width of a video frame and an offset for calculating a horizontal wrap-around position. In some embodiments, Figure 24 The method 24000 shown in Figure 4 can be performed by the apparatus 400 shown in Figure 24 The method 24000 shown in Figure 15 can be performed according to the syntax shown in Figure 16 or the semantics shown in Figure 24 The method 24000 shown in Figure 24 includes a wrap-around motion compensation process performed according to the VVC standard. In some embodiments,
[0175] In step S24010, a wrap-around motion compensation flag is received, where the wrap-around motion compensation flag is associated with an image. For example, as shown in Figure 15 or Figure 16 the wrap-around motion compensation flag can be the variable pps_ref_wraparound_enabled_flag. In some embodiments, the image is in the bitstream. In some embodiments, the image is part of a 360-degree video.
[0176] In step S24020, it is determined whether the wrap-around motion compensation flag is enabled. For example, as shown in Figure 15 it is determined whether the variable pps_ref_wraparound_enabled_flag is equal to 1.
[0177] In step S24030, in response to determining that the wrap-around motion compensation flag is enabled, the difference between the width of the image and an offset for determining a horizontal wrap-around position is received. For example, as shown in Figure 15As shown, when it is determined that the variable pps_ref_wraparound_enabled_flag is equal to 1, the variable pps_pic_width_minus_wraparound_offset can be received or signaled. In some embodiments, the difference is less than or equal to the width of the image divided by the size of the minimum luminance coding block minus the size of the luminance coding tree block divided by the size of the minimum luminance coding block minus 2.
[0178] In some embodiments, in S24030, receiving the difference includes receiving a wraparound offset type flag. For example, as Figure 17 and Figure 18 shown, the flag wraparound_offset_type can be signaled to indicate whether the value of the signaled wraparound offset is the original ERP image width or the difference between the coded image width and the original ERP image width. In some embodiments, as Figure 17 and Figure 18 shown, the value of the flag wraparound_offset_type can be 0 or 1.
[0179] In step S24040, motion compensation is performed on the image according to the wraparound motion compensation flag and the difference. For example, motion compensation can be performed on the image according to the variable pps_ref_wraparound_enabled_flag and pps_pic_width_minus_wraparound_offset shown in Figure 15 and Figure 16 . In some embodiments, performing the motion compensation further includes: determining a wraparound motion compensation offset according to the width of the image and the difference, and performing the motion compensation on the image according to the wraparound motion compensation offset. In some embodiments, the wraparound motion compensation offset can be determined as the width of the image divided by the size of the minimum luminance coding block minus the difference.
[0180] Figure 25 FIG. shows a flowchart of an example video coding method according to some embodiments of the present disclosure, the video coding method having a variable that signals the number of stripes in a video frame minus 2. In some embodiments, Figure 25 the method 25000 shown in Figure 4 can be executed by the apparatus 400 shown in Figure 25 . In some embodiments, the method 25000 shown in Figure 19 can be executed according to the syntax shown in Figure 25 or the semantics shown in FIG. 20. In some embodiments, the method 25000 shown in
[0181] In step S25010, an image for encoding is received. In some embodiments, the image may include one or more slices. In some embodiments, the image is in a bitstream. In some embodiments, the one or more slices are rectangular slices.
[0182] In step S25020, a variable indicating the number of slices in the image minus 2 is signaled in the picture parameter set of the image. For example, this variable may be Figure 19 or num_slices_in_pic_minus2 as shown in FIG. 20. In some embodiments, adding 2 to the value of the variable may specify the number of rectangular slices in each image. In some embodiments, the variable is part of the PPS. In some embodiments, similar to the semantics shown in FIG. 20, the variable num_slices_in_pic_minus2 may replace the variable num_slices_in_pic_minus_1.
[0183] Figure 26 A flowchart of an example video encoding method with a variable according to some embodiments of the present disclosure is shown, where the variable signals a variable indicating the number of slices in a video frame minus the number of sub - pictures in the video frame minus 1. In some embodiments, Figure 26 the method 26000 shown in Figure 4 may be executed by the apparatus 400 shown in Figure 26 The method 26000 shown in Figure 21 may be executed according to the syntax shown in Figure 26 or the semantics shown in FIG. 22. In some embodiments,
[0184] In step S26010, an image for encoding is received. In some embodiments, the image may include one or more slices and one or more sub - pictures. In some embodiments, the video frame is in a bitstream. In some embodiments, the one or more slices are rectangular slices.
[0185] In step S26020, a variable indicating the number of slices in the image minus the number of sub - pictures in the image minus 1 is signaled in the picture parameter set of the image. For example, the variable may be as Figure 21 and num_slices_in_pic_minus_subpic_num_minus1 shown in FIG. 22. In some embodiments, the minimum number of slices may be equal to the number of sub - pictures plus 1. In some embodiments, as Figure 21As shown in FIG. 22, the number of slices to be signaled can be the number of slices in the picture minus the number of sub - pictures minus 1 (e.g., num_slices_in_pic_minus_subpic_num_minus1), instead of the number of slices in the picture minus 1 (e.g., num_slices_in_pic_minus1), to reduce the number of bits to be signaled. In some embodiments, the variable is part of the PPS. In some embodiments, as shown in FIG. 22, the value of num_slices_in_pic_minus_subpic_num_minus1 plus the number of sub - pictures and plus 1 can specify the number of rectangular slices in each picture with reference to the PPS. In some embodiments, similar to the semantics shown in FIG. 20, the variable num_slices_in_pic_minus2 can replace the variable num_slices_in_pic_minus_1.
[0186] In some embodiments, as shown in step S26020, the variable indicating the number of slices in the picture can be determined according to the variable indicating the number of slices in the video frame minus the number of sub - pictures in the video frame minus 1. For example, as Figure 21 and FIG. 22 show, the flag or variable SliceNumInPic can be derived according to the variable num_slices_in_pic_minus_subpic_num_minus1.
[0187] Figure 27 The flowchart of an example video encoding method with a variable indicating the presence of an image header syntax structure in the slice header of a video frame according to some embodiments of the present disclosure is shown. In some embodiments, Figure 27 the method 27000 shown in Figure 4 can be executed by the apparatus 400 shown in Figure 27 The method 27000 shown in Figure 23 can be executed according to the syntax shown in Figure 27 In some embodiments, the method 27000 shown in
[0188] In step S27010, an image for encoding is received. The image includes one or more slices. In some embodiments, the image is in the bitstream. In some embodiments, the one or more slices are rectangular slices.
[0189] In step S27020, a variable indicating whether there is an image header syntax structure for the image in the slice header of one or more slices is signaled. For example, the variable can be as Figure 23The picture_header_in_slice_header_flag shown. In some embodiments, such as Figure 23 shown, when the PH syntax structure is present in the slice header, the picture may have only one slice. Therefore, there is no need to signal the slice address either. Thus, the signaling of the slice address can be conditional on the variable picture_header_in_slice_header_flag. In some embodiments, the variable is part of the PPS.
[0190] In some embodiments, a non-transitory computer-readable storage medium including instructions is also provided, and the instructions can be executed by a device (such as the disclosed encoder and decoder) for performing the above method. Common forms of non-transitory media include, for example, floppy disks, hard disks, solid state drives, magnetic tapes or any other magnetic data storage media, CD-ROMs, any other optical data storage media, any physical media with a hole pattern, RAM, PROM, and EPROM, FLASH-EPROM or any other flash memory, NVRAM, caches, registers, any other memory chips or cartridges, and their networked versions. The device may include one or more processors (CPUs), input / output interfaces, network interfaces, and / or memories.
[0191] It should be noted that the relational terms herein, such as "first" and "second", are only used to distinguish an entity or operation from another entity or operation, and do not require or imply any actual relationship or order between these entities or operations. In addition, the words "comprising", "having", "containing", and "including" and other similar forms are equivalent in meaning and are open-ended, because one or more items following any of these words do not mean an exhaustive list of such one or more items, or are limited to the listed one or more items.
[0192] As used herein, unless otherwise specifically stated, the term "or" includes all possible combinations, unless it is infeasible. For example, if it is stated that a database may include A or B, then unless otherwise explicitly stated or infeasible, the database may include A, or B, or A and B. As a second example, if it is stated that a database may include A, B, or C, then unless otherwise explicitly stated or infeasible, the database may include A, or B, or C, or A and B, or A and C, or B and C, or A, B, and C.
[0193] It should be understood that the above embodiments can be implemented by hardware, or software (program code), or a combination of hardware and software. If implemented by software, it can be stored in the above computer-readable medium. When executed by a processor, the software can execute the disclosed method. The computing units and other functional units described in the present disclosure can be implemented by hardware, or software, or a combination of hardware and software. Those of ordinary skill in the art will also understand that the above multiple modules / units can be combined into one module / unit, and each of the above modules / units can be further divided into multiple sub-modules / sub-units.
[0194] In the foregoing specification, embodiments have been described with reference to numerous specific details, which may vary with the implementation. Certain modifications and changes can be made to the described embodiments. By considering the specification and practice of the present invention disclosed herein, other embodiments will be apparent to those skilled in the art. The specification and embodiments are to be considered as exemplary only, and the true scope and spirit of the present invention are indicated by the appended claims. The step sequences shown in the drawings are for illustrative purposes only and are not intended to be limited to any particular step sequence. Thus, those skilled in the art will understand that these steps can be executed in a different order while implementing the same method.
[0195] The embodiments can be further described using the following clauses: 1. A video decoding method, comprising: Receiving a surrounding motion compensation flag; Determining whether to enable surrounding motion compensation based on the surrounding motion compensation flag; In response to determining that the surrounding motion compensation is enabled, receiving data indicating a difference between the width of an image and an offset for determining a horizontal surrounding position; and Performing motion compensation according to the surrounding motion compensation flag and the difference. 2. The method according to clause 1, wherein the difference is in units of the size of the smallest luma coding block. 3. The method according to clause 2, wherein the difference is less than or equal to (pps_pic_width_in_luma_samples / MinCbSizeY)-(CtbSizeY / MinCbSizeY)–2, where pps_pic_width_in_luma_samples is the width of the image in luma samples, MinCbSizeY is the size of the smallest luma coding block, and CtbSizeY is the size of the luma coding tree block. 4. The method according to clause 1, wherein performing the motion compensation further comprises: Determine a wrap-around motion compensation offset based on the width of the image and the difference; and Perform the motion compensation based on the wrap-around motion compensation offset. 5. The method according to clause 4, wherein determining the wrap-around motion compensation offset based on the width of the image and the difference further includes: Divide the width of the image in luminance samples by the size of the smallest luminance coding block to generate a first value; and Determine the wrap-around motion compensation offset to be equal to the first value minus the difference. 6. The method according to clause 1, wherein receiving the data indicating the difference further includes: Receive a wrap-around offset type flag; Determine whether the wrap-around offset type flag is equal to a first value or a second value; In response to determining that the wrap-around offset type flag is equal to the first value, receive data indicating the difference between the width of the image and the offset for calculating the horizontal wrap-around position; and In response to determining that the wrap-around offset type flag is equal to the second value, receive data indicating the offset for calculating the horizontal wrap-around position. 7. The method according to clause 6, wherein each of the first value and the second value is 0 or 1. 8. The method according to clause 1, wherein the motion compensation is performed according to a general video coding standard. 9. The method according to clause 1, wherein the image is part of a 360-degree video sequence. 10. The method according to clause 1, wherein the wrap-around motion compensation flag and the difference are signaled in an image parameter set (PPS). 11. A video decoding method, comprising: Signal a wrap-around motion compensation flag indicating whether wrap-around motion compensation is enabled; In response to the wrap-around motion compensation flag indicating that the wrap-around motion compensation is enabled, signal data indicating the difference between the width of the image and the offset for determining the horizontal wrap-around position; and Perform motion compensation based on the wrap-around motion compensation flag and the difference. 12. The method according to clause 11, wherein the difference is in units of the size of the smallest luminance coding block. 13. The method according to clause 12, wherein the difference is less than or equal to (pps_pic_width_in_luma_samples / MinCbSizeY)-(CtbSizeY / MinCbSizeY)–2, where pps_pic_width_in_luma_samples is the width of the image in luma samples, MinCbSizeY is the size of the smallest luma coding block, and CtbSizeY is the size of the luma coding tree block. 14. The method according to clause 11, wherein performing the motion compensation further comprises: Determining a wrap-around motion compensation offset based on the width of the image and the difference; and Performing the motion compensation based on the wrap-around motion compensation offset. 15. The method according to clause 14, wherein determining the wrap-around motion compensation offset based on the width of the image and the difference further comprises: Dividing the width of the image in luma samples by the size of the smallest luma coding block to generate a first value; and Determining the wrap-around motion compensation offset to be equal to the first value minus the difference. 16. The method according to clause 11, wherein signaling the data indicating the difference further comprises: Signaling a wrap-around offset type flag, where the value of the wrap-around offset type flag is a first value or a second value; In response to the value of the wrap-around offset type flag being equal to the first value, signaling data indicating the difference between the width of the image and the offset for calculating the horizontal wrap-around position; and In response to the value of the wrap-around offset type flag being equal to the second value, signaling data for calculating the offset for the horizontal wrap-around position. 17. The method according to clause 16, wherein each of the first value and the second value is 0 or 1. 18. The method according to clause 11, wherein the motion compensation is performed according to a common video coding standard. 19. The method according to clause 11, wherein the image is part of a 360-degree video sequence. 20. The method according to clause 11, wherein the wrap-around motion compensation flag and the difference are signaled in an image parameter set (PPS). 21. A video coding method, comprising: Receiving an image for encoding, wherein the image includes one or more slices; and In the picture parameter set of the said picture, signal a variable indicating the number of stripes in the video frame minus 2. 22. The method according to clause 21, wherein the said picture is in the bitstream. 23. The method according to clause 21, wherein the said picture is encoded according to the general video coding standard. 24. The method according to clause 21, wherein the one or more stripes are rectangular stripes. 25. A video coding method, comprising: Receiving a picture for encoding, wherein the said picture includes one or more stripes and one or more sub-pictures; and In the picture parameter set of the said picture, signal a variable indicating the number of stripes in the said picture minus the number of sub-pictures in the picture minus 1. 26. The method according to clause 25, wherein the said picture is in the bitstream. 27. The method according to clause 25, further comprising: Determining a variable indicating the number of stripes in the said picture according to the variable indicating the number of stripes in the said picture minus the number of sub-pictures in the picture minus 1. 28. The method according to clause 25, wherein the said picture is encoded according to the general video coding standard. 29. The method according to clause 25, wherein the one or more stripes are rectangular stripes. 30. A video coding method, comprising: Receiving a picture for encoding, wherein the said picture includes one or more stripes; Signaling a variable indicating whether the picture header syntax structure of the said picture exists in the stripe headers of the one or more stripes; and Signaling the stripe address according to the said variable. 31. The method according to clause 30, wherein the said picture is in the bitstream. 32. The method according to clause 30, wherein the said picture is encoded according to the general video coding standard. 33. The method according to clause 30, wherein the one or more stripes are rectangular. 34. A system for performing video data processing, the said system comprising: A memory storing an instruction set; and A processor configured to execute the said instruction set to cause the said system to perform the following operations: Receiving a loop motion compensation flag; Determining whether to enable loop motion compensation based on the said loop motion compensation flag; In response to determining that the wrap-around motion compensation is enabled, receive data indicating a difference between the width of the image and an offset for determining a horizontal wrap-around position; and Perform motion compensation based on the wrap-around motion compensation flag and the difference. 35. The system according to clause 34, wherein the difference is in units of the size of the minimum luma coded block. 36. The system according to clause 35, wherein the difference is less than or equal to (pps_pic_width_in_luma_samples / MinCbSizeY)-(CtbSizeY / MinCbSizeY)–2, where pps_pic_width_in_luma_samples is the width of the image in luma samples, MinCbSizeY is the size of the minimum luma coded block, and CtbSizeY is the size of the luma coding tree block. 37. The system according to clause 34, wherein, when performing the motion compensation, the processor is configured to execute the instruction set to cause the system to perform: Determine a wrap-around motion compensation offset based on the width of the image and the difference; and Perform the motion compensation based on the wrap-around motion compensation offset. 38. The system according to clause 37, wherein, when determining the wrap-around motion compensation offset based on the width of the image and the difference, the processor is configured to execute the instruction set to cause the system to perform: Divide the width of the image in luma samples by the size of the minimum luma coded block to generate a first value; and Determine the wrap-around motion compensation offset to be equal to the first value minus the difference. 39. The system according to clause 34, wherein, when receiving the data indicating the difference, the processor is configured to execute the instruction set to cause the system to perform: Receive a wrap-around offset type flag; Determine whether the wrap-around offset type flag is equal to a first value or a second value; In response to determining that the wrap-around offset type flag is equal to the first value, receive data indicating a difference between the width of the image and an offset for calculating a horizontal wrap-around position; and In response to determining that the wrap-around offset type flag is equal to the second value, receive data indicating an offset for calculating a horizontal wrap-around position. 40. The system according to clause 39, wherein each of the first value and the second value is 0 or 1. 41. The system according to clause 34, wherein the motion compensation is performed according to a general video coding standard. 42. The system according to clause 34, wherein the image is part of a 360-degree video sequence. 43. The system according to clause 34, wherein the wrap-around motion compensation flag and the difference are signaled in an Image Parameter Set (PPS). 44. A system for performing video data processing, the system comprising: a memory storing an instruction set; and a processor configured to execute the instruction set to cause the system to perform the following operations: signal a wrap-around motion compensation flag indicating whether wrap-around motion compensation is enabled; in response to the wrap-around motion compensation flag indicating that the wrap-around motion compensation is enabled, signal data indicating a difference between the width of the image and an offset for determining a horizontal wrap-around position; and perform motion compensation according to the wrap-around motion compensation flag and the difference. 45. The system according to clause 44, wherein the difference is in units of the size of a minimum luminance coded block. 46. The system according to clause 45, wherein the difference is less than or equal to (pps_pic_width_in_luma_samples / MinCbSizeY)-(CtbSizeY / MinCbSizeY)–2, where pps_pic_width_in_luma_samples is the width of the image in luminance samples, MinCbSizeY is the size of a minimum luminance coded block, and CtbSizeY is the size of a luminance coded tree block. 47. The system according to clause 44, wherein, when performing the motion compensation, the processor is configured to execute the instruction set to cause the system to perform: determine a wrap-around motion compensation offset according to the width of the image and the difference; and perform the motion compensation according to the wrap-around motion compensation offset. 48. The system according to clause 47, wherein, when determining the wrap-around motion compensation offset according to the width of the image and the difference, the processor is configured to execute the instruction set to cause the system to perform: divide the width of the image in luminance samples by the size of a minimum luminance coded block to generate a first value; and determine the wrap-around motion compensation offset to be equal to the first value minus the difference. 49. The system according to clause 44, wherein, when receiving the data indicating the difference, the processor is configured to execute the instruction set to cause the system to perform: Signal a wrap-around offset type flag, wherein the value of the wrap-around offset type flag is a first value or a second value; In response to the value of the wrap-around offset type flag being equal to the first value, signal data indicating the difference between the width of the image and the offset used to calculate the horizontal wrap-around position; and In response to the value of the wrap-around offset type flag being equal to the second value, signal data for calculating the offset of the horizontal wrap-around position. 50. The system according to clause 49, wherein each of the first value and the second value is 0 or 1. 51. The system according to clause 44, wherein the motion compensation is performed according to a general video coding standard. 52. The system according to clause 44, wherein the image is part of a 360-degree video sequence. 3. The system according to clause 44, wherein the wrap-around motion compensation flag and the difference are signaled in an image parameter set (PPS). 54. A system for performing video coding, the system comprising: A memory storing an instruction set; and A processor configured to execute the instruction set to cause the system to perform the following operations: Receive an image for encoding, wherein the image includes one or more stripes; and In the image parameter set of the image, signal a variable indicating the number of stripes in the video frame minus 2. 55. The system according to clause 54, wherein the image is in a bitstream. 56. The system according to clause 54, wherein the image is encoded according to a general video coding standard. 57. The system according to clause 54, wherein the one or more stripes are rectangular stripes. 58. A system for performing video coding, the system comprising: A memory storing an instruction set; and A processor configured to execute the instruction set to cause the system to perform the following operations: Receive an image for encoding, wherein the image includes one or more stripes and one or more sub-images; and In the image parameter set of the image, signal a variable indicating the number of stripes in the image minus the number of sub-images in the image minus 1. 59. The system according to clause 58, wherein the image is in the bitstream. 60. The system according to clause 58, wherein the processor is configured to execute the instruction set to cause the system to perform: Determine a variable indicating the number of stripes in the image according to a variable that subtracts the number of sub-images in the image minus 1 from the number of stripes in the image according to the indication. 61. The system according to clause 58, wherein the image is encoded according to a common video coding standard. 62. The system according to clause 58, wherein the one or more stripes are rectangular stripes. 63. A system for performing video coding, the system comprising: A memory storing an instruction set; and A processor configured to execute the instruction set to cause the system to perform the following operations: Receive an image for encoding, wherein the image includes one or more stripes; Signal a variable indicating whether an image header syntax structure of the image exists in a stripe header indicating the one or more stripes; and Signal a stripe address according to the variable. 64. The system according to clause 63, wherein the image is in the bitstream. 65. The system according to clause 63, wherein the image is encoded according to a common video coding standard. 66. The system according to clause 63, wherein the one or more stripes are rectangular. 67. A non-transitory computer-readable medium storing an instruction set, the instruction set being executable by one or more processors of a device to cause the device to initiate a method for performing video data processing, the method comprising: Receive a loop motion compensation flag; Determine whether loop motion compensation is enabled based on the loop motion compensation flag; In response to determining that the loop motion compensation is enabled, receive data indicating a difference between the width of the image and an offset for determining a horizontal loop position; and Perform motion compensation according to the loop motion compensation flag and the difference. 68. The non-transitory computer-readable medium according to clause 67, wherein the difference is in units of the size of a minimum luminance coding block. 69. The non-transitory computer-readable medium according to clause 68, wherein the difference is less than or equal to (pps_pic_width_in_luma_samples / MinCbSizeY)-(CtbSizeY / MinCbSizeY)–2, where pps_pic_width_in_luma_samples is the width of the image in luma samples, MinCbSizeY is the size of the smallest luma coding block, and CtbSizeY is the size of the luma coding tree block. 70. The non-transitory computer-readable medium according to clause 67, wherein performing the motion compensation further comprises: determining a wrap-around motion compensation offset based on the width of the image and the difference; and performing the motion compensation based on the wrap-around motion compensation offset. 71. The non-transitory computer-readable medium according to clause 70, wherein determining the wrap-around motion compensation offset based on the width of the image and the difference further comprises: dividing the width of the image in luma samples by the size of the smallest luma coding block to generate a first value; and determining the wrap-around motion compensation offset to be equal to the first value minus the difference. 72. The non-transitory computer-readable medium according to clause 67, wherein receiving data indicating the difference further comprises: receiving a wrap-around offset type flag; determining whether the wrap-around offset type flag is equal to a first value or a second value; in response to determining that the wrap-around offset type flag is equal to the first value, receiving data indicating the difference between the width of the image and the offset for calculating the horizontal wrap-around position; and in response to determining that the wrap-around offset type flag is equal to the second value, receiving data indicating the offset for calculating the horizontal wrap-around position. 73. The non-transitory computer-readable medium according to clause 72, wherein each of the first value and the second value is 0 or 1. 74. The non-transitory computer-readable medium according to clause 67, wherein the motion compensation is performed according to a general video coding standard. 75. The non-transitory computer-readable medium according to clause 67, wherein the image is part of a 360-degree video sequence. 76. The non-transitory computer-readable medium according to clause 67, wherein the wrap-around motion compensation flag and the difference are signaled in an image parameter set (PPS). 77. A non - transitory computer - readable medium storing an instruction set, the instruction set being executable by one or more processors of a device to cause the device to initiate a method for performing video data processing, the method comprising: Signaling a wrap - around motion compensation flag, the flag indicating whether wrap - around motion compensation is enabled; In response to the wrap - around motion compensation flag indicating that the wrap - around motion compensation is enabled, signaling data indicating a difference between the width of the image and an offset for determining a horizontal wrap - around position; and Performing motion compensation based on the wrap - around motion compensation flag and the difference. 78. The non - transitory computer - readable medium according to clause 77, wherein the difference is in units of the size of a minimum - luminance - coded block. 79. The non - transitory computer - readable medium according to clause 78, wherein the difference is less than or equal to (pps_pic_width_in_luma_samples / MinCbSizeY)-(CtbSizeY / MinCbSizeY)–2, where pps_pic_width_in_luma_samples is the width of the image in luminance samples, MinCbSizeY is the size of a minimum - luminance - coded block, and CtbSizeY is the size of a luminance - coded tree block. 80. The non - transitory computer - readable medium according to clause 77, wherein performing the motion compensation further comprises: Determining a wrap - around motion compensation offset based on the width of the image and the difference; and Performing the motion compensation based on the wrap - around motion compensation offset. 81. The non - transitory computer - readable medium according to clause 80, wherein determining the wrap - around motion compensation offset based on the width of the image and the difference further comprises: Dividing the width of the image in luminance samples by the size of a minimum - luminance - coded block to generate a first value; and Determining the wrap - around motion compensation offset to be equal to the first value minus the difference. 82. The non - transitory computer - readable medium according to clause 77, wherein receiving data indicating the difference further comprises: Signaling a wrap - around offset type flag, where the value of the wrap - around offset type flag is a first value or a second value; In response to the value of the wrap - around offset type flag being equal to the first value, signaling data indicating a difference between the width of the image and an offset for calculating a horizontal wrap - around position; and In response to the value of the wrap-around offset type flag being equal to the second value, signal data for calculating an offset of a horizontal wrap-around position. 83. The non-transitory computer-readable medium according to clause 82, wherein each of the first value and the second value is 0 or 1. 84. The non-transitory computer-readable medium according to clause 77, wherein the motion compensation is performed according to a general video coding standard. 85. The non-transitory computer-readable medium according to clause 77, wherein the image is part of a 360-degree video sequence. 86. The non-transitory computer-readable medium according to clause 77, wherein the wrap-around motion compensation flag and the difference are signaled in a picture parameter set (PPS). 87. A non-transitory computer-readable medium storing a set of instructions, the set of instructions being executable by one or more processors of a device to cause the device to initiate a method for performing video coding, the method comprising: receiving an image for encoding, wherein the image includes one or more slices; and in a picture parameter set of the image, signaling a variable indicating the number of slices in a video frame minus 2. 88. The non-transitory computer-readable medium according to clause 87, wherein the image is in a bitstream. 89. The non-transitory computer-readable medium according to clause 87, wherein the image is encoded according to a general video coding standard. 90. The non-transitory computer-readable medium according to clause 87, wherein the one or more slices are rectangular slices. 91. A non-transitory computer-readable medium storing a set of instructions, the set of instructions being executable by one or more processors of a device to cause the device to initiate a method for performing video coding, the method comprising: receiving an image for encoding, wherein the image includes one or more slices and one or more sub-images; and in a picture parameter set of the image, signaling a variable indicating the number of slices in the image minus the number of sub-images in the image minus 1. 92. The non-transitory computer-readable medium according to clause 91, wherein the image is in a bitstream. 93. The non-transitory computer-readable medium according to clause 91, further comprising: determining a variable indicating the number of slices in the image according to the variable indicating the number of slices in the image minus the number of sub-images in the image minus 1. 94. The non-transitory computer-readable medium according to clause 91, wherein the image is encoded according to a general video coding standard. 95. The non-transitory computer-readable medium according to clause 91, wherein the one or more stripes are rectangular stripes. 96. A non-transitory computer-readable medium storing an instruction set executable by one or more processors of a device to cause the device to initiate a method for performing video coding, the method comprising: receiving an image for encoding, wherein the image includes one or more stripes; signaling a variable of an image header syntax structure indicating whether an image header of the image is present in a stripe header of the one or more stripes; and signaling a stripe address according to the variable. 97. The non-transitory computer-readable medium according to clause 96, wherein the image is in a bitstream. 98. The non-transitory computer-readable medium according to clause 96, wherein the image is encoded according to a general video coding standard. 99. The non-transitory computer-readable medium according to clause 96, wherein the one or more stripes are rectangular.
[0196] In the drawings and the specification, exemplary embodiments have been disclosed. However, many variations and modifications can be made to these embodiments. Therefore, although specific terms are employed, they are used in a general and descriptive sense only and not for purposes of limitation.
Claims
1. A video encoding method, comprising: signaling a rectangular strip flag in a Picture Parameter Set (PPS); determining whether the rectangular strip flag is equal to a first value or a second value, wherein the rectangular strip flag being equal to the first value indicates that a raster scan strip mode is used for each picture referenced by the PPS, and the rectangular strip flag being equal to the second value indicates that a rectangular strip mode is used for each picture referenced by the PPS; in response to the rectangular strip flag being equal to the second value, signaling a parameter in the PPS, the parameter indicating the number of rectangular strips in each picture referenced by the PPS minus 1; and signaling a wrap-around motion compensation flag and a picture width in the PPS; determining a wrap-around motion compensation parameter according to the picture width; performing wrap-around motion compensation according to the wrap-around motion compensation flag and the wrap-around motion compensation parameter.
2. The video encoding method according to claim 1, wherein, each of the first value and the second value is 0 or 1.
3. The video encoding method according to claim 1, wherein, the motion compensation is performed according to a general video coding standard.
4. The video encoding method according to claim 1, wherein, the picture is part of a 360-degree video sequence.
5. A video decoding method, comprising: receiving a rectangular strip flag from a Picture Parameter Set (PPS); determining whether the rectangular strip flag is equal to a first value or a second value, wherein the rectangular strip flag being equal to the first value indicates that a raster scan strip mode is used for each picture referenced by the PPS, and the rectangular strip flag being equal to the second value indicates that a rectangular strip mode is used for each picture referenced by the PPS; in response to the rectangular strip flag being equal to the second value, receiving a parameter in the PPS, the parameter indicating the number of rectangular strips in each picture referenced by the PPS minus 1; and receiving a wrap-around motion compensation flag and a picture width from the PPS; determining a wrap-around motion compensation parameter according to the picture width; performing wrap-around motion compensation according to the wrap-around motion compensation flag and the wrap-around motion compensation parameter.
6. The video decoding method according to claim 5, wherein, each of the first value and the second value is 0 or 1.
7. The video decoding method according to claim 5, wherein, the motion compensation is performed according to a general video coding standard.
8. A non-transitory computer-readable storage medium storing a bitstream, the non-transitory computer-readable medium being part of a computing device, the bitstream being generated by execution of a set of instructions by one or more processors of the computing device, wherein, execution of the set of instructions causes the computing device to perform: signaling a rectangular strip flag in a Picture Parameter Set (PPS); determining whether the rectangular strip flag is equal to a first value or a second value, wherein the rectangular strip flag being equal to the first value indicates that a raster scan strip mode is used for each picture referenced by the PPS, and the rectangular strip flag being equal to the second value indicates that a rectangular strip mode is used for each picture referenced by the PPS; In response to the rectangular strip flag being equal to the second value, signal a parameter in the PPS, the parameter indicating the number of rectangular strips in each image referenced by the PPS minus one; and Signal the wrap-around motion compensation flag and the image width in the PPS; Determine wrap-around motion compensation parameters based on the image width; perform wrap-around motion compensation based on the wrap-around motion compensation flag and the wrap-around motion compensation parameters.
9. The non-transitory computer-readable storage medium according to claim 8, wherein, each of the first value and the second value is 0 or 1.
10. The non-transitory computer-readable storage medium according to claim 8, wherein, perform the motion compensation according to a general video coding standard.
Citation Information
Patent Citations
Encoding control apparatus, encoding control method, and storage medium
US20080298465A1
Video coding apparatus and video decoding apparatus
US20210136407A1
Video coding device and video decoding device
WO2019078169A1
Methods and apparatus for flexible grid regions
WO2020056247A1