Method of signaling video encoding data

By receiving the surround motion compensation flag and strip information, enabling surround motion compensation and processing strip addresses, the problem of insufficient coding efficiency in existing video coding standards is solved, and more efficient video data processing is achieved.

CN120075459BActive Publication Date: 2026-01-09HFI INNOVATION INC
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510317570.1
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Priority Date
2020-03-26
Filing Date
2021-03-26
Publication Date
2026-01-09
Estimated Expiration
2041-03-26

AI Technical Summary

Technical Problem

Existing video coding standards have not fully utilized surround motion compensation and strip layout information in high-efficiency video coding technologies, resulting in coding efficiency that needs to be improved.

Method used

By receiving the surround motion compensation flag and strip information, it determines whether surround motion compensation is enabled, and notifies the relevant variables in the image parameter set with signals to perform motion compensation and strip address processing.

Benefits of technology

It improves the coding efficiency of video encoding, enhances the flexibility and accuracy of video data processing, and adapts to the needs of different coding standards.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120075459B_ABST
    Figure CN120075459B_ABST
Patent Text Reader

Abstract

The present disclosure provides systems and methods for wraparound motion compensation. An example method includes receiving a wraparound motion compensation flag, determining whether wraparound motion compensation is enabled based on the wraparound motion compensation flag, in response to determining that the wraparound motion compensation is enabled, receiving data indicative of a difference between a width of a picture and an offset used to determine a horizontal wraparound position, and performing motion compensation according to the wraparound motion compensation flag and the difference.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] Cross-reference to related applications

[0002] This disclosure claims priority to and the benefits of priority to U.S. Provisional Patent Application No. 63 / 000,443, filed March 26, 2020. The provisional application is incorporated herein by reference in its entirety. Technical Field

[0003] This disclosure generally relates to video data processing, and more specifically, to methods and apparatus for signaling information regarding surround motion compensation, strip layout, and strip address. Background Technology

[0004] Video is a set of still images (or "frames") that capture visual information. To reduce storage memory and transmission bandwidth, video can be compressed before storage or transmission and then decompressed before display. The compression process is usually called encoding, and the decompression process is usually called decoding. There are various video coding formats that use standardized video coding techniques, the most common being prediction, transform, quantization, entropy coding, and in-loop filtering. Standardization organizations have developed video coding standards, such as the High Efficiency Video Coding (HEVC / H.265) standard, the Universal Video Coding (VVC / H.266) standard, and the AVS standard, specifying particular video coding formats. As more and more advanced video coding techniques are adopted in video standards, the coding efficiency of new video coding standards is becoming increasingly higher. Summary of the Invention

[0005] This disclosure provides a method for signaling video encoded data, the method comprising: receiving a surround motion compensation flag; determining whether to enable surround motion compensation based on the surround motion compensation flag; in response to determining that surround motion compensation is enabled, receiving data of the difference between the width of an indication image and an offset used to determine a horizontal surround position; and performing motion compensation based on the surround motion compensation flag and the difference.

[0006] This disclosure also provides a method for signaling video encoded data, the method comprising: receiving an image for encoding, wherein the image includes one or more stripes; and signaling a variable indicating the number of stripes in a video frame minus 2, within an image parameter set of the image.

[0007] This disclosure also provides a method for signaling video encoded data, the method comprising: receiving an image for encoding, wherein the image includes one or more stripes and one or more sub-images; and, in the image parameter set of the image, signaling a variable indicating the number of stripes in the image minus the number of sub-images in the image minus 1.

[0008] The embodiments of the present disclosure further provide a method for signaling video coding data, the method comprising: receiving an image for coding, wherein the image comprises one or more slices; signaling a variable indicating whether a picture header syntax structure of the image is present in a slice header of the one or more slices; and signaling a slice address according to the variable.

[0009] Embodiments of the present disclosure further provide a system for performing video data processing, the system comprising: a memory storing a set of instructions; and a processor configured to execute the set of instructions to cause the system to perform: receiving a wrap-around motion compensation flag; determining whether to enable wrap-around motion compensation based on the wrap-around motion compensation flag; in response to determining to enable the wrap-around motion compensation, receiving data indicating a difference between a width of a picture and an offset for determining a horizontal wrap-around position; and performing motion compensation according to the wrap-around motion compensation flag and the difference.

[0010] Embodiments of the present disclosure further provide a system for performing video data processing, the system comprising: a memory storing a set of instructions; and a processor configured to execute the set of instructions to cause the system to perform: receiving an image for coding, wherein the image comprises one or more slices; and in a picture parameter set of the image, signaling a variable indicating a number of slices in a video frame minus 2.

[0011] Embodiments of the present disclosure further provide a system for performing video data processing, the system comprising: a memory storing a set of instructions; and a processor configured to execute the set of instructions to cause the system to perform: receiving an image for coding, wherein the image comprises one or more slices and one or more sub-pictures; and in a picture parameter set of the image, signaling a variable indicating a number of slices in the image minus a number of sub-pictures in the image minus 1.

[0012] Embodiments of the present disclosure further provide a system for performing video data processing, the system comprising: a memory storing a set of instructions; and a processor configured to execute the set of instructions to cause the system to perform: receiving an image for coding, wherein the image comprises one or more slices; signaling a variable indicating whether a picture header syntax structure of the image is present in a slice header of the one or more slices; and signaling a slice address according to the variable.

[0013] Embodiments of the present disclosure also provide a non-transitory computer- readable medium storing a set of instructions, which is executable by one or more processors of an apparatus to cause the apparatus to initiate a method for performing video data processing, the method comprising: receiving a wrap-around motion compensation flag; determining whether wrap-around motion compensation is enabled based on the wrap-around motion compensation flag; in response to determining that the wrap-around motion compensation is enabled, receiving data indicating a difference between a width of a picture and an offset for determining a horizontal wrap-around position; and performing motion compensation according to the wrap-around motion compensation flag and the difference.

[0014] Embodiments of the present disclosure also provide a non-transitory computer- readable medium storing a set of instructions, which is executable by one or more processors of an apparatus to cause the apparatus to initiate a method for performing video data processing, the method comprising: receiving a picture for encoding, wherein the picture comprises one or more slices; and in a picture parameter set of the picture, signaling a variable indicating a number of slices in a video frame minus 2.

[0015] Embodiments of the present disclosure also provide a non-transitory computer- readable medium storing a set of instructions, which is executable by one or more processors of an apparatus to cause the apparatus to initiate a method for performing video data processing, the method comprising: receiving a picture for encoding, wherein the picture comprises one or more slices and one or more sub-pictures; and in a picture parameter set of the picture, signaling a variable indicating a number of slices in the picture minus a number of sub-pictures in the picture minus 1.

[0016] Embodiments of the present disclosure also provide a non-transitory computer- readable medium storing a set of instructions, which is executable by one or more processors of an apparatus to cause the apparatus to initiate a method for performing video data processing, the method comprising: receiving a picture for encoding, wherein the picture comprises one or more slices; signaling a variable indicating whether a picture header syntax structure of the picture is present in a slice header of the one or more slices; and signaling a slice address according to the variable. BRIEF DESCRIPTION OF DRAWINGS

[0017] Embodiments of the present disclosure and various aspects thereof are illustrated in the detailed description and the accompanying drawings. Various features shown in the drawings are not drawn to scale.

[0018] Figure 1 The structure of an example video sequence according to some embodiments of the present disclosure is shown.

[0019] Figure 2A A schematic diagram of an example encoding process according to some embodiments of the present disclosure is shown.

[0020] Figure 2BA diagram illustrating another example encoding process according to some embodiments of the disclosure is shown.

[0021] Figure 3A A diagram illustrating an example decoding process according to some embodiments of the disclosure is shown.

[0022] Figure 3B A diagram illustrating another example decoding process according to some embodiments of the disclosure is shown.

[0023] Figure 4 A block diagram illustrating an example apparatus for encoding or decoding a video according to some embodiments of the disclosure is shown.

[0024] Figure 5A A diagram illustrating an example blending operation for generating a reconstructed equirectangular projection according to some embodiments of the disclosure is shown.

[0025] Figure 5B A diagram illustrating an example cropping operation for generating a reconstructed equirectangular projection according to some embodiments of the disclosure is shown.

[0026] Figure 6A A diagram illustrating an example horizontal wrap-around motion compensation process for equirectangular projection according to some embodiments of the disclosure is shown.

[0027] Figure 6B A diagram illustrating an example horizontal wrap-around motion compensation process for padding equirectangular projection according to some embodiments of the disclosure is shown.

[0028] Figure 7 A syntax of an example advanced wrap-around offset according to some embodiments of the disclosure is shown.

[0029] Figure 8 A semantics of an example advanced wrap-around offset according to some embodiments of the disclosure is shown.

[0030] Figure 9 A diagram illustrating an example slice and subpicture partitioning of an image according to some embodiments of the disclosure is shown.

[0031] Figure 10 A diagram illustrating an example slice and subpicture partitioning of an image with different slices and subpictures according to some embodiments of the disclosure is shown.

[0032] Figure 11 A syntax of an example picture parameter set for tile mapping and slice layout according to some embodiments of the disclosure is shown.

[0033] A semantics of an example picture parameter set for tile mapping and slice layout according to some embodiments of the disclosure is shown.

[0034] Figure 13 Syntax of an example slice header is shown in accordance with some embodiments of the present disclosure.

[0035] Figure 14 Semantics of an example slice header is shown in accordance with some embodiments of the present disclosure.

[0036] Figure 15 Syntax of an example improved picture parameter set is shown in accordance with some embodiments of the present disclosure.

[0037] Figure 16 Semantics of an example improved picture parameter set is shown in accordance with some embodiments of the present disclosure.

[0038] Figure 17 Syntax of an example picture parameter set with variable wraparound_offset_type is shown in accordance with some embodiments of the present disclosure.

[0039] Figure 18 Semantics of an example improved picture parameter set with variable wraparound_offset_type is shown in accordance with some embodiments of the present disclosure.

[0040] Figure 19 Syntax of an example picture parameter set with variable num_slices_in_pic_minus2 is shown in accordance with some embodiments of the present disclosure.

[0041] Figure 20 shows semantics of an example improved picture parameter set with variable num_slices_in_pic_minus2 in accordance with some embodiments of the present disclosure.

[0042] Figure 21 Syntax of an example picture parameter set with variable num_slices_in_pic_minus_subpic_num_minus1 is shown in accordance with some embodiments of the present disclosure.

[0043] Figure 22 shows semantics of an example improved picture parameter set with variable num_slices_in_pic_minus_subpic_num_minus1 in accordance with some embodiments of the present disclosure.

[0044] Figure 23 Syntax of an example updated slice header is shown in accordance with some embodiments of the present disclosure.

[0045] Figure 24A flowchart of an example video encoding method having a variable that signals a difference between a width of a video frame and an offset used to calculate a horizontal wrap-around position is shown, in accordance with some embodiments of the disclosure.

[0046] Figure 25 A flowchart of an example video encoding method having a variable that signals a number of slices in a video frame minus 2 is shown, in accordance with some embodiments of the disclosure.

[0047] Figure 26 A flowchart of an example video encoding method having a variable that signals a number of slices in a video frame minus a number of sub-pictures in a video frame minus 1 is shown, in accordance with some embodiments of the disclosure.

[0048] Figure 27 A flowchart of an example video encoding method having a variable that indicates whether an image header syntax structure is present in a slice header of a video frame is shown, in accordance with some embodiments of the disclosure. DETAILED DESCRIPTION

[0049] Reference will now be made in detail to the example embodiments, examples of which are illustrated in the accompanying drawings. The following description refers to the accompanying drawings, in which the same numbers represent the same or similar elements between the several figures. The implementation set forth in the following description of example embodiments is not intended to be exhaustive or to be limited to any one implementation. Rather, it is intended to be illustrative only and to be read in conjunction with the appended claims to provide a full understanding of the aspects of the disclosure. The following detailed description is not meant to limit the aspects of the disclosure to the particular examples presented, as such examples are merely for illustrative purposes only. The terminology used herein is for the purpose of describing only the aspects of the disclosure and is not intended to be limiting. The terminology is intended to be interpreted broadly according to the principles of the disclosure.

[0050] The Joint Video Expert Team (JVET) of the ITU-T Video Coding Experts Group (ITU-T VCEG) and the ISO / IEC Moving Picture Experts Group (ISO / IEC MPEG) is currently developing the Versatile Video Coding (VVC / H.266) standard. The VVC standard aims to double the compression efficiency of its predecessor, High Efficiency Video Coding (HEVC / H.265) standard. In other words, the goal of VVC is to achieve the same subjective quality using half the bandwidth of HEVC / H.265.

[0051] To achieve the same subjective quality using half the bandwidth of HEVC / H.265, the JVET has been developing techniques beyond HEVC using the Joint Exploration Model (JEM) reference software. As coding techniques are incorporated into JEM, JEM achieves higher coding performance than HEVC.

[0052] The VVC standard is recently developed and continues to include more coding techniques that provide better compression performance. VVC is based on a hybrid video coding system that has been used in modern video compression standards such as HEVC, H.264 / AVC, MPEG2, H.263, etc.

[0053] A video is a set of still images (or “frames”) arranged in a temporal order to store visual information. Video capture devices (e.g., cameras) can be used to capture and store these images in a temporal order, and video playback devices (e.g., televisions, computers, smartphones, tablet computers, video players, or any end-user terminal with display functionality) can be used to display such images in a temporal sequence. Moreover, in some applications, video capture devices can transmit captured videos to video playback devices (e.g., computers with monitors) in real time, for example, for surveillance, conferencing, or live broadcasting.

[0054] To reduce the storage space and transmission bandwidth required for such applications, videos can be compressed before storage and transmission, and decompressed before display. Compression and decompression can be implemented by software executed by a processor (e.g., a processor of a general-purpose computer) or special-purpose hardware. Modules for compression are often referred to as “encoders”, and modules for decompression are often referred to as “decoders”. Encoders and decoders can be collectively referred to as “codecs”. Encoders and decoders can be implemented in any of a variety of suitable hardware, software, or combinations thereof. For example, hardware implementations of encoders and decoders can include circuitry such as one or more microprocessors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), discrete logic or any combination thereof. Software implementations of encoders and decoders can include program code, computer-executable instructions, firmware, or any suitable computer-implemented algorithms or processes fixed in a computer-readable medium. Video compression and decompression can be implemented by various algorithms or standards, such as MPEG-1, MPEG-2, MPEG-4, H.26x series, etc. In some applications, a codec can decompress a video from a first encoding standard and re-compress the decompressed video using a second encoding standard, in which case the codec can be referred to as a “transcoder”.

[0055] A video encoding process can identify and retain useful information that can be used to reconstruct images, and ignore unimportant reconstruction information. Such an encoding process can be referred to as “lossy” if the ignored unimportant information cannot be completely reconstructed. Otherwise, it can be referred to as “lossless”. Most encoding processes are lossy, which is a trade-off to reduce the required storage space and transmission bandwidth.

[0056] In many cases, useful information about an encoded image (referred to as the "current image") includes changes relative to a reference image (e.g., a previously encoded and reconstructed image). Such changes can include variations in pixel position, brightness, or color, with positional changes being the most important. The positional changes of a set of pixels representing an object can reflect the object's movement between the reference and current images.

[0057] An image encoded without referencing another image (i.e., it is its own reference image) is called an "I-image". An image encoded using a previous image as a reference image is called a "P-image", and an image encoded using both a previous image and a future image as reference images is called a "B-image" (the reference is "bidirectional").

[0058] Figure 1 The structure of an example video sequence 100 according to some embodiments of the present disclosure is shown. The video sequence 100 may be live video or video that has been captured and archived. The video 100 may be real-life video, computer-generated video (e.g., computer game video), or a combination of both (e.g., real video with augmented reality effects). The video sequence 100 may be input from a video capture device (e.g., a camera), a video archive containing previously captured video (e.g., a video file stored on a storage device), or a video feed interface (e.g., a video broadcast transceiver) receiving video from a video content provider.

[0059] like Figure 1 As shown, video sequence 100 may include a series of images arranged temporally along a timeline, including images 102, 104, 106, and 108. Images 102-106 are consecutive, with more images between images 106 and 108. Figure 1 In this diagram, image 102 is an I-image, and its reference image is image 102 itself. Image 104 is a P-image, and its reference image is image 102, as indicated by the arrow. Image 106 is a B-image, and its reference images are images 104 and 108, as indicated by the arrow. In some embodiments, the reference image of an image (e.g., image 104) may not immediately precede or follow the image. For example, the reference image of image 104 may be an image preceding image 102. It should be noted that the reference images of images 102-106 are merely examples, and this disclosure does not limit the scope to such cases. Figure 1 An example of the reference image shown.

[0060] Typically, due to the computational complexity of encoding and decoding tasks, video codecs do not encode or decode the entire image at once. Instead, they can segment the image into basic segments and encode or decode each segment sequentially. In this disclosure, such basic segments are referred to as basic processing units (“BPUs”). For example, Figure 1Structure 110 in FIG. 1 illustrates an example structure of a picture (e.g., any of pictures 102-108) of video sequence 100. In structure 110, the picture is divided into 4x4 basic processing units, the boundaries of which are shown as dashed lines. In some embodiments, a basic processing unit can be referred to as a “macroblock” in some video coding standards (e.g., the MPEG family, H.261, H.263, or H.264 / AVC), or as a “coding tree unit” (“CTU”) in some other video coding standards (e.g., H.265 / HEVC or H.266 / VVC). A basic processing unit can have a variable size in pixels, such as 128x128, 64x64, 32x32, 16x16, 4x8, 16x32, or any arbitrary shape and size. The size and shape of a basic processing unit can be selected for a picture based on a balance of coding efficiency and level of detail to be maintained in the basic processing unit.

[0061] A basic processing unit can be a logical unit that can include a set of different types of video data stored in computer memory (e.g., in a video frame buffer). For example, a basic processing unit of a color picture can include a luma component (Y) representing achromatic luminance information, one or more chroma components (e.g., Cb and Cr) representing color information, and associated syntax elements, where the luma and chroma components can have the same size as the basic processing unit. In some video coding standards (e.g., H.265 / HEVC or H.266 / VVC), the luma and chroma components can be referred to as “coding tree blocks” (“CTBs”). Any operation performed on a basic processing unit can be repeated on each of its luma and chroma components.

[0062] Video coding has multiple stages of operations, examples of which are shown in Figures 2A-2B and Figures 3A-3BAs shown. For each stage, the size of the basic processing unit can still be too large for processing, and thus can be further divided into segments, referred to in this disclosure as“basic processing subunits.” In some embodiments, the basic processing subunits can be referred to as“blocks” in some video coding standards (e.g., the MPEG family, H.261, H.263, or H.264 / AVC), or as“coding units” (“CUs”) in some other video coding standards (e.g., H.265 / HEVC or H.266 / VVC). The basic processing subunits can have the same size as the basic processing unit or have a smaller size than the basic processing unit. Similar to the basic processing unit, the basic processing subunit is also a logical unit, which can include a set of different types of video data (e.g., Y, Cb, Cr, and associated syntax elements) stored in a computer memory (e.g., in a video frame buffer). Any operations performed on the basic processing subunit can be repeated on each of its luma and chroma components. It should be noted that such division can be performed to further levels as needed for processing. It should also be noted that different stages can use different schemes to divide the basic processing unit.

[0063] For example, at the mode decision stage (examples of which are shown in Figure 2B The encoder can decide what prediction mode (e.g., intra prediction or inter prediction) to use for a basic processing unit that can be too large to make such a decision. The encoder can divide the basic processing unit into multiple basic processing subunits (e.g., as CUs in H.265 / HEVC or H.266 / VVC), and decide the prediction type for each individual basic processing subunit.

[0064] For another example, at the prediction stage (examples of which are shown in Figures 2A-2B The encoder can perform the prediction operation at the level of basic processing subunits (e.g., CUs). However, in some cases, the basic processing subunits can still be too large for processing. The encoder can further divide the basic processing subunits into smaller segments (e.g., referred to as“prediction blocks” or“PBs” in H.265 / HEVC or H.266 / VVC), at which level the prediction operation can be performed.

[0065] For another example, at the transform stage (examples of which are shown in Figures 2A-2BAs shown in the diagram, the encoder can perform transformation operations on residual basic processing subunits (e.g., CUs). However, in some cases, the basic processing subunits may still be too large to process. The encoder can further divide the basic processing subunits into smaller segments (e.g., referred to as "transform blocks" or "TBs" in H.265 / HEVC or H.266 / VVC), at which level transformation operations can be performed. It is important to note that the partitioning scheme of the same basic processing subunit can differ between the prediction and transformation phases. For example, in H.265 / HEVC or H.266 / VVC, the prediction blocks and transform blocks of the same CU can have different sizes and numbers.

[0066] exist Figure 1 In structure 110, the basic processing unit 112 is further divided into 3×3 basic processing sub-units, the boundaries of which are shown by dashed lines. Different basic processing units of the same image can be divided into basic processing sub-units in different schemes.

[0067] In some implementations, to provide parallel processing capabilities and fault tolerance for video encoding and decoding, an image can be divided into regions for processing, such that the encoding or decoding process for a given region of the image can be independent of information from any other region of the image. In other words, each region of the image can be processed independently. By doing so, the codec can process different regions of the image in parallel, thereby improving encoding efficiency. Furthermore, when data in one region is corrupted during processing or lost during network transmission, the codec can correctly encode or decode other regions of the same image without relying on the corrupted or lost data, thus providing fault tolerance. In some video coding standards, images can be divided into different types of regions. For example, H.265 / HEVC and H.266 / VVC provide two types of regions: “slices” and “tiles.” It should also be noted that different images in the video sequence 100 can have different partitioning schemes for dividing the image into regions.

[0068] For example, in Figure 1 In the diagram, structure 110 is divided into three regions 114, 116, and 118, whose boundaries are shown as solid lines within structure 110. Region 114 comprises four basic processing units. Regions 116 and 118 each comprise six basic processing units. It should be noted that... Figure 1 The basic processing unit, basic processing subunit, and structural region of 110 are merely examples, and this disclosure does not limit its embodiments.

[0069] Figure 2A A schematic diagram of an exemplary encoding process 200A according to an embodiment of the present disclosure is shown. For example, the encoding process 200A may be performed by an encoder. Figure 2AAs shown, the encoder can encode the video sequence 202 into a video bitstream 228 according to the process 200A. Similar to the structure 110 in Figure 1 , the video sequence 202 can include a set of pictures (referred to as “original pictures”) arranged in a temporal order. Similar to the structure 110 in Figure 1 , each original picture of the video sequence 202 can be divided by the encoder into basic processing units, basic processing sub-units, or regions for processing. In some embodiments, the encoder can perform the process 200A at the level of basic processing units for each original picture of the video sequence 202. For example, the encoder can perform the process 200A in an iterative manner, where the encoder can encode a basic processing unit in one iteration of the process 200A. In some embodiments, the encoder can perform the process 200A in parallel for regions (e.g., regions 114-118) of each original picture of the video sequence 202.

[0070] Referring to Figure 2A , the encoder can feed a basic processing unit (referred to as “original BPU”) of an original picture of the video sequence 202 to the prediction stage 204 to generate prediction data 206 and a predicted BPU 208. The encoder can subtract the predicted BPU 208 from the original BPU to generate a residual BPU 210. The encoder can feed the residual BPU 210 to the transform stage 212 and the quantization stage 214 to generate quantized transform coefficients 216. The encoder can feed the prediction data 206 and the quantized transform coefficients 216 to the binary encoding stage 226 to generate the video bitstream 228. The components 202, 204, 206, 208, 210, 212, 214, 216, 226, and 228 can be referred to as the “forward path.” During the process 200A, after the quantization stage 214, the encoder can feed the quantized transform coefficients 216 to the inverse quantization stage 218 and the inverse transform stage 220 to generate a reconstructed residual BPU 222. The encoder can add the reconstructed residual BPU 222 to the predicted BPU 208 to generate a prediction reference 224, which is used in the prediction stage 204 for the next iteration of the process 200A. The components 218, 220, 222, and 224 of the process 200A can be referred to as the “reconstruction path.” The reconstruction path can be used to ensure that both the encoder and the decoder use the same reference data for prediction.

[0071] The encoder can iteratively perform the process 200A to encode each original BPU of an original picture (in the forward path) and generate a prediction reference 224 for encoding the next original BPU of the original picture (in the reconstruction path). After encoding all original BPUs of an original picture, the encoder can proceed to encode the next picture in the video sequence 202.

[0072] Referring to process 200A, an encoder can receive a video sequence 202 generated by a video capture device (e.g., a camera). As used herein, the term “receive” can refer to any action that gets, obtains, retrieves, acquires, reads, accesses, or any action that is used to input data in any manner.

[0073] At a prediction stage 204, at a current iteration, the encoder can receive an original BPU and a prediction reference 224 and perform a prediction operation to generate prediction data 206 and a predicted BPU 208. The prediction reference 224 can be generated from the reconstruction path of a previous iteration of process 200A. The purpose of prediction stage 204 is to reduce information redundancy by extracting, from the prediction data 206 and the prediction reference 224, prediction data 206 that can be used to reconstruct the original BPU into the predicted BPU 208.

[0074] Ideally, the predicted BPU 208 can be identical to the original BPU. However, due to non-ideal prediction and reconstruction operations, the predicted BPU 208 is typically slightly different from the original BPU. To account for these differences, upon generating the predicted BPU 208, the encoder can subtract it from the original BPU to generate a residual BPU 210. For example, the encoder can subtract the values (e.g., grayscale values or RGB values) of the corresponding pixels of the predicted BPU 208 from the values of the pixels of the original BPU. Each pixel of the residual BPU 210 can have a residual value as a result of such subtraction between the corresponding pixels of the original BPU and the predicted BPU 208. The prediction data 206 and the residual BPU 210 can have a smaller number of bits compared to the original BPU, but they can be used to reconstruct the original BPU without a noticeable quality degradation. Thus, the original BPU is compressed.

[0075] To further compress the residual BPU 210, at a transform stage 212, the encoder can reduce its spatial redundancy by decomposing the residual BPU 210 into a set of two-dimensional “base patterns.” Each base pattern is associated with a “transform coefficient.” The base patterns can have the same size (e.g., the size of the residual BPU 210), and each base pattern can represent a component of the residual BPU 210 of a varying frequency (e.g., of luminance). None of the base patterns can be reproduced from any combination (e.g., linear combination) of any of the other base patterns. In other words, the decomposition can decompose the variations of the residual BPU 210 into the frequency domain. This decomposition is similar to a discrete Fourier transform of a function, where the base images are similar to the base functions (e.g., trigonometric functions) of the discrete Fourier transform, and the transform coefficients are similar to the coefficients associated with the base functions.

[0076] Different transform algorithms can use different basis patterns. Various transform algorithms can be used at transform stage 212, such as discrete cosine transform, discrete sine transform, etc. The transform at transform stage 212 is invertible. That is, the encoder can recover the residual BPU 210 through an inverse operation of the transform, called “inverse transform.” For example, to recover a pixel of the residual BPU 210, the inverse transform can be multiplying the values of the corresponding pixels of the basis pattern by the respective correlation coefficients and adding the products to produce a weighted sum. For video coding standards, both the encoder and the decoder can use the same transform algorithm (and thus have the same basis pattern). Thus, the encoder can record only the transform coefficients from which the decoder can reconstruct the residual BPU 210 without receiving the basis pattern from the encoder. The transform coefficients can have fewer bits compared to the residual BPU 210, but they can be used to reconstruct the residual BPU 210 without significant quality degradation. Thus, the residual BPU 210 is further compressed.

[0077] The encoder can further compress the transform coefficients at quantization stage 214. During the transform process, different basis patterns can represent different frequencies of variation (e.g., frequencies of luminance variation). Because the human eye is generally better at recognizing low-frequency variations, the encoder can ignore information of high-frequency variations without causing noticeable quality degradation in decoding. For example, at quantization stage 214, the encoder can generate quantized transform coefficients 216 by dividing each transform coefficient by an integer value, called “quantization parameter,” and rounding the quotient to its nearest integer. After such an operation, some transform coefficients of high-frequency basis patterns can be converted to zero, and transform coefficients of low-frequency basis patterns can be converted to smaller integers. The encoder can ignore the quantized transform coefficients 216 of zero value, whereby the transform coefficients are further compressed. This quantization process is also invertible, where the quantized transform coefficients 216 can be reconstructed to transform coefficients in an inverse operation of quantization, called “inverse quantization.”

[0078] Because the encoder ignores the remainder of the division in the rounding operation, quantization stage 214 can be lossy. Generally, quantization stage 214 can contribute the most information loss in process 200A. The greater the information loss, the fewer the number of bits required for the quantized transform coefficients 216. To obtain different levels of information loss, the encoder can use different quantization parameter values or any other parameters of the quantization process.

[0079] At the binarization stage 226, the encoder can binarize the prediction data 206 and the quantized transform coefficients 216 using a binarization technique, such as entropy coding, variable length coding, arithmetic coding, Huffman coding, context adaptive binary arithmetic coding, or any other lossless or lossy compression algorithm. In some embodiments, in addition to the prediction data 206 and the quantized transform coefficients 216, the encoder can binarize other information at the binarization stage 226, such as the prediction modes used at the prediction stage 204, parameters of the prediction operations, the type of transform at the transform stage 212, parameters of the quantization process (e.g., quantization parameters), encoder control parameters (e.g., bitrate control parameters), and the like. The encoder can use the output data of the binarization stage 226 to generate a video bitstream 228. In some embodiments, the video bitstream 228 can be further packetized for network transmission.

[0080] Referring to the reconstruction path of the process 200A, at the inverse quantization stage 218, the encoder can perform inverse quantization on the quantized transform coefficients 216 to generate reconstructed transform coefficients. At the inverse transform stage 220, the encoder can generate a reconstructed residual BPU 222 based on the reconstructed transform coefficients. The encoder can add the reconstructed residual BPU 222 to the predicted BPU 208 to generate a prediction reference 224 to be used in the next iteration of the process 200A.

[0081] It should be noted that other variants of the process 200A can be used to encode the video sequence 202. In some embodiments, the stages of the process 200A can be performed by the encoder in a different order. In some embodiments, one or more stages of the process 200A can be combined into a single stage. In some embodiments, a single stage of the process 200A can be split into multiple stages. For example, the transform stage 212 and the quantization stage 214 can be combined into a single stage. In some embodiments, the process 200A can include additional stages. In some embodiments, the process 200A can omit one or more stages of the process 200A. Figure 2A

[0082] Figure 2B A schematic diagram illustrating another example encoding process 200B is shown, in accordance with an embodiment of the disclosure. The process 200B can be a modification of the process 200A. For example, the process 200B can be used by an encoder conforming to a hybrid video coding standard (e.g., the H.26x family of standards). In comparison to the process 200A, the forward path of the process 200B further includes a mode decision stage 230 and splits the prediction stage 204 into a spatial prediction stage 2042 and a temporal prediction stage 2044, and the reconstruction path of the process 200B further additionally includes a loop filtering stage 232 and a buffer 234.

[0083] ​In general, prediction techniques can be divided into two types: spatial prediction and temporal prediction. Spatial prediction (e.g., intra-image prediction or “intra-prediction”) can use pixels from one or more already encoded neighboring BPU in the same image to predict a current BPU. That is, the prediction reference 224 in spatial prediction can include neighboring BPU. Spatial prediction can reduce spatial redundancy inherent in images. Temporal prediction (e.g., inter-image prediction or “inter-prediction”) can use regions from one or more already encoded images to predict a current BPU. That is, the prediction reference 224 in temporal prediction can include encoded images. Temporal prediction can reduce temporal redundancy inherent in images.

[0084] Referring to the process 200B, in the forward path, the encoder performs prediction operations at a spatial prediction stage 2042 and a temporal prediction stage 2044. For example, at the spatial prediction stage 2042, the encoder can perform intra-prediction. For an original BPU of an image being encoded, the prediction reference 224 can include one or more neighboring BPU in the same image that have been encoded (in the forward path) and reconstructed (in the reconstruction path). The encoder can generate a predicted BPU 208 by interpolating the neighboring BPU. Interpolation techniques can include, for example, linear interpolation or interpolation, polynomial interpolation or interpolation, etc. In some embodiments, the encoder can perform interpolation at a pixel level, e.g., by interpolating the value of a corresponding pixel for each pixel of the predicted BPU 208. The neighboring BPU used for interpolation can be located in various directions relative to the original BPU, e.g., in a vertical direction (e.g., at the top of the original BPU), a horizontal direction (e.g., at the left of the original BPU), a diagonal direction (e.g., at the lower left, lower right, upper left, or upper right of the original BPU), or any direction defined in the video coding standard used. For intra-prediction, the prediction data 206 can include, for example, the location (e.g., coordinates) of the neighboring BPU used, the size of the neighboring BPU used, parameters for interpolation, the direction of the neighboring BPU used relative to the original BPU, etc.

[0085] For another example, at the temporal prediction stage 2044, the encoder can perform inter prediction. For the original BPU of the current image, the prediction reference 224 can include one or more images (referred to as “reference images”) that have been encoded (in the forward path) and reconstructed (in the reconstruction path). In some embodiments, the reference images can be encoded and reconstructed on a BPU-by-BPU basis. For example, the encoder can add the reconstructed residual BPU 222 to the predicted BPU 208 to generate a reconstructed BPU. When all reconstructed BPUs of the same image are generated, the encoder can generate a reconstructed image as a reference image. The encoder can perform an operation of “motion estimation” to search for a matching region in a range (referred to as a “search window”) of the reference image. The location of the search window in the reference image can be determined based on the location of the original BPU in the current image. For example, the search window can be centered at a location in the reference image that has the same coordinates as the original BPU in the current image, and can extend outward by a predetermined distance. When the encoder identifies (e.g., by using a pel recursive algorithm, a block matching algorithm, etc.) a region in the search window that is similar to the original BPU, the encoder can determine such a region as a matching region. The matching region can have a different size (e.g., smaller, equal, larger, or have a different shape) than the original BPU. Because the reference image and the current image are separated in time on a timeline (e.g., as shown in FIG. 1), the matching region can be considered to “move” to the location of the original BPU over time. The encoder can record the direction and distance of such motion as a “motion vector.” When multiple reference images are used (e.g., as in FIG. 1), the encoder can search for matching regions and determine their associated motion vectors for each reference image. In some embodiments, the encoder can assign weights to the pixel values of the matching regions of the respective reference images. Figure 1 Figure 1

[0086] Motion estimation can be used to identify various types of motion, such as translation, rotation, scaling, etc. For inter prediction, the prediction data 206 can include, for example, the location (e.g., coordinates) of the matching region, the motion vector associated with the matching region, the number of reference images, the weights associated with the reference images, etc.

[0087] To generate the predicted BPU 208, the encoder can perform an operation of “motion compensation.” Motion compensation can be used to reconstruct the predicted BPU 208 based on the prediction data 206 (e.g., the motion vector) and the prediction reference 224. For example, the encoder can move the matching region of the reference image according to the motion vector, where the encoder can predict the original BPU of the current image. When multiple reference images are used (e.g., as in FIG. 1), the encoder can move the matching region of each reference image according to its associated motion vector, where the encoder can predict the original BPU of the current image. Figure 1 ​​In some embodiments, the encoder can move the matching region of the reference image according to the individual motion vectors and average pixel values of the matching region. In some embodiments, if the encoder has assigned weights to the pixel values of the matching region of the individual matching reference images, the encoder can add the weighted sum of the pixel values of the moved matching region.

[0088] In some embodiments, inter prediction can be uni-directional or bi-directional. Uni-directional inter prediction can use one or more reference images in the same temporal direction relative to the current image. For example, Figure 1 Image 104 in FIG. 1 is a uni-directional inter prediction image, where the reference image (i.e., image 102) precedes image 104. Bi-directional inter prediction can use one or more reference images in both temporal directions relative to the current image. For example, Figure 1 Image 106 in FIG. 1 is a bi-directional inter prediction image, where the reference images (i.e., images 104 and 108) precede image 106 in both temporal directions.

[0089] Still referring to the forward path of process 200B, after the spatial prediction 2042 and temporal prediction stage 2044, at a mode decision stage 230, the encoder can select a prediction mode (e.g., one of intra prediction or inter prediction) for the current iteration of process 200B. For example, the encoder can perform a rate-distortion optimization technique, where the encoder can select the prediction mode to minimize the value of a cost function according to the bit rate of the candidate prediction mode and the distortion of the reconstructed reference image under the candidate prediction mode. Depending on the selected prediction mode, the encoder can generate the corresponding predicted BPU 208 and prediction data 206.

[0090] In the reconstruction path of process 200B, if an intra prediction mode has been selected in the forward path, after generating the prediction reference 224 (e.g., the current BPU that has been encoded and reconstructed in the current picture), the encoder can feed the prediction reference 224 directly to the spatial prediction stage 2042 for later use (e.g., for interpolating the next BPU of the current picture). If an inter prediction mode has been selected in the forward path, after generating the prediction reference 224 (e.g., the current picture in which all the BPUs have been encoded and reconstructed), the encoder can feed the prediction reference 224 to the in-loop filter stage 232. At this stage, the encoder can apply in-loop filters to the prediction reference 224 to reduce or eliminate the distortion (e.g., blockiness artifacts) introduced by inter prediction. The encoder can apply various in-loop filter techniques at the in-loop filter stage 232, such as deblocking, sample adaptive offset, adaptive loop filter, etc. The in-loop filtered reference picture can be stored in the buffer 234 (or “decoded picture buffer”) for later use (e.g., as an inter prediction reference picture for future pictures of the video sequence 202). The encoder can store one or more reference pictures in the buffer 234 for use at the temporal prediction stage 2044. In some embodiments, the encoder can encode parameters of the in-loop filters (e.g., in-loop filter strength) at the binary encoding stage 226 along with the quantized transform coefficients 216, prediction data 206, and other information.

[0091] Figure 3A A schematic diagram illustrating an example decoding process 300A in accordance with embodiments of the present disclosure is shown. The process 300A can be a decompression process corresponding to the compression process 200A in Figure 2A some embodiments, the process 300A can be similar to the reconstruction path of the process 200A. A decoder can decode the video bitstream 228 into a video stream 304 according to the process 300A. The video stream 304 can be very similar to the video sequence 202. However, due to information loss in the compression and decompression processes (e.g., the quantization stage 214 in the process 200A in Figures 2A-2B some embodiments, the process 300A can be similar to the reconstruction path of the process 200A. A decoder can decode the video bitstream 228 into a video stream 304 according to the process 300A. The video stream 304 can be very similar to the video sequence 202. However, due to information loss in the compression and decompression processes (e.g., the quantization stage 214 in the process 200A in Figures 2A-2B similar to the processes 200A and 200B in

[0092] as described above with respect to the process 200A in Figure 3AAs shown, the decoder can feed a portion of the video bitstream 228 associated with a basic processing unit of the encoded image (referred to as an "encoded BPU") to a binarization decoding stage 302, where the decoder can decode the portion into prediction data 206 and quantized transform coefficients 216. The decoder can feed the quantized transform coefficients 216 to an inverse quantization stage 218 and an inverse transform stage 220 to generate a reconstructed residual BPU 222. The decoder can feed the prediction data 206 to a prediction stage 204 to generate a predicted BPU 208. The decoder can add the reconstructed residual BPU 222 to the predicted BPU 208 to generate a prediction reference 224. In some embodiments, the prediction reference 224 can be stored in a buffer (e.g., a decoded image buffer in computer memory). The decoder can feed the prediction reference 224 to the prediction stage 204 for performing a prediction operation in the next iteration of the process 300A.

[0093] The decoder can iteratively perform the process 300A to decode each encoded BPU of an encoded image and generate a prediction reference 224 for a next encoded BPU of the encoded image. After decoding all encoded BPUs of an encoded image, the decoder can output the image to a video stream 304 for display and continue decoding a next encoded image in the video bitstream 228.

[0094] At the binarization decoding stage 302, the decoder can perform an inverse operation of the binarization encoding technique used by the encoder (e.g., entropy encoding, variable length encoding, arithmetic encoding, Huffman encoding, context adaptive binary arithmetic encoding, or any other lossless compression algorithm). In some embodiments, in addition to the prediction data 206 and quantized transform coefficients 216, the decoder can decode other information at the binarization decoding stage 302, such as prediction modes, parameters of prediction operations, transform types, parameters of quantization processes (e.g., quantization parameters), encoder control parameters (e.g., bitrate control parameters), and the like. In some embodiments, if the video bitstream 228 is transmitted over a network in packets, the decoder can depacketize the video bitstream 228 before feeding it to the binarization decoding stage 302.

[0095] Figure 3B A schematic diagram illustrating another example decoding process 300B according to embodiments of the disclosure is shown. The process 300B can be a modification of the process 300A. For example, the process 300B can be used by a decoder conforming to a hybrid video coding standard (e.g., the H.26x family of standards). In comparison with the process 300A, the process 300B additionally divides the prediction stage 204 into a spatial prediction stage 2042 and a temporal prediction stage 2044, and additionally includes a loop filtering stage 232 and a buffer 234.

[0096] In process 300B, for a coded base processing unit (referred to as "current BPU") of a decoded coded picture (referred to as "current picture"), the prediction data 206 decoded by the decoder from the binary decoding stage 302 can include various types of data depending on what prediction mode is used by the encoder to code the current BPU. For example, if the current BPU is coded using intra prediction by the encoder, the prediction data 206 can include a prediction mode indicator (e.g., a flag value) indicating the intra prediction, parameters of the intra prediction operation, and the like. The parameters of the intra prediction operation can include, for example, locations (e.g., coordinates) of one or more neighboring BPUs used as references, sizes of the neighboring BPUs, parameters of interpolation, directions of the neighboring BPUs relative to the original BPU, and the like. For another example, if the current BPU is coded using inter prediction by the encoder, the prediction data 206 can include a prediction mode indicator (e.g., a flag value) indicating the inter prediction, parameters of the inter prediction operation, and the like. The parameters of the inter prediction operation can include, for example, a number of reference pictures associated with the current BPU, weights respectively associated with the reference pictures, locations (e.g., coordinates) of one or more matching regions in the respective reference pictures, one or more motion vectors respectively associated with the matching regions, and the like.

[0097] Based on the prediction mode indicator, the decoder can decide whether to perform spatial prediction (e.g., intra prediction) at the spatial prediction stage 2042 or temporal prediction (e.g., inter prediction) at the temporal prediction stage 2044, details of performing such spatial or temporal prediction are described in Figure 2B , which will not be repeated here. After performing such spatial or temporal prediction, the decoder can generate a predicted BPU 208, to which the decoder can add the reconstructed residual BPU 222 to generate a prediction reference 224, as described in Figure 3A .

[0098] In process 300B, the decoder can feed the prediction reference 224 to the spatial prediction stage 2042 or the temporal prediction stage 2044 for performing prediction operations in the next iteration of process 300B. For example, if the current BPU is decoded using intra prediction at the spatial prediction stage 2042, after generating the prediction reference 224 (e.g., the decoded current BPU), the decoder can feed the prediction reference 224 directly to the spatial prediction stage 2042 for later use (e.g., for interpolating the next BPU of the current picture). If the current BPU is decoded using inter prediction at the temporal prediction stage 2044, after generating the prediction reference 224 (e.g., the reference picture in which all BPUs are decoded), the encoder can feed the prediction reference 224 to the loop filter stage 232 to reduce or eliminate distortion (e.g., blocking artifacts). The decoder can feed the prediction reference 224 to the loop filter stage 232 as described in Figure 2BThe illustrated manner applies the in-loop filter to the prediction reference 224. The in-loop filtered reference picture can be stored in a buffer 234 (e.g., a decoded picture buffer in computer memory) for later use (e.g., as an inter-prediction reference picture for future encoded pictures of the video bitstream 228). The decoder can store one or more reference pictures in the buffer 234 for use at the temporal prediction stage 2044. In some embodiments, when the prediction mode indicator of the prediction data 206 indicates that inter-prediction is used to encode the current BPU, the prediction data can further include parameters of the in-loop filter (e.g., in-loop filter strength). The reconstructed picture from the buffer 234 can also be sent to a display, such as a TV, PC, smartphone, or tablet, for end-user viewing.

[0099] There can be four types of in-loop filters. For example, the in-loop filters can include a deblocking filter, a sample adaptive offset (“SAO”) filter, a luma mapping with chroma scaling (“LMCS”) filter, and an adaptive loop filter (“ALF”). The order of applying the four types of in-loop filters can be the LMCS filter, the deblocking filter, the SAO filter, and the ALF. The LMCS filter can include two main components. The first component is an in-loop mapping of the luma component based on an adaptive piecewise linear model. The second component can be used for the chroma component and can apply luma-dependent chroma residual scaling

[0100] Figure 4 is a block diagram of an example apparatus 400 for encoding or decoding a video according to embodiments of the present disclosure. As Figure 4 shown, the apparatus 400 can include a processor 402. When the processor 402 executes instructions described herein, the apparatus 400 can become a special-purpose machine for video encoding or decoding. The processor 402 can be any type of circuitry capable of manipulating or processing information. For example, the processor 402 can include any number of central processing units (or “CPUs”), graphics processing units (or “GPUs”), neural processing units (“NPUs”), microcontroller units (“MCUs”), optical processors, programmable logic controllers, microprocessors, digital signal processors, intellectual property (IP) cores, programmable logic arrays (PLAs), programmable array logic (PALs), generic array logic (GALs), complex programmable logic devices (CPLDs), a field-programmable gate array (FPGA), a system on a chip (SoC), an application-specific integrated circuit (ASIC), etc., in any combination. In some embodiments, the processor 402 can also be a group of processors grouped as a single logical component. For example, as Figure 4 shown, the processor 402 can include multiple processors, including a processor 402a, a processor 402b, and a processor 402n.

[0101] The device 400 may also include a memory 404 configured to store data (e.g., instruction sets, computer code, intermediate data, etc.). For example, such as Figure 4 As shown, the stored data may include program instructions (e.g., for implementing stages in processes 200A, 200B, 300A, or 300B) and data for processing (e.g., video sequence 202, video bitstream 228, or video stream 304). Processor 402 can access the program instructions and data for processing (e.g., via bus 410) and execute the program instructions to perform operations or manipulations on the data for processing. Memory 404 may include a high-speed random access memory device or a non-volatile memory device. In some embodiments, memory 404 may include any combination of any number of random access memories (RAM), read-only memories (ROM), optical discs, magnetic disks, hard disks, solid-state drives, flash drives, secure digital cards (SD cards), memory sticks, compact flash memory (CF cards), etc. Memory 404 may also be a group of memories grouped into single logical components. Figure 4 (Not shown in the image).

[0102] Bus 410 may be a communication device for transmitting data between components within device 400, such as an internal bus (e.g., CPU-memory bus), an external bus (e.g., a Universal Serial Bus port, a Peripheral Component Interconnect Fast Port), or the like.

[0103] For ease of explanation and to avoid ambiguity, the processor 402 and other data processing circuitry are collectively referred to as "data processing circuitry" in this disclosure. The data processing circuitry may be implemented entirely in hardware, or as a combination of software, hardware, or firmware. Furthermore, the data processing circuitry may be a single, separate module, or may be wholly or partially integrated into any other component of the device 400.

[0104] The device 400 may also include a network interface 406 to provide wired or wireless communication with a network (e.g., the Internet, intranet, local area network, mobile communication network, etc.). In some embodiments, the network interface 406 may include any combination of any number of network interface controllers (NICs), radio frequency (RF) modules, transceivers, transceivers, modems, routers, gateways, wired network adapters, wireless network adapters, Bluetooth adapters, infrared adapters, near field communication (“NFC”) adapters, cellular network chips, etc.

[0105] In some embodiments, optionally, the device 400 may further include a peripheral interface 408 to provide connectivity to one or more peripheral devices. Figure 4As shown, the peripherals can include, but are not limited to, a cursor control device (e.g., a mouse, a touchpad, or a touch screen), a keyboard, a display (e.g., a cathode ray tube display, a liquid crystal display, or a light-emitting diode display), a video input device (e.g., a camera or an input interface coupled to a video archive), and the like.

[0106] It should be noted that a video codec (e.g., a codec that performs process 200A, 200B, 300A, or 300B) can be implemented as any combination of any of the software or hardware modules in the apparatus 400. For example, some or all of the stages of process 200A, 200B, 300A, or 300B can be implemented as one or more software modules of the apparatus 400, such as a program instance that can be loaded into memory 404. For another example, some or all of the stages of process 200A, 200B, 300A, or 300B can be implemented as one or more hardware modules of the apparatus 400, such as a special-purpose data processing circuit (e.g., FPGA, ASIC, NPU, etc.).

[0107] In the quantization and inverse quantization function blocks (e.g., Figure 2A or Figure 2B quantization 214 and inverse quantization 218 of the encoder 202, Figure 3A or Figure 3B inverse quantization 218 of the decoder 204), a quantization parameter (QP) is used to determine the amount of quantization (and inverse quantization) applied to the prediction residual. An initial QP value for encoding an image or slice can be signaled at a higher level, e.g., using an init_qp_minus26 syntax element in a picture parameter set (PPS) and using a slice_qp_delta syntax element in a slice header. In addition, a delta QP value sent at the granularity of a quantization group can be used to adapt the QP value at a local level per CU.

[0108] Equirectangular projection (“ERP”) format is a commonly used projection format for representing 360-degree videos and images. The projection maps the meridians to vertically straight lines of constant pitch and the latitudinal circles to horizontally straight lines of constant pitch. ERP is one of the most common projections for 360-degree videos and images due to the particularly simple relationship between the location of an image pixel on the map and its corresponding geographical location on the sphere.

[0109] The algorithm description of projection format conversion and video quality metrics output by JVET gives the coordinate conversion and introduction between ERP and sphere. For 2D to 3D coordinate conversion, given a sample location (m, n), (u, v) can be calculated based on the following equations (1) and (2).

[0110] u = (m + 0.5) / W, 0 < m < W Equation (1)

[0111] v = (n + 0.5) / H, 0 < n < H Equation (2)

[0112] Then, longitude and latitude (φ, θ) in the sphere can be computed from (u, v) based on the following Equations (3) and (4).

[0113] φ = (u - 0.5) x (2 x π) Equation (3)

[0114] θ = (0.5 - v) x π Equation (4)

[0115] 3D coordinates (X, Y, Z) can be computed based on the following Equations (5)-(7).

[0116] X = cos(θ) cos(φ) Equation (5)

[0117] Y = sin(θ) Equation (6)

[0118] Z = -cos(θ) sin(φ) Equation (7)

[0119] For 3D to 2D coordinate conversion starting from (X, Y, Z), (φ, θ) can be computed based on the following Equations (8) and (9). Then, (u, v) is computed based on Equations (3) and (4). Finally, 2D coordinates (m, n) can be computed according to Equations (1) and (2).

[0120] φ = tan-1(-Z / X) Equation (8)

[0121] θ = sin-1(Y / (X2+Y2+Z2)1 / 2) Equation (9)

[0122] To reduce seam artifacts in the reconstructed viewport containing the left and right borders of the ERP image, a new format called padded equirectangular projection (“PERP”) is provided by padding samples on each of the left and right sides of the ERP image.

[0123] When using PERP to represent 360-degree video, the PERP image is encoded. After decoding, the reconstructed PERP is converted back to the reconstructed ERP by blending the copied samples or cropping the padded areas.

[0124] Figure 5A A schematic diagram showing example blending operations for generating a reconstructed equirectangular projection according to some embodiments of the present disclosure is shown. Unless otherwise specified, “recPERP” is used to denote the reconstructed PERP before post-processing, and “recERP” is used to denote the reconstructed ERP after post-processing. In Figure 5AIn some embodiments, A1 and B2 are the boundary regions in the ERP image, and B1 and A2 are the padding regions, where A2 is padded from A1 and B1 is padded from B2. As shown in FIG. 3, the replicated samples of recPERP can be blended by applying a distance-based weighted average operation. For example, region A can be generated by blending A1 with A2, and region B can be generated by blending B1 with B2. Figure 5A

[0125] In the following description, the width and height of the un-padded recERP are denoted as "W" and "H", respectively. The left and right padding widths are denoted as "P L " and "P R ", respectively. The total padding width is denoted as "P w ", which can be the sum of P L and P R . In some embodiments, recPERP can be converted to recERP by a blending operation. For example, for a sample recERP(j, i) in A, where (j, i) is a coordinate in the ERP image, i is in [0, P R - 1], and j is in [0, H - 1], recERP(j, i) can be determined according to the following equation.

[0126] A = w x A1 + (1 - w) x A2, where w is from PL / Pw to 1 Equation (10)

[0127] recERP(j, i) in A = (recPERP(j, i + PL) x (i + PL) + recPERP(j, i + PL + W) x (PR - i) + (PW » 1)) / PW Equation (11) where recPERP(y, x) is a sample on the reconstructed PERP image, where (y, x) is a coordinate of the sample in the PERP image.

[0128] In some embodiments, for a sample recERP(j, i) in B, where (j, i) is a coordinate in the ERP image, and i is in [W - P L , W - 1] and j is in [0, H - 1], recERP(j, i) can be generated according to the following equation.

[0129] B = k x B1 + (1 - k) x B2, where k is from 0 to PL / Pw Equation (12)

[0130] recERP(j, i) in B = (recPERP(j, i + PL) x (PR - i + W)

[0131] ​+ recPERP(j, i + PL - w) x (i - w + PL) + (Pw » 1)) / PW Equation (13) where recPERP(y, x) is a sample on the reconstructed PERP image, where (y, x) is the coordinate of the sample in the PERP image.

[0132] Figure 5B A schematic diagram showing an example cropping operation for generating a reconstructed equirectangular projection is shown in accordance with some embodiments of the present disclosure. In Figure 5B A1 and B2 are the boundary regions within the ERP image, and B1 and A2 are the padded regions, where A2 is padded from A1 and B1 is padded from B2. As Figure 5B shown, in the cropping process, the padded samples in recPERP can be discarded directly to obtain recERP. For example, the padded samples B1 and A2 can be discarded.

[0133] In some embodiments, horizontal wraparound motion compensation can be used to improve the coding performance of ERP. For example, horizontal wraparound motion compensation can be used in the VVC standard as a 360-specific coding tool designed to improve the visual quality of reconstructed 360-degree video in ERP format or PERP format. In conventional motion compensation, when a motion vector reference samples outside the image boundary of a reference image, a repetition padding is applied by copying from those nearest neighboring samples on the corresponding image boundary to derive the value of the out-of-bound sample. For 360-degree video, this repetition padding approach is not suitable and can cause visual artifacts known as “seam artifacts” in the reconstructed viewport video. Because 360-degree video is captured on a sphere and inherently has no “boundaries”, a reference sample outside the boundary of a reference image in the projection domain can be obtained from the neighboring sample in the spherical domain. For general projection formats, it can be difficult to derive the corresponding neighboring sample in the spherical domain because it involves 2D-to-3D and 3D-to-2D coordinate conversion, as well as sample interpolation for fractional sample positions. For the left and right boundaries of ERP or PERP projection formats, this problem can be solved because the spherical neighbor sample outside the left image boundary can be obtained from the sample within the right image boundary, and vice versa. Given the wide use of ERP or PERP projection formats and the relative ease of implementation, VVC adopted horizontal wraparound motion compensation to improve the visual quality of 360-degree video coded in ERP or PERP projection formats.

[0134] Figure 6A A schematic diagram showing an example horizontal wraparound motion compensation process for equirectangular projection is shown in accordance with some embodiments of the present disclosure. As Figure 6AAs shown, when a portion of the reference block lies outside the left (or right) boundary of the reference image in the projection domain, it is not repeated padding; the "outside the boundary" portion can be obtained from the corresponding spherical neighbor portion located within the reference image facing the right (or left) boundary in the projection domain. In some embodiments, repeated padding can be used for the top and bottom image boundaries.

[0135] Figure 6B A schematic diagram of an example horizontal surround motion compensation process for filling a rectangular projection according to some embodiments of the present disclosure is shown. Figure 6B As shown, horizontal wrap motion compensation can be combined with decanted padding methods commonly used in 360-degree video coding. In some embodiments, this is achieved by signaling a high-level syntax element to indicate a wrap motion compensation offset, which can be set to the ERP image width before padding. This syntax can be used to adjust the position of the horizontal wrap accordingly. In some embodiments, this syntax is unaffected by a specific amount of padding on the left or right image boundaries. As a result, this syntax naturally supports asymmetric padding of ERP images. In asymmetric padding of ERP images, the left and right padding can be different. In some embodiments, the wrap motion compensation can be determined according to the following formula:

[0136]

[0137] The offset can be a motion compensation offset signaled in the bitstream, picW can be the image width including the padding region before encoding, and pos... x It can be a reference position determined by the current block position and the motion vector, and the formula pos x The output of `_wrap` can be the actual reference position from the reference block in the surround motion compensation. To save on the signal notification overhead of the surround motion compensation offset, it can be in units of the smallest luminance coded block; therefore, the offset can be replaced with `offset`. w ×MinCbSizeY, where offset w This is the surround motion compensation offset in units of the minimum luminance coded block, which is signaled in the bitstream, and MinCbSizeY is the size of the minimum luminance coded block. Conversely, in conventional motion compensation, the actual reference position from which the reference block originates can be determined by shearing pos within the range of 0 to picW-1. x Export directly.

[0138] Horizontal wraparound motion compensation can provide more meaningful information for motion compensation when the reference sample is outside the left or right boundary of the reference picture. In 360 video common test conditions, this tool can improve compression performance not only in rate-distortion, but also in reducing seam artifacts and subjective quality of reconstructed 360 video. Horizontal wraparound motion compensation can also be used for other single-sided projection formats with constant sampling density in the horizontal direction, such as adjusted equirectangular projection.

[0139] In VVC (e.g., VVC Draft 8), wraparound motion compensation can be implemented by signaling a high-level syntax variable pps_ref_wraparound_offset to indicate the wraparound offset. Sometimes, the wraparound offset should be set to the ERP picture width before padding. This syntax can be used to adjust the position of horizontal wraparound motion compensation accordingly. This syntax can not be affected by the specific amount of padding on the left or right picture boundary. As a result, this syntax can naturally support asymmetric padding of ERP pictures (e.g., when the left padding and right padding are different). Horizontal wraparound motion compensation can provide more meaningful information for motion compensation when the reference sample is outside the left or right boundary of the reference picture.

[0140] Figure 7 An example high-level wraparound offset syntax according to some embodiments of the disclosure is shown. It should be understood that, Figure 7 The shown syntax can be used in VVC (e.g., VVC Draft 8). As Figure 7 shown, the wraparound offset pps_ref_wraparound_offset can be directly signaled when wraparound motion compensation is enabled (e.g., pps_ref_wraparound_enabled_flag == 1). As Figure 7 shown, the syntax elements or variables in the bitstream are shown in bold.

[0141] Figure 8 An example semantics of high-level wraparound offset according to some embodiments of the disclosure is shown. It should be understood that, Figure 8 The semantics shown in Figure 7 can correspond to Figure 8As shown, a value of 1 for pps_ref_wraparound_enabled_flag implies that horizontal wraparound motion compensation is applied in inter prediction. If the value of pps_ref_wraparound_enabled_flag is 0, then horizontal wraparound motion compensation is not applied. When the value of CtbSizeY / MinCbSizeY + 1 is greater than pic_width_in_luma_samples / MinCbSizeY - 1, the value of pps_ref_wraparound_enabled_flag shall be equal to 0. When sps_ref_wraparound_enabled_flag is equal to 0, the value of pps_ref_wraparound_enabled_flag shall be equal to 0. CtbSizeY is the size of a luma coding tree block.

[0142] In some embodiments, as Figure 8 As shown, a value of pps_ref_wraparound_offset + (CtbSizeY / MinCbSizeY) + 2 can specify an offset used to calculate a horizontal wraparound position in units of MinCbSizeY luma samples. The value of pps_ref_wraparound_offset can be in the range of 0 to (pic_width_in_luma_samples / MinCbSizeY) - (CtbSizeY / MinCbSizeY) - 2, inclusive. A variable PpsRefWraparoundOffset can be set equal to pps_ref_wraparound_offset + (CtbSizeY / MinCbSizeY) + 2. In some embodiments, the variable PpsRefWrapararoundOffset can be used to determine luma positions of sub-blocks (e.g., in VVC Draft 8).

[0143] In VVC (e.g., VVC Draft 8), a picture can be divided into one or more tile rows or one or more tile columns. A tile can be a sequence of coding tree units (“CTUs”) that cover a rectangular region of the picture. A slice can include an integer number of complete tiles or an integer number of consecutive complete CTU rows within a tile of a picture.

[0144] Two modes of slices can be supported, namely raster-scan slice mode and rectangular slice mode. In the raster-scan slice mode, a slice can include a complete tile sequence in a tile raster scan of the picture. In the rectangular slice mode, a slice can include multiple complete tiles that collectively form a rectangular region of the picture, or multiple consecutive complete CTU rows of one tile that collectively form a rectangular region of the picture. Tiles within a rectangular slice are scanned in tile raster scan order within the rectangular region corresponding to the slice.

[0145] A sub-picture can include one or more slices that collectively cover a rectangular region of the picture. Figure 9 A diagram showing example slice and sub-picture partitioning of a picture is shown in accordance with some embodiments of the disclosure. As shown, Figure 9 The picture is partitioned into 20 tiles with 5 tile columns and 4 tile rows. There are 12 tiles on the left, each covering a slice of 4x4 CTUs. There are 8 tiles on the right, each covering a slice of 2 vertically stacked 2x2 CTUs. There are 28 slices and 28 sub-pictures of different sizes in total (e.g., each slice can be one sub-picture).

[0146] Figure 10 A diagram showing example slice and sub-picture partitioning of a picture with different slices and sub-pictures is shown in accordance with some embodiments of the disclosure. As shown, Figure 10 The picture is partitioned into 20 tiles with 5 tile columns and 4 tile rows. There are 12 tiles on the left, each covering a slice of 4x4 CTUs. There are 8 tiles on the right, each covering a slice of 2 vertically stacked 2x2 CTUs. There are 28 slices in total. For the 12 slices on the left, each slice is a sub-picture. For the 16 slices on the right, each group of 4 slices forms a sub-picture. As a result, there are 16 sub-pictures of the same size in total.

[0147] In VVC (e.g., VVC Draft 8), information about slice layout can be signaled in a picture parameter set (“PPS”). In some embodiments, a picture parameter set is a syntax structure that includes syntax elements or variables that apply to zero or more entire coded pictures determined by syntax elements found in each picture header. Figure 11 Syntax of an example picture parameter set for tile mapping and slice layout is shown in accordance with some embodiments of the disclosure. It should be understood that Figure 11 The syntax shown can be used in VVC (e.g., VVC Draft 8). As shown, Figure 11 Syntax elements or variables in the bitstream are shown in bold. As shown, Figure 11As shown, if the number of tile in the current picture is greater than 1 and rectangular slice mode is used (e.g., rect_slice_flag == 1), a flag called single_slice_per_subpic_flag can be first signaled to indicate that each subpicture includes only one slice. In this case (e.g., single_slice_per_subpic_flag == 1), the slice layout information does not need to be further signaled because it can be the same as the subpicture layout that has already been signaled in the sequence parameter set (“SPS”). In some embodiments, the SPS is a syntax structure that includes syntax elements that apply to zero or more entire coded layer video sequences (“CLVS”), determined by the content of syntax elements found in the picture parameter set referenced by syntax elements found in each picture header. In some embodiments, the picture header is a syntax structure that includes syntax elements that apply to all slices of a coded picture. In some embodiments, as Figure 11 As shown, if the value of single_slice_per_subpic_flag is 0, the number of slices in the picture (e.g., num_slices_in_pic_minus1) can be first signaled, followed by the slice position and size information for each slice.

[0148] In some embodiments, to signal the number of slices, the number of slices minus 1 (e.g., num_slices_in_pic_minus1) can be signaled instead of directly signaling the number of slices because there is at least 1 slice in the picture. Generally, signaling a smaller positive value can cost fewer bits and improve the overall efficiency of performing video processing.

[0149] FIG. 12 illustrates semantics of an example picture parameter set for tile mapping and slice layout according to some embodiments of the present disclosure. It should be appreciated that the semantics 2 shown in FIG. 12 can correspond to the syntax shown in Figure 11 In some embodiments, the semantics shown in FIG. 12 correspond to VVC (e.g., VVC Draft 8), as shown.

[0150] In some embodiments, as shown in FIG. 12, the variable rect_slice_flag equal to 0 indicates that the tile within each slice is in a raster scan order and slice information is not signaled in the PPS. When the variable rect_slice_flag is equal to 1, the tile within each slice can cover a rectangular region of the picture and slice information can be signaled in the PPS. In some embodiments, the variable rect_slice_flag can be inferred to be equal to 1 when not present. In some embodiments, the value of rect_slice_flag shall be equal to 1 when the variable subpic_info_present_flag is equal to 1.

[0151] In some embodiments, as shown in FIG. 12, the variable single_slice_per_subpic_flag equal to 1 means that each subpicture can include one and only one rectangular slice. When the variable single_slice_per_subpic_flag is equal to 0, each subpicture can include one or more rectangular slices. In some embodiments, when the variable single_slice_per_subpic_flag is equal to 1, the variable num_slices_in_pic_minus1 can be inferred to be equal to the variable sps_num_subpics_minus1. In some embodiments, the value of single_slice_per_subpic_flag can be inferred to be equal to 0 when not present.

[0152] In some embodiments, as shown in FIG. 12, the variable num_slices_in_pic_minus1 plus 1 is the number of rectangular slices in each picture that refer to the PPS. In some embodiments, the value of num_slices_in_pic_minus1 can be in the range of 0 to MaxSlicesPerPicture - 1, inclusive. In some embodiments, when the variable no_pic_partition_flag is equal to 1, the value of num_slices_in_pic_minus1 can be inferred to be equal to 0.

[0153] In some embodiments, as shown in FIG. 12, if the variable tile_idx_delta_present_flag is equal to 0, the value of tile_idx_delta is not present in the PPS and all rectangular slices in the picture referring to the PPS are specified in raster order. In some embodiments, when the variable tile_idx_delta_present_flag is equal to 1, the value of tile_idx_delta can be present in the PPS and all rectangular slices in the picture referring to the PPS are specified in the order indicated by the value of tile_idx_delta. In some embodiments, when not present, the value of tile_idx_delta_present_flag can be inferred to be equal to 0.

[0154] In some embodiments, as shown in FIG. 12, the variable slice_width_in_tiles_minusl[i] plus 1 specifies the width of the i-th rectangular slice in tile units. In some embodiments, the value of slice_width_in_tiles_minusl[i] shall be in the range of 0 to NumTileColumns - 1, inclusive. In some embodiments, if slice_width_in_tiles_minusl[i] is not present, the following can apply: if NumTileColumns is equal to 1, the value of slice_width_in_tiles_minusl[i] can be inferred to be equal to 0; otherwise, the value of slice_width_in_tiles_minusl[i] can be inferred to be the value specified in the subclause in VVC Draft (e.g., VVC Draft 8).

[0155] In some embodiments, as shown in FIG. 12, the variable slice height in tiles minusl plus 1 specifies the height of the i-th rectangular slice in units of tile rows. In some embodiments, the value of slice height in tiles minusl [i] shall be in the range of 0 to NumTileRows - 1, inclusive. In some embodiments, when the variable slice height in tiles minusl [i] is not present, the following applies: if NumTileRows is equal to 1, or the variable tile idx delta present flag is equal to 0 and tileldx % NumTileColumns is greater than 0, the value of slice height in tiles minusl [i] is inferred to be equal to 0; otherwise (e.g., NumTileRows is not equal to 1, and tile idx delta present flag is equal to 1 or tileldx % NumTileColumns is equal to 0), the value of slice height in tiles minusl [i] can be inferred to be equal to slice height in tiles minusl [i - 1] when tile idx delta present flag is equal to 1 or tileldx % NumTileColumns is equal to 0.

[0156] In some embodiments, as shown in FIG. 12, the value of num exp slices in tile [i] specifies the number of explicitly provided slice heights in the current tile that includes multiple rectangular slices. In some embodiments, the value of num exp slices in tile [i] shall be in the range of 0 to RowHeight[tileY] - 1, inclusive, where tileY is the tile row index that includes the i-th slice. In some embodiments, when not present, the value of num exp slices in tile [i] can be inferred to be equal to 0. In some embodiments, when num exp slices in tile [i] is equal to 0, the value of the variable NumSliceInTile [i] is derived to be equal to 1.

[0157] In some embodiments, as shown in FIG. 12, the value of exp_slice_height_in_ctus_minus1[ j ] plus 1 specifies the height of the j-th rectangular slice in the current tile in CTU units. In some embodiments, the value of exp_slice_height_in_ctus_minus1[ j ] shall be in the range of 0 to RowHeight[ tileY ] - 1, inclusive, where tileY is the tile row index of the current tile.

[0158] In some embodiments, as shown in FIG. 12, when the variable num_exp_slices_in_tile[ i ] is greater than 0, the variable NumSlicesInTile[ i ] and SliceHeightInCtusMinus1[ i + k ] can be derived with k in the range of 0 to NumSlicesInTile[ i ] - 1. In some embodiments, as shown in FIG. 12, when the value of num_exp_slices_in_tile[ i ] is greater than 0, the values of NumSlicesInTile[ i ] and SliceHeightInCtusMinus1[ i + k ] can be derived. Figure 8

[0159] In some embodiments, as shown in FIG. 12, the value of tile_idx_delta[ i ] specifies the difference between the tile index of the first tile in the i-th rectangular slice and the tile index of the first tile in the (i+1)-th rectangular slice. The value of tile_idx_delta[ i ] shall be in the range of -NumTilesInPic+1 to NumTilesInPic-1, inclusive. In some embodiments, when not present, the value of tile_idx_delta[ i ] can be inferred to be equal to 0. In some embodiments, when present, the value of tile_idx_delta[ i ] can be inferred to be not equal to 0.

[0160] In VVC (e.g., VVC Draft 8), in order to locate each slice in a picture, one or more slice addresses can be signaled in the slice header. Figure 13 Syntax of an example slice header according to some embodiments of the present disclosure is shown. It should be understood that, Figure 13 Syntax 3 shown can be used in VVC (e.g., VVC Draft 8). As shown, Figure 11 variables or variables in the bitstream are shown in bold. As shown, Figure 13 If the variable picture_header_in_slice_header_flag is equal to 1, the picture header syntax structure is present in the slice header, as shown.

[0161] ​Figure 14 The semantics of an example slice header according to some embodiments of the disclosure is shown. It should be understood that Figure 14 The semantics shown in Figure 13 The syntax shown. In some embodiments, Figure 14 The semantics shown in correspond to VVC (e.g., VVC Draft 8).

[0162] In some embodiments, as Figure 14 shown, the requirement for bitstream conformance is that the value of picture_header_in_slice_header_flag shall be the same in all coded slices in a CLVS.

[0163] In some embodiments, as Figure 14 shown, when the variable picture_header_in_slice_header_flag is equal to 1 for a coded slice, the requirement for bitstream conformance is that there shall be no video coding layer (“VCL”) network abstraction layer (“NAL”) units with nal_unit_type equal to PH NUT in the CLVS.

[0164] In some embodiments, as Figure 14 shown, when picture_header_in_slice_header_flag is equal to 0, all coded slices in the current picture shall have picture_header_in_slice_header_flag equal to 0, and the current PU shall have a PH NAL unit.

[0165] In some embodiments, as Figure 14 shown, the variable slice_address can specify the slice address of the slice. In some embodiments, when not present, the value of slice_address can be inferred to be equal to 0. When the variable rect_slice_flag is equal to 1 and NumSliceInSubpic[CurrSubpicIdx] is equal to 1, the value of slice_address is inferred to be equal to 0.

[0166] In some embodiments, as Figure 14 shown, if the variable rect_slice_flag is equal to 0, the following can apply: the slice address can be the raster-scan tile index; the length of slice_address can be ceil(Log2(NumTilesInPic)) bits; and the value of slice_address shall be in the range of 0 to NumTilesInPic - 1, inclusive.

[0167] In some embodiments, as shown in Figure 14 The bitstream conformance requirement applies the following constraint: If the variable rect_slice_flag is equal to 0 or the variable subpic_info_present_flag is equal to 0, the value of slice_address shall not be equal to the value of slice_address of any other coded slice NAL unit of the same coded picture; otherwise, the pair of slice_subpic_id and slice_address values shall not be equal to the pair of slice_subpic_id and slice_address values of any other coded slice NAL unit of the same coded picture.

[0168] In some embodiments, as shown in Figure 15 The shape of a slice of a picture shall be such that each CTU shall have its entire left and top boundaries, when decoded, include picture boundaries or include boundaries of previously decoded CTUs.

[0169] In some embodiments, as shown in Figure 15 The variable slice_address can only be signaled if one of the following two conditions is met: rectangular slice mode is used and the number of slices in the current subpicture is greater than 1; or rectangular slice mode is not used and the number of tiles in the current picture is greater than 1. In some embodiments, if neither of the above two conditions is met, there is only one slice in the current subpicture or the current picture. In this case, since the entire subpicture or the entire picture is a single slice, there is no need to signal the slice address.

[0170] There are many problems in the current design of VVC. First, the wraparound offset pps_ref_wraparound_offset is signaled in the bitstream, and it should be set to the ERP picture width before padding. To save bits, the minimum value of the wraparound offset (e.g., CtbSizeY / MinCbSizeY)+2) is subtracted from the wraparound offset before signaling. However, the width of the padded region is much smaller than the width of the original ERP picture. This is especially true for coded ERP pictures where the padded region width can be 0. Given that the total width of the picture to be coded or decoded is known, it can cost more bits to signal the width of the original ERP portion than to signal the width of the padded region. As a result, the current signaling of the wraparound offset in MinCbSizeY units of the original ERP width is not very efficient.

[0171] Furthermore, the signaling of slice layout can also be improved. For example, subtract 1 from the number of slices before signaling the number of slices, because the number of slices in a picture is always greater than or equal to 1, and signaling a smaller positive value requires fewer bits. However, in the current VVC (e.g., VVC Draft 8), a subpicture contains an integer number of complete slices. As a result, the number of slices in a picture is greater than or equal to the number of subpictures in the picture. In the current VVC (e.g., VVC Draft 8), num_slices_in_pic_minus1 is signaled only when the variable single_slice_per_subpic_flag is 0, and the variable single_slice_per_subpic_flag equal to 0 indicates that at least one subpicture contains multiple slices. Therefore, in this case, the number of slices must be greater than the number of subpictures with a minimum value of 1. As a result, the minimum value of num_slices_in_pic_minus1 is greater than zero. It is inefficient to signal a non-negative value that does not start from zero.

[0172] Furthermore, there are other issues with slice address signaling. When there is only one slice in the current picture, it can not be necessary to signal the slice address because the entire subpicture or the entire picture is a single slice. The slice address can be inferred to be 0. However, the two conditions to skip slice address signaling are not complete. For example, in the raster scan slice mode, even if the number of tiles is greater than 1, there can be only one slice that includes all the tiles in the picture, in which case slice address signaling can also be avoided.

[0173] Embodiments of the present disclosure provide methods to solve the above problems. In some embodiments, since the width of the padding region is usually smaller than the width of the original ERP picture, which can be the same as the wraparound offset in wraparound motion compensation, it can be suggested to signal the difference between the coded picture width in the bitstream and the original ERP picture width. For example, it can be proposed to signal the difference between the coded picture width and the wraparound motion compensation offset in the bitstream, and perform derivation after parsing the signaled difference to obtain the wraparound offset at the decoder side. Since the difference between the coded picture width and the wraparound offset is usually smaller than the wraparound offset itself, this method can save bits for signaling.

[0174] Figure 7 An example improved picture parameter set syntax according to some embodiments of the present disclosure is shown. As Figure 15 shown, syntax elements or variables in the bitstream are shown in bold, and changes from the syntax shown in Figure 16 VVC Draft 8 are shown in italic type, where the proposed deletion syntax is further shown with deletion lines. As Figure 16As shown, a new variable pps_pic_width_minus_wraparound_offset can be created. The variable

[0175] The value of pps_pic_width_minus_wraparound_offset can be signaled according to

[0176] The value of pps_ref_wraparound_enabled_flag is signaled. In some embodiments, the new variable pps_pic_width_minus_wraparound_offset can replace the variable pps_ref_wraparound_offset in the original VVC (e.g., VVC Draft 8).

[0177] Figure 8 An example improved picture parameter set semantics according to some embodiments of the disclosure is shown. As shown, the changes to the previous VVC (e.g., the semantics shown in Figure 16 Figure 15 The semantics shown in Figure 16 Figure 16 The semantics shown in Figure 8

[0178] As shown, the semantics of the new variable pps_pic_width_minus_wraparound_offset is different from the variable pps_ref_wraparound_offset (e.g., as shown in Figure 16 Figure 17 The variable pps_pic_width_minus_wraparound_offset can specify the difference between the picture width and the offset used to calculate the horizontal wrap-around position in units of MinCbSizeY luma samples. In some embodiments, as shown in Figure 17 Figure 17 The value of pps_pic_width_minus_wraparound_offset shall be less than or equal to (pic_width_in_luma_samples / MinCbSizeY) - (CtbSizeY / MinCbSizeY) - 2. In some embodiments, the variable PpsRefWraparoundOffset can be set equal to pic_width_in_luma_samples / MinCbSizeY - pps_pic_width_minus_wraparound_offset.​​​

[0179] In some embodiments, a flag wraparound_offset_type can be signaled to indicate whether the signaled wraparound offset is the original ERP picture width or the difference between the coded picture width and the original ERP picture width. The encoder can choose the smaller one from the two values and signal it in the bitstream, which can further reduce the signaling overhead. Figure 7 A syntax of an example picture parameter set with variable wraparound_offset_type according to some embodiments of the present disclosure is shown. As shown, Figure 17 The syntax elements or variables in the bitstream are shown in boldface, and the changes from the previous VVC (e.g., the syntax shown in Figure 18 ) are shown in italic type. As shown, Figure 18 A new variable pps_ref_wraparound_offset can be added to specify the value used to determine PpsRefWraparoundOffset.

[0180] Figure 8 A semantics of an example improved picture parameter set with variable wraparound_offset_type according to some embodiments of the present disclosure is shown. As shown, Figure 18 The changes from the previous VVC (e.g., the semantics shown in Figure 17 ) are shown in italic type, and the proposed deletion semantics are further shown in strikeout. It should be understood that, Figure 18 The semantics shown in Figure 18 may correspond to the syntax shown in Figure 18 In some embodiments, the semantics shown in correspond to VVC (e.g., VVC Draft 8).

[0181] In some embodiments, as shown in Figure 19 A new variable wraparound_offset_type can be added to specify the type of variable pps_ref_wraparound_offset. The value of pps_ref_wraparound_offset shall be in the range of 0 to ((pps_pic_width_in_luma_samples / MinCbSizeY) - (CtbSizeY / MinCbSizeY) - 2) / 2.

[0182] In some embodiments, as shown in Figure 19As shown, the value of the variable PpsRefWraparoundOffset can be derived. For example, when the variable wraparound_offset_type is equal to 0, the signaled wrapound offset is the original ERP picture width. As a result, the variable PpsRefWraparoundOffset is equal to pps_ref_wraparound_offset + (CrbSizeY / MinCbSizeY) + 2. When the variable wraparound_offset_type is not equal to 0, the signaled wrapound offset is the difference between the coded picture width and the original ERP picture width. As a result, the variable PpsRefWraparoundOffset is equal to

[0183] pps_pic_width_in_luma_samples / MinCbSizeY - pps_ref_wraparound_offset.

[0184] In the previous VVC (e.g., VVC Draft 8), when the variable single_slice_per_subpic_flag is 0, the number of slices is signaled in PPS, which means there is at least one subpicture containing more than one slice. In this case, considering that each subpicture should contain one or more complete slices, the number of slices needs to be greater than the number of subpictures, which thus results in the minimum number of slices being 2. This is because there is at least one subpicture in an image.

[0185] Embodiments of the present disclosure provide a method for improving the signaling of the number of slices. Figure 11 An example picture parameter set with the variable num_slices_in_pic_minus2 according to some embodiments of the present disclosure is shown. As shown, the syntax elements or variables in the bitstream are shown in bold, and the changes from the previous VVC (e.g., the syntax shown in Figure 19 Figure 19 As shown, one new variable num_slices_in_pic_minus2 can be added to specify the value for determining PpsRefWraparoundOffset. Figure 21

[0186] ​​FIG. 20 illustrates an example modified semantics of a picture parameter set with variable num_slices_in_pic_minus2, according to some embodiments of the disclosure. As shown in FIG. 20, changes from previous VVC (e.g., the semantics shown in FIG. 12) are shown in italic type, and the proposed deleted syntax is further shown in strikeout. It will be appreciated that the semantics shown in FIG. 20 can correspond to Figure 21 the syntax shown. In some embodiments, the semantics shown in FIG. 20 correspond to VVC (e.g., VVC Draft 8).

[0187] In some embodiments, as shown in FIG. 20, the value of variable num_slices_in_pic_minus2 plus 2 can specify the number of rectangular slices in each picture referring to the PPS. In some embodiments, as shown in FIG. 20, variable num_slices_in_pic_minus2 can replace variable num_slices_in_pic_minus_1. In some embodiments, the value of num_slices_in_pic_minus2 plus 2 shall be in the range of 0 to MaxSlicesPerPicture - 2, inclusive, where MaxSlicesPerPicture is specified in VVC (e.g., VVC Draft 8). When variable no_pic_partition_flag is equal to 1, the value of num_slices_in_pic_minus2 is inferred to be equal to -1.

[0188] In some embodiments, the number of slices can be signaled using a variable that is the number of slices minus the number of subpictures and then minus 1 (e.g., num_slices_in_pic_minus_subpic_num_minus1). Figure 11 FIG. 21 illustrates an example syntax of a picture parameter set with variable num_slices_in_pic_minus_subpic_num_minus1, according to some embodiments of the disclosure. As shown in FIG. 21, the syntax elements or variables in the bitstream are shown in bold, and changes from previous VVC (e.g., the syntax shown in FIG. 18) are shown in italic type, where the proposed deleted syntax is further shown in strikeout. Figure 21 Figure 21

[0189] ​​FIG. 22 illustrates semantics of an example modified picture parameter set with variable num_slices_in_pic_minus_subpic_num_minus1 according to some embodiments of the present disclosure. As shown in FIG. 22, changes from previous VVC (e.g., semantics shown in FIG. 12) are shown in italic type, and the proposed deletion of syntax is further shown with a strike-through line. It should be understood that the semantics shown in FIG. 22 can correspond to Figure 23 the syntax shown. In some embodiments, the semantics shown in FIG. 22 correspond to VVC (e.g., VVC Draft 8).

[0190] In VVC (e.g., VVC Draft 8), the number of slices is signaled in PPS only when variable single_slice_per_subpic_flag is equal to 0, which can be equivalent to stating that there is at least one subpicture containing more than one slice. Therefore, the number of slices should be greater than the number of subpictures, since a subpicture should contain one or more complete slices. As a result, the minimum number of slices is equal to the number of subpictures plus 1. As Figure 23 and shown in FIG. 22, the number of slices minus the number of subpictures minus 1 (e.g., num_slices_in_pic_minus_subpic_num_minus1) is signaled instead of the number of slices minus 1 (e.g., num_slices_in_pic_minus1) to reduce the number of bits signaled.

[0191] num_slices_in_pic_minus_subpic_num_minus1) instead of the number of slices minus 1 (e.g., num_slices_in_pic_minus1) to reduce the number of bits signaled.

[0192] In some embodiments, as shown in FIG. 22, the value of num_slices_in_pic_minus_subpic_num_minus1 plus the number of sub-pictures plus 1 can specify the number of rectangular slices in each picture referring to the PPS. In some embodiments, the value of num_slices_in_pic_minus_subpic_num_minus1 shall be in the range of 0 to MaxSlicesPerPicture minus sps_num_subpics_minus1 minus 2, inclusive, where MaxSlicesPerPicture can be specified in VVC (e.g., VVC Draft 8). In some embodiments, when the variable no_pic_partition_flag is equal to 1, the value of num_slices_in_pic_minus_subpic_num_minus1 can be inferred to be equal to (sps_num_subpics_minus1 + 1). In some embodiments, the variable SliceNumInPic can be derived as SliceNumInPic = num_slices_in_pic_minus_subpic_num_minus1 + sps_num_subpics_minus1 + 2.

[0193] Embodiments of the present disclosure also provide a new way of signaling slice addresses. Figure 13 An example updated slice header syntax according to some embodiments of the present disclosure is shown. As shown in Figure 23 , syntax elements or variables in the bitstream are shown in bold, and changes from previous VVC (e.g., syntax shown in Figure 23 ) are shown in italic type, and the proposed deleted syntax is further shown with strikeout.

[0194] In some embodiments, as shown in Figure 24 , the variable picture_header_in_slice_header_flag can be signaled in the slice header to indicate whether the PH syntax structure is present within the slice header. In VVC (e.g., VVC Draft 8), there are constraints on the presence or absence of the PH syntax structure in the slice header and the number of slices in a picture. When the PH syntax structure is present in the slice header, the picture should have only one slice. Therefore, there is also no need to signal the slice address. Thus, the signaling of the slice address can be conditioned on the picture_header_in_slice_header_flag. As shown in Figure 24As shown, the value of picture_header_in_slice_header_flag can be used as another condition to decide whether to signal the variable slice_address. When picture_header_in_slice_header_flag is equal to 1, the signaling of the variable slice_address is skipped. In some embodiments, when the variable slice_address is not signaled, it can be inferred to be 0.

[0195] Embodiments of the present disclosure also provide methods for performing video encoding. Figure 4 A flowchart of an example video encoding method according to some embodiments of the present disclosure is shown, which has a variable that signals a difference between a width of a video frame and an offset used to calculate a horizontal wrap-around position. In some embodiments, Figure 24 The method 24000 shown in FIG. 24 can be performed by Figure 15 The apparatus 400 shown. In some embodiments, Figure 16 The method 24000 shown can be performed according to Figure 24 The syntax shown or Figure 24 The semantics shown. In some embodiments, Figure 15 The method 24000 shown in FIG. 24 includes a wrap-around motion compensation process performed according to the VVC standard. In some embodiments, Figure 16 The method 24000 shown in FIG. 24 can be performed with a 360-degree video sequence as input.

[0196] In step S24010, a wrap-around motion compensation flag is received, where the wrap-around motion compensation flag is associated with a picture. For example, as shown in Figure 15 or Figure 15 The wrap-around motion compensation flag can be the variable pps_ref_wraparound_enabled_flag. In some embodiments, the picture is in a bitstream. In some embodiments, the picture is part of a 360-degree video.

[0197] In step S24020, it is determined whether the wrap-around motion compensation flag is enabled. For example, as shown in Figure 17 It is determined whether the variable pps_ref_wraparound_enabled_flag is equal to 1.

[0198] In step S24030, in response to determining that the wrap-around motion compensation flag is enabled, a difference between a width of the picture and an offset used to determine a horizontal wrap-around position is received. For example, as shown in Figure 18As shown, when it is determined that the variable pps_ref_wraparound_enabled_flag is equal to 1, the variable pps_pic_width_minus_wraparound_offset can be received or signaled. In some embodiments, the difference is less than or equal to the width of the picture divided by the size of the minimum luma coding block minus the size of the luma coding tree block divided by the size of the minimum luma coding block minus 2.

[0199] In some embodiments, in S24030, receiving the difference includes receiving a wraparound offset type flag. For example, as shown in Figure 17 and Figure 18 The flag wraparound_offset_type can be signaled to indicate whether the signaled value of the wraparound offset is the original ERP picture width or the difference between the coded picture width and the original ERP picture width. In some embodiments, as shown in Figure 15 and Figure 16 The value of the flag wraparound_offset_type can be 0 or 1.

[0200] In S24040, the picture is motion compensated according to the wraparound motion compensation flag and the difference. For example, the picture can be motion compensated according to the variables pps_ref_wraparound_enabled_flag and pps_pic_width_minus_wraparound_offset shown in Figure 25 and Figure 25 In some embodiments, performing the motion compensation further includes determining a wraparound motion compensation offset according to the width of the picture and the difference, and performing the motion compensation on the picture according to the wraparound motion compensation offset. In some embodiments, the wraparound motion compensation offset can be determined as the width of the picture divided by the size of the minimum luma coding block minus the difference.

[0201] Figure 4 A flowchart of an example video encoding method according to some embodiments of the present disclosure is shown, which has a variable signaling the number of slices in a video frame minus 2. In some embodiments, the method 25000 shown in Figure 25 may be performed by the apparatus 400 shown in Figure 19 In some embodiments, the method 25000 shown in Figure 25 may be performed according to the syntax shown in Figure 19 or the semantics shown in FIG. 20. In some embodiments, the method 25000 shown in Figure 26 may be performed according to the VVC standard.

[0202] In step S25010, an image for encoding is received. In some embodiments, the image can include one or more slices. In some embodiments, the image is in a bitstream. In some embodiments, the one or more slices are rectangular slices.

[0203] In step S25020, a variable is signaled in a picture parameter set of the image, the variable indicating a number of slices in the picture minus 2. For example, the variable can be Figure 26 num_slices_in_pic_minus2 shown in FIG. 20. In some embodiments, a value of the variable plus 2 can specify a number of rectangular slices in each picture. In some embodiments, the variable is part of a PPS. In some embodiments, the variable num_slices_in_pic_minus2 can replace the variable num_slices_in_pic_minus_1 similar to the semantics shown in FIG. 20.

[0204] Figure 4 A flowchart of an example video encoding method having a variable signaling a number of slices in a video frame minus a number of sub-pictures in the video frame minus 1 is shown, according to some embodiments of the disclosure. In some embodiments, Figure 26 The method 26000 shown in FIG. 26 can be performed by Figure 21 the apparatus 400 shown in FIG. 4. In some embodiments, Figure 26 The method 26000 shown in FIG. 26 can be performed according to Figure 21 the syntax shown in FIG. 22 or the semantics shown in FIG. 22. In some embodiments, Figure 21 The method 26000 shown in FIG. 26 can be performed according to the VVC standard.

[0205] In step S26010, an image for encoding is received. In some embodiments, the image can include one or more slices and one or more sub-pictures. In some embodiments, the video frame is in a bitstream. In some embodiments, the one or more slices are rectangular slices.

[0206] In step S26020, a variable is signaled in a picture parameter set of the image, the variable indicating a number of slices in the picture minus a number of sub-pictures in the picture minus 1. For example, the variable can be as Figure 21 num_slices_in_pic_minus_subpic_num_minus1 shown in FIG. 22. In some embodiments, a minimum number of slices can be equal to a number of sub-pictures plus 1. In some embodiments, the variable num_slices_in_pic_minus_subpic_num_minus1 can replace the variable num_slices_in_pic_minus_1 similar to the semantics shown in FIG. 22. Figure 27As shown in FIG. 22, instead of the number of slices minus 1 (e.g., num_slices_in_pic_minus1), the number of slices minus the number of sub-pictures minus 1 (e.g., num_slices_in_pic_minus_subpic_num_minus1) can be signaled to reduce the number of bits signaled. In some embodiments, the variable is part of the PPS. In some embodiments, as shown in FIG. 22, the value of num_slices_in_pic_minus_subpic_num_minus1 plus the number of sub-pictures plus 1 can specify the number of rectangular slices in each picture referring to the PPS. In some embodiments, similar to the semantics shown in FIG. 20, the variable num_slices_in_pic_minus2 can replace the variable num_slices_in_pic_minus_1.

[0207] In some embodiments, as shown in step S26020, a variable indicating the number of slices in a picture can be determined according to a variable indicating the number of slices in a video frame minus the number of sub-pictures in the video frame minus 1. For example, as shown in FIG. 26, the variable num_slices_in_pic_minus_subpic_num_minus1 can be derived according to the variable num_slices_in_pic_minus_subpic_num_minus1. Figure 27 As shown in FIG. 22, a flag or variable SliceNumInPic can be derived according to the variable num_slices_in_pic_minus_subpic_num_minus1.

[0208] Figure 4 A flowchart of an example video encoding method having a variable indicating whether a picture header syntax structure is present within slice headers of a video frame is shown according to some embodiments of the present disclosure. In some embodiments, the method 27000 can be performed by a video encoder. Figure 27 The method 27000 shown in FIG. 27 can be performed by a video encoder. Figure 23 In some embodiments, the method 27000 shown in FIG. 27 can be performed according to the VVC standard. Figure 27 The method 27000 shown in FIG. 27 can be performed according to the VVC standard. Figure 23 In some embodiments, the method 27000 shown in FIG. 27 can be performed according to the syntax shown in FIG. 22. Figure 23 The method 27000 shown in FIG. 27 can be performed according to the VVC standard.

[0209] In step S27010, a picture for encoding is received. The picture includes one or more slices. In some embodiments, the picture is in a bitstream. In some embodiments, the one or more slices are rectangular slices.

[0210] In step S27020, a variable indicating whether a picture header syntax structure for the picture is present in slice headers of the one or more slices is signaled. For example, the variable can be signaled as ​picture_header_in_slice_header_flag. In some embodiments, as shown in FIG. 3, the picture can have only one slice when the PH syntax structure is present in the slice header. Thus, there is no need to signal the slice address either. Consequently, the signaling of the slice address can be conditioned on the variable picture_header_in_slice_header_flag. In some embodiments, the variable is part of the PPS. ​

[0211] In some embodiments, non-transitory computer-readable storage medium comprising instructions is also provided, and the instructions can be executed by a device, such as the disclosed encoders and decoders, for performing the above-described methods. Common forms of non-transitory media include, for example, a floppy disk, a flexible disk, a hard disk, a solid-state drive, a magnetic tape, or any other magnetic data storage medium, a CD-ROM, any other optical data storage medium, any physical medium with patterns of holes, a RAM, a PROM, and EPROM, a FLASH-EPROM or any other flash memory, NVRAM, a cache, a register, any other memory chip or cartridge, and a networked version of any of the foregoing. The device can include one or more processors (CPUs), input / output interfaces, network interfaces, and / or memory.

[0212] It should be noted that relational terms, such as“first” and“second,” are used herein solely to distinguish one entity or action from another entity or action, without necessarily requiring or implying any actual relationship or order between such entities or actions. Moreover, the words“comprises,”“has,”“includes,” and“containing” and other similar forms are intended to be equivalent in meaning and be open ended in that an item or items following any one of these words is not meant to be an exhaustive listing of such item or items, or meant to be limited to only the listed item or items.

[0213] As used herein, the term“or” includes all possible combinations, unless otherwise specifically indicated, unless infeasible. For example, if a database is said to include A or B, then unless specifically stated otherwise or infeasible, the database can include A, or B, or A and B. As a second example, if a database is said to include A, B, or C, then unless specifically stated otherwise or infeasible, the database can include A, or B, or C, or A and B, or A and C, or B and C, or A, B, and C.

[0214] ​It should be understood that the above-described embodiments can be implemented by hardware, or software (program code), or a combination of hardware and software. If implemented by software, it can be stored in the above-mentioned computer readable medium. The software can perform the disclosed method when executed by the processor. The computing units and other functional units described in the present disclosure can be implemented by hardware, or software, or a combination of hardware and software. Those of ordinary skill in the art will also understand that the above-mentioned multiple modules / units can be combined into one module / unit, and each of the above-mentioned modules / units can be further divided into multiple sub-modules / sub-units.

[0215] In the foregoing specification, embodiments have been described with reference to numerous specific details that can vary from embodiment to embodiment. Certain modifications and changes can be made to the described embodiments, including but not limited to those made neither to the description nor to the examples understood from the specification and / or practices. Other embodiments will be apparent from consideration of the specification and practice of the application disclosed herein. It is intended that the specification and examples be considered as exemplary only, with the true scope and spirit of the application being indicated by the following claims. The sequence of steps shown in the accompanying drawings is for illustrative purposes only and is not intended to be limiting to any particular sequence of steps. It will be appreciated by those skilled in the art that these steps can be performed in different orders while implementing the same method.

[0216] Embodiments can be further described using the following clauses:

[0217] 1. A method of video decoding, comprising:

[0218] receiving a wrap-around motion compensation flag;

[0219] determining whether wrap-around motion compensation is enabled based on the wrap-around motion compensation flag;

[0220] in response to determining that the wrap-around motion compensation is enabled, receiving data indicating a difference between a width of a picture and an offset used to determine a horizontal wrap-around position; and

[0221] performing motion compensation according to the wrap-around motion compensation flag and the difference.

[0222] 2. The method of clause 1, wherein the difference is in units of a size of a minimum luma coding block.

[0223] 3. The method of clause 2, wherein the difference is less than or equal to (pps_pic_width_in_luma_samples / MinCbSizeY) - (CtbSizeY / MinCbSizeY) - 2, where pps_pic_width_in_luma_samples is a width of the picture in luma samples, MinCbSizeY is a size of a minimum luma coding block, and CtbSizeY is a size of a luma coding tree block.

[0224] 4. The method of clause 1, wherein performing the motion compensation further comprises:

[0225] determining a wrap-around motion compensation offset based on the width of the picture and the difference; and

[0226] performing the motion compensation based on the wrap-around motion compensation offset.

[0227] 5. The method of clause 4, wherein determining the wrap-around motion compensation offset based on the width of the picture and the difference further comprises:

[0228] dividing a width of a picture in luma samples by a size of a minimum luma coding block to generate a first value; and

[0229] determining the wrap-around motion compensation offset to be equal to the first value minus the difference.

[0230] 6. The method of clause 1, wherein receiving the data indicative of the difference further comprises:

[0231] receiving a wrap-around offset type flag;

[0232] determining whether the wrap-around offset type flag is equal to a first value or a second value;

[0233] in response to determining that the wrap-around offset type flag is equal to the first value, receiving data indicative of a difference between a width of the picture and an offset used to calculate a horizontal wrap-around position; and

[0234] in response to determining that the wrap-around offset type flag is equal to the second value, receiving data indicative of an offset used to calculate a horizontal wrap-around position.

[0235] 7. The method of clause 6, wherein each of the first value and the second value is 0 or 1.

[0236] 8. The method of clause 1, wherein the motion compensation is performed according to a Versatile Video Coding standard.

[0237] 9. The method of clause 1, wherein the picture is part of a 360-degree video sequence.

[0238] 10. The method of clause 1, wherein the wrap-around motion compensation flag and the difference are signaled in a picture parameter set (PPS).

[0239] 11. A method of video decoding, comprising:

[0240] signaling a wrap-around motion compensation flag, the flag indicating whether wrap-around motion compensation is enabled;

[0241] in response to the wrap-around motion compensation flag indicating that the wrap-around motion compensation is enabled, signaling data indicating a difference between a width of the picture and an offset used to determine a horizontal wrap-around position; and

[0242] performing motion compensation according to the wrap-around motion compensation flag and the difference.

[0243] 12. The method of clause 11, wherein the difference is in units of a size of a minimum luma coding block.

[0244] 13. The method of clause 12, wherein the difference is less than or equal to (pps_pic_width_in_luma_samples / MinCbSizeY) - (CtbSizeY / MinCbSizeY) - 2, where pps_pic_width_in_luma_samples is a width of the picture in luma samples, MinCbSizeY is a size of a minimum luma coding block, and CtbSizeY is a size of a luma coding tree block.

[0245] 14. The method of clause 11, wherein performing the motion compensation further comprises:

[0246] determining a wrap-around motion compensation offset according to the width of the picture and the difference; and

[0247] performing the motion compensation according to the wrap-around motion compensation offset.

[0248] 15. The method of clause 14, wherein determining the wrap-around motion compensation offset according to the width of the picture and the difference further comprises:

[0249] dividing a width of the picture in luma samples by a size of a minimum luma coding block to generate a first value; and

[0250] determining the wrap-around motion compensation offset to be equal to the first value minus the difference.

[0251] 16. The method of clause 11, wherein signaling the data indicative of the difference further comprises:

[0252] signaling a wrap-around offset type flag, wherein a value of the wrap-around offset type flag is a first value or a second value;

[0253] in response to the value of the wrap-around offset type flag being equal to the first value, signaling data indicative of a difference between a width of the picture and an offset used to calculate a horizontal wrap-around position; and

[0254] in response to the value of the wrap-around offset type flag being equal to the second value, signaling data used to calculate an offset used to calculate a horizontal wrap-around position.

[0255] 17. The method of clause 16, wherein each of the first value and the second value is 0 or 1.

[0256] 18. The method of clause 11, wherein the motion compensation is performed according to a Versatile Video Coding standard.

[0257] 19. The method of clause 11, wherein the picture is part of a 360-degree video sequence.

[0258] 20. The method of clause 11, wherein the wrap-around motion compensation flag and the difference are signaled in a picture parameter set (PPS).

[0259] 21. A method of video encoding, comprising:

[0260] receiving a picture for encoding, wherein the picture comprises one or more slices; and

[0261] in a picture parameter set of the picture, signaling a variable indicative of a number of slices in a video frame minus 2.

[0262] 22. The method of clause 21, wherein the picture is in a bitstream.

[0263] 23. The method of clause 21, wherein the picture is encoded according to a Versatile Video Coding standard.

[0264] 24. The method of clause 21, wherein the one or more slices are rectangular slices.

[0265] 25. A method of video encoding, comprising:

[0266] receiving a picture for encoding, wherein the picture comprises one or more slices and one or more sub-pictures; and

[0267] In a picture parameter set of the picture, a variable is signaled that indicates a number of slices in the picture minus a number of sub-pictures in the picture minus 1.

[0268] 26. The method of clause 25, wherein the picture is in a bitstream.

[0269] 27. The method of clause 25, further comprising:

[0270] determining a variable that indicates a number of slices in the picture based on the variable that indicates a number of slices in the picture minus a number of sub-pictures in the picture minus 1.

[0271] 28. The method of clause 25, wherein the picture is encoded according to a Versatile Video Coding standard.

[0272] 29. The method of clause 25, wherein the one or more slices are rectangular slices.

[0273] 30. A method of video encoding, comprising:

[0274] receiving a picture for encoding, wherein the picture comprises one or more slices;

[0275] signaling a variable that indicates whether a picture header syntax structure of the picture is present in a slice header of the one or more slices; and

[0276] signaling a slice address based on the variable.

[0277] 31. The method of clause 30, wherein the picture is in a bitstream.

[0278] 32. The method of clause 30, wherein the picture is encoded according to a Versatile Video Coding standard.

[0279] 33. The method of clause 30, wherein the one or more slices are rectangular.

[0280] 34. A system for performing video data processing, the system comprising:

[0281] a memory that stores a set of instructions; and

[0282] a processor configured to execute the set of instructions to cause the system to perform the following operations:

[0283] receiving a wrap-around motion compensation flag;

[0284] determining whether to enable wrap-around motion compensation based on the wrap-around motion compensation flag;

[0285] in response to determining that the wraparound motion compensation is enabled, receiving data indicative of a difference between a width of a picture and an offset for determining a horizontal wraparound position; and

[0286] performing motion compensation according to the wraparound motion compensation flag and the difference.

[0287] 35. The system of clause 34, wherein the difference is in units of a size of a minimum luma coding block.

[0288] 36. The system of clause 35, wherein the difference is less than or equal to (pps_pic_width_in_luma_samples / MinCbSizeY) - (CtbSizeY / MinCbSizeY) - 2, where pps_pic_width_in_luma_samples is a width of the picture in luma samples, MinCbSizeY is a size of a minimum luma coding block, and CtbSizeY is a size of a luma coding tree block.

[0289] 37. The system of clause 34, wherein, in performing the motion compensation, the processor is configured to execute the instruction set to cause the system to perform:

[0290] determining a wraparound motion compensation offset according to the width of the picture and the difference; and

[0291] performing the motion compensation according to the wraparound motion compensation offset.

[0292] 38. The system of clause 37, wherein, in determining the wraparound motion compensation offset according to the width of the picture and the difference, the processor is configured to execute the instruction set to cause the system to perform:

[0293] dividing a width of a picture in luma samples by a size of a minimum luma coding block to generate a first value; and

[0294] determining the wraparound motion compensation offset to be equal to the first value minus the difference.

[0295] 39. The system of clause 34, wherein, in receiving the data indicative of the difference, the processor is configured to execute the instruction set to cause the system to perform:

[0296] receiving a wraparound offset type flag;

[0297] determining whether the wraparound offset type flag is equal to a first value or a second value;

[0298] in response to determining that the wrap-around offset type flag is equal to the first value, receiving data indicative of a difference between a width of the picture and an offset amount used to calculate a horizontal wrap-around position; and

[0299] in response to determining that the wrap-around offset type flag is equal to the second value, receiving data indicative of an offset amount used to calculate a horizontal wrap-around position.

[0300] 40. The system of clause 39, wherein each of the first value and the second value is 0 or 1.

[0301] 41. The system of clause 34, wherein the motion compensation is performed in accordance with a Versatile Video Coding standard.

[0302] 42. The system of clause 34, wherein the picture is part of a 360-degree video sequence.

[0303] 43. The system of clause 34, wherein the wrap-around motion compensation flag and the difference value are signaled in a picture parameter set (PPS).

[0304] 44. A system for performing video data processing, the system comprising:

[0305] a memory storing a set of instructions; and

[0306] a processor configured to execute the set of instructions to cause the system to perform the following operations:

[0307] signal a wrap-around motion compensation flag, the flag indicating whether wrap-around motion compensation is enabled;

[0308] in response to the wrap-around motion compensation flag indicating that the wrap-around motion compensation is enabled, signal data indicative of a difference between a width of the picture and an offset amount used to determine a horizontal wrap-around position; and

[0309] perform motion compensation in accordance with the wrap-around motion compensation flag and the difference value.

[0310] 45. The system of clause 44, wherein the difference value is in units of a size of a smallest luma coding block.

[0311] 46. The system of clause 45, wherein the difference is less than or equal to (pps_pic_width_in_luma_samples / MinCbSizeY) - (CtbSizeY / MinCbSizeY) - 2, where pps_pic_width_in_luma_samples is a width of the picture in luma samples, MinCbSizeY is a size of a minimum luma coding block, and CtbSizeY is a size of a luma coding tree block.

[0312] 47. The system of clause 44, wherein, in performing the motion compensation, the processor is configured to execute the instruction set to cause the system to perform:

[0313] determining a wrap-around motion compensation offset based on the width of the picture and the difference; and

[0314] performing the motion compensation based on the wrap-around motion compensation offset.

[0315] 48. The system of clause 47, wherein, in determining the wrap-around motion compensation offset based on the width of the picture and the difference, the processor is configured to execute the instruction set to cause the system to perform:

[0316] dividing the width of the picture in luma samples by a size of a minimum luma coding block to generate a first value; and

[0317] determining the wrap-around motion compensation offset to be equal to the first value minus the difference.

[0318] 49. The system of clause 44, wherein, in receiving the data indicative of the difference, the processor is configured to execute the instruction set to cause the system to perform:

[0319] signaling a wrap-around offset type flag, wherein a value of the wrap-around offset type flag is a first value or a second value;

[0320] in response to the value of the wrap-around offset type flag being equal to the first value, signaling data indicative of a difference between a width of the picture and an offset used to calculate a horizontal wrap-around position; and

[0321] in response to the value of the wrap-around offset type flag being equal to the second value, signaling data used to calculate an offset for a horizontal wrap-around position.

[0322] 50. The system of clause 49, wherein each of the first value and the second value is 0 or 1.

[0323] 51. The system of clause 44, wherein the motion compensation is performed according to a Versatile Video Coding standard.

[0324] 52. The system of clause 44, wherein the image is part of a 360-degree video sequence.

[0325] 3. The system of clause 44, wherein the wrap-around motion compensation flag and the difference value are signaled in a picture parameter set (PPS).

[0326] 54. A system for performing video encoding, the system comprising:

[0327] a memory storing a set of instructions; and

[0328] a processor configured to execute the set of instructions to cause the system to perform the following operations:

[0329] receive an image for encoding, wherein the image comprises one or more slices; and

[0330] signal, in a picture parameter set of the image, a variable indicating a number of slices in a video frame minus 2.

[0331] 55. The system of clause 54, wherein the image is in a bitstream.

[0332] 56. The system of clause 54, wherein the image is encoded according to a Versatile Video Coding standard.

[0333] 57. The system of clause 54, wherein the one or more slices are rectangular slices.

[0334] 58. A system for performing video encoding, the system comprising:

[0335] a memory storing a set of instructions; and

[0336] a processor configured to execute the set of instructions to cause the system to perform the following operations:

[0337] receive an image for encoding, wherein the image comprises one or more slices and one or more sub-pictures; and

[0338] signal, in a picture parameter set of the image, a variable indicating a number of slices in the image minus a number of sub-pictures in the image minus 1.

[0339] 59. The system of clause 58, wherein the image is in a bitstream.

[0340] 60. The system of clause 58, wherein the processor is configured to execute the set of instructions to cause the system to perform:

[0341] determine a variable indicating a number of slices in the picture according to the indication of the number of slices in the picture minus the number of sub-pictures in the picture minus 1.

[0342] 61. The system of clause 58, wherein the picture is encoded according to a Versatile Video Coding standard.

[0343] 62. The system of clause 58, wherein the one or more slices are rectangular slices.

[0344] 63. A system for performing video encoding, the system comprising:

[0345] a memory storing a set of instructions; and

[0346] a processor configured to execute the set of instructions to cause the system to perform:

[0347] receive a picture for encoding, wherein the picture comprises one or more slices;

[0348] signal a variable indicating whether a picture header syntax structure of the picture is present in a slice header of the one or more slices; and

[0349] signal a slice address according to the variable.

[0350] 64. The system of clause 63, wherein the picture is in a bitstream.

[0351] 65. The system of clause 63, wherein the picture is encoded according to a Versatile Video Coding standard.

[0352] 66. The system of clause 63, wherein the one or more slices are rectangular.

[0353] 67. A non-transitory computer-readable medium storing a set of instructions executable by one or more processors of an apparatus to cause the apparatus to initiate a method for performing video data processing, the method comprising:

[0354] receiving a wrap-around motion compensation flag;

[0355] determining whether wrap-around motion compensation is enabled based on the wrap-around motion compensation flag;

[0356] in response to determining that the wrap-around motion compensation is enabled, receiving data indicating a difference between a width of a picture and an offset amount used to determine a horizontal wrap-around position; and

[0357] performing motion compensation according to the wrap-around motion compensation flag and the difference value.

[0358] 68. The non-transitory computer-readable medium of clause 67, wherein the difference value is in units of a size of a minimum luma coding block.

[0359] 69. The non-transitory computer-readable medium of clause 68, wherein the difference value is less than or equal to (pps_pic_width_in_luma_samples / MinCbSizeY) - (CtbSizeY / MinCbSizeY) - 2, where pps_pic_width_in_luma_samples is a width of the picture in luma samples, MinCbSizeY is a size of a minimum luma coding block, and CtbSizeY is a size of a luma coding tree block.

[0360] 70. The non-transitory computer-readable medium of clause 67, wherein performing the motion compensation further comprises:

[0361] determining a wrap-around motion compensation offset according to the width of the picture and the difference value; and

[0362] performing the motion compensation according to the wrap-around motion compensation offset.

[0363] 71. The non-transitory computer-readable medium of clause 70, wherein determining the wrap-around motion compensation offset according to the width of the picture and the difference value further comprises:

[0364] dividing a width of a picture in luma samples by a size of a minimum luma coding block to generate a first value; and

[0365] determining the wrap-around motion compensation offset to be equal to the first value minus the difference value.

[0366] 72. The non-transitory computer-readable medium of clause 67, wherein receiving data indicative of a difference value further comprises:

[0367] receiving a wrap-around offset type flag;

[0368] determining whether the wrap-around offset type flag is equal to a first value or a second value;

[0369] in response to determining that the wrap-around offset type flag is equal to the first value, receiving data indicative of a difference value between a width of the picture and an offset used to calculate a horizontal wrap-around position; and

[0370] in response to determining that the wrap-around offset type flag is equal to the second value, receiving data indicative of a difference between a width of the picture and an offset used to determine a horizontal wrap-around position.

[0371] 73. The non-transitory computer-readable medium of clause 72, wherein each of the first value and the second value is 0 or 1.

[0372] 74. The non-transitory computer-readable medium of clause 67, wherein the motion compensation is performed in accordance with a Versatile Video Coding standard.

[0373] 75. The non-transitory computer-readable medium of clause 67, wherein the picture is part of a 360-degree video sequence.

[0374] 76. The non-transitory computer-readable medium of clause 67, wherein the wrap-around motion compensation flag and the difference are signaled in a picture parameter set (PPS).

[0375] 77. A non-transitory computer-readable medium storing a set of instructions executable by one or more processors of an apparatus to cause the apparatus to initiate a method for performing video data processing, the method comprising:

[0376] signaling a wrap-around motion compensation flag, the flag indicating whether wrap-around motion compensation is enabled;

[0377] in response to the wrap-around motion compensation flag indicating that the wrap-around motion compensation is enabled, signaling data indicative of a difference between a width of the picture and an offset used to determine a horizontal wrap-around position; and

[0378] performing motion compensation in accordance with the wrap-around motion compensation flag and the difference.

[0379] 78. The non-transitory computer-readable medium of clause 77, wherein the difference is in units of a size of a minimum luma coding block.

[0380] 79. The non-transitory computer-readable medium of clause 78, wherein the difference is less than or equal to (pps_pic_width_in_luma_samples / MinCbSizeY) - (CtbSizeY / MinCbSizeY) - 2, where pps_pic_width_in_luma_samples is a width of the picture in luma samples, MinCbSizeY is a size of a minimum luma coding block, and CtbSizeY is a size of a luma coding tree block.

[0381] 80. The non-transitory computer-readable medium of clause 77, wherein performing the motion compensation further comprises:

[0382] determining a wrap-around motion compensation offset based on the width of the picture and the difference; and

[0383] performing the motion compensation based on the wrap-around motion compensation offset.

[0384] 81. The non-transitory computer-readable medium of clause 80, wherein determining the wrap-around motion compensation offset based on the width of the picture and the difference further comprises:

[0385] dividing the width of the picture in units of luma samples by a size of a minimum luma coding block to generate a first value; and

[0386] determining the wrap-around motion compensation offset to be equal to the first value minus the difference.

[0387] 82. The non-transitory computer-readable medium of clause 77, wherein receiving data indicative of a difference further comprises:

[0388] signaling a wrap-around offset type flag, wherein a value of the wrap-around offset type flag is a first value or a second value;

[0389] in response to the value of the wrap-around offset type flag being equal to the first value, signaling data indicative of a difference between a width of the picture and an offset used to calculate a horizontal wrap-around position; and

[0390] in response to the value of the wrap-around offset type flag being equal to the second value, signaling data used to calculate an offset for a horizontal wrap-around position.

[0391] 83. The non-transitory computer-readable medium of clause 82, wherein each of the first value and the second value is 0 or 1.

[0392] 84. The non-transitory computer-readable medium of clause 77, wherein the motion compensation is performed according to a Versatile Video Coding standard.

[0393] 85. The non-transitory computer-readable medium of clause 77, wherein the picture is part of a 360-degree video sequence.

[0394] 86. The non-transitory computer-readable medium of clause 77, wherein the wrap-around motion compensation flag and the difference are signaled in a picture parameter set (PPS).

[0395] 87. A non-transitory computer-readable medium storing a set of instructions executable by one or more processors of an apparatus to cause the apparatus to initiate a method for performing video encoding, the method comprising:

[0396] receiving an image for encoding, wherein the image comprises one or more slices; and

[0397] in a picture parameter set of the image, signaling a variable indicating a number of slices in a video frame minus 2.

[0398] 88. The non-transitory computer-readable medium of clause 87, wherein the image is in a bitstream.

[0399] 89. The non-transitory computer-readable medium of clause 87, wherein the image is encoded according to a Versatile Video Coding standard.

[0400] 90. The non-transitory computer-readable medium of clause 87, wherein the one or more slices are rectangular slices.

[0401] 91. A non-transitory computer-readable medium storing a set of instructions executable by one or more processors of an apparatus to cause the apparatus to initiate a method for performing video encoding, the method comprising:

[0402] receiving an image for encoding, wherein the image comprises one or more slices and one or more sub-pictures; and

[0403] in a picture parameter set of the image, signaling a variable indicating a number of slices in the image minus a number of sub-pictures in the image minus 1.

[0404] 92. The non-transitory computer-readable medium of clause 91, wherein the image is in a bitstream.

[0405] 93. The non-transitory computer-readable medium of clause 91, further comprising:

[0406] determining, according to the variable indicating a number of slices in the image minus a number of sub-pictures in the image minus 1, a variable indicating a number of slices in the image.

[0407] 94. The non-transitory computer-readable medium of clause 91, wherein the image is encoded according to a Versatile Video Coding standard.

[0408] 95. The non-transitory computer-readable medium of clause 91, wherein the one or more slices are rectangular slices.

[0409] 96. A non-transitory computer-readable medium storing a set of instructions executable by one or more processors of an apparatus to cause the apparatus to initiate a method for performing video encoding, the method comprising:

[0410] receiving an image for encoding, wherein the image comprises one or more slices;

[0411] signaling a variable indicating whether a picture header syntax structure of the image is present in a slice header of the one or more slices; and

[0412] signaling a slice address according to the variable.

[0413] 97. The non-transitory computer-readable medium of clause 96, wherein the image is in a bitstream.

[0414] 98. The non-transitory computer-readable medium of clause 96, wherein the image is encoded according to a Versatile Video Coding standard.

[0415] 99. The non-transitory computer-readable medium of clause 96, wherein the one or more slices are rectangular.

[0416] In the drawings and specification, there have been disclosed exemplary embodiments. However, many variations and modifications can be made to these embodiments. Consequently, it is intended that the scope of the application be limited only by the broadest interpretation of the appended claims to be accorded under applicable law, as much as by the embodiments disclosed above and as appropriately interpreted now or in the future.

Claims

1. A method of video coding, comprising: signaling a rectangular slice flag in a picture parameter set (PPS) ; determining whether the rectangular slice flag is equal to a first value or a second value, wherein the rectangular slice flag being equal to the first value indicates that a raster scan slice mode is used in each picture referring to the PPS, and the rectangular slice flag being equal to the second value indicates that a rectangular slice mode is used in each picture referring to the PPS; in response to the rectangular slice flag being equal to the second value, signaling in the PPS a parameter indicating a number of rectangular slices in each picture referring to the PPS minus 1; and signaling in the PPS a wraparound motion compensation flag and a picture width; determining a wraparound motion compensation parameter according to the picture width and a difference between the picture width and an offset used to determine a horizontal wraparound position; and performing wraparound motion compensation according to the wraparound motion compensation flag and the wraparound motion compensation parameter.

2. The video coding method of claim 1, wherein, The first value is 1 and the second value is 0.

3. The video coding method of claim 1, wherein, The motion compensation is performed according to a general video coding standard.

4. The video coding method of claim 1, wherein, The picture is part of a 360-degree video sequence. 5.A method of video decoding, comprising: receiving a rectangular slice flag from a picture parameter set (PPS) ; determining whether the rectangular slice flag is equal to a first value or a second value, wherein the rectangular slice flag being equal to the first value indicates that a raster scan slice mode is used in each picture referring to the PPS, and the rectangular slice flag being equal to the second value indicates that a rectangular slice mode is used in each picture referring to the PPS; in response to the rectangular slice flag being equal to the second value, receiving in the PPS a parameter indicating a number of rectangular slices in each picture referring to the PPS minus 1; and receiving in the PPS a wraparound motion compensation flag and a picture width; determining a wraparound motion compensation parameter according to the picture width and a difference between the picture width and an offset used to determine a horizontal wraparound position; and performing wraparound motion compensation according to the wraparound motion compensation flag and the wraparound motion compensation parameter.

6. The video decoding method of claim 5, wherein, The first value is 1 and the second value is 0.

7. The video decoding method of claim 5, wherein, The motion compensation is performed according to a general video coding standard.

8. A non-transitory computer-readable storage medium storing a bitstream, the non-transitory computer-readable medium being part of a computing device, the bitstream being generable by one or more processors of the computing device executing a set of instructions, wherein, Execution of the set of instructions causes the computing device to perform: signaling a rectangular slice flag in a picture parameter set (PPS) ; determining whether the rectangular slice flag is equal to a first value or a second value, wherein the rectangular slice flag being equal to the first value indicates that a raster scan slice mode is used in each picture referring to the PPS, and the rectangular slice flag being equal to the second value indicates that a rectangular slice mode is used in each picture referring to the PPS; in response to the rectangular slice flag being equal to the second value, signaling in the PPS a parameter indicating a number of rectangular slices in each picture referring to the PPS minus 1; and signaling in the PPS a wraparound motion compensation flag and a picture width; A wrap-around motion compensation parameter is determined according to the picture width and a difference between the picture width and an offset used to determine a horizontal wrap-around position; and wrap-around motion compensation is performed according to the wrap-around motion compensation flag and the wrap-around motion compensation parameter.

9. The non-transitory computer-readable storage medium of claim 8, wherein, The first value is 1 and the second value is 0.

10. The non-transitory computer-readable storage medium of claim 8, wherein, The motion compensation is performed according to a general video coding standard.

Citation Information

Patent Citations

  • Encoding control apparatus, encoding control method, and storage medium

    US20080298465A1

  • Methods and apparatus for flexible grid regions

    WO2020056247A1