Method of signaling video encoded data
By receiving the surrounding motion compensation flag and strip information, and enabling the surrounding motion compensation and strip address processing, the problem of insufficient encoding efficiency in the existing video encoding standards is solved, and more efficient video data processing is achieved.
Patent Information
- Application Number
- CN202510318494.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2020-03-26
- Filing Date
- 2021-03-26
- Publication Date
- 2025-08-29
AI Technical Summary
In the efficient video encoding technology, the existing video encoding standards have not fully utilized the surrounding motion compensation and strip layout information, resulting in insufficient encoding efficiency.
By receiving the surround motion compensation flag and strip information, it is determined whether surround motion compensation is enabled and signal related variables in the image parameter set to perform motion compensation and strip address processing.
Improves the efficiency of video encoding, reduces the bandwidth required for storage and transmission, and enhances the parallel processing and fault tolerance of encoder and decoder.
Smart Images

Figure CN120568071A_ABST
Abstract
Description
CROSS-REFERENCE TO RELATED APPLICATIONS
[0001] This disclosure claims priority to and the benefit of U.S. Provisional Patent Application No. 63 / 000,443, filed on March 26, 2020. The provisional application is incorporated herein by reference in its entirety. Technical Field
[0002] The present disclosure relates generally to video data processing, and more particularly, to methods and apparatus for signaling information regarding surround motion compensation, slice layout, and slice addresses. Background Art
[0003] A video is a set of static images (or "frames") that capture visual information. To reduce storage memory and transmission bandwidth, the video can be compressed before storage or transmission, and then decompressed before display. The compression process is usually called encoding, and the decompression process is usually called decoding. There are many video coding formats that use standardized video coding techniques, the most common of which are based on prediction, transform, quantization, entropy coding, and in-loop filtering. Standardization organizations have developed video coding standards, such as the High Efficiency Video Coding (HEVC / H.265) standard, the Versatile Video Coding (VVC / H.266) standard, and the AVS standard, which specify specific video coding formats. As more and more advanced video coding technologies are adopted in video standards, the coding efficiency of new video coding standards is getting higher and higher. Summary of the Invention
[0004] An embodiment of the present disclosure provides a method for notifying video encoding data with a signal, the method comprising: receiving a surround motion compensation flag; determining whether to enable surround motion compensation based on the surround motion compensation flag; in response to determining that the surround motion compensation is enabled, receiving data indicating the difference between the width of an image and an offset used to determine a horizontal surround position; and performing motion compensation based on the surround motion compensation flag and the difference.
[0005] An embodiment of the present disclosure also provides a method for signaling video encoding data, the method comprising: receiving an image for encoding, wherein the image comprises one or more slices; and in an image parameter set of the image, signaling a variable indicating the number of slices in a video frame minus 2.
[0006] An embodiment of the present disclosure also provides a method for signaling video encoding data, the method comprising: receiving an image for encoding, wherein the image comprises one or more slices and one or more sub-images; and in an image parameter set of the image, signaling a variable indicating the number of slices in the image minus the number of sub-images in the image minus 1.
[0007] An embodiment of the present disclosure also provides a method for signaling video encoding data, the method comprising: receiving an image for encoding, wherein the image comprises one or more slices; signaling a variable indicating whether a slice header syntax structure of the image exists in a slice header of the one or more slices; and signaling a slice address based on the variable.
[0008] An embodiment of the present disclosure also provides a system for performing video data processing, the system comprising: a memory storing an instruction set; and a processor configured to execute the instruction set to cause the system to: receive a surround motion compensation flag; determine whether to enable surround motion compensation based on the surround motion compensation flag; in response to determining that the surround motion compensation is enabled, receive data indicating the difference between the width of an image and an offset used to determine a horizontal surround position; and perform motion compensation based on the surround motion compensation flag and the difference.
[0009] An embodiment of the present disclosure also provides a system for performing video data processing, the system comprising: a memory storing a set of instructions; and a processor configured to execute the set of instructions so that the system performs: receiving an image for encoding, wherein the image comprises one or more slices; and in an image parameter set of the image, signaling a variable indicating the number of slices in the video frame minus 2.
[0010] An embodiment of the present disclosure also provides a system for performing video data processing, the system comprising: a memory storing an instruction set; and a processor configured to execute the instruction set so that the system performs: receiving an image for encoding, wherein the image comprises one or more slices and one or more sub-images; and in an image parameter set of the image, signaling a variable indicating the number of slices in the image minus the number of sub-images in the image minus 1.
[0011] An embodiment of the present disclosure also provides a system for performing video data processing, the system comprising: a memory storing an instruction set; and a processor configured to execute the instruction set so that the system performs: receiving an image for encoding, wherein the image comprises one or more slices; signaling a variable indicating whether a picture header syntax structure of the image exists in a slice header of the one or more slices; and signaling a slice address based on the variable.
[0012] Embodiments of the present disclosure also provide a non-transitory computer-readable medium storing an instruction set, which can be executed by one or more processors of a device to cause the device to initiate a method for performing video data processing, the method comprising: receiving a surround motion compensation flag; determining whether to enable surround motion compensation based on the surround motion compensation flag; in response to determining that the surround motion compensation is enabled, receiving data indicating the difference between the width of the image and the offset used to determine the horizontal surround position; and performing motion compensation based on the surround motion compensation flag and the difference.
[0013] Embodiments of the present disclosure also provide a non-transitory computer-readable medium storing an instruction set, wherein the instruction set can be executed by one or more processors of a device to cause the device to initiate a method for performing video data processing, the method comprising: receiving an image for encoding, wherein the image includes one or more slices; and in an image parameter set of the image, signaling a variable indicating the number of slices in the video frame minus 2.
[0014] Embodiments of the present disclosure also provide a non-transitory computer-readable medium storing an instruction set, wherein the instruction set can be executed by one or more processors of a device to cause the device to initiate a method for performing video data processing, the method comprising: receiving an image for encoding, wherein the image comprises one or more slices and one or more sub-images; and in an image parameter set for the image, signaling a variable indicating the number of slices in the image minus the number of sub-images in the image minus 1.
[0015] Embodiments of the present disclosure also provide a non-transitory computer-readable medium storing an instruction set, wherein the instruction set can be executed by one or more processors of a device to cause the device to initiate a method for performing video data processing, the method comprising: receiving an image for encoding, wherein the image includes one or more slices; signaling a variable indicating whether a picture header syntax structure of the image exists in a slice header of the one or more slices; and signaling a slice address based on the variable. BRIEF DESCRIPTION OF THE DRAWINGS
[0016]
[0011] Embodiments and aspects of the present disclosure are illustrated in the following detailed description and accompanying drawings.The various features shown in the drawings are not drawn to scale.
[0017] Figure 1 The structure of an example video sequence according to some embodiments of the present disclosure is shown.
[0018] Figure 2A A schematic diagram of an exemplary encoding process according to some embodiments of the present disclosure is shown.
[0019] Figure 2BA schematic diagram illustrating another example encoding process according to some embodiments of the present disclosure is shown.
[0020] Figure 3A A schematic diagram illustrating an exemplary decoding process according to some embodiments of the present disclosure is shown.
[0021] Figure 3B A schematic diagram illustrating another example decoding process according to some embodiments of the present disclosure is shown.
[0022] Figure 4 A block diagram of an example apparatus for encoding or decoding video according to some embodiments of the present disclosure is shown.
[0023] Figure 5A A schematic diagram illustrating an example blending operation for generating a reconstructed equirectangular projection according to some embodiments of the present disclosure is shown.
[0024] Figure 5B A schematic diagram illustrating an example cropping operation for generating a reconstructed equirectangular projection according to some embodiments of the present disclosure is shown.
[0025] Figure 6A A schematic diagram illustrating an exemplary horizontal surround motion compensation process for equirectangular projection according to some embodiments of the present disclosure is shown.
[0026] Figure 6B A schematic diagram illustrating an exemplary horizontal surround motion compensation process for filling an equirectangular projection according to some embodiments of the present disclosure is shown.
[0027] Figure 7 The syntax of an example advanced wraparound offset is shown in accordance with some embodiments of the present disclosure.
[0028] Figure 8 The semantics of example high-level wrap-around offsets according to some embodiments of the present disclosure are shown.
[0029] Figure 9 Schematic diagram showing example stripe and sub-image partitioning of an image according to some embodiments of the present disclosure.
[0030] Figure 10 Schematic diagram showing example stripe and sub-image partitioning of an image with different stripes and sub-images according to some embodiments of the present disclosure.
[0031] Figure 11 The syntax of an example image parameter set for tile mapping and stripe layout according to some embodiments of the present disclosure is shown.
[0032] Figure 12A and Figure 12BThe semantics of example image parameter sets for tile mapping and stripe layout are shown according to some embodiments of the present disclosure.
[0033] Figure 13 The syntax of an example slice header according to some embodiments of the present disclosure is shown.
[0034] Figure 14 The semantics of an example slice header according to some embodiments of the present disclosure are shown.
[0035] Figure 15 An example improved syntax of an image parameter set according to some embodiments of the present disclosure is shown.
[0036] Figure 16 The semantics of an example improved image parameter set according to some embodiments of the present disclosure are shown.
[0037] Figure 17 The syntax of an example image parameter set with the variable wraparound_offset_type is shown according to some embodiments of the present disclosure.
[0038] Figure 18 The semantics of an example improved picture parameter set with the variable wraparound_offset_type according to some embodiments of the present disclosure are shown.
[0039] Figure 19 Shown is the syntax of an example picture parameter set with the variable num_slices_in_pic_minus2, according to some embodiments of the present disclosure.
[0040] Figure 20A and Figure 20B The semantics of an example improved picture parameter set with the variable num_slices_in_pic_minus2 is shown according to some embodiments of the present disclosure.
[0041] Figure 21 Shown is the syntax of an example picture parameter set with the variable num_slices_in_pic_minus_subpic_num_minus1 according to some embodiments of the present disclosure.
[0042] Figure 22A and Figure 22B An example improved picture parameter set semantics with the variable num_slices_in_pic_minus_subpic_num_minus1 is shown according to some embodiments of the present disclosure.
[0043] Figure 23 The syntax of an example updated slice header according to some embodiments of the present disclosure is shown.
[0044] Figure 24 A flow chart is shown of an example video encoding method having a variable that signals the difference between the width of a video frame and an offset used to calculate horizontal surround position, according to some embodiments of the present disclosure.
[0045] Figure 25 A flow chart is shown of an example video encoding method with a variable signaling the number of slices in a video frame minus two, according to some embodiments of the present disclosure.
[0046] Figure 26 A flow chart of an example video encoding method with a variable that signals a variable indicating the number of slices in a video frame minus the number of sub-pictures in the video frame minus one is shown, according to some embodiments of the present disclosure.
[0047] Figure 27 A flow chart illustrating an example video encoding method with a variable indicating whether a picture header syntax structure is present in a slice header of a video frame according to some embodiments of the present disclosure is shown. DETAILED DESCRIPTION
[0048] Reference will now be made in detail to exemplary embodiments, examples of which are illustrated in the accompanying drawings. The following description refers to the accompanying drawings, in which the same numbers in different figures represent the same or similar elements unless otherwise indicated. The embodiments set forth in the following description of the exemplary embodiments do not represent all embodiments consistent with the present disclosure. Instead, they are merely examples of devices and methods consistent with aspects related to the present disclosure as described in the appended claims. Specific aspects of the present disclosure are described in more detail below. In the event of a conflict with terms and / or definitions incorporated by reference, the terms and definitions provided herein shall prevail.
[0049] The Joint Video Experts Team (JVET) of the ITU-T Video Coding Experts Group (ITU-T VCEG) and the ISO / IEC Moving Picture Experts Group (ISO / IEC MPEG) is currently developing the Versatile Video Coding (VVC / H.266) standard. The VVC standard aims to double the compression efficiency of its predecessor, the High Efficiency Video Coding (HEVC / H.265) standard. In other words, VVC aims to achieve the same subjective quality as HEVC / H.265 using half the bandwidth.
[0050] To achieve the same subjective quality as HEVC / H.265 using half the bandwidth, JVET has been developing technologies beyond HEVC using the Joint Exploration Model (JEM) reference software. As coding technologies are incorporated into JEM, JEM achieves higher coding performance than HEVC.
[0051] The VVC standard was developed recently and continues to include more coding techniques that provide better compression performance. VVC is based on the same hybrid video coding system that has been used in modern video compression standards such as HEVC, H.264 / AVC, MPEG2, and H.263.
[0052] A video is a set of static images (or "frames") arranged in time to store visual information. These images can be captured and stored in time using a video capture device (e.g., a camera), and displayed in time using a video playback device (e.g., a television, computer, smartphone, tablet, video player, or any end-user terminal with a display). Furthermore, in some applications, the video capture device can send the captured video to a video playback device (e.g., a computer with a monitor) in real time, for example, for monitoring, conferencing, or live broadcasting.
[0053] To reduce the storage space and transmission bandwidth required for such applications, video can be compressed before storage and transmission, and decompressed before display. Compression and decompression can be implemented by software executed by a processor (e.g., a processor of a general-purpose computer) or dedicated hardware. The module used for compression is generally referred to as an "encoder," and the module used for decompression is generally referred to as a "decoder." Encoders and decoders can be collectively referred to as "codecs." Encoders and decoders can be implemented as any of various suitable hardware, software, or combinations thereof. For example, hardware implementations of encoders and decoders can include circuits such as one or more microprocessors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), discrete logic, or any combination thereof. Software implementations of encoders and decoders can include program code, computer-executable instructions, firmware, or any suitable computer-implemented algorithm or process fixed in a computer-readable medium. Video compression and decompression can be implemented using various algorithms or standards, such as MPEG-1, MPEG-2, MPEG-4, the H.26x series, and the like. In some applications, a codec can decompress video from a first coding standard and recompress the decompressed video using a second coding standard. In this case, the codec can be referred to as a "transcoder."
[0054] The video encoding process can identify and retain useful information that can be used to reconstruct the image, while ignoring unimportant reconstruction information. If ignoring unimportant information cannot be fully reconstructed, such an encoding process can be called "lossy." Otherwise, it can be called "lossless." Most encoding processes are lossy, as a trade-off to reduce the required storage space and transmission bandwidth.
[0055] In many cases, useful information about the image being coded (referred to as the "current image") includes changes relative to a reference image (e.g., a previously coded and reconstructed image). Such changes can include changes in pixel position, brightness, or color, with position changes being of primary interest. A change in the position of a group of pixels representing an object can reflect the object's motion between the reference image and the current image.
[0056] A picture that is coded without reference to another picture (i.e., it is its own reference picture) is called an "I-picture." A picture coded using a previous picture as a reference is called a "P-picture," and a picture coded using both a previous picture and a future picture as reference is called a "B-picture" (the reference is "bidirectional").
[0057] Figure 1 The structure of an example video sequence 100 according to some embodiments of the present disclosure is shown. The video sequence 100 can be live video or video that has been captured and archived. The video 100 can be real-life video, computer-generated video (e.g., computer game video), or a combination of the two (e.g., real video with augmented reality effects). The video sequence 100 can be input from a video capture device (e.g., a camera), an archive containing previously captured video (e.g., a video file stored on a storage device), or a video feed interface (e.g., a video broadcast transceiver) that receives video from a video content provider.
[0058] like Figure 1 As shown, video sequence 100 may include a series of images arranged temporally along a timeline, including images 102, 104, 106, and 108. Images 102-106 are consecutive, with more images between images 106 and 108. Figure 1 , picture 102 is an I-picture, and its reference picture is picture 102 itself. Picture 104 is a P-picture, and its reference picture is picture 102, as indicated by the arrow. Picture 106 is a B-picture, and its reference pictures are pictures 104 and 108, as indicated by the arrow. In some embodiments, the reference picture of a picture (e.g., picture 104) may not be immediately before or after the picture. For example, the reference picture of picture 104 may be the picture before picture 102. It should be noted that the reference pictures of pictures 102-106 are merely examples, and the present disclosure is not limited to such pictures. Figure 1 Example of a reference image shown.
[0059] Typically, video codecs do not encode or decode an entire image at once due to the computational complexity of the encoding and decoding tasks. Instead, they can divide the image into basic segments and encode or decode the image segments one by one. In this disclosure, such a basic segment is referred to as a basic processing unit ("BPU"). For example, Figure 1Structure 110 in FIG. 1 shows an example structure of an image of video sequence 100 (e.g., any of images 102-108). In structure 110, the image is divided into 4×4 basic processing units, whose boundaries are shown as dashed lines. In some embodiments, the basic processing unit may be referred to as a "macroblock" in some video coding standards (e.g., the MPEG family, H.261, H.263, or H.264 / AVC), or as a "coding tree unit" ("CTU") in some other video coding standards (e.g., H.265 / HEVC or H.266 / VVC). The basic processing unit can have variable sizes within the image, such as 128×128, 64×64, 32×32, 16×16, 4×8, 16×32, or pixels of arbitrary shapes and sizes. The size and shape of the basic processing unit for the image can be selected based on a balance between coding efficiency and the level of detail to be maintained in the basic processing unit.
[0060] A basic processing unit may be a logical unit that may include a set of different types of video data stored in a computer memory (e.g., in a video frame buffer). For example, a basic processing unit of a color image may include a luma component (Y) representing achromatic luminance information, one or more chroma components (e.g., Cb and Cr) representing color information, and associated syntax elements, where the luma and chroma components may have the same size as the basic processing unit. In some video coding standards (e.g., H.265 / HEVC or H.266 / VVC), the luma and chroma components may be referred to as "coding tree blocks" ("CTBs"). Any operation performed on a basic processing unit may be repeated for each of its luma and chroma components.
[0061] Video encoding has multiple stages of operation, examples of which are Figures 2A-2B and Figures 3A-3BAs shown. For each stage, the size of the basic processing unit may still be too large for processing, so it can be further divided into segments referred to as "basic processing sub-units" in this disclosure. In some embodiments, the basic processing sub-unit may be referred to as a "block" in some video coding standards (e.g., MPEG family, H.261, H.263 or H.264 / AVC), or as a "coding unit" ("CU") in some other video coding standards (e.g., H.265 / HEVC or H.266 / VVC). The basic processing sub-unit may have the same size as the basic processing unit or a smaller size than the basic processing unit. Similar to the basic processing unit, the basic processing sub-unit is also a logical unit, which may include a set of different types of video data (e.g., Y, Cb, Cr and associated syntax elements) stored in a computer memory (e.g., in a video frame buffer). Any operation performed on the basic processing sub-unit can be repeated for each of its luminance and chrominance components. It should be noted that this division can be performed to a further level according to processing needs. It should also be noted that different stages can use different schemes to divide the basic processing units.
[0062] For example, in the mode decision phase (an example of which is Figure 2B As shown in FIG, the encoder can decide what prediction mode (e.g., intra prediction or inter prediction) to use for a basic processing unit, which may be too large to make such a decision. The encoder can divide the basic processing unit into multiple basic processing sub-units (e.g., CUs in H.265 / HEVC or H.266 / VVC) and decide the prediction type for each individual basic processing sub-unit.
[0063] For another example, in the prediction phase (the example is Figures 2A-2B ), the encoder can perform prediction operations at the level of basic processing sub-units (e.g., CUs). However, in some cases, the basic processing sub-units may still be too large to process. The encoder can further divide the basic processing sub-units into smaller segments (e.g., called "prediction blocks" or "PBs" in H.265 / HEVC or H.266 / VVC), at which prediction operations can be performed.
[0064] For another example, in the transformation phase (the example is Figures 2A-2B), the encoder can perform transform operations on the residual basic processing sub-unit (e.g., CU). However, in some cases, the basic processing sub-unit may still be too large to process. The encoder can further divide the basic processing sub-unit into smaller segments (e.g., called "transform blocks" or "TBs" in H.265 / HEVC or H.266 / VVC), at which level transform operations can be performed. It should be noted that the division scheme of the same basic processing sub-unit can be different in the prediction stage and the transform stage. For example, in H.265 / HEVC or H.266 / VVC, the prediction blocks and transform blocks of the same CU can have different sizes and numbers.
[0065] exist Figure 1 In the structure 110, the basic processing unit 112 is further divided into 3×3 basic processing sub-units, whose boundaries are shown by dotted lines. Different basic processing units of the same image can be divided into basic processing sub-units in different schemes.
[0066] In some embodiments, in order to provide parallel processing capabilities and error resilience for video encoding and decoding, the image can be divided into regions for processing so that the encoding or decoding process for a region of the image does not depend on information from any other region of the image. In other words, each region of the image can be processed separately. By doing so, the codec can process different regions of the image in parallel, thereby improving coding efficiency. In addition, when the data of a region is damaged during processing or lost during network transmission, the codec can correctly encode or decode other regions of the same image without relying on the damaged or lost data, thereby providing error resilience. In some video coding standards, the image can be divided into different types of regions. For example, H.265 / HEVC and H.266 / VVC provide two types of regions: "slices" and "tiles". It should also be noted that different images in the video sequence 100 can have different partitioning schemes for dividing the image into regions.
[0067] For example, in Figure 1 In FIG, the structure 110 is divided into three regions 114, 116 and 118, whose boundaries are shown as solid lines inside the structure 110. Region 114 includes four basic processing units. Regions 116 and 118 each include six basic processing units. It should be noted that Figure 1 The basic processing units, basic processing sub-units, and structural areas in 110 are merely examples, and the present disclosure does not limit the embodiments thereof.
[0068] Figure 2A Schematic diagram of an exemplary encoding process 200A according to an embodiment of the present disclosure is shown. For example, the encoding process 200A may be performed by an encoder. Figure 2AAs shown, the encoder may encode the video sequence 202 into a video bitstream 228 according to process 200A. Figure 1 The video sequence 100 in FIG. 2 may include a set of images (referred to as “original images”) arranged in time sequence. Figure 1 In the structure 110 in FIG. 1 , each original image of the video sequence 202 can be divided by the encoder into basic processing units, basic processing sub-units, or regions for processing. In some embodiments, the encoder can perform process 200A at the level of a basic processing unit for each original image of the video sequence 202. For example, the encoder can perform process 200A in an iterative manner, wherein the encoder can encode a basic processing unit in one iteration of process 200A. In some embodiments, the encoder can perform process 200A in parallel for each region (e.g., regions 114-118) of the original image of the video sequence 202.
[0069] refer to Figure 2A , the encoder may feed the basic processing units of the original images of the video sequence 202 (referred to as "original BPUs") to the prediction stage 204 to generate prediction data 206 and prediction BPUs 208. The encoder may subtract the predicted BPUs 208 from the original BPUs to generate residual BPUs 210. The encoder may feed the residual BPUs 210 to the transform stage 212 and the quantization stage 214 to generate quantized transform coefficients 216. The encoder may feed the prediction data 206 and the quantized transform coefficients 216 to the binary encoding stage 226 to generate a video bitstream 228. Components 202, 204, 206, 208, 210, 212, 214, 216, 226, and 228 may be referred to as the "forward path." During process 200A, after the quantization stage 214, the encoder may feed the quantized transform coefficients 216 to the inverse quantization stage 218 and the inverse transform stage 220 to generate a reconstructed residual BPU 222. The encoder can add the reconstructed residual BPU 222 to the predicted BPU 208 to generate a prediction reference 224, which is used in the prediction phase 204 of the next iteration of process 200A. Components 218, 220, 222, and 224 of process 200A can be referred to as a "reconstruction path." The reconstruction path can be used to ensure that both the encoder and decoder use the same reference data for prediction.
[0070] The encoder may iteratively perform process 200A to encode each original BPU of the original image (in the forward path) and generate a prediction reference 224 for encoding the next original BPU of the original image (in the reconstruction path). After encoding all the original BPUs of the original image, the encoder may proceed to encode the next image in the video sequence 202.
[0071] Referring to process 200A, an encoder may receive a video sequence 202 generated by a video capture device (e.g., a camera). As used herein, the term "receive" may refer to any action of receiving, inputting, acquiring, retrieving, obtaining, reading, accessing, or inputting data in any manner.
[0072] In the prediction phase 204, at the current iteration, the encoder may receive the original BPU and the prediction reference 224 and perform a prediction operation to generate the prediction data 206 and the predicted BPU 208. The prediction reference 224 may be generated from the reconstruction path of the previous iteration of the process 200A. The purpose of the prediction phase 204 is to reduce information redundancy by extracting the prediction data 206 that can be used to reconstruct the original BPU into the predicted BPU 208 from the prediction data 206 and the prediction reference 224.
[0073] Ideally, the predicted BPU 208 would be identical to the original BPU. However, due to non-ideal prediction and reconstruction operations, the predicted BPU 208 typically differs slightly from the original BPU. To account for these differences, after generating the predicted BPU 208, the encoder may subtract it from the original BPU to generate a residual BPU 210. For example, the encoder may subtract the value of the corresponding pixel of the predicted BPU 208 (e.g., grayscale value or RGB value) from the value of the corresponding pixel of the original BPU. Each pixel of the residual BPU 210 may have a residual value as a result of this subtraction between the corresponding pixel of the original BPU and the predicted BPU 208. Compared to the original BPU, the predicted data 206 and the residual BPU 210 may have fewer bits, but they can be used to reconstruct the original BPU without a significant loss in quality. Therefore, the original BPU is compressed.
[0074] To further compress the residual BPU 210, in the transform stage 212, the encoder can reduce its spatial redundancy by decomposing the residual BPU 210 into a set of two-dimensional "base patterns". Each base pattern is associated with a "transform coefficient". The base patterns can have the same size (e.g., the size of the residual BPU 210), and each base pattern can represent a frequency-varying (e.g., frequency of brightness variation) component of the residual BPU 210. None of the base patterns can be reproduced from any combination (e.g., linear combination) of any other base patterns. In other words, the decomposition can decompose the variation of the residual BPU 210 into the frequency domain. This decomposition is similar to the discrete Fourier transform of a function, where the base images are similar to the basis functions of the discrete Fourier transform (e.g., trigonometric functions), and the transform coefficients are similar to the coefficients associated with the basis functions.
[0075] Different transform algorithms can use different basic patterns. Various transform algorithms can be used in the transform stage 212, such as discrete cosine transform, discrete sine transform, etc. The transform at the transform stage 212 is reversible. That is, the encoder can restore the residual BPU 210 by performing the inverse operation of the transform (referred to as an "inverse transform"). For example, to restore the pixels of the residual BPU 210, the inverse transform may be multiplying the values of corresponding pixels of the basic pattern by the corresponding associated coefficients and adding the products to produce a weighted sum. For video coding standards, both the encoder and decoder can use the same transform algorithm (and therefore the same basic pattern). Therefore, the encoder can only record the transform coefficients, from which the decoder can reconstruct the residual BPU 210 without receiving the basic pattern from the encoder. Compared to the residual BPU 210, the transform coefficients may have fewer bits, but they can be used to reconstruct the residual BPU 210 without significant quality degradation. As a result, the residual BPU 210 is further compressed.
[0076] The encoder can further compress the transform coefficients in the quantization stage 214. During the transform process, different basis patterns can represent different frequencies of change (e.g., the frequency of brightness changes). Because the human eye is generally better at detecting low-frequency changes, the encoder can ignore information about high-frequency changes without causing noticeable quality degradation in decoding. For example, in the quantization stage 214, the encoder can generate quantized transform coefficients 216 by dividing each transform coefficient by an integer value (called a "quantization parameter") and rounding the quotient to the nearest integer. After this operation, some transform coefficients of the high-frequency basis patterns can be converted to zero, and transform coefficients of the low-frequency basis patterns can be converted to smaller integers. The encoder can ignore quantized transform coefficients 216 with zero values, thereby further compressing the transform coefficients. This quantization process is also reversible, where the quantized transform coefficients 216 can be reconstructed as transform coefficients in the inverse operation of quantization (called "inverse quantization").
[0077] Because the encoder ignores the remainder of this division during rounding operations, the quantization stage 214 can be lossy. Generally, the quantization stage 214 can contribute the most information loss in process 200A. The greater the information loss, the fewer bits are required to quantize the transform coefficients 216. To achieve different levels of information loss, the encoder can use different values for the quantization parameter or any other parameters of the quantization process.
[0078] In the binary encoding stage 226, the encoder may encode the prediction data 206 and the quantized transform coefficients 216 using a binary encoding technique, such as entropy coding, variable length coding, arithmetic coding, Huffman coding, context-adaptive binary arithmetic coding, or any other lossless or lossy compression algorithm. In some embodiments, in addition to the prediction data 206 and the quantized transform coefficients 216, the encoder may encode other information in the binary encoding stage 226, such as the prediction mode used in the prediction stage 204, the parameters of the prediction operation, the transform type at the transform stage 212, the parameters of the quantization process (e.g., quantization parameters), encoder control parameters (e.g., bit rate control parameters), etc. The encoder may use the output data of the binary encoding stage 226 to generate a video bitstream 228. In some embodiments, the video bitstream 228 may be further packaged for network transmission.
[0079] Referring to the reconstruction path of process 200A, in the inverse quantization stage 218, the encoder can perform inverse quantization on the quantized transform coefficients 216 to generate reconstructed transform coefficients. In the inverse transform stage 220, the encoder can generate a reconstructed residual BPU 222 based on the reconstructed transform coefficients. The encoder can add the reconstructed residual BPU 222 to the prediction BPU 208 to generate a prediction reference 224 to be used in the next iteration of process 200A.
[0080] It should be noted that other variations of process 200A may be used to encode video sequence 202. In some embodiments, the stages of process 200A may be performed by the encoder in a different order. In some embodiments, one or more stages of process 200A may be combined into a single stage. In some embodiments, a single stage of process 200A may be divided into multiple stages. For example, transform stage 212 and quantization stage 214 may be combined into a single stage. In some embodiments, process 200A may include additional stages. In some embodiments, process 200A may be omitted. Figure 2A one or more stages in a process.
[0081] Figure 2B A schematic diagram of another example encoding process 200B according to an embodiment of the present disclosure is shown. Process 200B can be modified from process 200A. For example, process 200B can be used by an encoder that complies with a hybrid video coding standard (e.g., the H.26x series). Compared to process 200A, the forward path of process 200B also includes a mode decision stage 230 and divides the prediction stage 204 into a spatial prediction stage 2042 and a temporal prediction stage 2044. The reconstruction path of process 200B also includes a loop filter stage 232 and a buffer 234.
[0082] In general, prediction techniques can be divided into two types: spatial prediction and temporal prediction. Spatial prediction (e.g., intra-frame image prediction or "intra-frame prediction") can use pixels from one or more already encoded adjacent BPUs in the same image to predict the current BPU. That is, the prediction reference 224 in spatial prediction can include adjacent BPUs. Spatial prediction can reduce the spatial redundancy inherent in the image. Temporal prediction (e.g., inter-image prediction or "inter-frame prediction") can use regions from one or more already encoded images to predict the current BPU. That is, the prediction reference 224 in temporal prediction can include encoded images. Temporal prediction can reduce the temporal redundancy inherent in the image.
[0083] Referring to process 200B, in the forward path, the encoder performs prediction operations in a spatial prediction stage 2042 and a temporal prediction stage 2044. For example, in the spatial prediction stage 2042, the encoder may perform intra-frame prediction. For an original BPU of the image being encoded, the prediction reference 224 may include one or more neighboring BPUs that have been encoded (in the forward path) and reconstructed (in the reconstruction path) in the same image. The encoder may generate a predicted BPU 208 by interpolating the neighboring BPUs. Interpolation techniques may include, for example, linear interpolation or interpolation, polynomial interpolation or interpolation, etc. In some embodiments, the encoder may perform interpolation at the pixel level, for example, by interpolating the values of corresponding pixels for each pixel of the predicted BPU 208. The neighboring BPUs used for interpolation may be located in various directions relative to the original BPU, such as vertically (e.g., on top of the original BPU), horizontally (e.g., to the left of the original BPU), diagonally (e.g., below left, below right, above left, or above right of the original BPU), or in any direction defined in the video coding standard being used. For intra prediction, the prediction data 206 may include, for example, the location (eg, coordinates) of the used neighboring BPUs, the size of the used neighboring BPUs, interpolation parameters, the direction of the used neighboring BPUs relative to the original BPU, and the like.
[0084] As another example, during the temporal prediction stage 2044, the encoder may perform inter-frame prediction. For the original BPU of the current image, the prediction reference 224 may include one or more images (referred to as "reference images") that have been encoded (in the forward path) and reconstructed (in the reconstruction path). In some embodiments, the reference images may be encoded and reconstructed BPU by BPU. For example, the encoder may add the reconstructed residual BPU 222 to the prediction BPU 208 to generate a reconstructed BPU. When all reconstructed BPUs for the same image have been generated, the encoder may generate a reconstructed image as a reference image. The encoder may perform a "motion estimation" operation to search for a matching region within a range of reference images (referred to as a "search window"). The position of the search window in the reference image may be determined based on the position of the original BPU in the current image. For example, the search window may be centered at a location in the reference image with the same coordinates as the original BPU in the current image and may extend outward by a predetermined distance. When the encoder identifies (e.g., using a PEL recursive algorithm, a block matching algorithm, etc.) an area similar to the original BPU in the search window, the encoder may determine such an area as a matching region. The matching region may have a different size than the original BPU (e.g., smaller, equal, larger, or a different shape). Because the reference image and the current image are temporally separated on the timeline (e.g., as Figure 1 ), so the matching area can be considered to "move" to the position of the original BPU over time. The encoder can record the direction and distance of this movement as a "motion vector". When using multiple reference images (e.g., Figure 1 The encoder may search for matching regions and determine the motion vector associated with each reference image. In some embodiments, the encoder may assign weights to the pixel values of the matching regions of the respective matching reference images.
[0085] Motion estimation can be used to identify various types of motion, such as translation, rotation, scaling, etc. For inter-frame prediction, the prediction data 206 may include, for example, the location (e.g., coordinates) of the matching region, the motion vector associated with the matching region, the number of reference images, weights associated with the reference images, etc.
[0086] To generate the predicted BPU 208, the encoder may perform a "motion compensation" operation. Motion compensation may be used to reconstruct the predicted BPU 208 based on the prediction data 206 (e.g., motion vectors) and the prediction reference 224. For example, the encoder may move a matching area of the reference image according to the motion vector, where the encoder may predict the original BPU of the current image. When multiple reference images are used (e.g., Figure 1In some embodiments, if the encoder has assigned weights to the pixel values of the matching regions of the reference image, the encoder may add the weighted sum of the pixel values of the moved matching regions.
[0087] In some embodiments, inter-frame prediction can be unidirectional or bidirectional. Unidirectional inter-frame prediction can use one or more reference images in the same temporal direction relative to the current image. For example, Figure 1 The picture 104 in is a unidirectional inter-frame predicted picture, where the reference picture (i.e., picture 102) precedes picture 04. Bidirectional inter-frame prediction can use one or more reference pictures in both temporal directions relative to the current picture. For example, Figure 1 Picture 106 in is a bidirectional inter-predicted picture, where the reference pictures (ie, pictures 104 and 08) are relative to picture 104 in both temporal directions.
[0088] Still referring to the forward path of process 200B, after the spatial prediction 2042 and temporal prediction stages 2044, in the mode decision stage 230, the encoder can select a prediction mode (e.g., one of intra-frame prediction and inter-frame prediction) for the current iteration of process 200B. For example, the encoder can perform a rate-distortion optimization technique, in which the encoder can select a prediction mode to minimize the value of a cost function based on the bit rate of the candidate prediction mode and the distortion of the reconstructed reference image under the candidate prediction mode. Based on the selected prediction mode, the encoder can generate a corresponding prediction BPU 208 and prediction data 206.
[0089] In the reconstruction path of process 200B, if intra-prediction mode has been selected in the forward path, after generating the prediction reference 224 (e.g., the current BPU that has been encoded and reconstructed in the current image), the encoder can feed the prediction reference 224 directly to the spatial prediction stage 2042 for later use (e.g., for interpolating the next BPU of the current image). If inter-prediction mode has been selected in the forward path, after generating the prediction reference 224 (e.g., the current image in which all BPUs have been encoded and reconstructed), the encoder can feed the prediction reference 224 to the loop filter stage 232. At this stage, the encoder can apply a loop filter to the prediction reference 224 to reduce or eliminate distortion (e.g., blocking artifacts) introduced by inter-prediction. The encoder can apply various loop filter techniques at the loop filter stage 232, such as deblocking, sample adaptive compensation, adaptive loop filtering, etc. The loop filtered reference pictures may be stored in a buffer 234 (or "decoded picture buffer") for later use (e.g., as inter-frame prediction reference pictures for future pictures of the video sequence 202). The encoder may store one or more reference pictures in the buffer 234 for use at the temporal prediction stage 2044. In some embodiments, the encoder may encode the parameters of the loop filter (e.g., loop filter strength) along with the quantized transform coefficients 216, the prediction data 206, and other information at the binary encoding stage 226.
[0090] Figure 3A A schematic diagram of an exemplary decoding process 300A according to an embodiment of the present disclosure is shown. Process 300A may correspond to Figure 2A In some embodiments, process 300A may be similar to the reconstruction path of process 200A. The decoder may decode the video bitstream 228 into a video stream 304 according to process 300A. Video stream 304 may be very similar to video sequence 202. However, due to information loss during compression and decompression (e.g., Figures 2A-2B quantization stage 214 in), typically, the video stream 304 is different from the video sequence 202. Figures 2A-2B 200A and 200B, the decoder may perform process 300A at a basic processing unit (BPU) level for each picture encoded in the video bitstream 228. For example, the decoder may perform process 300A in an iterative manner, where the decoder may decode a basic processing unit in one iteration of process 300A. In some embodiments, the decoder may perform process 300A in parallel for each region (e.g., regions 114-118) of each picture encoded in the video bitstream 228.
[0091] like Figure 3AAs shown, the decoder may feed a portion of the video bitstream 228 associated with a basic processing unit (referred to as a "coding BPU") for an encoded picture to a binary decoding stage 302, where the decoder may decode the portion into prediction data 206 and quantized transform coefficients 216. The decoder may feed the quantized transform coefficients 216 to an inverse quantization stage 218 and an inverse transform stage 220 to generate a reconstructed residual BPU 222. The decoder may feed the prediction data 206 to the prediction stage 204 to generate a prediction BPU 208. The decoder may add the reconstructed residual BPU 222 to the prediction BPU 208 to generate a prediction reference 224. In some embodiments, the prediction reference 224 may be stored in a buffer (e.g., a decoded picture buffer in computer memory). The decoder may feed the prediction reference 224 to the prediction stage 204 for performing a prediction operation in the next iteration of process 300A.
[0092] The decoder may iteratively perform process 300A to decode each coded BPU of a coded picture and generate a prediction reference 224 for the next coded BPU of the coded picture. After decoding all coded BPUs of a coded picture, the decoder may output the picture to a video stream 304 for display and continue decoding the next coded picture in the video bitstream 228.
[0093] In the binary decoding stage 302, the decoder can perform the inverse of the binary coding technique used by the encoder (e.g., entropy coding, variable length coding, arithmetic coding, Huffman coding, context-adaptive binary arithmetic coding, or any other lossless compression algorithm). In some embodiments, in addition to the prediction data 206 and the quantized transform coefficients 216, the decoder can decode other information in the binary decoding stage 302, such as the prediction mode, parameters of the prediction operation, the transform type, parameters of the quantization process (e.g., quantization parameters), encoder control parameters (e.g., bit rate control parameters), etc. In some embodiments, if the video bitstream 228 is transmitted in packet form over the network, the decoder can depacketize the video bitstream 228 before feeding it to the binary decoding stage 302.
[0094] Figure 3B A schematic diagram of another example decoding process 300B according to an embodiment of the present disclosure is shown. Process 300B can be modified from process 300A. For example, process 300B can be used by a decoder that complies with a hybrid video coding standard (e.g., the H.26x series). Compared to process 300A, process 300B additionally divides the prediction stage 204 into a spatial prediction stage 2042 and a temporal prediction stage 2044, and additionally includes a loop filter stage 232 and a buffer 234.
[0095] In process 300B, for a coded basic processing unit (referred to as a "current BPU") of a decoded coded image (referred to as a "current image"), the prediction data 206 decoded by the decoder from the binary decoding stage 302 may include various types of data, depending on what prediction mode the encoder used to encode the current BPU. For example, if the encoder used intra-frame prediction to encode the current BPU, the prediction data 206 may include a prediction mode indicator (e.g., a flag value) indicating intra-frame prediction, parameters of the intra-frame prediction operation, etc. The parameters of the intra-frame prediction operation may include, for example, the location (e.g., coordinates) of one or more neighboring BPUs used as a reference, the size of the neighboring BPUs, parameters of interpolation, the direction of the neighboring BPU relative to the original BPU, etc. For another example, if the encoder used inter-frame prediction to encode the current BPU, the prediction data 206 may include a prediction mode indicator (e.g., a flag value) indicating inter-frame prediction, parameters of the inter-frame prediction operation, etc. The parameters of the inter-frame prediction operation may include, for example, the number of reference images associated with the current BPU, the weights associated with the reference images respectively, the positions (e.g., coordinates) of one or more matching regions in the corresponding reference images, one or more motion vectors associated with the matching regions respectively, and the like.
[0096] Based on the prediction mode indicator, the decoder can decide whether to perform spatial prediction (e.g., intra-frame prediction) in the spatial prediction stage 2042 or temporal prediction (e.g., inter-frame prediction) in the temporal prediction stage 2044. The details of performing such spatial prediction or temporal prediction are described in detail in Figure 2B After performing such spatial prediction or temporal prediction, the decoder may generate a predicted BPU 208, which may be added to the predicted BPU 208 and the reconstructed residual BPU 222 to generate a prediction reference 224, as shown in FIG. Figure 3A As described in.
[0097] In process 300B, the decoder may feed the prediction reference 224 to the spatial prediction stage 2042 or the temporal prediction stage 2044 for performing a prediction operation in the next iteration of process 300B. For example, if the current BPU is decoded using intra-frame prediction in the spatial prediction stage 2042, then after generating the prediction reference 224 (e.g., the decoded current BPU), the decoder may feed the prediction reference 224 directly to the spatial prediction stage 2042 for later use (e.g., for interpolating the next BPU of the current picture). If the current BPU is decoded using inter-frame prediction in the temporal prediction stage 2044, then after generating the prediction reference 224 (e.g., the reference picture in which all BPUs are decoded), the encoder may feed the prediction reference 224 to the loop filter stage 232 to reduce or eliminate distortion (e.g., blocking artifacts). The decoder may Figure 2BThe loop filter is applied to the prediction reference 224 in the manner shown. The loop-filtered reference picture can be stored in a buffer 234 (e.g., a decoded picture buffer in computer memory) for later use (e.g., as an inter-prediction reference picture for a future coded picture in the video bitstream 228). The decoder can store one or more reference pictures in the buffer 234 for use at the temporal prediction stage 2044. In some embodiments, when the prediction mode indicator of the prediction data 206 indicates that inter-frame prediction was used to encode the current BPU, the prediction data can further include parameters of the loop filter (e.g., loop filter strength). The reconstructed picture from the buffer 234 can also be sent to a display, such as a TV, PC, smartphone, or tablet, for viewing by the end user.
[0098] There can be four types of loop filters. For example, the loop filters can include a deblocking filter, a sample adaptive offset ("SAO") filter, a luma mapping with chroma scaling ("LMCS") filter, and an adaptive loop filter ("ALF"). The order in which the four types of loop filters are applied can be the LMCS filter, the deblocking filter, the SAO filter, and the ALF. The LMCS filter can include two main components. The first component is an in-loop mapping of the luma component based on an adaptive piecewise linear model. The second component can be used for chroma components and can apply luma-dependent chroma residual scaling.
[0099] Figure 4 FIG is a block diagram of an example apparatus 400 for encoding or decoding a video according to an embodiment of the present disclosure. Figure 4 As shown, the device 400 may include a processor 402. When the processor 402 executes the instructions described herein, the device 400 may become a special-purpose machine for video encoding or decoding. The processor 402 may be any type of circuit capable of manipulating or processing information. For example, the processor 402 may include any number of central processing units (or "CPUs"), graphics processing units (or "GPUs"), neural processing units ("NPUs"), microcontroller units ("MCUs"), optical processors, programmable logic controllers, microprocessors, digital signal processors, intellectual property (IP) cores, programmable logic arrays (PLAs), programmable array logic (PALs), general array logic (GALs), complex programmable logic devices (CPLDs), a field programmable gate array (FPGA), a system on a chip (SoC), an application-specific integrated circuit (ASIC), and the like. In some embodiments, the processor 402 may also be a group of processors grouped into a single logical component. For example, as Figure 4 As shown, processor 402 may include multiple processors, including processor 402a, processor 402b, and processor 402n.
[0100] The apparatus 400 may further include a memory 404 configured to store data (eg, instruction sets, computer code, intermediate data, etc.). Figure 4 As shown, the stored data may include program instructions (e.g., for implementing stages in process 200A, 200B, 300A, or 300B) and data for processing (e.g., video sequence 202, video bitstream 228, or video stream 304). Processor 402 may access program instructions and data for processing (e.g., via bus 410) and execute the program instructions to perform operations or manipulations on the data for processing. Memory 404 may include a high-speed random access memory device or a non-volatile memory device. In some embodiments, memory 404 may include any number of random access memories (RAM), read-only memories (ROM), optical disks, magnetic disks, hard drives, solid-state drives, flash drives, secure digital (SD) cards, memory sticks, compact flash (CF) cards, etc. Memory 404 may also be a group of memories grouped into a single logical component ( Figure 4 not shown).
[0101] The bus 410 may be a communication device that transmits data between components within the apparatus 400 , such as an internal bus (eg, a CPU-memory bus), an external bus (eg, a Universal Serial Bus port, a Peripheral Component Interconnect Express port), or the like.
[0102] For ease of explanation and to avoid ambiguity, in this disclosure, processor 402 and other data processing circuitry are collectively referred to as "data processing circuitry." The data processing circuitry may be implemented entirely in hardware, or as a combination of software, hardware, or firmware. Furthermore, the data processing circuitry may be a single, separate module, or may be fully or partially integrated into any other component of device 400.
[0103] The device 400 may also include a network interface 406 to provide wired or wireless communication with a network (e.g., the Internet, an intranet, a local area network, a mobile communication network, etc.). In some embodiments, the network interface 406 may include any number of network interface controllers (NICs), radio frequency (RF) modules, transponders, transceivers, modems, routers, gateways, wired network adapters, wireless network adapters, Bluetooth adapters, infrared adapters, near field communication ("NFC") adapters, cellular network chips, etc.
[0104] In some embodiments, the apparatus 400 may optionally further include a peripheral interface 408 to provide a connection to one or more peripheral devices. Figure 4As shown, the peripheral device may include, but is not limited to, a cursor control device (e.g., a mouse, a touchpad, or a touch screen), a keyboard, a display (e.g., a cathode ray tube display, a liquid crystal display, or a light emitting diode display), a video input device (e.g., a camera or an input interface coupled to a video archive), etc.
[0105] It should be noted that a video codec (e.g., the codec that executes processes 200A, 200B, 300A, or 300B) may be implemented as any combination of any software or hardware modules in device 400. For example, some or all stages of processes 200A, 200B, 300A, or 300B may be implemented as one or more software modules of device 400, such as program instances that may be loaded into memory 404. For another example, some or all stages of processes 200A, 200B, 300A, or 300B may be implemented as one or more hardware modules of device 400, such as dedicated data processing circuits (e.g., FPGA, ASIC, NPU, etc.).
[0106] In the quantization and inverse quantization functional blocks (e.g., Figure 2A or Figure 2B quantization 214 and inverse quantization 218 of Figure 3A or Figure 3B inverse quantization 218 of , the quantization parameter (QP) is used to determine the amount of quantization (and inverse quantization) applied to the prediction residual. The initial QP value used to encode an image or a slice may be signaled at a higher level, e.g., using the init_qp_minus26 syntax element in the picture parameter set (PPS) and using the slice_qp_delta syntax element in the slice header. In addition, the QP value may be adapted at the local level of each CU using incremental QP values sent in the granularity of quantization groups.
[0107] Rectangular projection (“ERP”) formats such as
[0108] are common projection formats for representing 360-degree videos and images. The projection maps the meridians to vertical lines with a constant spacing and the latitude circles to horizontal lines with a constant spacing. Since the relationship between the position of an image pixel on the map and its corresponding geographical location on the sphere is particularly simple, ERP is one of the most common projections for 360-degree videos and images. u = (m + 0.5) / W, 0 ≤ m < W Equation (1) v = (n + 0.5) / H, 0 ≤ n < H, Equation (2)
[0109] Then, the longitude and latitude (φ, θ) in the sphere can be calculated from (u, v) based on the following equations (3) and (4). φ = (u - 0.5) × (2 × π), Equation (3) θ = (0.5 - v) × π, Equation (4)
[0110] The 3D coordinates (X, Y, Z) can be calculated based on the following equations (5)-(7). X = cos(θ)cos(φ), Equation (5) Y = sin(θ), Equation (6) Z = -cos(θ)sin(φ), Equation (7)
[0111] For the 3D to 2D coordinate conversion starting from (X, Y, Z), (φ, θ) can be calculated based on the following equations (8) and (9). Then, (u, v) are calculated based on equations (3) and (4). Finally, the 2D coordinates (m, n) can be calculated according to equations (1) and (2). φ = tan-1(-Z / X), Equation (8) θ = sin-1(Y / (X2 + Y2 + Z2)1 / 2), Equation (9)
[0112] To reduce the seam artifacts in the reconstructed viewport that includes the left and right boundaries of the ERP image, a new format called padded equirectangular projection ("PERP") is provided by padding samples on each of the left and right sides of the ERP image.
[0113] When representing a 360-degree video using PERP, the PERP image is encoded. After decoding, the reconstructed PERP is converted back to the reconstructed ERP by blending the replicated samples or cropping the padded regions.
[0114] Figure 5A A schematic diagram showing an example blending operation for generating a reconstructed equirectangular projection according to some embodiments of the present disclosure is shown. Unless otherwise specified, "recPERP" is used to represent the reconstructed PERP before post-processing, and "recERP" is used to represent the reconstructed ERP after post-processing. In Figure 5A where A1 and B2 are the boundary regions in the ERP image, and B1 and A2 are the padded regions, where A2 is padded from A1 and B1 is padded from B2. As Figure 5AAs shown, the replicated samples of recPERP can be mixed by applying a distance-based weighted average operation. For example, region A can be generated by mixing region A1 with A2, while region B can be generated by mixing region B1 with b2.
[0115] In the following description, the width and height of the unfilled recERP are denoted as "W" and "H" respectively. The left and right padded widths are denoted as "P" respectively. L ” and “P R The total fill width is expressed as "P w ”, which can be P L and P R In some embodiments, recPERP can be converted to recERP by a blending operation. For example, for a sample in A, recERP(j, i), where (j, i) is the coordinate in the ERP image, i is in [0, P R -1], j is in [0, H-1], recERP(j, i) can be determined according to the following formula. A=w×A1+(1–w)×A2, where w is PL / Pw to 1 Formula (10) recERP(j,i) in A=(recPERP(j,i+PL)×(i+PL)+recPERP(j,i+PL+W)×(PR-i)+(PW>>1)) / PW Formula (11) where recPERP(y, x) is the sample on the reconstructed PERP image, where (y, x) are the coordinates of the sample in the PERP image.
[0116] In some embodiments, for a sample in B, recERP(j, i), where (j, i) is the coordinate in the ERP image, and i is in [WP L , W-1] and j is in [0, H-1], recERP(j, i) can be generated according to the following equation: B=k×B1+(1–k)×B2, where k ranges from 0 to PL / Pw Formula (12) recERP(j,i)=(recPERP(j,i+PL)×(PR-i+W)+recPERP(j,i+PL-w)×(i–W+PL)+(Pw>>1)) / PW in B Formula (13) where recPERP(y, x) is the sample on the reconstructed PERP image, where (y, x) are the coordinates of the sample in the PERP image.
[0117] Figure 5BA schematic diagram of an example cropping operation for generating a reconstructed equirectangular projection according to some embodiments of the present disclosure is shown. Figure 5B In , A1 and B2 are the boundary areas within the ERP image, and B1 and A2 are the filling areas, where A2 is filled from A1 and B1 is filled from B2. Figure 5B As shown in Figure 2, during the clipping process, the filler samples in recPERP can be directly discarded to obtain recERP. For example, the filler samples B1 and A2 can be discarded.
[0118] In some embodiments, horizontal surround motion compensation can be used to improve the coding performance of ERP. For example, horizontal surround motion compensation can be used as a 360-degree-specific coding tool in the VVC standard, designed to improve the visual quality of reconstructed 360-degree video in ERP or PERP formats. In conventional motion compensation, when a motion vector references samples beyond the image boundaries of a reference image, padding is applied by copying the nearest neighboring samples from the corresponding image boundaries to derive the value of the out-of-bounds sample. For 360-degree video, this padding approach is inappropriate and can result in visual artifacts known as "seams" in the reconstructed viewport video. Because 360-degree video is captured on a sphere and inherently has no "boundaries," reference samples outside the boundaries of the reference image in the projected domain can be obtained from neighboring samples in the spherical domain. For general projected formats, deriving corresponding neighboring samples in the spherical domain can be difficult because it involves 2D-to-3D and 3D-to-2D coordinate conversions, as well as sample interpolation at fractional sample positions. For the left and right boundaries of the ERP or PERP projection format, this problem can be solved because the spherical neighbor samples outside the left image boundary can be obtained from the samples within the right image boundary, and vice versa. Given the widespread use of the ERP or PERP projection format and its relative ease of implementation, VVC uses horizontal surround motion compensation to improve the visual quality of 360-degree videos encoded in the ERP or PERP projection format.
[0119] Figure 6A Schematic diagram showing an example horizontal surround motion compensation process for equirectangular projection according to some embodiments of the present disclosure. Figure 6A As shown, when a portion of the reference block is outside the left (or right) boundary of the reference image in the projection domain, instead of repeated padding, the "out-of-bounds" portion can be obtained from the corresponding spherical neighbor portion located within the reference image towards the right (or left) boundary in the projection domain. In some embodiments, repeated padding can be used for the top and bottom image boundaries.
[0120] Figure 6B Schematic diagram showing an example horizontal surround motion compensation process for filled equirectangular projection according to some embodiments of the present disclosure. Figure 6BAs shown, horizontal surround motion compensation can be combined with the non-standard padding method commonly used in 360-degree video encoding. In some embodiments, this is achieved by signaling a high-level syntax element to indicate a surround motion compensation offset, which can be set to the ERP image width before padding. This syntax can be used to adjust the position of the horizontal surround accordingly. In some embodiments, the syntax is not affected by the specific amount of padding on the left or right image boundary. As a result, this syntax can naturally support asymmetric padding of ERP images. In asymmetric padding of ERP images, the padding on the left and right can be different. In some embodiments, the surround motion compensation can be determined according to the following formula: where offset may be the surround motion compensation offset signaled in the bitstream, picW may be the picture width including padding before encoding, and pos x It can be the reference position determined by the current block position and motion vector, and the formula pos x The output of _wrap can be the actual reference position of the reference block from the wrap motion compensation. To save the signaling overhead of the wrap motion compensation offset, it can be in units of the minimum luma coding block, so the offset can be replaced by offset w ×MinCbSizeY, where offset w is the surround motion compensation offset in units of the smallest luma coding block that is signaled in the bitstream, and MinCbSizeY is the size of the smallest luma coding block. In contrast, in conventional motion compensation, the actual reference position from which the reference block comes can be obtained by clipping pos in the range 0 to picW-1. x Export directly.
[0121] Horizontal surround motion compensation provides more meaningful information for motion compensation when the reference samples are outside the left and right boundaries of the reference image. Under common test conditions for 360-degree video, this tool can improve compression performance not only in terms of rate-distortion, but also in terms of reducing seam artifacts and improving the subjective quality of the reconstructed 360-degree video. Horizontal surround motion compensation can also be used with other single-sided projection formats with constant sampling density in the horizontal direction, such as the adjusted equal-area projection.
[0122] In VVC, (e.g., VVC draft 8), wraparound motion compensation can be implemented by signaling the high-level syntax variable pps_ref_wraparound_offset to indicate the wraparound offset. Sometimes, the wraparound offset should be set to the ERP image width before padding. This syntax can be used to adjust the position of horizontal wraparound motion compensation accordingly. This syntax may not be affected by the specific amount of padding on the left or right image boundary. As a result, the syntax can naturally support asymmetric padding of ERP images (e.g., when the left and right padding are different). When the reference samples are outside the left and right boundaries of the reference image, horizontal wraparound motion compensation can provide more meaningful information for motion compensation.
[0123] Figure 7 The syntax of an example high-level wraparound offset according to some embodiments of the present disclosure is shown. It should be understood that Figure 7 The syntax shown can be used in VVC (e.g., VVC draft 8). Figure 7 As shown, when wraparound motion compensation is enabled (eg, pps_ref_wraparound_enabled_flag == 1), the wraparound offset pps_ref_wraparound_offset may be directly signaled. Figure 7 As shown, syntax elements or variables in the bitstream are shown in bold.
[0124] Figure 8 The semantics of example advanced wraparound offsets according to some embodiments of the present disclosure are shown. It should be understood that Figure 8 The semantics shown in can correspond to Figure 7 In some embodiments, such as Figure 8 As shown, the value of pps_ref_wraparound_enabled_flag is 1, which means that horizontal wrap motion compensation is applied in inter prediction. If the value of pps_ref_wraparound_enabled_flag is 0, horizontal wrap motion compensation is not applied. When the value of CtbSizeY / MinCbSizeY+1 is greater than pic_width_in_luma_samples / MinCbSizeY-1, The value of pps_ref_wraparound_enabled_flag shall be equal to 0. When sps_ref_wraparound_enabled_flag is equal to 0, the value of pps_ref_wraparound_enabled_flag shall be equal to 0. CtbSizeY is the size of the luma coding tree block.
[0125] In some embodiments, as Figure 8 As shown, the value of pps_ref_wraparound_offset plus (CtbSizeY / MinCbSizeY)+2 can specify the offset used to calculate the horizontal wrap position in units of MinCbSizeY luma samples. The value of pps_ref_wraparound_offset can be in the range of 0 to (pic_width_in_luma_samples / MinCbSizeY)-(CtbSizeY / MinCbSizeY)-2, inclusive. The variable PpsRefWraparoundOffset can be set equal to pps_ref_wraparound_offset+(CtbSizeY / MinCbSizeY)+2. In some embodiments, the variable PpsRefWraparoundOffset can be used to determine the luma position of the sub-block (e.g., in VVC draft 8).
[0126] In VVC (e.g., VVC draft 8), a picture can be divided into one or more tile rows or one or more tile columns. A tile can be a sequence of coding tree units ("CTUs") covering a rectangular area of the picture. A slice can include an integer number of complete tile slices or an integer number of consecutive complete CTU rows within a tile of the picture.
[0127] Two modes of striping are supported: raster scan striping and rectangular striping. In raster scan striping, a strip can include a sequence of complete tiles in a raster scan of a tile of an image. In rectangular striping, a strip can include multiple complete tiles that together form a rectangular region of the image, or multiple consecutive complete CTU rows that together form a tile of a rectangular region of the image. Tiles within a rectangular strip are scanned in tile raster scan order within the rectangular region corresponding to the strip.
[0128] A sub-image may comprise one or more strips that together cover a rectangular area of the image. Figure 9 Schematic diagram showing example strips and sub-image partitioning of an image according to some embodiments of the present disclosure. Figure 9 As shown, the image is divided into 20 tiles, with 5 tile columns and 4 tile rows. There are 12 tiles on the left, each tile covering a 4×4 CTU strip. There are 8 tiles on the right, each tile covering 2 vertically stacked 2×2 CTU strips. There are a total of 28 slices and 28 sub-images of different sizes (for example, each slice can be a sub-image).
[0129] Figure 10 Schematic diagrams showing example stripe and sub-image partitioning with different stripes and sub-images according to some embodiments of the present disclosure are shown. Figure 10 As shown, the image is divided into 20 tiles, with 5 tile columns and 4 tile rows. There are 12 tiles on the left, each covering a 4×4 CTU strip. There are 8 tiles on the right, each covering two vertically stacked 2×2 CTU strips. There are 28 strips in total. For the 12 strips on the left, each strip is a sub-image. For the 16 strips on the right, each of 4 strips constitutes a sub-image. As a result, there are 16 sub-images of the same size.
[0130] In VVC (e.g., VVC draft 8), information about slice layout can be signaled in a picture parameter set ("PPS"). In some embodiments, a picture parameter set is a syntax structure that includes syntax elements or variables that apply to zero or more entire coded pictures determined by syntax elements found in each picture header. Figure 11 The syntax of an example image parameter set for tile mapping and stripe layout according to some embodiments of the present disclosure is shown. It should be understood that Figure 11 The syntax shown can be used in VVC (e.g., VVC draft 8). Figure 11 As shown in the figure, syntax elements or variables in the bitstream are shown in bold. Figure 11 As shown, if the number of slices in the current picture is greater than 1 and rectangular slice mode is used (e.g., rect_slice_flag == 1), a flag called single_slice_per_subpic_flag may first be signaled to indicate that each sub-picture includes only one slice. In this case (e.g., single_slice_per_subpic_flag == 1), there is no need to further signal the layout information of the slices, as it may be the same as the sub-picture layout already signaled in the sequence parameter set ("SPS"). In some embodiments, an SPS is a syntax structure that includes syntax elements that apply to zero or more entire coded layer video sequences ("CLVS"), as determined by the contents of syntax elements found in a picture parameter set referenced by syntax elements found in each picture header. In some embodiments, a picture header is a syntax structure that includes syntax elements that apply to all slices of a coded picture. In some embodiments, as Figure 11 As shown, if the value of single_slice_per_subpic_flag is 0, the number of slices in the picture (eg, num_slices_in_pic_minus1) may be signaled first, followed by the slice position and size information of each slice.
[0131] In some embodiments, to signal the number of slices, the number of slices minus 1 (e.g., num_slices_in_pic_minus1) may be signaled instead of directly signaling the number of slices, since there is at least 1 slice in the picture. In general, signaling a smaller positive value may consume fewer bits and improve the overall efficiency of performing video processing.
[0132] Figure 12A and Figure 12B The semantics of an example image parameter set for tile mapping and stripe layout according to some embodiments of the present disclosure are shown. It should be understood that Figure 12A and Figure 12B The semantics shown in 2 can correspond to Figure 11 In some embodiments, such as Figure 12A and Figure 12B The semantics shown in corresponds to VVC (eg, VVC draft 8).
[0133] In some embodiments, as Figure 12A and Figure 12B As shown, the variable rect_slice_flag equal to 0 indicates that the tiles within each slice are in raster scan order and that slice information is not signaled in the PPS. When the variable rect_slice_flag is equal to 1, the tiles within each slice may cover a rectangular area of the image and slice information may be signaled in the PPS. In some embodiments, when not present, the variable rect_slice_flag may be inferred to be equal to 1. In some embodiments, when the variable subpic_info_present_flag is equal to 1, the value of rect_slice_flag shall be equal to 1.
[0134] In some embodiments, as Figure 12A and Figure 12B As shown, the variable single_slice_per_subpic_flag equal to 1 means that each sub-image can include one and only one rectangular slice. When the variable single_slice_per_subpic_flag is equal to 0, each sub-image can include one or more rectangular slices. In some embodiments, when the variable single_slice_per_subpic_flag is equal to 1, the variable num_slices_in_pic_minus1 can be inferred to be equal to the variable sps_num_subpics_minus1. In some embodiments, when the variable single_slice_per_subpic_flag is not present, the value of single_slice_per_subpic_flag can be inferred to be equal to 0.
[0135] In some embodiments, as Figure 12A and Figure 12B As shown, the variable num_slices_in_pic_minus1 plus 1 is the number of rectangular slices in each picture that reference the PPS. In some embodiments, the value of num_slices_in_pic_minus1 can be in the range of 0 to MaxSlicesPerPicture-1, inclusive. In some embodiments, when the variable no_pic_partition_flag is equal to 1, the value of num_slices_in_pic_minus1 can be inferred to be equal to 0.
[0136] In some embodiments, as Figure 12A and Figure 12B As shown, if the variable tile_idx_delta_present_flag is equal to 0, then the value of tile_idx_delta is not present in the PPS, and all rectangular strips in the picture referenced by the PPS are specified in raster order. In some embodiments, when the variable tile_idx_delta_present_flag is equal to 1, the value of tile_idx_delta may be present in the PPS, and all rectangular strips in the picture referenced by the PPS are specified in the order indicated by the value of tile_idx_delta. In some embodiments, when not present, the value of tile_idx_delta_present_flag may be inferred to be equal to 0.
[0137] In some embodiments, as Figure 12A and Figure 12B As shown, the variable slice_width_in_tiles_minus1[i] plus 1 specifies the width of the i-th rectangular strip in units of tiles. In some embodiments, the value of slice_width_in_tiles_minus1[i] should be in the range of 0 to NumTileColumns-1, inclusive. In some embodiments, if slice_width_in_tiles_minus1[i] does not exist, the following may apply: if NumTileColumns is equal to 1, then the value of slice_width_in_tiles_minus1[i] may be inferred to be equal to 0; otherwise, the value of slice_width_in_tiles_minus_1[i] may be inferred to be the value specified in the clause in the VVC draft (e.g., VVC draft 8).
[0138] In some embodiments, as Figure 12A and Figure 12BAs shown, the variable slice_height_in_tiles_minus1 plus 1 specifies the height of the i-th rectangular strip in tile rows. In some embodiments, the value of slice_height_in_tiles_minus1[i] should be in the range of 0 to NumTileRows-1, inclusive. In some embodiments, when the variable slice_height_in_tiles_minus1[i] does not exist, the following applies: if NumTileRows is equal to 1, or the variable tile_idx_delta_present_flag is equal to 0 and tileIdx%NumTileColumns is greater than 0, then the value of slice_height_in_tiles_minus1[i] is inferred to be equal to 0; otherwise (for example, NumTileRows is not equal to 1, and tile_idx_delta_present_flag is equal to 1 or tileIdx%NumTileColumns is equal to 0), when tile_idx_delta_present_flag is equal to 1 or tileIdx%NumTileColumns is equal to 0, the value of slice_height_in_tiles_minus1[i] can be inferred to be equal to slice_height_in_tiles_minus1[i-1].
[0139] In some embodiments, as Figure 12A and Figure 12B As shown, the value of num_exp_slices_in_tile[i] specifies the number of explicitly provided strip heights in the current tile that includes multiple rectangular strips. In some embodiments, the value of num_exp_slices_in_tile[i] should be in the range of 0 to RowHeight[tileY]-1, inclusive, where tileY is the tile row index that includes the i-th strip. In some embodiments, when not present, the value of num_exp_slices_in_tile[i] can be inferred to be equal to 0. In some embodiments, when num_exp_slices_in_tile[i] is equal to 0, the value of the variable NumSliceInTile[i] is inferred to be equal to 1.
[0140] In some embodiments, as Figure 12A and Figure 12BAs shown, the value of exp_slice_height_in_ctus_minus1[j] plus 1 specifies the height of the j-th rectangular slice in the current tile in units of CTU rows. In some embodiments, the value of exp_slice_height_in_ctus_minus1[j] should be in the range of 0 to RowHeight[tileY]-1, inclusive, where tileY is the tile row index of the current tile.
[0141] In some embodiments, as Figure 12A and Figure 12B As shown, when the variable num_exp_slices_in_tile[i] is greater than 0, the variables NumSlicesInTile[i] and SliceHeightInCtusMinus1[i+k] with k in the range of 0 to NumSlicesInTile[i]-1 can be obtained. In some embodiments, as Figure 8 As shown, when the value of num_exp_slices_in_tile[i] is greater than 0, the values of NumSlicesInTile[i] and SliceHeightInCtusMinus1[i+k] can be obtained.
[0142] In some embodiments, as Figure 12A and Figure 12B As shown, the value of tile_idx_delta[i] specifies the difference between the tile index of the first tile in the i-th rectangular strip and the tile index of the first tile in the (i+1)-th rectangular strip. The value of tile_idx_delta[i] should be in the range of -NumTilesInPic+1 to NumTilesInPic-1, inclusive. In some embodiments, when not present, the value of tile_idx_delta[i] can be inferred to be equal to 0. In some embodiments, when present, the value of tile_idx_delta[i] can be inferred to be non-zero.
[0143] In VVC (eg, VVC draft 8), to locate each slice in a picture, one or more slice addresses may be signaled in a slice header. Figure 13 FIGURE 1 shows the syntax of an example slice header according to some embodiments of the present disclosure. It should be understood that Figure 13 The syntax 3 shown can be used in VVC (e.g., VVC draft 8). Figure 11 As shown in the figure, syntax elements or variables in the bitstream are shown in bold. Figure 13As shown, if the variable picture_header_in_slice_header_flag is equal to 1, the picture header syntax structure exists in the slice header.
[0144] Figure 14 The semantics of an example slice header according to some embodiments of the present disclosure are shown. It should be understood that Figure 14 The semantics shown in can correspond to Figure 13 In some embodiments, Figure 14 The semantics shown in corresponds to VVC (eg, VVC draft 8).
[0145] In some embodiments, as Figure 14 As shown, the bitstream conformance requirement is that the value of picture_header_in_slice_header_flag shall be the same in all coded slices in CLVS.
[0146] In some embodiments, as Figure 14 As shown, when the variable picture_header_in_slice_header_flag is equal to 1 for a coded slice, a bitstream conformance requirement is that no video coding layer ("VCL") network abstraction layer ("NAL") units with nal_unit_type equal to PH_NUT shall be present in the CLVS.
[0147] In some embodiments, as Figure 14 As shown, when picture_header_in_slice_header_flag is equal to 0, all coded slices in the current picture shall have picture_header_in_slice_header_flag equal to 0, and the current PU shall have a PH NAL unit.
[0148] In some embodiments, as Figure 14 As shown, the variable slice_address may specify the slice address of the slice. In some embodiments, when not present, the value of slice_address may be inferred to be equal to 0. When the variable rect_slice_flag is equal to 1 and NumSliceInSubpic[CurrSubpicIdx] is equal to 1, the value of slice_address is inferred to be equal to 0.
[0149] In some embodiments, as Figure 14As shown, if the variable rect_slice_flag is equal to 0, the following may apply: the strip address may be a raster scan tile index; the length of slice_address may be ceil(Log2(NumTilesInPic)) bits; and the value of slice_address shall be in the range of 0 to NumTilesInPic-1, inclusive.
[0150] In some embodiments, as Figure 14 As shown, bitstream conformance requires that the following constraints apply: If the variable rect_slice_flag is equal to 0 or the variable subpic_info_present_flag is equal to 0, then the value of slice_address shall not be equal to the value of slice_address of any other coded slice NAL unit of the same coded picture; otherwise, the slice_subpic_id and slice_address value pair shall not be equal to the slice_subpic_id and slice_address value pair of any other coded slice NAL unit of the same coded picture.
[0151] In some embodiments, as Figure 14 As shown, the shape of the slices of the picture should be such that each CTU, when decoded, should have its entire left border and entire top border including the picture border or including the border of the previously decoded CTU.
[0152] In some embodiments, as Figure 14 As shown, the variable slice_address can be signaled only when one of the following two conditions is met: rectangular strip mode is used and the number of strips in the current sub-image is greater than one; or rectangular strip mode is not used and the number of tiles in the current image is greater than one. In some embodiments, if neither of the above conditions is met, there is only one strip in the current sub-image or current image. In this case, since the entire sub-image or the entire image is a single strip, there is no need to signal the strip address.
[0153] There are many problems with the current design of VVC. First, the wraparound offset pps_ref_wraparound_offset is signaled in the bitstream and should be set to the ERP image width before padding. To save bits, the minimum value of the wraparound offset (e.g., CtbSizeY / MinCbSizeY)+2) is subtracted from the wraparound offset before signaling. However, the width of the padded area is much smaller than the width of the original ERP image. This is especially true for coded ERP images where the padded area width can be 0. Assuming that the total width of the image to be encoded or decoded is known, signaling the width of the original ERP part may cost more bits than signaling the width of the padded area. As a result, the current signaling of the wraparound offset of the original ERP width in units of MinCbSizeY is not very efficient.
[0154] Furthermore, the signaling of the slice layout can be improved. For example, 1 can be subtracted from the number of slices before signaling the number of slices, since the number of slices in a picture is always greater than or equal to 1, and signaling a smaller positive value requires fewer bits. However, in current VVC (e.g., VVC draft 8), a sub-picture includes an integer number of complete slices. As a result, the number of slices in a picture is greater than or equal to the number of sub-pictures in the picture. In current VVC (e.g., VVC draft 8), num_slices_in_pic_minus1 is signaled only when the variable single_slice_per_subpic_flag is 0, and the variable single_slice_per_subpic_flag being equal to 0 indicates that at least one sub-picture contains multiple slices. Therefore, in this case, the number of slices must be greater than the number of sub-pictures, which has a minimum value of 1. As a result, the minimum value of num_slices_in_pic_minus1 is greater than zero. Signaling non-negative values whose range does not start at zero is inefficient.
[0155] Furthermore, there are other issues with slice address signaling. When there is only one slice in the current image, there may be no need to signal the slice address, as the entire sub-image or the entire image is a single slice. The slice address can be inferred to be 0. However, the two conditions for skipping slice address signaling are not complete. For example, in raster scan strip mode, even if the number of tiles is greater than one, there may be only one strip that includes all the tiles in the image, in which case slice address signaling can be avoided.
[0156] Embodiments of the present disclosure provide methods for addressing the aforementioned issues. In some embodiments, since the width of the padded region is typically smaller than the width of the original ERP image, the width of the original ERP image can be the same as the wraparound offset used in wraparound motion compensation. It is therefore advisable to signal the difference between the coded image width and the original ERP image width in the bitstream. For example, it is advisable to signal the difference between the coded image width and the wraparound motion compensation offset in the bitstream, and then perform a derivation on the decoder side to obtain the wraparound offset after parsing the signaled difference. Since the difference between the coded image width and the wraparound offset is typically smaller than the wraparound offset itself, this approach can save signaled bits.
[0157] Figure 15 FIGURE 1 shows an example improved syntax of an image parameter set according to some embodiments of the present disclosure. Figure 15 As shown, syntax elements or variables in the bitstream are shown in bold, and changes to previous VVCs (e.g., Figure 7 ), where the suggested deletion syntax is further shown in strikethrough. Figure 15 As shown, a new variable pps_pic_width_minus_wraparound_offset can be created. The value of the variable pps_pic_width_minus_wraparound_offset can be signaled according to the value of pps_ref_wraparound_enabled_flag. In some embodiments, the new variable pps_pic_width_minus_wraparound_offset can replace the variable pps_ref_wraparound_offset in the original VVC (e.g., VVC draft 8).
[0158] Figure 16 The semantics of an example improved image parameter set according to some embodiments of the present disclosure are shown. Figure 16 As shown, changes to the previous VVC (e.g., Figure 8 The semantics shown in () are shown in italic type, and the suggested deletion syntax is further shown in strikethrough. It should be understood that Figure 16 The semantics shown in can correspond to Figure 15 In some embodiments, Figure 16 The semantics shown correspond to VVC (eg, VVC draft 8).
[0159] like Figure 16As shown, the semantics of the new variable pps_pic_width_minus_wraparound_offset is different from that of the variable pps_ref_wraparound_offset (e.g., Figure 8 The variable pps_pic_width_minus_wraparound_offset may specify the difference between the picture width and the offset used to calculate the horizontal wraparound position in units of MinCbSizeY luma samples. In some embodiments, as Figure 16 As shown, the value of pps_pic_width_minus_wraparound_offset should be less than or equal to (pic_width_in_luma_samples / MinCbSizeY)-(CtbSizeY / MinCbSizeY)-2. In some embodiments, the variable PpsRefWraparoundOffset may be set equal to pic_width_in_luma_samples / MinCbSizeY-pps_pic_width_minus_wraparound_offset.
[0160] In some embodiments, a flag wraparound_offset_type can be signaled to indicate whether the wraparound offset signaled is the original ERP picture width or the difference between the coded picture width and the original ERP picture width. The encoder can choose the smaller of these two values and signal it in the bitstream, which can further reduce signaling overhead. Figure 17 FIGURE 1 shows the syntax of an example image parameter set with a variable wraparound_offset_type according to some embodiments of the present disclosure. Figure 17 As shown, syntax elements or variables in the bitstream are shown in bold and are changes to previous VVCs (e.g., Figure 7 The syntax shown in ) is shown in italic type. Figure 17 As shown, a new variable pps_ref_wraparound_offset can be added to specify the value used to determine PpsRefWraparoundOffset.
[0161] Figure 18 The semantics of an example improved image parameter set with the variable wraparound_offset_type according to some embodiments of the present disclosure are shown. Figure 18 As shown, changes to the previous VVC (e.g., Figure 8The semantics shown in () are shown in italic type, and the suggested deletion syntax is further shown in strikethrough. It should be understood that Figure 18 The semantics shown can correspond to Figure 17 In some embodiments, Figure 18 The semantics shown correspond to VVC (eg, VVC draft 8).
[0162] In some embodiments, as Figure 18 As shown in the figure, a new variable wraparound_offset_type can be added to specify the type of the variable pps_ref_wraparound_offset. The value of pps_ref_wraparound_offset should be in the range of 0 to ((pps_pic_width_in_luma_samples / MinCbSizeY)–(CtbSizeY / MinCbSizeY)–2) / 2.
[0163] In some embodiments, as Figure 18 As shown, the value of the variable PpsRefWraparoundOffset can be derived. For example, when the variable wraparound_offset_type is equal to 0, the wraparound offset signaled is the original ERP picture width. As a result, the variable PpsRefWraparoundOffset is equal to pps_ref_wraparound_offset + (CrbSizeY / MinCbSizeY) + 2. When the variable wraparound_offset_type is not equal to 0, the wraparound offset signaled is the difference between the coded picture width and the original ERP picture width. As a result, the variable PpsRefWraparoundOffset is equal to pps_pic_width_in_luma_samples / MinCbSizeY – pps_ref_wraparound_offset.
[0164] In previous VVC (e.g., VVC draft 8), when the variable single_slice_per_subpic_flag is 0, the number of slices is signaled in the PPS, which means that there is at least one sub-picture containing more than one slice. In this case, considering that each sub-picture should contain one or more complete slices, the number of slices needs to be greater than the number of sub-pictures, which therefore results in a minimum number of slices of 2. This is because there is at least one sub-picture in a picture.
[0165] Embodiments of the present disclosure provide methods for improving signaling of the number of stripes. Figure 19FIGURE 2 shows the syntax of an example picture parameter set with the variable num_slices_in_pic_minus2 according to some embodiments of the present disclosure. Figure 19 As shown, syntax elements or variables in the bitstream are shown in bold and are changes to previous VVCs (e.g., Figure 11 The syntax shown in ) is shown in italic type, with the suggested deletion syntax further shown in strikethrough. Figure 19 As shown, a new variable num_slices_in_pic_minus2 can be added to specify the value used to determine PpsRefWraparoundOffset.
[0166] Figure 20A and Figure 20B The semantics of an example improved picture parameter set with the variable num_slices_in_pic_minus2 according to some embodiments of the present disclosure are shown. Figure 20A and Figure 20B As shown, changes to the previous VVC (e.g., Figure 12A and Figure 12B The semantics shown in () are shown in italic type, and the suggested deletion syntax is further shown in strikethrough. It should be understood that Figure 20A and Figure 20B The semantics shown can correspond to Figure 19 In some embodiments, Figure 20A and Figure 20B The semantics shown correspond to VVC (eg, VVC draft 8).
[0167] In some embodiments, as Figure 20A and Figure 20B As shown, the value of the variable num_slices_in_pic_minus2 plus 2 can specify the number of rectangular slices in each picture of the reference PPS. Figure 20A and Figure 20B As shown, the variable num_slices_in_pic_minus2 can replace the variable num_slices_in_pic_minus_1. In some embodiments, the value of num_slices_in_pic_minus2 plus 2 should be in the range of 0 to MaxSlicesPerPicture-2, inclusive, where MaxSlicesPerPicture is specified in VVC (e.g., VVC draft 8). When the variable no_pic_partition_flag is equal to 1, the value of num_slices_in_pic_minus2 is inferred to be equal to -1.
[0168] In some embodiments, the number of slices may be signaled using a variable that is the number of slices minus the number of sub-pictures minus 1 (eg, num_slices_in_pic_minus_subpic_num_minus1). Figure 21 : shows the syntax of an example picture parameter set with the variable num_slices_in_pic_minus_subpic_num_minus1 according to some embodiments of the present disclosure. Figure 21 As shown, syntax elements or variables in the bitstream are shown in bold and are changes to previous VVCs (e.g., Figure 11 The syntax shown in ) is shown in italic type, with suggested deletion syntax further shown in strikethrough.
[0169] Figure 22A and Figure 22B The semantics of an example improved picture parameter set with the variable num_slices_in_pic_minus_subpic_num_minus1 according to some embodiments of the present disclosure are shown. Figure 22A and Figure 22B As shown, changes to the previous VVC (e.g., Figure 12A and Figure 12B The semantics shown in () are shown in italic type, and the suggested deletion syntax is further shown in strikethrough. It should be understood that Figure 22A and Figure 22B The semantics shown in can correspond to Figure 21 In some embodiments, Figure 22A and Figure 22B The semantics shown correspond to VVC (eg, VVC draft 8).
[0170] In VVC (e.g., VVC draft 8), the number of slices is signaled in the PPS only when the variable single_slice_per_subpic_flag is equal to 0, which can be equivalent to stating that there is at least one sub-image containing more than one slice. Therefore, the number of slices should be greater than the number of sub-images, because a sub-image should contain one or more complete slices. As a result, the minimum number of slices is equal to the number of sub-images plus 1. Figure 21 、 Figure 22A and Figure 22B As shown, the number of slices minus the number of sub-pictures minus 1 is signaled (e.g., num_slices_in_pic_minus_subpic_num_minus1) instead of the number of slices minus 1 (e.g., num_slices_in_pic_minus1) to reduce the number of bits signaled.
[0171] In some embodiments, as Figure 22A and Figure 22B As shown, the value of num_slices_in_pic_minus_subpic_num_minus1 plus the number of sub-pictures plus 1 can specify the number of rectangular slices in each picture that references the PPS. In some embodiments, the value of num_slices_in_pic_minus_subpic_num_minus1 should be in the range of 0 to MaxSlicesPerPicture minus sps_num_subpics_minus1 minus 2, inclusive, where MaxSlicesPerPicture may be specified in VVC (e.g., VVC draft 8). In some embodiments, when the variable no_pic_partition_flag is equal to 1, the value of num_slices_in_pic_minus_subpic_num_minus1 can be inferred to be equal to (sps_num_subpics_minus1+1). In some embodiments, the variable SliceNumInPic can be derived as SliceNumInPic=num_slices_in_pic_minus_subpic_num_minus1+sps_num_subpics_minus1+2.
[0172] Embodiments of the present disclosure also provide a new way to signal stripe addresses. Figure 23 FIGURE 2 shows the syntax of an example updated slice header according to some embodiments of the present disclosure. Figure 23 As shown, syntax elements or variables in the bitstream are shown in bold, and changes to the previous VVC (e.g., Figure 13 The syntax shown in ) is shown in italic type, and the suggested deletion syntax is further shown in strikethrough.
[0173] In some embodiments, as Figure 23 As shown, the variable picture_header_in_slice_header_flag can be signaled in the slice header to indicate whether the PH syntax structure is present in the slice header. In VVC (e.g., VVC draft 8), there are constraints on the presence or absence of the PH syntax structure in the slice header and the number of slices in a picture. When the PH syntax structure is present in the slice header, the picture should have only one slice. Therefore, there is no need to signal the slice address. Therefore, the signaling of the slice address can be conditional on the picture_header_in_slice_header_flag. Figure 23As shown, the value of picture_header_in_slice_header_flag can be used as another condition to decide whether to signal the variable slice_address. When picture_header_in_slice_header_flag is equal to 1, signaling of the variable slice_address is skipped. In some embodiments, when the variable slice_address is not signaled, it can be inferred to be 0.
[0174] Embodiments of the present disclosure also provide a method for performing video encoding. Figure 24 A flow chart illustrating an example video encoding method having a variable that signals the difference between the width of a video frame and an offset used to calculate horizontal surround position according to some embodiments of the present disclosure is shown. In some embodiments, Figure 24 The method 24000 shown in Figure 4 The apparatus 400 shown in FIG. 4 is executed. In some embodiments, Figure 24 The method 24000 shown can be based on Figure 15 The syntax shown or Figure 16 In some embodiments, Figure 24 The method 24000 shown in includes a surround motion compensation process performed according to the VVC standard. In some embodiments, Figure 24 The method 24000 shown in can be performed with a 360-degree video sequence as input.
[0175] In step S24010, a surround motion compensation flag is received, wherein the surround motion compensation flag is associated with the image. Figure 15 or Figure 16 As shown, the wraparound motion compensation flag may be the variable pps_ref_wraparound_enabled_flag. In some embodiments, the image is in a bitstream. In some embodiments, the image is part of a 360-degree video.
[0176] In step S24020, it is determined whether the surround motion compensation flag is enabled. Figure 15 As shown, it is determined whether the variable pps_ref_wraparound_enabled_flag is equal to 1.
[0177] In step S24030, in response to determining that the surround motion compensation flag is enabled, a difference between the width of the image and the offset used to determine the horizontal surround position is received. Figure 15As shown, the variable pps_pic_width_minus_wraparound_offset may be received or signaled when the variable pps_ref_wraparound_enabled_flag is determined to be equal to 1. In some embodiments, the difference is less than or equal to the width of the picture divided by the size of the smallest luma coding block minus the size of the luma coding tree block divided by the size of the smallest luma coding block minus 2.
[0178] In some embodiments, in S24030, receiving the difference value includes receiving a wraparound offset type flag. Figure 17 and Figure 18 As shown, a flag wraparound_offset_type may be signaled to indicate whether the value of the signaled wraparound offset is the original ERP picture width or the difference between the coded picture width and the original ERP picture width. Figure 17 and Figure 18 As shown, the value of the flag wraparound_offset_type can be 0 or 1.
[0179] In step S24040, motion compensation is performed on the image according to the surround motion compensation flag and the difference. Figure 15 and Figure 16 In some embodiments, the motion compensation is performed on an image using the variables pps_ref_wraparound_enabled_flag and pps_pic_width_minus_wraparound_offset shown in
[0014] . In some embodiments, performing the motion compensation further comprises determining a wrap motion compensation offset based on the width of the image and the difference value, and performing the motion compensation on the image based on the wrap motion compensation offset. In some embodiments, the wrap motion compensation offset may be determined as the width of the image divided by the size of the minimum luma coding block minus the difference value.
[0180] Figure 25 A flow chart illustrating an example video encoding method with a variable that signals the number of slices in a video frame minus 2, according to some embodiments of the present disclosure. In some embodiments, Figure 25 The method 25000 shown in FIG. 25000 may be performed by Figure 4 The apparatus 400 shown in FIG. 4 is executed. In some embodiments, Figure 25 The method 25000 shown can be based on Figure 19 The syntax shown or Figure 20A and Figure 20B In some embodiments, Figure 25 The method 25000 shown in FIG. 2 may be performed according to the VVC standard.
[0181] In step S25010, an image for encoding is received. In some embodiments, the image may include one or more slices. In some embodiments, the image is in a bitstream. In some embodiments, the one or more slices are rectangular slices.
[0182] In step S25020, a variable indicating the number of stripes in the image minus 2 is signaled in the image parameter set for the image. For example, the variable may be Figure 19 or Figure 20A and Figure 20B In some embodiments, the value of the variable plus 2 can specify the number of rectangular slices in each image. In some embodiments, the variable is part of the PPS. In some embodiments, similar to Figure 20A and Figure 20B The variable num_slices_in_pic_minus2 can replace the variable num_slices_in_pic_minus_1 according to the semantics shown.
[0183] Figure 26 A flow chart illustrating an example video encoding method with a variable that signals a variable indicating the number of slices in a video frame minus the number of sub-pictures in the video frame minus 1, according to some embodiments of the present disclosure. In some embodiments, Figure 26 The method 26000 shown in Figure 4 The apparatus 400 shown in FIG. 4 is executed. In some embodiments, Figure 26 The method 26000 shown can be based on Figure 21 The syntax shown or Figure 22A and Figure 22B In some embodiments, Figure 26 The method 26000 shown in FIG. 2 may be performed according to the VVC standard.
[0184] In step S26010, an image for encoding is received. In some embodiments, the image may include one or more slices and one or more sub-images. In some embodiments, the video frame is in a bitstream. In some embodiments, the one or more slices are rectangular slices.
[0185] In step S26020, a variable indicating the number of stripes in the image minus the number of sub-images in the image minus 1 is signaled in the image parameter set for the image. For example, the variable may be Figure 21 、 Figure 22A and Figure 22BIn some embodiments, the minimum number of slices may be equal to the number of sub-pictures plus 1. Figure 21 、 Figure 22A and Figure 22B As shown, the number of slices minus the number of sub-pictures minus 1 (e.g., num_slices_in_pic_minus_subpic_num_minus1) may be signaled instead of the number of slices minus 1 (e.g., num_slices_in_pic_minus1) to reduce the number of bits signaled. In some embodiments, the variable is part of the PPS. In some embodiments, as Figure 22A and Figure 22B As shown, the value of num_slices_in_pic_minus_subpic_num_minus1 plus the number of sub-pictures plus 1 can specify the number of rectangular slices in each picture of the reference PPS. Figure 20A and Figure 20B The variable num_slices_in_pic_minus2 can replace the variable num_slices_in_pic_minus_1 according to the semantics shown.
[0186] In some embodiments, as shown in step S26020, a variable indicating the number of stripes in an image may be determined based on a variable indicating the number of stripes in a video frame minus the number of sub-images in the video frame minus 1. For example, Figure 21 、 Figure 22A and Figure 22B As shown, the flag or variable SliceNumInPic can be derived from the variable num_slices_in_pic_minus_subpic_num_minus1.
[0187] Figure 27 A flow chart of an example video encoding method with a variable indicating whether a picture header syntax structure is present within a slice header of a video frame according to some embodiments of the present disclosure is shown. In some embodiments, Figure 27 The method 27000 shown in FIG. 27000 may be performed by Figure 4 The apparatus 400 shown in FIG. 4 is executed. In some embodiments, Figure 27 The method 27000 shown can be performed according to Figure 23 In some embodiments, Figure 27 The method 27000 shown in FIG. 27000 may be performed according to the VVC standard.
[0188] In step S27010, an image for encoding is received. The image includes one or more slices. In some embodiments, the image is in a bitstream. In some embodiments, the one or more slices are rectangular slices.
[0189] In step S27020, a variable indicating whether a picture header syntax structure for the picture is present in the slice headers of one or more slices is signaled. For example, the variable may be Figure 23 In some embodiments, as shown in picture_header_in_slice_header_flag Figure 23 As shown, when the PH syntax structure is present in the slice header, the picture may have only one slice. Therefore, there is no need to signal the slice address. Therefore, the signaling of the slice address can be conditioned on the variable picture_header_in_slice_header_flag. In some embodiments, the variable is part of the PPS.
[0190] In some embodiments, a non-transitory computer-readable storage medium including instructions is also provided, and the instructions can be executed by a device (such as the disclosed encoder and decoder) for performing the above method. Common forms of non-transitory media include, for example, floppy disks, hard disks, solid-state drives, tapes or any other magnetic data storage media, CD-ROMs, any other optical data storage media, any physical media with a hole pattern, RAM, PROM and EPROM, FLASH-EPROM or any other flash memory, NVRAM, cache, registers, any other memory chip or cartridge storage, and networked versions thereof. The device may include one or more processors (CPUs), input / output interfaces, network interfaces, and / or memories.
[0191] It should be noted that relational terms in this document, such as "first" and "second", are used only to distinguish an entity or operation from another entity or operation, and do not require or imply any actual relationship or order between these entities or operations. In addition, the words "include", "have", "include" and "including" and other similar forms are equivalent in meaning and are open-ended, in that one or more items following any of these words is not intended to be an exhaustive list of such one or more items, or to be limited to the listed one or more items.
[0192] As used herein, unless specifically stated otherwise, the term "or" includes all possible combinations, unless otherwise specified. For example, if it is stated that a database may include A or B, then unless otherwise specified or not applicable, the database may include A, or B, or A and B. As a second example, if it is stated that a database may include A, B, or C, then unless otherwise specified or not applicable, the database may include A, or B, or C, or A and B, or A and C, or B and C, or A, B, and C.
[0193] It should be understood that the above embodiments can be implemented by hardware, or software (program code), or a combination of hardware and software. If implemented by software, it can be stored in the above-mentioned computer-readable medium. The software can perform the disclosed method when executed by a processor. The computing units and other functional units described in this disclosure can be implemented by hardware, or software, or a combination of hardware and software. Those of ordinary skill in the art will also understand that the above-mentioned multiple modules / units can be combined into one module / unit, and each module / unit in the above-mentioned modules / units can be further divided into multiple sub-modules / sub-units.
[0194] In the foregoing description, embodiments have been described with reference to numerous specific details, which may vary from implementation to implementation. Certain modifications and variations may be made to the described embodiments. Other embodiments will be apparent to those skilled in the art upon consideration of the specification and practice of the invention disclosed herein. This description and embodiments are to be considered exemplary only, with the true scope and spirit of the invention being indicated by the appended claims. The sequence of steps shown in the accompanying drawings is for illustrative purposes only and is not intended to be limiting to any particular sequence of steps. Therefore, it will be understood by those skilled in the art that these steps may be performed in a different order while implementing the same method.
[0195] The embodiments may be further described using the following terms: 1. A video decoding method, comprising: receiving a surround motion compensation flag; determining whether to enable surround motion compensation based on the surround motion compensation flag; In response to determining that the surround motion compensation is enabled, receiving data indicating a difference between a width of an image and an offset used to determine a horizontal surround position; and Motion compensation is performed based on the surround motion compensation flag and the difference value. 2. The method of clause 1, wherein the difference value is in units of the size of a minimum luma coding block. 3. A method according to clause 2, wherein the difference is less than or equal to (pps_pic_width_in_luma_samples / MinCbSizeY)-(CtbSizeY / MinCbSizeY)-2, where pps_pic_width_in_luma_samples is the width of the picture in luma samples, MinCbSizeY is the size of the minimum luma coding block, and CtbSizeY is the size of the luma coding tree block. 4. The method of clause 1, wherein performing the motion compensation further comprises: determining a surround motion compensation offset based on the width of the image and the difference; and The motion compensation is performed according to the surround motion compensation offset. 5. The method of clause 4, wherein determining the surround motion compensation offset based on the width of the image and the difference further comprises: dividing a width of the image in units of luma samples by a size of a minimum luma coding block to generate a first value; and The surround motion compensation offset is determined to be equal to the first value minus the difference value. 6. The method of clause 1, wherein receiving the data indicative of the difference further comprises: Receives the wraparound offset type flag; determining whether the wraparound offset type flag is equal to a first value or a second value; In response to determining that the wrap offset type flag is equal to the first value, receiving data indicating a difference between a width of the image and an offset used to calculate a horizontal wrap position; and In response to determining that the wrap offset type flag is equal to the second value, data indicating an offset for calculating a horizontal wrap position is received. 7. The method of clause 6, wherein each of the first value and the second value is 0 or 1. 8. The method of clause 1, wherein the motion compensation is performed according to a general video coding standard. 9. The method of clause 1, wherein the image is part of a 360-degree video sequence. 10. The method of clause 1, wherein the surround motion compensation flag and the difference value are signaled in a picture parameter set (PPS). 11. A video decoding method, comprising: signaling a surround motion compensation flag, the flag indicating whether surround motion compensation is enabled; In response to the surround motion compensation flag indicating that the surround motion compensation is enabled, signaling data indicating a difference between a width of the image and an offset used to determine a horizontal surround position; and Motion compensation is performed based on the surround motion compensation flag and the difference value. 12. A method according to clause 11, wherein the difference value is in units of the size of a minimum luma coding block. 13. A method according to clause 12, wherein the difference is less than or equal to (pps_pic_width_in_luma_samples / MinCbSizeY)-(CtbSizeY / MinCbSizeY)-2, where pps_pic_width_in_luma_samples is the width of the picture in luma samples, MinCbSizeY is the size of the minimum luma coding block, and CtbSizeY is the size of the luma coding tree block. 14. The method of clause 11, wherein performing the motion compensation further comprises: determining a surround motion compensation offset based on the width and the difference of the image; and The motion compensation is performed according to the surround motion compensation offset. 15. The method of clause 14, wherein determining the surround motion compensation offset based on the width of the image and the difference further comprises: dividing a width of the image in units of luma samples by a size of a minimum luma coding block to generate a first value; and The surround motion compensation offset is determined to be equal to the first value minus the difference value. 16. The method of clause 11, wherein signaling the data indicative of the difference further comprises: signaling a wraparound offset type flag, wherein a value of the wraparound offset type flag is a first value or a second value; In response to the value of the wrap offset type flag being equal to the first value, signaling data indicating a difference between a width of the image and an offset used to calculate a horizontal wrap position; and In response to the value of the wrap offset type flag being equal to the second value, data for calculating an offset of a horizontal wrap position is signaled. 17. The method of clause 16, wherein each of the first value and the second value is 0 or 1. 18. The method of clause 11, wherein the motion compensation is performed according to a general video coding standard. 19. The method of clause 11, wherein the image is part of a 360-degree video sequence. 20. The method of clause 11, wherein the surround motion compensation flag and the difference value are signaled in a picture parameter set (PPS). 21. A video encoding method, comprising: receiving a picture for encoding, wherein the picture comprises one or more slices; and In the picture parameter set for that picture, a variable indicating the number of slices in the video frame minus 2 is signaled. 22. The method of clause 21, wherein the image is in a bitstream. 23. The method of clause 21, wherein the image is encoded according to a common video coding standard. 24. The method of clause 21, wherein the one or more strips are rectangular strips. 25. A video encoding method, comprising: receiving a picture for encoding, wherein the picture comprises one or more slices and one or more sub-pictures; and In a picture parameter set for the picture, a variable indicating the number of slices in the picture minus the number of sub-pictures in the picture minus one is signaled. 26. The method of clause 25, wherein the image is in a bitstream. 27. The method according to clause 25, further comprising: A variable indicating the number of strips in the image is determined based on the variable indicating the number of strips in the image minus the number of sub-images in the image minus one. 28. The method of clause 25, wherein the image is encoded according to a common video coding standard. 29. The method of clause 25, wherein the one or more strips are rectangular strips. 30. A video encoding method, comprising: receiving an image for encoding, wherein the image comprises one or more slices; signaling a variable indicating whether a picture header syntax structure of the picture is present in a slice header of the one or more slices; and The stripe address is signaled according to the variable. 31. The method of clause 30, wherein the image is in a bitstream. 32. The method of clause 30, wherein the image is encoded according to a common video coding standard. 33. The method of clause 30, wherein the one or more strips are rectangular. 34. A system for performing video data processing, the system comprising: a memory storing an instruction set; and a processor configured to execute the set of instructions to cause the system to: receiving a surround motion compensation flag; determining whether to enable surround motion compensation based on the surround motion compensation flag; In response to determining that the surround motion compensation is enabled, receiving data indicating a difference between a width of an image and an offset used to determine a horizontal surround position; and Motion compensation is performed based on the surround motion compensation flag and the difference value. 35. The system of clause 34, wherein the difference value is in units of a minimum luma coding block size. 36. A system according to clause 35, wherein the difference is less than or equal to (pps_pic_width_in_luma_samples / MinCbSizeY)-(CtbSizeY / MinCbSizeY)-2, where pps_pic_width_in_luma_samples is the width of the picture in luma samples, MinCbSizeY is the size of a minimum luma coding block, and CtbSizeY is the size of a luma coding tree block. 37. The system of clause 34, wherein, in performing the motion compensation, the processor is configured to execute the set of instructions to cause the system to: determining a surround motion compensation offset based on the width of the image and the difference; and The motion compensation is performed according to the surround motion compensation offset. 38. The system of clause 37, wherein, in determining the surround motion compensation offset based on the width and the disparity of the image, the processor is configured to execute the set of instructions to cause the system to: dividing a width of the image in units of luma samples by a size of a minimum luma coding block to generate a first value; and The surround motion compensation offset is determined to be equal to the first value minus the difference value. 39. The system of clause 34, wherein upon receiving the data indicative of the difference, the processor is configured to execute the set of instructions to cause the system to: Receives the wraparound offset type flag; determining whether the wraparound offset type flag is equal to a first value or a second value; In response to determining that the wrap offset type flag is equal to the first value, receiving data indicating a difference between a width of the image and an offset used to calculate a horizontal wrap position; and In response to determining that the wrap offset type flag is equal to the second value, data indicating an offset for calculating a horizontal wrap position is received. 40. The system of clause 39, wherein each of the first value and the second value is 0 or 1. 41. The system of clause 34, wherein the motion compensation is performed according to a common video coding standard. 42. The system of clause 34, wherein the image is part of a 360-degree video sequence. 43. The system of clause 34, wherein the surround motion compensation flag and the difference value are signaled in a picture parameter set (PPS). 44. A system for performing video data processing, the system comprising: a memory storing an instruction set; and a processor configured to execute the set of instructions to cause the system to: signaling a surround motion compensation flag, the flag indicating whether surround motion compensation is enabled; In response to the surround motion compensation flag indicating that the surround motion compensation is enabled, signaling data indicating a difference between a width of the image and an offset used to determine a horizontal surround position; and Motion compensation is performed based on the surround motion compensation flag and the difference value. 45. The system of clause 44, wherein the difference value is in units of a minimum luma coding block size. 46. A system according to clause 45, wherein the difference is less than or equal to (pps_pic_width_in_luma_samples / MinCbSizeY)-(CtbSizeY / MinCbSizeY)-2, where pps_pic_width_in_luma_samples is the width of the picture in luma samples, MinCbSizeY is the size of a minimum luma coding block, and CtbSizeY is the size of a luma coding tree block. 47. The system of clause 44, wherein, in performing the motion compensation, the processor is configured to execute the set of instructions to cause the system to: determining a surround motion compensation offset based on the width and the difference of the image; and The motion compensation is performed according to the surround motion compensation offset. 48. The system of clause 47, wherein, in determining the surround motion compensation offset based on the width of the image and the difference, the processor is configured to execute the set of instructions to cause the system to: dividing a width of the image in units of luma samples by a size of a minimum luma coding block to generate a first value; and The surround motion compensation offset is determined to be equal to the first value minus the difference value. 49. The system of clause 44, wherein upon receiving the data indicative of the difference, the processor is configured to execute the set of instructions to cause the system to: signaling a wraparound offset type flag, wherein a value of the wraparound offset type flag is a first value or a second value; In response to the value of the wrap offset type flag being equal to the first value, signaling data indicating a difference between a width of the image and an offset used to calculate a horizontal wrap position; and In response to the value of the wrap offset type flag being equal to the second value, data for calculating an offset of a horizontal wrap position is signaled. 50. The system of clause 49, wherein each of the first value and the second value is 0 or 1. 51. The system of clause 44, wherein the motion compensation is performed according to a common video coding standard. 52. The system of clause 44, wherein the image is part of a 360-degree video sequence. 3. The system of clause 44, wherein the surround motion compensation flag and the difference value are signaled in a picture parameter set (PPS). 54. A system for performing video encoding, the system comprising: a memory storing an instruction set; and a processor configured to execute the set of instructions to cause the system to: receiving a picture for encoding, wherein the picture comprises one or more slices; and In the picture parameter set for that picture, a variable indicating the number of slices in the video frame minus 2 is signaled. 55. The system of clause 54, wherein the image is in a bitstream. 56. The system of clause 54, wherein the image is encoded according to a common video coding standard. 57. The system of clause 54, wherein the one or more strips are rectangular strips. 58. A system for performing video encoding, the system comprising: a memory storing an instruction set; and a processor configured to execute the set of instructions to cause the system to: receiving a picture for encoding, wherein the picture comprises one or more slices and one or more sub-pictures; and In a picture parameter set for the picture, a variable indicating the number of slices in the picture minus the number of sub-pictures in the picture minus one is signaled. 59. The system of clause 58, wherein the image is in a bitstream. 60. The system of clause 58, wherein the processor is configured to execute the set of instructions to cause the system to: A variable indicating the number of strips in the image is determined based on the variable indicating the number of strips in the image minus the number of sub-images in the image minus one. 61. The system of clause 58, wherein the image is encoded according to a common video coding standard. 62. The system of clause 58, wherein the one or more strips are rectangular strips. 63. A system for performing video encoding, the system comprising: a memory storing an instruction set; and a processor configured to execute the set of instructions to cause the system to: receiving an image for encoding, wherein the image comprises one or more slices; signaling a variable indicating whether a picture header syntax structure of the picture is present in a slice header of the one or more slices; and The stripe address is signaled according to the variable. 64. The system of clause 63, wherein the image is in a bitstream. 65. The system of clause 63, wherein the image is encoded according to a common video coding standard. 66. The system of clause 63, wherein the one or more strips are rectangular. 67. A non-transitory computer-readable medium storing a set of instructions executable by one or more processors of a device to cause the device to initiate a method for performing video data processing, the method comprising: receiving a surround motion compensation flag; determining whether to enable surround motion compensation based on the surround motion compensation flag; In response to determining that the surround motion compensation is enabled, receiving data indicating a difference between a width of an image and an offset used to determine a horizontal surround position; and Motion compensation is performed based on the surround motion compensation flag and the difference value. 68. The non-transitory computer-readable medium of clause 67, wherein the difference value is in units of a minimum luma coding block size. 69. A non-transitory computer-readable medium according to clause 68, wherein the difference is less than or equal to (pps_pic_width_in_luma_samples / MinCbSizeY)-(CtbSizeY / MinCbSizeY)-2, where pps_pic_width_in_luma_samples is the width of the image in units of luma samples, MinCbSizeY is the size of a minimum luma coding block, and CtbSizeY is the size of a luma coding tree block. 70. The non-transitory computer-readable medium of clause 67, wherein performing the motion compensation further comprises: determining a surround motion compensation offset based on the width of the image and the difference; and The motion compensation is performed according to the surround motion compensation offset. 71. The non-transitory computer-readable medium of clause 70, wherein determining the surround motion compensation offset based on the width of the image and the difference further comprises: dividing a width of the image in units of luma samples by a size of a minimum luma coding block to generate a first value; and The surround motion compensation offset is determined to be equal to the first value minus the difference value. 72. The non-transitory computer-readable medium of clause 67, wherein receiving data indicative of a difference further comprises: Receives the wraparound offset type flag; determining whether the wraparound offset type flag is equal to a first value or a second value; In response to determining that the wrap offset type flag is equal to the first value, receiving data indicating a difference between a width of the image and an offset used to calculate a horizontal wrap position; and In response to determining that the wrap offset type flag is equal to the second value, data indicating an offset for calculating a horizontal wrap position is received. 73. The non-transitory computer-readable medium of clause 72, wherein each of the first value and the second value is 0 or 1. 74. The non-transitory computer-readable medium of clause 67, wherein the motion compensation is performed according to a general video coding standard. 75. The non-transitory computer-readable medium of clause 67, wherein the image is part of a 360-degree video sequence. 76. The non-transitory computer-readable medium of clause 67, wherein the surround motion compensation flag and the difference value are signaled in a picture parameter set (PPS). 77. A non-transitory computer-readable medium storing a set of instructions executable by one or more processors of a device to cause the device to initiate a method for performing video data processing, the method comprising: signaling a surround motion compensation flag, the flag indicating whether surround motion compensation is enabled; In response to the surround motion compensation flag indicating that the surround motion compensation is enabled, signaling data indicating a difference between a width of the image and an offset used to determine a horizontal surround position; and Motion compensation is performed based on the surround motion compensation flag and the difference value. 78. The non-transitory computer-readable medium of clause 77, wherein the difference value is in units of a minimum luma coding block size. 79. A non-transitory computer-readable medium according to clause 78, wherein the difference is less than or equal to (pps_pic_width_in_luma_samples / MinCbSizeY)-(CtbSizeY / MinCbSizeY)-2, where pps_pic_width_in_luma_samples is the width of the image in units of luma samples, MinCbSizeY is the size of a minimum luma coding block, and CtbSizeY is the size of a luma coding tree block. 80. The non-transitory computer-readable medium of clause 77, wherein performing the motion compensation further comprises: determining a surround motion compensation offset based on the width and the difference of the image; and The motion compensation is performed according to the surround motion compensation offset. 81. The non-transitory computer-readable medium of clause 80, wherein determining the surround motion compensation offset based on the width of the image and the difference further comprises: dividing a width of the image in units of luma samples by a size of a minimum luma coding block to generate a first value; and The surround motion compensation offset is determined to be equal to the first value minus the difference value. 82. The non-transitory computer-readable medium of clause 77, wherein receiving data indicating a difference further comprises: signaling a wraparound offset type flag, wherein a value of the wraparound offset type flag is a first value or a second value; In response to the value of the wrap offset type flag being equal to the first value, signaling data indicating a difference between a width of the image and an offset used to calculate a horizontal wrap position; and In response to the value of the wrap offset type flag being equal to the second value, data for calculating an offset of a horizontal wrap position is signaled. 83. The non-transitory computer-readable medium of clause 82, wherein each of the first value and the second value is 0 or 1. 84. The non-transitory computer-readable medium of clause 77, wherein the motion compensation is performed according to a general video coding standard. 85. The non-transitory computer-readable medium of clause 77, wherein the image is part of a 360-degree video sequence. 86. The non-transitory computer-readable medium of clause 77, wherein the surround motion compensation flag and the difference value are signaled in a picture parameter set (PPS). 87. A non-transitory computer-readable medium storing a set of instructions executable by one or more processors of a device to cause the device to initiate a method for performing video encoding, the method comprising: receiving a picture for encoding, wherein the picture comprises one or more slices; and In the picture parameter set for that picture, a variable indicating the number of slices in the video frame minus 2 is signaled. 88. The non-transitory computer-readable medium of clause 87, wherein the image is in a bitstream. 89. The non-transitory computer-readable medium of clause 87, wherein the image is encoded according to a common video coding standard. 90. The non-transitory computer-readable medium of clause 87, wherein the one or more strips are rectangular strips. 91. A non-transitory computer-readable medium storing a set of instructions executable by one or more processors of a device to cause the device to initiate a method for performing video encoding, the method comprising: receiving a picture for encoding, wherein the picture comprises one or more slices and one or more sub-pictures; and In a picture parameter set for the picture, a variable indicating the number of slices in the picture minus the number of sub-pictures in the picture minus one is signaled. 92. The non-transitory computer-readable medium of clause 91, wherein the image is in a bitstream. 93. The non-transitory computer-readable medium of clause 91, further comprising: A variable indicating the number of strips in the image is determined based on the variable indicating the number of strips in the image minus the number of sub-images in the image minus one. 94. The non-transitory computer-readable medium of clause 91, wherein the image is encoded according to a common video coding standard. 95. The non-transitory computer-readable medium of clause 91, wherein the one or more strips are rectangular strips. 96. A non-transitory computer-readable medium storing a set of instructions executable by one or more processors of a device to cause the device to initiate a method for performing video encoding, the method comprising: receiving an image for encoding, wherein the image comprises one or more slices; signaling a variable indicating whether a picture header syntax structure of the picture is present in a slice header of the one or more slices; and The stripe address is signaled according to the variable. 97. The non-transitory computer-readable medium of clause 96, wherein the image is in a bitstream. 98. The non-transitory computer-readable medium of clause 96, wherein the image is encoded according to a common video coding standard. 99. The non-transitory computer-readable medium of clause 96, wherein the one or more strips are rectangular.
[0196] In the drawings and the specification, exemplary embodiments have been disclosed. However, many variations and modifications may be made to these embodiments. Therefore, although specific terms are employed, they are used in a generic and descriptive sense only and not for purposes of limitation.
Claims
1. A video encoding method, comprising: signaling a surround motion compensation flag indicating whether surround motion compensation is enabled; responsive to the surround motion compensation flag indicating that the surround motion compensation is enabled, signaling data indicating a difference between a width of an image and a value used to determine a horizontal surround position offset; in, The surround motion compensation flag and the difference value are used in a motion compensation operation.
2. The method according to claim 1, wherein The difference is in units of the size of the minimum luma coding block.
3. The method of claim 1 , wherein the difference is less than or equal to (pps_pic_width_in_luma_samples / MinCbSizeY)-(CtbSizeY / MinCbSizeY)-2, where pps_pic_width_in_luma_samples is the width of the image in units of luma samples, MinCbSizeY is the size of a minimum luma coding block, and CtbSizeY is the size of a luma coding tree block.
4. The method according to claim 1, wherein In motion compensation operation, determining a surround motion compensation offset based on the width of the image and the difference; and The surround motion compensation offset is used to perform the motion compensation.
5. The method according to claim 4, wherein When determining the surround motion compensation offset according to the width of the image and the difference, dividing a width of the image in units of luma samples by a size of a minimum luma coding block to generate a first value; and The surround motion compensation offset is determined to be equal to the first value minus the difference value.
6. The method according to claim 1, wherein Signaling data indicative of the difference further comprises: Receives the wraparound offset type flag; determining whether the wraparound offset type flag is equal to a first value or a second value; In response to determining that the wrap offset type flag is equal to the first value, receiving data indicating a difference between a width of the image and an offset used to calculate a horizontal wrap position; and In response to determining that the wrap offset type flag is equal to the second value, data indicating an offset for calculating a horizontal wrap position is received.
7. The method according to claim 6, wherein: Each of the first value and the second value is 0 or 1.
8. The method according to claim 1, wherein The method further comprises: A picture for encoding is received, wherein the picture includes one or more slices.
9. The method according to claim 1, wherein: The images are in a bitstream.
10. The method according to claim 1, wherein The surround motion compensation flag and the difference value are signaled in a picture parameter set (PPS).
11. A system for performing video encoding, the system comprising: a memory storing an instruction set; and one or more processors configured to execute the set of instructions to cause the system to perform: signaling a surround motion compensation flag indicating whether surround motion compensation is enabled; responsive to the surround motion compensation flag indicating that the surround motion compensation is enabled, signaling data indicating a difference between a width of an image and a value used to determine a horizontal surround position offset; in, The surround motion compensation flag and the difference value are used in a motion compensation operation.
12. The system according to claim 11, wherein The difference is in units of the size of the minimum luma coding block.
13. The system according to claim 11, wherein: The difference is less than or equal to (pps_pic_width_in_luma_samples / MinCbSizeY)-(CtbSizeY / MinCbSizeY)-2, where pps_pic_width_in_luma_samples is the width of the image in units of luma samples, MinCbSizeY is the size of the minimum luma coding block, and CtbSizeY is the size of the luma coding tree block.
14. The system according to claim 11, wherein: In motion compensation operation, determining a surround motion compensation offset based on the width of the image and the difference; and The surround motion compensation offset is used to perform the motion compensation.
15. The system according to claim 14, wherein: When determining the surround motion compensation offset according to the width of the image and the difference, dividing a width of the image in units of luma samples by a size of a minimum luma coding block to generate a first value; and The surround motion compensation offset is determined to be equal to the first value minus the difference value.
16. The system according to claim 11, wherein signaling data indicative of the difference, the one or more processors being configured to execute the set of instructions to cause the system to further: Receives the wraparound offset type flag; determining whether the wraparound offset type flag is equal to a first value or a second value; In response to determining that the wrap offset type flag is equal to the first value, receiving data indicating a difference between a width of the image and an offset used to calculate a horizontal wrap position; and In response to determining that the wrap offset type flag is equal to the second value, data indicating an offset for calculating a horizontal wrap position is received.
17. The system according to claim 16, wherein: Each of the first value and the second value is 0 or 1.
18. The system according to claim 11, wherein: The image is encoded according to a common video coding standard.
19. The system according to claim 11, wherein: The surround motion compensation flag and the difference value are signaled in a picture parameter set (PPS).
20. A non-transitory computer-readable medium storing a bitstream, the non-transitory computer-readable medium being part of a computing device, the bitstream being generated by one or more processors of the computing device executing a set of instructions, wherein: Execution of the set of instructions causes the computing device to perform operations for video encoding, the operations comprising: signaling a surround motion compensation flag indicating whether surround motion compensation is enabled; responsive to the surround motion compensation flag indicating that the surround motion compensation is enabled, signaling data indicating a difference between a width of an image and a value used to determine a horizontal surround position offset; The surround motion compensation flag and the difference value are used in a motion compensation operation.
Citation Information
Patent Citations
Method for encoding and decoding images based on constrained offset compensation and loop filter, and apparatus therefor
CN103959794A
Adaptive loop filtering on deblocking filter results in video coding
US20190238845A1
Systems and methods for performing motion compensation for coding of video data
US20190273943A1
An apparatus, a method and a computer program for video coding and decoding
WO2019211522A2