Method for wrap-around padding for omnidirectional media coding

JP2024133577A5Active Publication Date: 2025-05-15TENCENT AMERICA LLC
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
JP2024107340
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Priority Date
2019-12-11
Filing Date
2024-07-03
Publication Date
2025-05-15
Estimated Expiration
2039-12-27

AI Technical Summary

Technical Problem

Existing video encoding and decoding technologies face challenges in efficiently handling 360-degree omnidirectional media due to seam visual artifacts resulting from conventional 2D image reprojection, which complicates the parsing and decoding of NAL units in video bitstreams.

Method used

The method involves applying wraparound padding and iterative decoding techniques to subregions of 360-degree images, using image segmentation information to determine the need for wraparound padding and applying appropriate padding methods to reduce seam artifacts during the decoding process.

Benefits of technology

This approach effectively reduces seam artifacts in omnidirectional media decoding, improving the accuracy and efficiency of video reconstruction by optimizing the parsing and decoding of NAL units in video bitstreams.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 00000000_0000_ABST
    Figure 00000000_0000_ABST
Patent Text Reader

Abstract

To provide a method of reconstructing a coded current picture for decoding a video.SOLUTION: A method includes the steps of: decoding picture partitioning information corresponding to a current picture; determining whether padding includes wrap-around padding, using the picture partitioning information, based on determining that the padding is applied; based on determining that the padding does not include wrap-around padding, applying repetition padding to sub-regions, and decoding the sub-regions using the repetition padding; based on determining that the padding includes wrap-around padding, applying the wrap-around padding to the sub-regions, and decoding the sub-regions using the wrap-around padding; and reconstructing the current picture based on the decoded sub-regions.SELECTED DRAWING: Figure 8
Need to check novelty before this filing date? Find Prior Art

Description

[Technical field]

[0001] CROSS-REFERENCE TO RELATED APPLICATIONS This application claims priority under 35 U.S.C. §119 to U.S. Provisional Patent Application No. 62 / 787,063, filed in the U.S. Patent and Trademark Office on December 31, 2018, and U.S. Patent Application No. 16 / 710,936, filed in the U.S. Patent and Trademark Office on December 11, 2019, the disclosures of which are incorporated herein by reference in their entireties.

[0002] The disclosed subject matter relates to video encoding and decoding, and more particularly, to including wraparound padding for 360-degree omnidirectional media encoding. [Background technology]

[0003] Examples of video encoding and decoding using inter-image prediction with motion compensation have been known for several decades. Uncompressed digital video may consist of a sequence of images, each image having spatial dimensions of, for example, 1920x1080 luminance samples and associated chrominance samples. The sequence of images may have a fixed or variable image rate (also informally known as frame rate) of, for example, 60 images per second, or 60 Hz. Uncompressed video has significant bitrate requirements. For example, 1080p60 4:2:0 video (1920x1080 luminance sample resolution at a 60 Hz frame rate) with 8 bits per sample requires a bandwidth of nearly 1.5 Gbits / s. One hour of such video requires more than 600 Gbytes of storage space.

[0004] Video encoding and decoding may aim at reducing redundancy in the input video signal through compression. Compression may help to reduce the aforementioned bandwidth or storage space requirements, possibly by a factor of 100 or more. Both lossless and lossy compression, as well as combinations of these, may be used. Lossless compression refers to techniques that allow a perfect copy of the original signal to be reconstructed from the compressed original signal. With lossy compression, the reconstructed signal may not be identical to the original signal, but the distortion between the original and reconstructed signals is small enough that the reconstructed signal is sufficiently useful for the intended application. For video, lossy compression is widely used. The amount of distortion is tolerable depending on the application, for example, users of some consumer streaming applications may tolerate higher order distortion than users of applications that contribute to television. The achievable compression ratio may reflect that the higher the possible / acceptable distortion, the higher the compression ratio.

[0005] Video encoders and decoders can use several broad categories of techniques, including, for example, motion compensation, transform, quantization, and entropy coding, some of which are introduced below.

[0006] The division of coded video bitstreams into packets for transmission over packet networks has been used for several decades. Early on, video coding standards and technologies were mostly optimized for bot-oriented transmission and defined bitstreams. The packetization occurring at the system layer interface was specified, for example, in the Real-time Transport Protocol (RTP) payload format. With the emergence of Internet connections suitable for the mass use of video over the Internet, video coding standards reflected their prominent use cases by making a conceptual distinction between the Video Coding Layer (VCL) and the Network Abstraction Layer (NAL). NAL units were introduced in H.264 in 2003 and have been retained in several video coding standards and technologies since then, with only minor modifications.

[0007] A NAL unit can often be considered as the smallest entity that a decoder can act on without having to decode all previous NAL units of a coded video sequence. To this extent, NAL units enable several error resilience techniques, as well as several bitstream manipulation techniques, to include bitstream pruning, by a Media Aware Network Element (MANE), such as a Selective Forwarding Unit (SFU) or a Multipoint Control Unit (MCU). Summary of the Invention [Problem to be solved by the invention]

[0008] Figure 1 shows the relevant parts of the parsing diagram of the NAL unit header according to H.264 (101) and H.265 (102), in both cases without the respective extensions. In both cases, forbidden_zero_bit is a zero bit used for start code emulation prevention in some system layer environments. The nal_unit_type syntax element indicates the type of data the NAL unit holds, which may be, for example, one of several slice types, parameter setting types, supplemental enhancement information (SEI-) messages, etc. The H.265 NAL unit header further includes nuh_layer_id and nuh_temporal_id_plus1, which indicate the spatial / SNR and temporal layer of the coded image to which the NAL unit belongs.

[0009] It can be observed that NAL unit headers contain only easily parsable fixed-length codewords, which do not have any parsing dependencies on other data in the bitstream, e.g., other NAL unit headers, parameter sets, etc. Since NAL unit headers are the first octets within a NAL unit, a MANE can easily extract, parse, and act on them. In contrast, other higher level syntax elements, e.g., slice or tile headers, are less accessible to a MANE, since they would require maintaining parameter set context or processing variable-length or arithmetically coded codepoints.

[0010] It may be further observed that the NAL unit header shown in FIG. 1 does not contain information that allows associating the NAL unit with a coded image consisting of multiple NAL units (e.g., including multiple tiles or slices, at least some of which are packetized in individual NAL units).

[0011] Some transmission technologies, such as RTP (RFC 3550), the MPEG Systems standard, and the ISO file formats, may contain some information, often in the form of timing information such as presentation time (for MPEG and ISO file formats) or capture time (for RTP), which the MANE can easily access and help it to associate its respective transmission unit with a coded picture. However, the semantics of this information may differ for each transmission / storage technology and may not be directly related to the picture structure used for the video coding. Therefore, such information is only heuristic and may not be particularly well suited to identify whether NAL units in a NAL unit stream belong to the same coded picture. [Means for solving the problem]

[0012] In an embodiment, a method is provided for reconstructing an encoded current image for video decoding using at least one processor, the method including the steps of: decoding image partitioning information corresponding to the current image; using the image partitioning information to determine whether padding is applied to multiple sub-regions of the current image; based on a determination that padding is not applied, decoding the multiple sub-regions without padding the multiple sub-regions; based on a determination that padding is applied, using the image partitioning information to determine whether the padding includes wrap-around padding; based on a determination that the padding does not include wrap-around padding, applying repeated padding to the multiple sub-regions and decoding the multiple sub-regions using the repeated padding; based on a determination that the padding includes wrap-around padding, applying wrap-around padding to the multiple sub-regions and decoding the multiple sub-regions using the wrap-around padding; and reconstructing the current image based on the decoded multiple sub-regions.

[0013] In an embodiment, an apparatus for reconstructing an encoded current image for decoding a video comprises at least one memory configured to store program code, and at least one processor configured to read the program code and to operate as instructed by the program code, the program code comprising: a first decoding code configured to cause the at least one processor to decode image segmentation information corresponding to the current image; a first decision code configured to cause the at least one processor to determine whether padding is applied to a plurality of sub-regions of the current image using the image segmentation information; a second decoding code configured to cause the at least one processor to decode the plurality of sub-regions without padding the plurality of sub-regions based on a determination that padding is not applied; a second decision code configured to determine whether the padding includes wrap-around padding using the image partition information based on a determination that the padding does not include wrap-around padding; a first iterative code configured to cause the at least one processor to apply repeated padding to the plurality of sub-regions and decode the plurality of sub-regions using the repeated padding based on a determination that the padding includes wrap-around padding; a second iterative code configured to cause the at least one processor to apply wrap-around padding to the plurality of sub-regions and decode the plurality of sub-regions using the wrap-around padding based on a determination that the padding includes wrap-around padding; and a reconstruction code configured to cause the at least one processor to reconstruct a current image based on the decoded plurality of sub-regions.

[0014] In an embodiment, a non-transitory computer-readable medium is provided that stores instructions, the instructions including one or more instructions that, when executed by one or more processors of an apparatus that reconstructs a current image encoded to decode video, cause the one or more processors to decode image segmentation information corresponding to the current image, use the image segmentation information to determine whether padding is applied to a plurality of sub-regions of the current image, decode the plurality of sub-regions without padding based on a determination that padding is not applied, use the image segmentation information to determine whether the padding includes wrap-around padding based on a determination that padding is applied, apply repeat padding to the plurality of sub-regions and decode the plurality of sub-regions using the repeat padding, apply wrap-around padding to the plurality of sub-regions and decode the plurality of sub-regions using the wrap-around padding based on a determination that the padding includes wrap-around padding, and reconstruct the current image based on the decoded plurality of sub-regions.

[0015] Further features, nature and various advantages of the subject matter of the present disclosure will become more apparent from the following detailed description and the accompanying drawings. [Brief description of the drawings]

[0016] [Figure 1] 1 is a schematic diagram of a NAL unit header according to H.264 and H.265. [Diagram 2] FIG. 1 is a schematic diagram of a simplified block diagram of a communication system, according to an embodiment. [Diagram 3] FIG. 1 is a schematic diagram of a simplified block diagram of a communication system, according to an embodiment. [Figure 4] FIG. 2 is a schematic diagram of a simplified block diagram of a decoder according to an embodiment; [Diagram 5] FIG. 2 is a schematic diagram of a simplified block diagram of an encoder according to an embodiment; [Figure 6]FIG. 2 is a schematic diagram of syntax elements for offset signal transmission according to an embodiment. [Figure 7] FIG. 1 is a schematic diagram of a syntax element for an encoder to signal padding width according to an embodiment. [Figure 8] 4 is a flowchart of an exemplary process for reconstructing a current encoded image for decoding video, according to an embodiment. [Figure 9] FIG. 1 is a schematic diagram of a computer system, according to an embodiment. DETAILED DESCRIPTION OF THE PREFERRED EMBODIMENTS

[0017] Problem to be solved by the invention The 360-degree video is mapped onto a 2D video using a 3D-2D projection method such as equirectangular projection (ERP). The projected video is encoded and decoded by a conventional 2D video encoder and rendered by reprojecting the 2D video onto a 3D surface. The encoded regions are then stitched together separately, resulting in visual seam artifacts from the reprojection process.

[0018] Detailed Description FIG. 2 shows a simplified block diagram of a communication system (200) according to an embodiment of the present disclosure. The system (200) may include at least two terminals (210, 220) interconnected via a network (250). In the case of a one-way transmission of data, a first terminal (210) may locally encode video data for transmission over the network (250) to the other terminal (220). The second terminal (220) may receive the other terminal's encoded video data from the network (250), decode the encoded data, and display the recovered video data. One-way data transmission may be common in media distribution applications, etc.

[0019] 2 shows a second pair of terminals (230, 240) provided to support bidirectional transmission of encoded video, such as may occur during a video conference. For bidirectional transmission of data, each terminal (230, 240) may encode video data captured at a local location for transmission over a network (250) to the other terminal. Each terminal (230, 240) may receive the encoded video data transmitted by the other terminal, may decode the encoded data, and may display the recovered video data on a local display device.

[0020] In FIG. 2, the terminals (210-240) may be depicted as servers, personal computers, and smartphones, although the principles of the present disclosure may not be so limited. The embodiments of the present disclosure also apply to notebook computers, tablet computers, media players, and / or dedicated video conferencing equipment. The network (250) represents any number of networks, including wired and / or wireless communication networks, that convey encoded video data between the terminals (210-240). The communication network (250) may exchange data over circuit-switched and / or packet-switched channels. Representative networks include telecommunications networks, local area networks, wide area networks, and / or the Internet. For purposes of this discussion, the architecture and topology of the network (250) may be irrelevant to the operation of the present disclosure, unless otherwise described below.

[0021] Figure 3 shows the arrangement of video encoders and decoders in a streaming environment as an example application of the disclosed subject matter, which is equally applicable to other video uses including, for example, video conferencing, digital television, and storage of compressed video on digital media such as CDs, DVDs, memory sticks, etc.

[0022] The streaming system may include a capture subsystem (313), which may include, for example, a video source (301), such as a digital camera, that creates an uncompressed video sample stream (302). The sample stream (302), shown in bold to emphasize its large amount of data compared to an encoded video bitstream, may be processed by an encoder (303) coupled to the camera (301). As described in more detail below, the encoder (303) may include hardware, software, or a combination thereof to enable or implement aspects of the disclosed subject matter. The encoded video bitstream (304), shown in thin to emphasize its small amount of data compared to the sample stream, may be stored on a streaming server (305) for later use. One or more streaming clients (306, 308) may access the streaming server (305) to retrieve copies (307, 309) of the encoded video bitstream (304). The client (306) may include a video decoder (310) that decodes a copy (307) of the incoming encoded video bitstream to generate an outgoing video sample stream (311) that may be displayed on a display device (312) or other display device (not shown). In some streaming systems, the video bitstreams (304, 307, 309) may be encoded according to some video encoding / compression standard. Examples of such standards include ITU-T Recommendation H.265. A video encoding standard, informally known as Versatile Video Coding (VVC), is under development. The disclosed subject matter may be used in conjunction with VVC.

[0023] FIG. 4 is a functional block diagram of a video decoder (310) according to an embodiment of the present disclosure.

[0024] The receiver (410) may receive one or more codec video sequences to be decoded by the decoder (310), or in the same or another embodiment, may receive one coded video sequence at a time, with the decoding of each coded video sequence being independent of the other coded video sequences. The coded video sequences may be received from a channel (412), which may be hardware / software coupled with a storage device that stores the coded video data. The receiver (410) may receive the coded video data along with other data, such as coded audio data and / or auxiliary data streams, which may be forwarded to respective use entities (not shown). The receiver (410) may separate the coded video sequences from other data. To combat network jitter, a buffer memory (415) may be coupled between the receiver (410) and the entropy decoder / parser (420) (hereinafter the "parser"). When the receiver 410 is receiving data from a storage / transmission device with sufficient bandwidth and controllability, or from an isosynchronous network, the buffer 415 may not be needed or may be small. The buffer 415 may be needed for use with best-effort packet networks such as the Internet, and may be relatively large and preferably of an adaptive size.

[0025] The video decoder (310) may include a parser (420) to reconstruct symbols (421) from the entropy coded video sequence. As shown in FIG. 3, such classification of symbols includes information used to manage the operation of the decoder (310) and potential information for controlling a display device, such as a display device (312) that is not an integral part of the decoder but may be coupled to it. The control information for the display device(s) may be in the form of Supplementary Enhancement Information (SEI) messages or Video Usability Information (VUI) parameter set fragments (not shown). The parser (420) may parse / entropy decode the received coded video sequence. The coding of the coded video sequence may be in accordance with a video coding technique or standard and may be in accordance with principles well known to those skilled in the art, including variable length coding, Huffman coding, context-dependent or non-context-dependent arithmetic coding, etc. The parser (420) may extract a set of subgroup parameters for at least one of the subgroups of pixels in the video decoder based on at least one parameter corresponding to the group from the coded video sequence. The subgroup may include a group of pictures (GOP), a picture, a subpicture, a tile, a slice, a brick, a macroblock, a coding tree unit (CTU), a coding unit (CU), a block, a transform unit (TU), a prediction unit (PU), etc. A tile may refer to a rectangular region of a CU / CTU in a particular tile column and row in an image. A brick may refer to a rectangular region of a CU / CTU row in a particular tile. A slice may refer to one or more bricks of an image, which are contained in a NAL unit. A subpicture may refer to a rectangular region of one or more slices in an image. The entropy decoder / parser may also extract information such as transform coefficients, quantizer parameter values, motion vectors, etc. from the coded video sequence.

[0026] The parser (420) may perform entropy decoding / parsing operations on the video sequence received from the buffer (415) to generate symbols (421).

[0027] The reconstruction of the symbols (421) can include a number of different units depending on the type of coded video or portions thereof (e.g., inter and intra pictures, inter and intra blocks) and other factors. Which units are included and how can be controlled by subgroup control information parsed from the coded video sequence by the parser (420). The flow of such subgroup control information between the parser (420) and the following units is not shown for clarity.

[0028] In addition to the functional blocks already mentioned, the decoder 310 may be conceptually subdivided into a number of functional units, as described below. In an actual implementation operating under commercial constraints, many of such units may interact closely with each other and may be at least partially integrated with each other. However, for purposes of describing the disclosed subject matter, a conceptual subdivision into the following functional units is appropriate:

[0029] The first unit is a scalar / inverse transform unit (451), which receives quantized transform coefficients as well as control information, including the transform to use, block size, quantization factor, quantization scaling matrix, etc., as symbol(s) (421) from the parser (420). The scalar / inverse transform unit (451) can output a block containing sample values, which can be input to an aggregator (455).

[0030] In some cases, the output samples of the scalar / inverse transform (451) may relate to intra-coded blocks, i.e., blocks that do not use prediction information from a previously reconstructed image may use prediction information from a previously reconstructed part of the current image. Such prediction information may be provided by an intra-image prediction unit (452). In some cases, the intra-image prediction unit (452) uses surrounding already reconstructed information taken from the current (partially reconstructed) image (458) to generate blocks of the same size and shape as the block being reconstructed. The aggregation device (455) may add the prediction information generated by the intra-prediction unit (452) on a sample-by-sample basis to the output sample information provided by the scalar / inverse transform unit (451).

[0031] In other cases, the output samples of the scalar / inverse transform unit (451) may relate to an inter-coded and potentially motion-compensated block. In such cases, the motion compensation prediction unit (453) may access the reference picture memory (457) to retrieve samples to use for prediction. After motion compensating the retrieved samples according to the symbols (421) associated with the block, these samples may be added to the output of the scalar / inverse transform unit (called residual samples or residual signals in this case) by the aggregation device (455) to generate output sample information. The addresses in the reference picture memory from which the motion compensation unit retrieves the prediction samples may be controlled by a motion vector, available to the motion compensation unit in the form of a symbol (421), and may have, for example, X, Y, and reference picture components. Motion compensation may also include interpolation of sample values ​​retrieved from the reference picture memory when sub-sample accurate motion vectors are used, motion vector prediction mechanisms, etc.

[0032] The output samples of the aggregator (455) can be subjected to various loop filtering techniques in a loop filter unit (456). Video compression techniques can include in-loop filter techniques controlled by parameters contained in the encoded video bitstream and made available to the loop filter unit (456) as symbols (421) from the parser (420), but can also be responsive to meta-information obtained during decoding of previous (in decoding order) portions of the encoded image or encoded video sequence, as well as to previously reconstructed loop filtered sample values.

[0033] The output of the loop filter unit (456) may be a sample stream that can be output to a display device (312) and / or stored in a reference picture memory for use in future inter-picture prediction.

[0034] Some coded pictures, once fully reconstructed, can be used as reference pictures for future predictions. Once a coded picture is fully reconstructed and the coded picture has been identified as a reference picture (e.g., by the parser (420)), the current reference picture (458) can become part of the reference picture buffer (457) and new current picture memory can be reallocated before starting the reconstruction of the subsequent coded picture.

[0035] The video decoder 420 may perform decoding operations according to a given video compression technique, which may be described in a standard such as ITU-T Rec. H.265. An encoded video sequence is said to comply with the syntax specified by the video compression technique or standard used in the sense that it adheres to the syntax of the video compression technique or standard, as specified in the video compression technique document or standard, and specifically in the profiles described therein. A further requirement for compliance may be that the complexity of the encoded video sequence be within a range prescribed by the level of the video compression technique or standard. In some cases, the level limits the maximum picture size, maximum frame rate, maximum reconstruction sample rate (e.g., measured in megasamples / second), maximum reference picture size, etc. The limits set by the level may be further limited in some cases by a Hypothetical Reference Decoder (HRD) specification and HRD buffer management metadata signaled in the encoded video sequence.

[0036] In an embodiment, the receiver (410) may receive additional (redundant) data along with the encoded video. The additional data may be included as part of the encoded video sequence(s). The additional data may be used by the video decoder (420) to properly decode the data and / or to more accurately reconstruct the original video data. The additional data may be in the form of, for example, temporal, spatial, or SNR enhancement layers, redundant slices, redundant images, forward error correction codes, etc.

[0037] FIG. 5 is a functional block diagram of a video encoder (303) according to an embodiment of the present disclosure.

[0038] The encoder (303) may receive video samples from a video source (301) (not part of the encoder) that may capture the video(s) to be encoded by the encoder (303).

[0039] The video source (301) may provide a source video sequence to be encoded by the encoder (303) in the form of a digital video sample stream that may be of any suitable bit depth (8-bit, 10-bit, 12-bit, etc.), any color space (BT.601 Y CrCB, RGB, etc.), and any suitable sampling structure (Y CrCb 4:2:0, Y CrCb 4:4:4, etc.). In a media delivery system, the video source (301) may be a storage device that stores previously prepared video. In a video conferencing system, the video source (303) may be a camera that captures local image information as a video sequence. The video data may be provided as a number of individual images that convey motion when viewed in sequence. The images themselves may be organized as a spatial array of pixels, each of which may contain one or more samples depending on the sampling structure, color space, etc., in use. Those skilled in the art will readily appreciate the relationship between pixels and samples. The following description focuses on samples.

[0040] According to an embodiment, the encoder (303) may encode and compress images of a source video sequence into an encoded video sequence (543) in real time or under other time constraints required by the application. Providing an appropriate encoding rate is one function of the controller (550). The controller controls and is operatively coupled to other functional units as described below. For clarity, coupling is not shown. Parameters set by the controller may include rate control related parameters (e.g., picture skip, quantization, lambda values ​​for rate-distortion optimization techniques), picture size, group of pictures (GOP) layout, maximum motion vector search range, etc. Those skilled in the art can easily identify other functions of the controller (550) as they may be relevant to a video encoder (303) optimized for some system designs.

[0041] Some video encoders operate in what is easily recognized by those skilled in the art as a "coding loop." In an oversimplified explanation, the coding loop may consist of an encoding section of the encoder (530) (hereafter "source encoder") (responsible for generating symbols based on the input image to be encoded and the reference image(s)), and a (local) decoder (533) embedded in the encoder (303) that reconstructs the symbols to generate sample data that the (remote) decoder also generates (since the compression between the symbols and the encoded video bitstream is lossless in the video compression techniques contemplated in the disclosed subject matter). The reconstructed sample stream is input to a reference image memory (534). As the decoding of the symbol stream results in bit-exact, regardless of the location of the decoder (local or remote), the contents of the reference image buffer are also bit-perfect between the local and remote encoders. In other words, the predictor section of the encoder "sees" the reference image samples as exactly the same sample values ​​that the decoder "sees" when using prediction during decoding. This basic principle of reference image synchronicity (and the resulting drift if synchronicity cannot be maintained, for example due to channel errors) is well known to those skilled in the art.

[0042] The operation of the "local" decoder (533) may be the same as the "remote" decoder (310), which has already been described in detail above in relation to Figure 4. However, referring also momentarily to Figure 4, because symbols are available and can be losslessly encoded / decoded into an encoded video sequence by the entropy encoder (545) and parser (420), the entropy decoding portion of the decoder (310), including the channel (412), receiver (410), buffer (415), and parser (420), need not be entirely implemented in the local decoder (533).

[0043] It is currently believed that any decoder techniques, other than parsing / entropy decoding, present in the decoder will naturally need to be present in the corresponding encoder in approximately the same functional form. For this reason, the disclosed subject matter focuses on the operation of the decoder. A description of the encoder techniques can be omitted, since they are the inverse of the decoder techniques, which are described generically. Only in a few areas are more detailed descriptions required and are described below.

[0044] The source encoder (530) may perform motion compensated predictive encoding as part of its operation, predictively encoding an input frame with respect to one or more previously encoded frames from the video sequence designated as “reference frames.” In this method, the encoding engine (532) encodes differences between pixel blocks of the input frame and pixel blocks of reference frame(s) that may be selected as predictive reference(s) for the input frame.

[0045] The local video decoder (533) may decode the encoded video data of frames that may be designated as reference frames based on symbols generated by the source encoder (530). The operation of the encoding engine (532) may preferably be a lossy process. When the encoded video data may be decoded by a video decoder (not shown in FIG. 5), the reconstructed video sequence may be a copy of the source video sequence, usually with some errors. The local video decoder (533) replicates the decoding process that may be performed on the reference frames by the video decoder and may cause the reconstructed reference frames to be stored in a reference picture cache (534). In this manner, the encoder (303) may locally store copies of reconstructed reference frames that have common content with the reconstructed reference frames obtained by the far-end video decoder (free of transmission errors).

[0046] The predictor (535) may perform a prediction search for the coding engine (532). That is, for a new frame to be coded, the predictor (535) may search for sample data (as candidate reference pixel blocks) from the reference image memory (534) or some metadata such as reference image motion vectors, block shapes, etc., which serve as suitable prediction references for the new image. The predictor (535) may operate on a sample by sample basis to find suitable prediction references. In some cases, as determined by the search results obtained by the predictor (535), the input image may have prediction references drawn from multiple reference images stored in the reference image memory (534).

[0047] The controller (550) may manage the encoding operations of the video encoder (530), including, for example, setting parameters and subgroup parameters used to encode the video data.

[0048] The output of all the aforementioned functional units may be entropy coded in an entropy encoder (545) that converts the symbols produced by the various functional units into an encoded video sequence by losslessly compressing the symbols with techniques known to those skilled in the art as Huffman coding, variable length coding, arithmetic coding, etc.

[0049] The transmitter (540) may buffer the encoded video sequence(s) as they are generated by the entropy encoder (545) in preparation for transmission over the communication channel (560), which may be a hardware / software association with a storage device that stores the encoded video data. The transmitter (540) may merge the encoded video data of the video encoder (530) with other data to be transmitted, such as encoded audio data and / or auxiliary data streams (sources not shown).

[0050] A controller (550) may manage the operation of the encoder (303). During encoding, the controller (550) may assign several encoding image types to each of the encoded images, which may affect the encoding technique that may be applied to each image. For example, images are often assigned to one of the following frame types:

[0051] An Intra picture (I-picture) is one that can be coded and decoded without using other frames in a sequence as a source of prediction. Some video codecs allow different kinds of Intra pictures, including, for example, Independent Decoder Refresh pictures. Those skilled in the art are aware of such variations of I-pictures, as well as their respective uses and characteristics.

[0052] A predicted image (P picture) can be coded and decoded using intra- or inter-prediction, using at most one motion vector and reference index to predict the sample values ​​of each block.

[0053] Bidirectionally predicted images (B-pictures) are those that can be coded and decoded using intra- or inter-prediction, using up to two motion vectors and reference indices to predict the sample values ​​of each block. Similarly, multi-predictive images can use more than two reference images and associated metadata to reconstruct a block.

[0054] A source image may be coded block by block, usually spatially subdivided into multiple sample blocks (e.g., blocks of 4x4, 8x8, 4x8, or 16x16 samples each). Blocks may be predictively coded with reference to other (already coded) blocks as determined by the code assignment applied to each image of the block. For example, blocks of I-pictures may be non-predictively coded or predictively coded with reference to already coded blocks of the same image (spatial or intra prediction). Pixel blocks of P-pictures may be non-predictively coded by spatial prediction with reference to one previously coded reference image or by temporal prediction. Blocks of B-pictures may be non-predictively coded by spatial prediction with reference to one or two previously coded reference images or by temporal prediction.

[0055] The video encoder (303) may perform encoding operations in accordance with a given video encoding technique or standard, such as ITU-T Rec. H.265. During its operation, the video encoder (303) may perform various compression operations, including predictive encoding operations that exploit temporal and spatial redundancy in the input video sequence. The encoded video data may therefore conform to a syntax specified by the video encoding technique or standard used.

[0056] In an embodiment, the transmitter (540) may transmit additional data along with the encoded video. The video encoder (530) may include such data as part of the encoded video sequence. The additional data may include other forms of redundant data, such as temporal / spatial / SNR enhancement layers, redundant pictures and slices, Supplementary Enhancement Information (SEI) messages, Visual Usability Information (VUI) parameter set fragments, etc.

[0057] 6-7, in an embodiment, 360-degree video is captured by a set of cameras, or a camera device with multiple lenses. The cameras may cover omnidirectionally around a center point of the camera set. Images of the same time instance are stitched together, possibly rotated, projected, and mapped onto the image. The packed images are encoded as encoded into the coded video bitstream and delivered according to a specific media container file format. The file includes metadata such as projection and packing information.

[0058] In an embodiment, the 360-degree image may be projected onto the 2D image using equirectangular projection (ERP). ERP projection may cause seam artifacts. The padded ERP (PERP) format may effectively reduce seam artifacts in the reconstructed viewport surrounding the left and right boundaries of the ERP image. However, padding and blending may not be sufficient to completely solve the seam problem.

[0059] In an embodiment, horizontal shape padding may be applied to the ERP or PERP to reduce seam artifacts. The process of padding for PERP may be the same as ERP, except that the offset may be based on the unpadded ERP width rather than the image width to account for the size of the padded region. If a reference block is outside the left (right) reference image boundary, it may be replaced with a "wrap-around" reference block that is shifted right (left) by the ERP width. Vertically, conventional repeated padding may be used. Blending of left and right padded regions does not involve looping as a post-processing operation.

[0060] In an embodiment, syntax such as seq_parameter_set_rbsp( 601 ) that enables horizontal shape padding of reference images for ERP and PERP formats is shown in FIG. 6 .

[0061] In an embodiment, sps_ref_wraparound_enabled_flag (602) specifies that horizontal wraparound motion compensation is used for inter prediction when it is set to 1. In an embodiment, sps_ref_wraparound_enabled_flag (602) specifies that this motion compensation method is not applied when it is set to 0.

[0062] In an embodiment, ref_wraparound_offset (603) specifies the offset of luma samples used to calculate the horizontal wraparound position. In an embodiment, ref_wraparound_offset (603) shall be greater than pic_width_in_luma_samples-1, not greater than pic_width_in_luma_samples, and an integer multiple of MinCbSizeY.

[0063] In an embodiment, syntax such as seq_parameter_set_rbsp( 701 ) that enables horizontal shape padding of reference images for ERP and PERP formats is shown in FIG.

[0064] In an embodiment, sps_ref_wraparound_enabled_flag (702) specifies that horizontal wraparound motion compensation is used for inter prediction when it is set to 1. sps_ref_wraparound_enabled_flag (702) specifies that this motion compensation method is not applied when it is set to 0.

[0065] In an embodiment, left_wraparound_padding_width (703) specifies the width of the left padding area in luma samples. In an embodiment, ref_wraparound_offset shall be greater than or equal to 0, not greater than pic_width_in_luma_samples / 2, and an integer multiple of MinCbSizeY.

[0066] In an embodiment, right_wraparound_padding_width (704) specifies the width of the right padding area in luma samples. In an embodiment, ref_wraparound_offset shall be greater than or equal to 0, not greater than pic_width_in_luma_samples / 2, and an integer multiple of MinCbSizeY.

[0067] In an embodiment, the wraparound offset value may be obtained by the following derivation process: if ref_wraparound_offset is present wrapAroundOffset=ref_wraparound_offset else if left_wraparound_padding_width and right_wraparound_padding_width are present wrapAroundOffset=pic_width_in_luma_samples-(left_wraparound_padding_width+right_wraparound_padding_width) else wrapAroundOffset=pic_width_in_luma_samples

[0068] In an embodiment, the luma and chroma sample interpolation process may be modified to allow for horizontal geometry padding of reference images in ERP and PERP formats.

number

number

[0069] An example of a luma sample interpolation process according to an embodiment and an example of a chroma sample interpolation process according to an embodiment are described below. Luma Sample Interpolation Process The inputs to this process are: -Complete sample unit (xInt L ,yInt L ) Luma position in -Fractional sample units (xFrac L ,yFrac L ) Luma position in - Luma reference sample array refPicLX L The output of this process is the predicted luma sample value predSampleLX. L The variables shift1, shift2, and shift3 are derived as follows: - Variable shift1 is Min(4,BitDepth Y -8), variable shift2 is set to 6, and variable shift3 is set to Max(2,14-BitDepth Y ) The variable picW is set to pic_width_in_luma_samples and the variable picH is set to pic_height_in_luma_samples. - The variable xOffset is set to wrapAroundOffset. xFrac L or yFrac L The luma interpolation filter coefficients f for each 1 / 16 fractional sample position p are equal to L [p] is specified below. Predicted luma sample value predSampleLX L is derived as follows: -xFrac L and yFrac L are both zero, then the following applies: - If sps_ref_wraparound_enabled_flag is 0, predSampleLX L The value of is derived as follows: predSampleLX L =refPicLX L [Clip3(0,picW-1,xIntL )][Clip3(0,picH-1,yInt L )]< <shift3 -or predSampleLX L The value of is derived as follows: predSampleLX L =refPicLX L [ClipH(xOffset,picW,xInt L )][Clip3(0,picH-1,yInt L )]< <shift3 - or xFrac L is not 0, and yFrac L If is zero, the following applies: -yPos L The value of is derived as follows: yPos L =Clip3(0,picH-1,yInt L ) - If sps_ref_wraparound_enabled_flag is 0, predSampleLX L The value of is derived as follows: predSampleLX L =(f L [xFrac L ][0]*refPicLX L [Clip3(0,picW-1,xInt L -3)][yPos L ]+ f L [xFrac L ][1]*refPicLX L [Clip3(0,picW-1,xInt L -2)][yPos L ]+ f L [xFrac L ][2]*refPicLX L [Clip3(0,picW-1,xInt L -1)][yPos L ]+ f L [xFrac L[3]*refPicLX L [Clip3(0,picW-1,xInt L )][yPos L + f L [xFrac L [4]*refPicLX L [Clip3(0,picW-1,xInt L +1)][yPos L + f L [xFrac L [5]*refPicLX L [Clip3(0,picW-1,xInt L +2)][yPos L + f L [xFrac L [6]*refPicLX L [Clip3(0,picW-1,xInt L +3)][yPos L + f L [xFrac L [7]*refPicLX L [Clip3(0,picW-1,xInt L +4)][yPos L )>>shift1 - Or predSampleLX L The value of is derived as follows. predSampleLX L =(f L [xFrac L [0]*refPicLX L [ClipH(xOffset,picW,xInt L -3)][yPos L + f L [xFrac L [1]*refPicLX L [ClipH(xOffset,picW,xInt L -2)][yPos L + f L [xFrac L][2]*refPicLX L [ClipH(xOffset,picW,xInt L -1)][yPos L ]+ f L [xFrac L ][3]*refPicLX L [ClipH(xOffset,picW,xInt L )][yPos L ]+ f L [xFrac L ][4]*refPicLX L [ClipH(xOffset,picW,xInt L +1)][yPos L ]+ f L [xFrac L ][5]*refPicLX L [ClipH(xOffset,picW,xInt L +2)][yPos L ]+ f L [xFrac L ][6]*refPicLX L [ClipH(xOffset,picW,xInt L +3)][yPos L ]+ f L [xFrac L ][7]*refPicLX L [ClipH(xOffset,picW,xInt L +4)][yPos L ])>>shift1 - or xFrac L is 0, and yFrac L If is not 0, predSampleLX L The value of is derived as follows: - If sps_ref_wraparound_enabled_flag is 0, then xPos L The value of is derived as follows: xPos L =Clip3(0,picW-1,xIntL ) -or- xPos L The value of is derived as follows: xPos L =ClipH(xOffset,picW,xInt L ) - predicted luma sample value predSampleLX L is derived as follows: predSampleLX L =(f L [yFrac L ][0]*refPicLX L [xPos L ][Clip3(0,picH-1,yInt L -3)]+ f L [yFrac L ][1]*refPicLX L [xPos L ][Clip3(0,picH-1,yInt L -2)]+ f L [yFrac L ][2]*refPicLX L [xPos L ][Clip3(0,picH-1,yInt L -1)]+ f L [yFrac L ][3]*refPicLX L [xPos L ][Clip3(0,picH-1,yInt L )]+ f L [yFrac L ][4]*refPicLX L [xPos L ][Clip3(0,picH-1,yInt L +1)]+ f L [yFrac L ][5]*refPicLX L [xPos L ][Clip3(0,picH-1,yInt L +2)]+ f L [yFrac L ][6]*refPicLX L [xPos L ][Clip3(0,picH-1,yInt L +3)]+ f L [yFrac L ][7]*refPicLX L [xPos L ][Clip3(0,picH-1,yInt L +4)])>>shift1 - or xFrac L is not 0, and yFrac L If is not 0, predSampleLX L The value of is derived as follows: If -sps_ref_wraparound_enabled_flag is 0, the sample array temp[n] for n=0 to 7 is derived as follows: yPos L =Clip3(0,picH-1,yInt L +n-3) temp[n]=(f L [xFrac L ][0]*refPicLX L [Clip3(0,picW-1,xInt L -3)][yPos L ]+ f L [xFrac L ][1]*refPicLX L [Clip3(0,picW-1,xInt L -2)][yPos L ]+ f L [xFrac L ][2]*refPicLX L [Clip3(0,picW-1,xInt L -1)][yPos L ]+ f L [xFrac L ][3]*refPicLX L[Clip3(0,picW-1,xInt L )][yPos L ]+ f L [xFrac L ][4]*refPicLX L [Clip3(0,picW-1,xInt L +1)][yPos L ]+ f L [xFrac L ][5]*refPicLX L [Clip3(0,picW-1,xInt L +2)][yPos L ]+ f L [xFrac L ][6]*refPicLX L [Clip3(0,picW-1,xInt L +3)][yPos L ]+ f L [xFrac L ][7]*refPicLX L [Clip3(0,picW-1,xInt L +4)][yPos L ])>>shift1 Alternatively, the sample array temp[n] for n=0 to 7 is derived as follows: yPos L =Clip3(0,picH-1,yInt L +n-3) temp[n]=(f L [xFrac L ][0]*refPicLX L [ClipH(xOffset,picW,xInt L -3)][yPos L ]+ f L [xFrac L ][1]*refPicLX L [ClipH(xOffset,picW,xInt L -2)][yPos L ]+ f L[xFrac L ][2]*refPicLX L [ClipH(xOffset,picW,xInt L -1)][yPos L ]+ f L [xFrac L ][3]*refPicLX L [ClipH(xOffset,picW,xInt L )][yPos L ]+ f L [xFrac L ][4]*refPicLX L [ClipH(xOffset,picW,xInt L +1)][yPos L ]+ f L [xFrac L ][5]*refPicLX L [ClipH(xOffset,picW,xInt L +2)][yPos L ]+ f L [xFrac L ][6]*refPicLX L [ClipH(xOffset,picW,xInt L +3)][yPos L ]+ f L [xFrac L ][7]*refPicLX L [ClipH(xOffset,picW,xInt L +4)][yPos L ])>>shift1 - predicted luma sample value predSampleLX L is derived as follows: predSampleLX L =(f L [yFrac L ][0]*temp[0]+ f L [yFrac L ][1]*temp[1]+ fL [yFrac L ][2]*temp[2]+ f L [yFrac L ][3]*temp[3]+ f L [yFrac L ][4]*temp[4]+ f L [yFrac L ][5]*temp[5]+ f L [yFrac L ][6]*temp[6]+ f L [yFrac L ][7]*temp[7])>>shift2 Chroma Sample Interpolation Process The inputs to this process are: -Complete sample unit (xInt C ,yInt C ) Chroma position -1 / 32 fractional sample unit (xFrac C ,yFrac C ) Chroma position - Chroma reference sample array refPicLX C It is. The output of this process is the predicted chroma sample value, predSampleLX. C The variables shift1, shift2, and shift3 are derived as follows: - Variable shift1 is Min(4,BitDepth C -8), variable shift2 is set to 6, and variable shift3 is set to Max(2,14-BitDepth C ) -Variable picW C is set to pic_width_in_luma_samples / SubWidthC, and the variable picH C is set to pic_height_in_luma_samples / SubHeightC. -Variable xOffset C is set to wrapAroundOffset / SubWidthC. xFrac C or yFrac C The luma interpolation filter coefficients f for each 1 / 32 fractional sample position p are equal to C [p] is specified below. Predicted chroma sample value predSampleLX C is derived as follows: -xFrac C and yFrac C are both zero, then the following applies: - If sps_ref_wraparound_enabled_flag is 0, predSampleLX C The value of is derived as follows: predSampleLX C =refPicLX C [Clip3(0,picW C -1,xInt C )][Clip3(0,picH C -1,yInt C )]< <shift3 -or predSampleLX C The value of is derived as follows: predSampleLX C =refPicLX C [ClipH(xOffset C ,picW C ,xInt C )][Clip3(0,picH C -1,yInt C )]< <shift3 - or xFrac C is not 0, and yFrac C If is 0 then the following applies: -yPos C The value of is derived as follows: yPos C =Clip3(0,picH C -1,yInt C ) - If sps_ref_wraparound_enabled_flag is 0, predSampleLX C The value of is derived as follows: predSampleLX C =(f C [xFrac C ][0]*refPicLX C [Clip3(0,picW C -1,xInt C -1)][yInt C ]+ f C [xFrac C ][1]*refPicLX C [Clip3(0,picW C -1,xInt C )][yInt C ]+ f C [xFrac C ][2]*refPicLX C [Clip3(0,picW C -1,xInt C +1)][yInt C ]+ f C [xFrac C ][3]*refPicLX C [Clip3(0,picW C -1,xInt C +2)][yInt C ])>>shift1 -or predSampleLX C The value of is derived as follows: predSampleLX C =(f C [xFrac C ][0]*refPicLX C [ClipH(xOffset C ,picW C ,xInt C -1)][yPos C ]+ f C [xFrac C ][1]*refPicLXC [ClipH(xOffset C ,picW C ,xInt C )][yPos C ]+ f C [xFrac C ][2]*refPicLX C [ClipH(xOffset C ,picW C ,xInt C +1)][yPos C ]+ f C [xFrac C ][3]*refPicLX C [ClipH(xOffset C ,picW C ,xInt C +2)][yPos C ])>>shift1 - or xFrac C is 0, and yFrac C If predSampleLX is not 0, C The value of is derived as follows: - If sps_ref_wraparound_enabled_flag is 0, then xPos C The value of is derived as follows: xPos C =Clip3(0,picW C -1,xInt C ) -or- xPos C The value of is derived as follows: xPos C =ClipH(xOffset C ,picW C ,xInt C ) -Predicted chroma sample value predSampleLX C is derived as follows: predSampleLX C =(f C [yFrac C ][0]*refPicLX C [xPosC ][Clip3(0,picH C -1,yInt C -1)]+ f C [yFrac C ][1]*refPicLX C [xPos C ][Clip3(0,picH C -1,yInt C )]+ f C [yFrac C ][2]*refPicLX C [xPos C ][Clip3(0,picH C -1,yInt C +1)]+ f C [yFrac C ][3]*refPicLX C [xPos C ][Clip3(0,picH C -1,yInt C +2)])>>shift1 - or xFrac C is not 0 and yFrac C If predSampleLX is not 0, C The value of is derived as follows: If -sps_ref_wraparound_enabled_flag is 0, the sample array temp[n] for n=0 to 3 is derived as follows: yPos C =Clip3(0,picH C -1,yInt C +n-1) temp[n]=(f C [xFrac C ][0]*refPicLX C [Clip3(0,picW C -1,xInt C -1)][yPos C ]+ f C [xFrac C ][1]*refPicLX C[Clip3(0,picW C -1,xInt C )][yPos C ]+ f C [xFrac C ][2]*refPicLX C [Clip3(0,picW C -1,xInt C +1)][yPos C ]+ f C [xFrac C ][3]*refPicLX C [Clip3(0,picW C -1,xInt C +2)][yPos C ])>>shift1 Alternatively, the sample array temp[n] for n=0 to 3 is derived as follows: yPos C =Clip3(0,picH C -1,yInt C +n-1) temp[n]=(f C [xFrac C ][0]*refPicLX C [ClipH(xOffset C ,picW C ,xInt C -1)][yPos C ]+ f C [xFrac C ][1]*refPicLX C [ClipH(xOffset C ,picW C ,xInt C )][yPos C ]+ f C [xFrac C ][2]*refPicLX C [ClipH(xOffset C ,picW C ,xInt C +1)][yPos C ]+ f C[xFrac C ][3]*refPicLX C [ClipH(xOffset C ,picW C ,xInt C +2)][yPos C ])>>shift1 -Predicted chroma sample value predSampleLX C is derived as follows: predSampleLX C =(f C [yFrac C ][0]*temp[0]+ f C [yFrac C ][1]*temp[1]+ f C [yFrac C ][2]*temp[2]+ f C [yFrac C ][3]*temp[3])>>shift2

[0070] In an embodiment, if sps_ref_wraparound_enabled_flag (601) is 0 or is not present, conventional repeat padding may be applied. Alternatively, wraparound padding may be applied.

[0071] In embodiments, wrap-around padding may be applied to both horizontal and vertical boundaries. A flag in a high-level syntax structure may indicate that wrap-around padding has been applied both horizontally and vertically.

[0072] In embodiments, wrap-around padding may be applied at brick, tile, slice, or sub-image boundaries. In embodiments, wrap-around padding may be applied at tile group boundaries. A flag in the high-level syntax structure may indicate that wrap-around padding is applied both horizontally and vertically. In an embodiment, the reference picture may be the same as the current picture for motion compensated prediction. When the current picture is the reference picture, wrap-around padding may be applied to the boundaries of the current picture.

[0073] Figure 8 is a flowchart of an example process 800 for generating a merge candidate list using intermediate candidates. In some implementations, one or more process blocks of Figure 8 may be performed by the decoder 310. In some implementations, one or more process blocks of Figure 8 may be performed by a device separate from the decoder 310, such as the encoder 303, or a collection of devices that includes the decoder 310.

[0074] As shown in FIG. 8, process 800 may include decoding image segmentation information corresponding to a current image (block 810).

[0075] As further shown in FIG. 8, process 800 may include using the image segmentation information to determine whether padding is applied to multiple sub-regions of the current image (block 820).

[0076] 8, based on a determination that padding is not applied, process 800 may include decoding the sub-regions without padding the sub-regions (block 830). Process 800 may then proceed to reconstructing the current image based on the decoded sub-regions (block 870).

[0077] As further shown in FIG. 8, process 800 may include, based on a determination that padding is applied, determining whether the padding includes wrap-around padding (block 840) using the image segmentation information.

[0078] 8, based on a determination that the padding does not include wrap-around padding, process 800 may include applying repeated padding to the multiple sub-regions and decoding the multiple sub-regions using the repeated padding (block 850). Process 800 may then proceed to reconstructing the current image based on the decoded multiple sub-regions (block 870).

[0079] 8, process 800 may include applying wrap-around padding to the multiple sub-regions based on a determination that the padding includes wrap-around padding and decoding the multiple sub-regions using the wrap-around padding (block 860). Process 800 may then proceed to reconstructing the current image based on the decoded multiple sub-regions (block 870).

[0080] In an embodiment, the image segmentation information may be included in the image parameter set corresponding to the current image.

[0081] In an embodiment, the image segmentation information includes at least one flag included in the image parameter set.

[0082] In an embodiment, the multiple sub-regions include at least one of: bricks, tiles, slices, tile groups, sub-images, or sub-layers.

[0083] In an embodiment, padding may be applied to boundaries of sub-regions among multiple sub-regions.

[0084] In an embodiment, the boundary may be a vertical boundary of the sub-region.

[0085] In an embodiment, the boundary may be a horizontal boundary of the sub-region.

[0086] In an embodiment, padding may be applied to vertical boundaries of a sub-region of the plurality of sub-regions and to horizontal boundaries of the sub-regions.

[0087] In an embodiment, the image division information may indicate an offset value for the wrap-around padding.

[0088] In an embodiment, the image division information may indicate left padding width information and right padding width information.

[0089] Although Figure 8 illustrates example blocks of process 800, in some implementations, process 800 may include additional blocks, fewer blocks, different blocks, or blocks arranged differently than those illustrated in Figure 8. Additionally or alternatively, two or more blocks of process 800 may be performed in parallel.

[0090] The proposed methods may also be implemented by a processing circuit (e.g., one or more processors or one or more integrated circuits). In one example, the one or more processors execute a program stored on a non-transitory computer-readable medium to perform one or more of the proposed methods.

[0091] The techniques described above can be implemented as computer software using computer readable instructions and physically stored on one or more computer readable media. For example, Figure 9 illustrates a computer system 900 suitable for implementing some embodiments of the disclosed subject matter.

[0092] Computer software can be encoded using any suitable machine code or computer language, which may follow mechanisms such as assembly, compilation, linking, etc. to produce code containing instructions that can be executed by a computer central processing unit (CPU), graphics processing unit (GPU), etc., directly, or via interpretation, microcode execution, etc.

[0093] The instructions may be executed in various types of computers or components thereof, such as personal computers, tablet computers, servers, smartphones, gaming consoles, and Internet of Things (IoT) devices.

[0094] 9 are exemplary in nature and are not intended to suggest any limitation to the scope of use or functionality of the computer software implementing the embodiments of the present disclosure. The arrangement of components should not be interpreted as having any dependency or requirement regarding any one or combination of components illustrated in the exemplary embodiment of computer system 900.

[0095] The computer system 900 may include several human interface input devices. Such human interface input devices may be responsive to input by one or more users, for example, by tactile input (pressing a key, swiping, moving a data glove, etc.), audio input (voice, clapping hands, etc.), visual input (gestures, etc.), olfactory input (not shown). The human interface devices may further be used to capture several media that do not necessarily involve direct conscious input by a person, such as audio (speech, music, ambient sounds, etc.), images (scanned images, photographic images captured by a still camera, etc.), video (two-dimensional video, three-dimensional video including stereoscopic video, etc.).

[0096] The input human interface devices may include one or more of a keyboard 901, a mouse 902, a trackpad 903, a touch screen 910, an associated graphics adapter 950, a data glove 1204, a joystick 905, a microphone 906, a scanner 907, and a camera 908 (only one of each is shown).

[0097] The computer system 900 may include a number of human interface output devices. Such human interface output devices may stimulate one or more of the user's senses, such as haptic output, sound, light, and smell / taste. Such human interface output devices may include haptic output devices (e.g., haptic feedback via a touch screen 910, data gloves 1204, or joystick 905, although there may also be haptic feedback devices that do not function as input devices), audio output devices (speakers 909, headphones (not shown)), visual output devices (such as screens 910 including cathode ray tube (CRT) screens, liquid crystal display (LCD) screens, plasma screens, organic light emitting diode (OLED) screens, each with or without touch screen input capability, each with or without haptic feedback capability, some of which may be capable of outputting output in more than three dimensions by means of two-dimensional video output, or stereoscopic output, etc., VR glasses (not shown), holographic displays, smoke tanks (not shown), and printers (not shown)).

[0098] The computer system 900 may further include human-accessible storage devices and associated media therewith, such as CD / DVD ROM / RW 920, including media 921 such as CDs / DVDs, thumb-drives 922, removable hard drives or solid-state drives 923, legacy magnetic media such as tapes and floppy disks (not shown), optical media including specialized ROM / ASIC / PLD based devices (not shown) such as security dongles.

[0099] Those skilled in the art should further appreciate that the term "computer-readable medium" as used with respect to the subject matter disclosed herein does not include transmission media, carrier waves or other transitory signals.

[0100] The computer system 900 may further comprise an interface to one or more communication networks 955. The networks may be, for example, wireless, wired, optical. The networks may further be local, wide area, metropolitan, in-vehicle and industrial, real-time, delay tolerant, etc. Examples of networks include local area networks such as Ethernet, wireless LANs, mobile communication networks including Global System for Mobile Communications (GSM), third generation (3G), fourth generation (4G), fifth generation (5G), long-term evolution (LTE), etc., wired or wireless wide area digital networks including cable television, satellite television, and terrestrial television, in-vehicle and industrial networks including CANBus, etc. Some networks typically require an external network interface adapter (954) attached to some general-purpose data port or peripheral bus (949) (e.g., a Universal Serial Bus (USB) port of the computer system 900), while other networks are typically integrated into the core of the computer system 900 by attaching to the system bus (e.g., an Ethernet interface to a PC computer system, or a mobile communication network interface to a smartphone computer system) as described below. As an example, a network 955 may be connected to the peripheral bus 949 using a network interface 954. Using any such network, the computer system 900 may communicate with other entities. Such communication may be one-way communication, receive-only communication (e.g., a television broadcast), one-way transmit-only communication (e.g., a device transmitting from a CANbus to a CANbus), or bidirectional communication to other computer systems, for example, using a local or wide area digital network. As previously described, several protocols and protocol stacks may be used for each such network and each network interface (954).

[0101] The human interface devices, human accessible storage, and network interfaces described above may be attached to a core 940 of the computer system 900 .

[0102] The core 940 may include one or more central processing units (CPU) 941, graphic processing units (GPU) 942, specialized programmable processing units in the form of FPGAs (Field Programmable Gate Areas) 943, hardware accelerators for some tasks 944, etc. Such devices may be connected via a system bus 1248, along with read only memory (ROM) 945, random access memory (RAM) 946, internal mass storage such as an internal non-user accessible hard drive, solid state drive (SSD) 947, etc. In some computer systems, the system bus 1248 may be accessible in the form of one or more physical plugs to allow expansion with additional CPUs, GPUs, etc. Peripheral devices may be attached directly to the core's system bus 1248 or via a peripheral bus 949. Peripheral bus architectures include peripheral component interconnect (PCI), USB, etc.

[0103] The CPU 941, GPU 942, FPGA 943, and accelerator 944 may execute a combination of instructions that may make up the computer code described above. The computer code may be stored in ROM 945 or RAM 946. Transient data may also be stored in RAM 946, whereas persistent data may be stored, for example, in internal mass storage device 947. A cache memory may be used to allow quick storage and retrieval of any memory device, which may be closely associated with one or more of the CPU 941, GPU 942, mass storage device 947, ROM 945, RAM 946, etc.

[0104] The computer-readable medium can bear computer code for performing various computer-implemented operations. The medium and computer code can be those specially designed and constructed for the purposes of the present disclosure, or they can be of the kind well known and available to those skilled in the computer software arts.

[0105] By way of example and not of limitation, a computer system having the architecture 900, and in particular the core 940, can provide functionality as a result of processor(s) (CPU, GPU, FPGA, accelerator, etc.) executing software embodied in one or more tangible computer-readable media. Such computer-readable media may be user-accessible mass storage devices as introduced above, as well as media associated with some storage devices of the core 940 that are non-transitory in nature, such as the core internal mass storage device 947 or ROM 945. Software implementing various embodiments of the present disclosure can be stored in such devices and executed by the core 940. The computer-readable media can include one or more memory devices or chips according to particular needs. The software can cause the core 940, and in particular the processors therein (including the CPU, GPU, FPGA, etc.) to perform certain operations, or certain portions of certain operations, described herein, including the definition of data structures stored in the RAM 946 and the modification of such data structures according to the software-defined operations. Additionally or alternatively, the computer system may provide functionality as a result of logic hardwired into or embodied in circuitry (e.g., accelerator 944), which may operate in place of or in conjunction with software to perform certain operations, or portions of certain operations, described herein. References to software may encompass logic, where appropriate, and vice versa. References to computer-readable media may encompass circuitry (such as integrated circuits (ICs)) that store software for execution, where appropriate, circuitry embodying logic for execution, or both. The present disclosure encompasses any suitable combination of hardware and software.

[0106] While this disclosure has described several exemplary embodiments, there are alterations, substitutions, and various substitute equivalents, which are within the scope of this disclosure. Thus, it will be appreciated that those skilled in the art will be able to devise many systems and methods that embody the principles of the present disclosure and are therefore within its scope, even if not explicitly shown or described herein. [Explanation of symbols]

[0107] 101 H.264 NAL unit header parsing diagram 102 H.265 NAL unit header parsing diagram 200 Communication Systems 210 First Terminal 220 Second Terminal 230 Terminals 240 Terminals 250 Communication Network 301 Video source (camera) 302 Video Sample Stream 303 Encoder 304 Video Bitstream 305 Streaming Server 306 Streaming Client 307 Copy of video bitstream 308 Streaming Client 309 Copy of video bitstream 310 Decoder 311 Video Sample Stream 312 Display device 313 Acquisition Subsystem 410 Receiver 412 Channels 415 Buffer Memory 420 Parser (Video Decoder) 421 Symbols 451 Scaler / Descaler Unit 452 Intra-Image Prediction Unit 453 Motion Compensation Prediction Unit 455 Aggregation Device 456 Loop Filter Unit 457 Reference Image Buffer, Reference Image Memory 458 Current Reference Images 530 Source Encoder, Video Encoder 532 encoding engine 533 Local Video Decoder 534 Reference Image Memory, Reference Image Cache 535 Predictors 540 Transmitter 543 Video Sequence 545 Entropy Encoder 550 Controller 560 Communication Channels 601 seq_parameter_set_rbsp() 602 sps_ref_wraparound_enabled_flag 603 ref_wraparound_offset 701 seq_parameter_set_rbsp() 702 sps_ref_wraparound_enabled_flag 703 left_wraparound_padding_width 704 right_wraparound_padding_width 800 processes 900 Computer Systems 901 Keyboard 902 Mouse 903 Trackpad 905 Joystick 906 Mike 907 Scanner 908 Camera 909 Audio output device 910 Touchscreen 920 Optical media 921 CD / DVD and other media 922 USB memory 923 Solid State Drive 940 cores 941 Central Processing Unit (CPU) 942 Graphics Processing Unit (GPU) 943 FPGA 944 Accelerator 945 Read-Only Memory (ROM) 946 Random Access Memory (RAM) 947 Mass storage device 949 Surrounding Bus 950 Graphics Adapter 954 Network Interface 955 Communication Network 1248 System Bus

Claims

1. A method for generating a video bitstream, performed by an encoder, comprising: making a second determination based on the first determination indicating whether padding is applied to a plurality of sub-regions of the current image that the padding includes wrap-around padding; generating image segmentation information based on the first determination and the second determination; encoding the current image based on the plurality of sub-regions and the image division information; generating a video bitstream including the image segmentation information and the encoded current image; a pixel location for motion compensated prediction in a reference image is determined by performing clipping based on a syntax element corresponding to the wrap-around padding; method.

2. prior to the step of encoding the current image, the method further comprises the step of encoding the plurality of sub-regions based on the wrap-around padding based on the second determination indicating that the padding includes the wrap-around padding; The step of encoding the current image based on the plurality of sub-regions and the image division information includes: The method of claim 2 , further comprising: encoding the current image based on the encoded sub-regions and the image segmentation information.

3. A method as described in claim 1 or 2, wherein the image splitting information includes an offset value, the offset value specifying an offset of a luma sample used to calculate a wrap-around position for motion compensation in inter prediction.

4. A method according to any one of claims 1 to 3, wherein the image segmentation information is included in a sequence parameter set corresponding to the current image.

5. The method of claim 4, wherein the image segmentation information includes at least one flag included in the sequence parameter set.

6. A method according to any one of claims 1 to 5, wherein the plurality of sub-regions comprises at least one of a brick, a tile, a slice, a tile group, a sub-image, or a sub-layer.

7. A method according to any one of claims 1 to 6, wherein the padding is applied to boundaries of a subregion among a plurality of subregions.

8. The method of claim 7, wherein the boundary is a vertical boundary of the subregion.

9. The method of claim 7, wherein the boundary is a horizontal boundary of the subregion.

10. A method according to any one of claims 1 to 9, wherein the padding is applied to vertical boundaries of a subregion and to horizontal boundaries of the subregion among a plurality of subregions.

11. A method according to any one of claims 1 to 3, wherein the image splitting information indicates an offset value for wraparound padding.

12. A method according to any one of claims 1 to 3, wherein the image division information indicates left padding width information and right padding width information.

13. An apparatus configured to perform a method according to any one of claims 1 to 12.

14. A computer program for causing an encoder to execute a method according to any one of claims 1 to 12.