Method, device, and computer program for wraparound motion compensation with reference picture resampling
Patent Information
- Application Number
- JP2025043044
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2020-10-06
- Filing Date
- 2025-03-18
- Publication Date
- 2025-07-18
- Estimated Expiration
- 2040-11-09
AI Technical Summary
Wrap-around motion compensation in video encoding and decoding fails to function correctly when the reference picture width differs from the current picture width, leading to increased implementation and computational complexity.
Disable wrap-around motion compensation when the layer of the current picture is a dependent layer or when reference picture resampling is enabled, and during interpolation processes for motion compensation if the reference picture width is different from the current picture width.
Simplifies the implementation and reduces computational complexity by ensuring wrap-around motion compensation is only used when the current layer is an independent layer and reference picture resampling is disabled, thereby maintaining accurate motion compensation.
Smart Images

Figure 00000000_0000_ABST
Abstract
Description
Technical Field
[0001] Cross - Reference to Related Applications This application claims priority from U.S. Provisional Patent Application No. 62 / 955,520, filed on December 31, 2019, and U.S. Patent Application No. 17 / 064,172, filed on October 6, 2020, the entire contents of which are incorporated herein by reference.
[0002] Field The disclosed subject matter relates to video encoding and decoding, and more particularly, to enabling and disabling wraparound motion compensation.
Background Art
[0003] Video encoding and decoding using inter - picture prediction with motion compensation is known. Uncompressed digital video can be composed of a series of pictures, and each picture has spatial dimensions of, for example, 1920×1080 luminance samples and related chrominance samples. A series of pictures can have a fixed or variable picture rate (informally also known as the frame rate), for example, 60 pictures per second or 60 Hz. Uncompressed video has significant bit - rate requirements. For example, 1080p60 4:2:0 video with 8 bits per sample (1920×1080 luminance sample resolution at a frame rate of 60 Hz) requires a bandwidth close to 1.5 Gbit / s. One hour of such video requires a storage space of more than 600 GB.
[0004] One purpose of video encoding and decoding can be the reduction of redundancy in the input video signal by compression. Compression can help reduce the aforementioned bandwidth or memory space requirements, sometimes by more than two orders of magnitude. Both reversible compression and irreversible compression, as well as combinations thereof, can be used. Reversible compression refers to a technique in which an exact copy of the original signal can be reconstructed from the compressed original signal. When using irreversible compression, the reconstructed signal may not be identical to the original signal, but the distortion between the original signal and the reconstructed signal is small enough to make the reconstructed signal useful for its intended purpose. In the case of video, irreversible compression is widely used. The amount of acceptable distortion depends on the application. For example, users of certain consumer streaming applications may tolerate higher distortion than users of television distribution applications. The achievable compression ratio can reflect the fact that higher acceptable / tolerable distortion can result in a higher compression ratio.
[0005] Video encoders and decoders can utilize techniques from several broad categories, including, for example, motion compensation, transformation, quantization, and entropy encoding. Some of these are introduced below.
[0006] Historically, video encoders and decoders tended to operate on a given picture size, which was typically defined for a coded video sequence (CVS), a group of pictures (GOP), or a similar multi-picture temporal frame and remained constant. For example, in MPEG-2, the system design is known to vary the horizontal resolution (and thus the picture size) depending on factors such as scene activity, but only in I pictures and thus typically for a GOP. Resampling of reference pictures for use within a CVS at different resolutions is known, for example, from ITU-T Rec. H.263 Annex P. However, here the picture size does not change, and only the reference pictures are resampled, such that potentially only a portion of the picture canvas is used (in the case of downsampling) or only a portion of the scene is captured (in the case of upsampling). Further, H.263 Annex Q allows resampling that doubles or halves (in each dimension) of individual macroblocks. Again, the picture size remains the same. The size of the macroblocks is fixed in H.263 and thus does not need to be signaled.
[0007] Changes in picture size in predicted pictures have become more mainstream in modern video coding. For example, VP9 allows resampling and resolution changes of reference pictures for the entire picture. Similarly, certain proposals for VVC (including, for example, Non-Patent Document 1, which is entirely incorporated herein) allow resampling the entire reference picture to a different - higher or lower - resolution. In that document, it is proposed that different candidate resolutions be coded in the sequence parameter set and referenced by per-picture syntax elements in the picture parameter set.
Non-Patent Document 1
Summary of the Invention
Means for Solving the Problems
[0008] In one embodiment, a method is provided for generating an encoded video bitstream using at least one processor. The method includes making a first determination as to whether the current layer of the current picture is an independent layer; making a second determination as to whether reference picture resampling is enabled for the current layer; disabling wrap-around compensation for the current layer based on the first determination and the second determination; and encoding the current layer without wrap-around compensation.
[0009] In one embodiment, an apparatus for generating an encoded video bitstream is provided, including at least one memory configured to store program code, and at least one processor configured to read the program code and operate as instructed by the program code. The program code includes: first determination code configured to cause the at least one processor to make a first determination as to whether the current layer of the current picture is an independent layer; second determination code configured to cause the at least one processor to make a second determination as to whether reference picture resampling is enabled for the current layer; disabling code configured to cause the at least one processor to disable wrap-around compensation for the current layer based on the first determination and the second determination; and encoding code configured to cause the at least one processor to encode the current layer without wrap-around compensation.
[0010] In one embodiment, a non-transitory computer-readable medium storing instructions is provided. When the instructions are executed by one or more processors of an apparatus that generates an encoded video bitstream, the one or more processors are caused to: make a first determination as to whether a current layer of a current picture is an independent layer; make a second determination as to whether reference picture resampling is enabled for the current layer; based on the first determination and the second determination, disable wrap-around compensation for the current layer; and encode the current layer without wrap-around compensation. The instructions include one or more instructions for causing the execution.
Brief Description of the Drawings
[0011] Further features, properties, and various advantages of the disclosed subject matter will become more apparent from the following detailed description and the accompanying drawings.
[0012]
Figure 1
[0013]
Figure 2
[0014]
Figure 3
[0015]
Figure 4
[0016]
Figure 5
[0017]
Figure 6
[0018]
Figure 7
[0019]
Figure 8
[0020]
Figure 9
DETAILED DESCRIPTION OF THE INVENTION
[0021] FIG. 1 shows a simplified block diagram of a communication system (100) according to an embodiment of the present disclosure. The system (100) may include at least two terminals (110 to 120) interconnected via a network (150). For one-way data transmission, the first terminal (110) can encode video data at a local location for transmission to the other terminal (120) via the network (150). The second terminal (120) can receive the encoded video data of the other terminal from the network (150), decode the encoded data, and display the recovered video data. One-way data transmission may be common in media providing applications and the like.
[0022] FIG. 1 shows a second pair of terminals (130, 140) provided to support two-way transmission of encoded video that may occur during a video conference. For two-way transmission of data, each terminal (130, 140) can encode video data captured at a local location for transmission to the other terminal via the network (150). Each terminal (130, 140) can also receive encoded video data transmitted by the other terminal, decode the encoded data, and display the recovered video data on a local display device.
[0023] In FIG. 1, the terminals (110-140) may be illustrated as servers, personal computers, and smartphones, but the principles of the present disclosure may not be limited thereto. Embodiments of the present disclosure find application with laptop computers, tablet computers, media players, and / or dedicated video conferencing facilities. The network (150) represents any number of networks that transmit encoded video data between the terminals (110-140), including, for example, wired and / or wireless communication networks. The communication network (150) can exchange data in circuit-switched and / or packet-switched channels. Representative networks include telecommunications networks, local area networks, wide area networks, and / or the Internet. For the purposes of the present discussion, the architecture and topology of the network (150) may not be important to the operation of the present disclosure, unless otherwise described below.
[0024] FIG. 2 shows the placement of video encoders and decoders in a streaming environment as an example of the application of the disclosed subject matter. The disclosed subject matter may be equally applicable to other video-enabled applications, including, for example, storage of compressed video on digital media, including video conferences, digital TVs, CDs, DVDs, memory sticks, and the like.
[0025] The streaming system can include a video source (201), such as a digital camera, and may include a capture subsystem (213) that generates, for example, an uncompressed video sample stream (202). The sample stream (202) is drawn in bold to emphasize its high data volume when compared to an encoded video bitstream and can be processed by an encoder (203) coupled to the camera (201). The encoder (203) can include hardware, software, or a combination thereof to enable or implement aspects of the disclosed subject matter, as will be described in more detail below. The encoded video bitstream (204), drawn in thin lines to emphasize its lower data volume when compared to the sample stream, can be stored in a streaming server (205) for future use. One or more streaming clients (206, 208) can access the streaming server (205) to retrieve copies (207, 209) of the encoded video bitstream (204). The client (206) can include a video decoder (210). The video decoder decodes an incoming copy of the encoded video bitstream (207) and generates an outgoing video sample stream (211) that can be rendered on a display (212) or other rendering device (not shown). In some streaming systems, the video bitstreams (204, 207, 209) can be encoded according to certain video encoding / compression standards. Examples of these standards include ITU-T Recommendation H.265. A video encoding standard, informally known as Versatile Video Coding or VVC, is also under development. The disclosed subject matter may be used in the context of VVC.
[0026] FIG. 3 may be a functional block diagram of a video decoder (210) according to an embodiment of the present disclosure.
[0027] The receiver (310) may receive one or more encoded video sequences to be decoded by the decoder (210); in the same or another embodiment, it is one encoded video sequence at a time, and the decoding of each encoded video sequence is independent of other encoded video sequences. The encoded video sequence may be received from a channel (312), which may be a hardware / software link to a storage device storing the encoded video data. The receiver (310) may receive the encoded video data together with other data, such as encoded audio data and / or auxiliary data streams, and these data may be transferred to their respective usage entities (not shown). The receiver (310) can separate the encoded video sequence from other data. As a countermeasure against network jitter, a buffer memory (315) may be coupled between the receiver (310) and the entropy decoder / parser (320) (hereinafter referred to as "parser"). If the receiver (310) is receiving data from a storage / transfer device with sufficient bandwidth and controllability, or from an isochronous network, the buffer (315) may not be required or may be small. For use in a best-effort packet network such as the Internet, a buffer (315) may be required, may be relatively large, and advantageously may be of an adaptable size.
[0028] Video decoder (210) may include a parser (320) for reconstructing symbols (321) from an entropy-coded video sequence. The categories of these symbols include information used to manage the operation of the decoder (210) and potentially information for controlling a rendering device such as a display (212) that is not an integral part of the decoder but can be coupled to the decoder as shown in FIG. 3. The control information for the rendering device(s) may be in the form of Supplementary Enhancement Information (SEI messages) or Video Usability Information (VUI) parameter set fragments (not shown). The parser (320) can parse / entropy-decode the received encoded video sequence. The encoding of the encoded video sequence can follow video encoding techniques or standards and can follow various well-known principles including variable length encoding, Huffman encoding, arithmetic encoding with or without context sensitivity, etc. The parser (320) can extract a set of subgroup parameters for at least one of the subgroups of pixels in the video decoder from the encoded video sequence based on at least one parameter corresponding to the group. Subgroups can include Group of Pictures (GOP), picture, subpicture, tile, slice, macroblock, Coding Tree Unit (CTU), Coding Unit (CU), block, Transform Unit (TU), Prediction Unit (PU), etc. A tile can indicate a rectangular region of CUs / CTUs within a specific tile column and row in a picture. A brick can indicate a rectangular region of a CU / CTU row within a specific tile. A slice may indicate one or more bricks included in a NAL unit of a picture.A sub-picture may indicate a rectangular area of one or more slices within a picture. The entropy decoder / parser can also extract information such as transform coefficients, quantizer parameter values, motion vectors, etc. from the encoded video sequence.
[0029] The parser (320) can perform an entropy decode / parse operation on the video sequence received from the buffer (315), thereby generating symbols (321).
[0030] The reconstruction of the symbols (321) can involve multiple different units depending on the type (e.g., intra-block) of the encoded video picture or its parts and other factors. How each unit is involved can be controlled by subgroup control information parsed by the parser (320) from the encoded video sequence. Such a flow of subgroup control information between the parser (320) and the multiple units below is not depicted for clarity.
[0031] In addition to the functional blocks already described, the decoder 210 can conceptually be divided into several functional units as described below. In an actual implementation operating under commercial constraints, many of these units can interact closely with each other and can be at least partially integrated with each other. However, for the purpose of describing the disclosed subject matter, the conceptual subdivision into the following functional units is appropriate.
[0032] The first unit is the scaler / inverse transform unit (351). The scaler / inverse transform unit (351) receives the quantized transform coefficients and control information as symbols (singular or plural) (321) from the parser (320). The control information includes which transform to use, block size, quantization factor, quantization scaling matrix, etc. The scaler / inverse transform unit can output a block containing sample values that can be input to the aggregator (355).
[0033] In some cases, the output samples of the scaler / inverse transform (351) may relate to intra-coded blocks, i.e., blocks that do not use prediction information from a previously reconstructed picture, but can use prediction information from a previously reconstructed part of the current picture. Such prediction information can be provided by an intra-picture prediction unit (352). In some cases, the intra-picture prediction unit (352) may use surrounding already reconstructed information taken from the current (partially reconstructed) picture (358) to generate a block of the same size and shape as the block being reconstructed. The aggregator (355) may, in some cases, add, sample by sample, the prediction information generated by the intra prediction unit (352) to the output sample information provided by the scaler / inverse transform unit (351).
[0034] In other cases, the output samples of the scaler / inverse transform unit (351) may relate to inter-coded and potentially motion-compensated blocks. In such cases, the motion compensation prediction unit (353) can access the reference picture memory (357) to fetch the samples used for prediction. After motion compensating the fetched samples according to the symbols (321) related to the block, these samples can be added by the aggregator (355) to the output of the scaler / inverse transform unit (in this case, called the residual samples or residual signal), thereby generating output sample information. The address in the reference picture memory from which the motion compensation unit fetches the prediction samples can be controlled by the motion vectors available to the motion compensation unit in the form of symbols (321). The symbols can have, for example, X, Y, and reference picture components. Motion compensation can include interpolation of the sample values fetched from the reference picture memory when exact motion vectors below the sample level are used, a motion vector prediction mechanism, etc.
[0035] The output samples of the aggregator (355) can undergo various loop filtering techniques within the loop filter unit (356). Video compression techniques can include in-loop filter techniques. The in-loop filter techniques are controlled by parameters included in the encoded video bitstream and made available to the loop filter unit (356) as symbols (321) from the parser (320), but also in response to meta information obtained during the decoding of the previous part (in decode order) of the encoded picture or encoded video sequence, and can also respond to previously reconstructed and loop filtered sample values.
[0036] The output of the loop filter unit (356) can be a sample stream, which can be output to the renderer device (212) and can also be stored in the reference picture memory for use in future inter-picture prediction.
[0037] Once a certain encoded picture is completely reconstructed, it can be used as a reference picture for future prediction. For example, when the encoded picture is completely reconstructed and the encoded picture is identified as a reference picture (e.g., by the parser (320)), the current reference picture (358) can become part of the reference picture buffer (357), and a fresh current picture memory can be reallocated before starting the reconstruction of subsequent encoded pictures.
[0038] The video decoder (210) can perform a decoding operation according to a predetermined video compression technique that may be documented in a standard such as ITU-T Recommendation H.265. The encoded video sequence can conform to the syntax specified by the video compression technique or standard being used, in the sense that it follows the syntax of the video compression technique or standard specified in the documentation or standard of the video compression technique, particularly in the profile document therein. For compliance, it may also be necessary that the complexity of the encoded video sequence is within the range defined by the level of the video compression technique or standard. In some cases, the level restricts the maximum picture size, maximum frame rate, maximum reconstruction sample rate (e.g., measured in megasamples per second), maximum reference picture size, etc. The limits set by the level can, in some cases, be further restricted through the hypothetical reference decoder (HRD) specifications and metadata signaled in the encoded video sequence for HRD buffer management.
[0039] In one embodiment, the receiver (310) may receive additional (redundant) data along with the encoded video. The additional data may be included as part of the encoded video sequence(s). The additional data may be used by the video decoder (210) to properly decode the data and / or to more accurately reconstruct the original video data. The additional data can be in the form of, for example, temporal, spatial, or SNR enhancement layers, redundant slices, redundant pictures, forward error correction codes, etc.
[0040] FIG. 4 can be a functional block diagram of a video encoder (203) according to an embodiment of the present disclosure.
[0041] The encoder (203) can receive video samples from a video source (201) (which is not part of the encoder) that can capture the video image to be encoded by the encoder (203).
[0042] The video source (201) can provide the source video sequence to be encoded by the encoder (203) in the form of a digital video sample stream that can be of any suitable bit depth (e.g., 8-bit, 10-bit, 12-bit, etc.), any color space (e.g., BT.601 YCrCB, RGB, etc.) and any suitable sampling structure (e.g., YCrCb 4:2:0, YCrCb 4:4:4). In a media service system, the video source (201) may be a storage device storing pre-prepared video. In a video conferencing system, the video source (203) may be a camera that captures local image information as a video sequence. The video data may be provided as a plurality of individual pictures that impart motion when viewed in sequence. Each picture itself may be organized as a spatial array of pixels, and each pixel can contain one or more samples depending on the sampling structure, color space, etc. in use. Those skilled in the art can easily understand the relationship between pixels and samples. The following description focuses on samples.
[0043] According to one embodiment, the encoder (203) can encode and compress pictures of the source video sequence in real time or under any other temporal constraints required by the application to obtain an encoded video sequence (443). Implementing an appropriate encoding speed is one function of the controller (450). The controller controls other functional units as described below and is functionally coupled to those units. Such couplings are not drawn for clarity. The parameters set by the controller can include parameters related to rate control (picture skip, quantizer, lambda value of rate-distortion optimization techniques, …), picture size, picture group (GOP) layout, maximum motion vector search range, etc. Those skilled in the art can readily identify other functions of the controller (450) that may be relevant to a video encoder (203) optimized for a certain system design.
[0044] Some video encoders operate in what those skilled in the art would readily recognize as an “encoding loop.” As a grossly simplified explanation, in one example, the encoding loop can consist of an encoding section of an encoder (430) (hereinafter, the “source coder”) that is responsible for generating symbols based on an input picture and reference picture(s) to be encoded, and a (local) decoder (433) embedded in the encoder (203). The decoder reconstructs the symbols to generate sample data that a (remote) decoder would also generate (in the video compression techniques contemplated by the disclosed subject matter, any compression between the symbols and the encoded video bitstream is lossless). The reconstructed sample stream is input into a reference picture memory (434). Since the decoding of the symbol stream yields bit-exact results regardless of the decoder position (local or remote), the contents of the reference picture buffer are also bit-exact between the local encoder and the remote encoder. In other words, the prediction section of the encoder “sees” the same sample values as reference picture samples that the decoder “sees” when the decoder uses prediction during decoding. This basic principle of reference picture synchronization (and the resulting drift if, for example, synchronization cannot be maintained due to channel errors) is well known to those skilled in the art.
[0045] The operation of the “local” decoder (433) may be the same as the operation of the “remote” decoder (210) already described in detail above in connection with FIG. 3. However, referring also momentarily to FIG. 4, since the symbols are available and the encoding / decoding of the symbols into the encoded video sequence by the entropy encoder (445) and the parser (320) can be reversible, the entropy decoding section of the decoder (210) including the channel (312), the receiver (310), the buffer (315), and the parser (320) may not be fully implemented in the local decoder (433).
[0046] The observations that can be made at this point are that any decoder technology, except for the parse / entropy decoding that exists within the decoder, must necessarily exist in substantially the same functional form within the corresponding encoder. For this reason, the disclosed subject matter focuses on decoder operations. The description of encoder technology can be abbreviated because it is the reverse of the decoder technology described comprehensively. More detailed explanations are necessary only in certain areas and are provided below.
[0047] As part of its operation, the source encoder (430) can perform motion-compensated predictive encoding that predictively encodes an input frame by referring to one or more previously encoded frames from a video sequence designated as "reference frames". In this way, the encoding engine (432) encodes the difference between a pixel block of the input frame and pixel blocks of the reference frame(s) that can be selected as a predictive reference for the input frame.
[0048] The local video decoder (433) can decode the encoded video data of a frame that can be designated as a reference frame based on the symbols generated by the source encoder (430). The operation of the encoding engine (432) can advantageously be a lossy process. When the encoded video data can be decoded by a video decoder (not shown in FIG. 4), the reconstructed video sequence can typically be a copy of the source video sequence with some errors. The local video decoder (433) can replicate the decoding process that can be performed on the reference frame by the video decoder and cause the reconstructed reference frame to be stored in the reference picture cache (434). In this way, the encoder (203) can locally store a copy of the reconstructed reference frame that has (in the absence of transmission errors) the common content as the reconstructed reference frame that would be obtained by a remote video decoder.
[0049] The predictor (435) can perform a prediction search for the encoding engine (432). That is, for a new frame to be encoded, the predictor (435) can search the reference picture memory (434) for sample data (as candidate reference pixel blocks) or certain metadata, such as reference picture motion vectors, block shapes, etc., that can serve as appropriate prediction references for the new picture. The predictor (435) can operate on a sample block-by-pixel block basis to find an appropriate prediction reference. Depending on what is determined by the search results obtained by the predictor (435), the input picture can have prediction references drawn from a plurality of reference pictures stored in the reference picture memory (434).
[0050] The controller (450) may manage the encoding operation of the video encoder (430), including, for example, setting parameters and subgroup parameters used to encode video data.
[0051] The outputs of all of the above functional units can undergo entropy encoding in the entropy encoder (445). The entropy encoder converts the symbols generated by the various functional units into an encoded video sequence by losslessly compressing the symbols according to techniques known to those skilled in the art, such as Huffman coding, variable length coding, arithmetic coding, etc.
[0052] The transmitter (440) can buffer the encoded video sequence generated by the entropy encoder (445) and prepare it for transmission via the communication channel (460). The communication channel may be a hardware / software link to a storage device that stores the encoded video data. The transmitter (440) can merge the encoded video data from the video encoder (430) with other data to be transmitted, such as encoded audio data and / or an auxiliary data stream (source not shown).
[0053] The controller (450) may manage the operation of the encoder (203). During encoding, the controller (450) can assign a certain encoded picture type to each encoded picture. The encoded picture type can affect the encoding technique applicable to each picture. For example, a picture may often be assigned as one of the following frame types.
[0054] An intra picture (I picture) can be encoded and decoded without using other pictures in the sequence as a prediction source. Some video codecs allow different types of intra pictures, including, for example, Independent Decoder Refresh pictures. Those skilled in the art will recognize these variations of I pictures, as well as their respective uses and characteristics.
[0055] A predicted picture (P picture) can be encoded and decoded using intra prediction or inter prediction that uses at most one motion vector and a reference index to predict the sample values of each block.
[0056] Bidirectional prediction pictures (B pictures) can be encoded and decoded using intra prediction or inter prediction that uses up to two motion vectors and reference indices to predict the sample values of each block. Similarly, multi-prediction pictures can use three or more reference pictures and associated metadata for the reconstruction of a single block.
[0057] The source picture is typically spatially divided into a plurality of sample blocks (e.g., blocks of 4×4, 8×8, 4×8, or 16×16 samples each) and can be encoded block by block. The blocks can be predictively encoded with reference to other (already encoded) blocks, as determined by the encoding assignment applied to each picture of the block. For example, blocks of an I picture may be encoded non-predictively or predictively with reference to already encoded blocks of the same picture (spatial prediction or intra prediction). Pixel blocks of a P picture may be predictively encoded via spatial prediction or via temporal prediction with reference to one previously encoded reference picture. Blocks of a B picture may be predictively encoded via spatial prediction or via temporal prediction with reference to one or two previously encoded reference pictures.
[0058] The video encoder (203) can perform an encoding operation according to a predetermined video encoding technology or standard such as ITU-T Recommendation H.265. In that operation, the video encoder (203) can perform various compression operations, including a predictive encoding operation that exploits the temporal and spatial redundancy in the input video sequence. Thus, the encoded video data can conform to the syntax specified by the video encoding technology or standard used.
[0059] In one embodiment, the transmitter (440) may transmit additional data along with the encoded video. The video encoder (430) may include such data as part of the encoded video sequence. The additional data may include temporal / spatial / SNR enhancement layers, other forms of redundant data such as redundant pictures and slices, supplementary enhancement information (SEI) messages, visual user information (VUI) parameter set fragments, and the like.
[0060] In recent years, there has been some attention paid to the aggregation and extraction in the compression domain of a single video picture of multiple semantically independent picture parts. In particular, for example, in the context of 360 encoding or certain surveillance applications, multiple semantically independent source pictures (e.g., the six cube surfaces of a cube-projected 360 scene, or the individual camera inputs in the case of a multi-camera surveillance setup) may require different adaptive resolution settings to handle the activities for different scenes at a given point in time. In other words, the encoder can choose to use different resampling factors for different semantically independent pictures that make up the 360 or the surveillance scene at a given point in time. When combined into a single picture, for this purpose, reference picture resampling is performed, and it is necessary that adaptive resolution encoded signaling be available for the various parts of the encoded picture.
[0061] In the following, some terms referred to in the remainder of this article are introduced.
[0062] A sub-picture may, in some cases, refer to a rectangular arrangement of samples, blocks, macro-blocks, coding units, or similar entities that are grouped and can be independently coded at a modified resolution. One or more sub-pictures may form a picture. One or more coded sub-pictures may form a coded picture. One or more sub-pictures may be assembled into one picture, and one or more sub-pictures may be extracted from one picture. In certain environments, one or more coded sub-pictures may be assembled into a coded picture in the compressed domain without transcoding to the sample level, and in the same or other cases, one or more coded sub-pictures may be extracted from a coded picture in the compressed domain.
[0063] Adaptive Resolution Change (ARC) may refer to a mechanism that allows for a change in the resolution of a picture or sub-picture within a coded video sequence, for example, by reference picture resampling. Hereinafter, ARC parameters refer to the control information required to perform adaptive resolution change and may include, for example, filter parameters, scaling factors, the resolution of the output and / or reference pictures, various control flags, etc.
[0064] In various embodiments, encoding and decoding may be performed on a single semantically independent coded video picture. Before describing the implications of encoding / decoding of multiple sub-pictures with independent ARC parameters and the additional complexity that implies, various options for signaling ARC parameters are described.
[0065] Referring to FIGS. 5A-5E, several embodiments for signaling ARC parameters are shown. As described for each of the embodiments, they may have certain advantages and certain disadvantages from the perspectives of coding efficiency, complexity, and architecture. A video coding standard or technology may select one or more of these embodiments or options known from related technologies for signaling ARC parameters. It is contemplated that those embodiments may not be mutually exclusive and may be interchangeable based on the needs of the application, the relevant standard technology, or the choice of encoder.
[0066] The class of ARC parameters may include the following.
[0067] · Separate or combined upsampling / downsampling factors in the X and Y dimensions.
[0068] · Adding a time dimension indicating a constant rate of zoom in / out for a given number of pictures to the upsampling / downsampling factor.
[0069] · Either of the above two may include the encoding of one or more, perhaps short syntax elements, that may point within a table containing the factor(s).
[0070] · The resolution in the X or Y dimension of the input picture, output picture, reference picture, encoded picture, in units of samples, blocks, macroblocks, coding units (CUs), or at any other suitable granularity. If there are multiple resolutions (e.g., one for the input picture and one for the reference picture), in some cases, one set of values may be inferred from another set of values. It can be gated, for example, by the use of a flag. For more detailed examples, see below.
[0071] · As described above, "warping" coordinates similar to those used in Annex P of H.263 at a suitable granularity. Annex P of H.263 defines one efficient way to encode such warping coordinates, but it is conceivable that other, potentially more efficient ways can be devised. For example, the variable-length reversible "Huffman"-type encoding of warping coordinates in Annex P can be replaced with a binary encoding of a suitable length. Here, the length of the binary codeword can be derived, for example, from the maximum picture size, possibly multiplied by a factor at the maximum picture size to allow "warping" outside the boundaries of the maximum picture size and offset by a certain value.
[0072] · Upsampling or downsampling filter parameters. In some embodiments, there may be only a single filter for upsampling and / or downsampling. However, in some embodiments, it is desirable to allow greater flexibility in filter design, which may require signaling of filter parameters. Such parameters may be selected through an index in a list of possible filter designs, the filter may be fully specified (e.g., through a list of filter coefficients, using a suitable entropy coding technique), or the filter may be implicitly selected through the upsampling / downsampling ratio signaled according to any of the mechanisms described above.
[0073] Hereinafter, this document assumes the encoding of a finite set of upsampling / downsampling factors (the same factor is used in both the X and Y dimensions) indicated through a codeword. The codeword may be variable-length encoded, for example, using an Ext-Golomb code common to certain syntax elements in video coding specifications such as H.264 and H.265. One suitable mapping of values to upsampling / downsampling factors may be, for example, according to Table 1:
Table 1
[0074] Many similar mappings can be devised according to the needs of the application and the capabilities of the upscaling and downscaling mechanisms available in video compression technologies or standards. This table can be extended to more values. The values may be represented by entropy coding mechanisms other than the Ext-Golomb code, using, for example, binary coding. This can have certain advantages, for example, by the MANE, when the resampling factor is of interest outside the video processing engine (encoder and decoder first) itself. For situations where no resolution change is required, a short Ext-Golomb code can be selected, and it should be noted that in the above table, it is only 1 bit. This may be more efficient in coding than using binary coding for the most common cases.
[0075] The number of items in the table and their meaning content may be fully or partially configurable. For example, the basic outline of the table may be transmitted in a "higher-level" parameter set such as a sequence or decoder parameter set. In various embodiments, one or more such tables may be defined in a video coding technology or standard and may be selected, for example, through a decoder or sequence parameter set.
[0076] The following describes how the upsampling / downsampling factor (ARC information) encoded as described above can be included in a video coding technology or standard syntax. Similar considerations can apply to one or a few codewords that control the up / downsampling filter. For discussions in cases where a relatively large amount of data is required for filters and other data structures, see below.
[0077] As shown in FIG. 5A, Annex P of H.263 includes ARC information (502) in the form of four distortion coordinates in the picture header (501), particularly in the H.263 PLUS PTYPE (503) header extension. This can be a reasonable design choice when a) there is an available picture header and b) frequent changes to the ARC information are expected. However, the overhead when using H.263-style signaling can be very high, and since the picture header may be of a temporary nature, the scaling factors may not be related between picture boundaries.
[0078] As shown in FIG. 5B, JVC ET-M135-v1 includes ARC reference information (505) (index) located within the picture parameter set (504), which indexes a table (506) that includes the target resolution located within the sequence parameter set (507). Placing the possible resolutions in the table (506) within the sequence parameter set (507) can be justified, according to oral statements made by the authors, by using the SPS as an interoperability trade-off point during capability exchange. The resolution can vary for each picture, within the limits set by the values in the table (506), by referring to the appropriate picture parameter set (504).
[0079] Referring to FIGS. 5C - 5E, the following embodiments may exist for transmitting ARC information in a video bitstream. Each of these options has certain advantages over the embodiments described above. The embodiments may co-exist simultaneously within the same video coding technology or standard.
[0080] In various embodiments, such as the embodiment shown in FIG. 5C, ARC information (509), such as a resampling (zoom) factor, may be in a slice header, GOP header, tile header, or tile group header. FIG. 5C shows an embodiment in which a tile group header (508) is used. This may be sufficient when the ARC information is small, for example, a single variable length ue(v) or a fixed length codeword of a few bits as shown above. Having the ARC information directly within the tile group header has the additional advantage that the ARC information may be applicable to a sub-picture represented by the tile group rather than the whole picture, for example. See also below. Further, even if the video compression technology or standard assumes only picture-wide adaptive resolution changes (as opposed to, for example, tile group-based adaptive resolution changes), putting the ARC information in the tile group header has certain advantages from the perspective of error resilience compared to putting it in an H.263 style picture header.
[0081] In various embodiments, such as the embodiment shown in FIG. 5D, the ARC information (512) itself may be present in a suitable parameter set, such as a picture parameter set, header parameter set, tile parameter set, adaptation parameter set, etc. FIG. 5D shows an embodiment in which an adaptation parameter set (511) is used. The scope of the parameter set may advantageously not be larger than the picture, for example, the tile group. The use of the ARC information is made implicitly through the activation of the relevant parameter set. For example, if the video coding technology or standard considers only picture-based ARC, a picture parameter set or equivalent may be appropriate.
[0082] In an embodiment, for example, the embodiment shown in FIG. 5E, the ARC reference information (513) may be present within the tile group header (514) or a similar data structure. The reference information (513) can refer to a subset of the ARC information (515) available within a parameter set (516) having a scope that extends beyond a single picture, for example, a sequence parameter set or a decoder parameter set.
[0083] Activation implied by additional levels of indirection of the PPS from the tile group header, PPS, SPS used in JVET-M0135-v1 seems unnecessary. This is because the picture parameter set can be used for capability negotiation or announcements (and can be present in certain standards such as RFC3984) similar to the sequence parameter set. However, if the ARC information is to be applicable to sub-pictures represented, for example, by a tile group, a parameter set with an activation scope limited to the tile group, such as an adaptation parameter set or a header parameter set, might be a better choice. Also, if the ARC information is large enough to be non-ignorable, for example, contains filter control information such as a large number of filter coefficients, the parameters can be a good choice from the perspective of coding efficiency rather than directly using the header (508). This is because those settings can be reusable by future pictures or sub-pictures by referring to the same parameter set.
[0084] When using a sequence parameter set or another higher-level parameter set having a scope that spans multiple pictures, certain considerations may apply.
[0085] 1. The parameters set to store the ARC information table (516) can, in some cases, be a sequence parameter set, but in other cases, advantageously, can be a decoder parameter set. The decoder parameter set can have a plurality of CVSs, i.e., encoded video streams, i.e., an activation scope for all the encoded video bits from session start to session end. Possible ARC factors may be decoder functions implemented in hardware, and such a scope may be more appropriate because the hardware functions tend not to change for any CVS (in at least some entertainment systems, the CVS is a picture group with a length of less than 1 second). That being said, putting the table into the sequence parameter set is explicitly included in the arrangement options described herein, particularly in relation to point 2 below.
[0086] 2. The ARC reference information (513) may preferably be placed directly in the picture / slice tile / GOP / tile group header, for example the tile group header (514), rather than in the picture parameter set as in JVCET-M0135-v1. For example, if the encoder wants to change a single value within the picture parameter set, such as the ARC reference information, a new PPS has to be created and the new PPS has to be referred to. Assume that only the ARC reference information changes and other information such as quantization matrix information in the PPS does not change. Such information may be of a fairly large size and may have to be resent to complete the new PPS. Since the ARC reference information (513) may be a single codeword such as an index to a table, which would be the only value changing, it would be cumbersome and wasteful to resent, for example, all of the quantization matrix information. Therefore, as proposed in JVET-M0135-v1, it may be quite excellent from the viewpoint of coding efficiency in order to avoid indirect reference through the PPS. Similarly, putting the ARC reference information into the PPS has an additional drawback that the ARC information referred to by the ARC reference information (513) may be applied to the whole picture rather than to a sub-picture because the activation scope of the picture parameter set by the ARC reference information (513) is the picture.
[0087] In the same or another embodiment, the signaling of the ARC parameters can follow the detailed example outlined in FIGS. 6A-6B. FIGS. 6A-6B show a syntax diagram in a representation format using a notation that generally follows C-style programming, as used, for example, in video coding standards since at least 1993. The thick lines indicate syntax elements present in the bitstream, and the non-thick lines often indicate control flow and variable settings.
[0088] As shown in FIG. 6A, the tile group header (601) can optionally include a variable length Exp-Golomb coded syntax element dec_pic_size_idx (602) (shown in bold) as an exemplary syntax structure of a header applicable to a (possibly rectangular) part of a picture. The presence of this syntax element in the tile group header can be gated using an adaptive resolution (603) - the value of a flag not shown in bold here. This means that the flag is present in the bitstream at the point where it appears in the syntax diagram. Whether the adaptive resolution is used for this picture or a part thereof can be signaled in any high level syntax structure inside or outside the bitstream. In the example shown, it is signaled in the sequence parameter set as outlined below.
[0089] Referring to FIG. 6B, an excerpt of the sequence parameter set (610) is also shown. The first syntax element shown is the adaptive_pic_resolution_change_flag (611). When true, the flag can indicate the use of an adaptive resolution, which may require some control information. In this example, such control information is conditionally present based on the value of the flag, based on an if() statement in the parameter set (612) and the tile group header (601).
[0090] When using adaptive resolution, in this example, the output resolution in terms of sample units (613) is encoded. The code 613 refers to both output_pic_width_in_luma_samples and output_pic_height_in_luma_samples, which together can define the resolution of the output picture. In other places in the video coding technology or standard, certain constraints can be defined for either value. For example, the level definition can limit the total number of output samples that may be the product of the values of these two syntax elements. Also, certain video coding technologies or standards, or external technologies or standards such as system specifications, for example, may limit the numbering range (e.g., one or both dimensions must be divisible by a power of two) or the aspect ratio (e.g., the width and height must be in a relationship such as 4:3 or 16:9). Such constraints may be introduced to facilitate hardware implementation or for other reasons and are well-known in the art.
[0091] In some applications, rather than implicitly assuming that its size is the output picture size, it may be desirable for the encoder to instruct the decoder to use a certain reference picture size. In this example, the syntax element reference_pic_size_present_flag (614) gates the conditional presence of the reference picture dimensions (615) (where again the numbers refer to both width and height).
[0092] Finally, a table of possible decoded picture widths and heights is shown. Such a table can be represented, for example, by the table indication (num_dec_pic_size_in_luma_samples_minus1) (616). "minus1" [subtracted by 1] can refer to the interpretation of the value of that syntax element. For example, if the encoded value is zero, there is one table entry, and if the value is 5, there are six table entries. For each "row" in the table, the width and height of the decoded picture are included in the syntax (617).
[0093] The presented table entries (617) can be indexed using the syntax element dec_pic_size_idx (602) in the tile group header, thereby allowing different decoded sizes - in effect, zoom factors - for each tile group.
[0094] In the implementation of related technologies of VVC, there may be a problem that the wrap-around motion compensation may not function correctly when the reference picture width is different from the current picture width. In some embodiments, the wrap-around motion compensation may be disabled in the high-level syntax when the layer of the current picture is a dependent layer or when the RPR is effective for the current layer. In some embodiments, when the reference picture width is different from the current picture width, the wrap-around process may be disabled during the interpolation process for motion compensation.
[0095] Wrap-around motion compensation can be a useful feature for encoding, for example, 360 projection pictures with an equirectangular projection (ERP) format. This can reduce some visual artifacts at the seams and improve the coding gain. In the current VVC specification draft JVET-P2001 (editorially updated by JVET-Q0041), sps_ref_wraparound_offset_minus1 in the SPS specifies the offset used for calculating the horizontal wrap-around position.
[0096] The problem can occur that the wrap-around offset value is determined in relation to the picture width. If the picture width of the reference picture is different from the current picture width, the wrap-around offset value should be changed in proportion to the scaling ratio between the current picture and the reference picture. However, in practice, adjusting the offset value according to the picture width of each reference picture can significantly increase the implementation and computational complexity compared to the advantages of wrap-around motion compensation. Inter-layer prediction and reference picture resampling (RPR) with different picture sizes can result in a daunting variety of combinations of different picture resolutions across layers and temporal pictures.
[0097] Embodiments can address this problem. For example, in an embodiment, when the layer of the current picture is a dependent layer, or when RPR is enabled for the current layer, wrap-around motion compensation may be disabled by setting sps_ref_wraparound_enabled_flag equal to 0. Thus, wrap-around motion compensation can be used only when the current layer is an independent layer and RPR is disabled. Under this condition, the reference picture size is equal to the current picture size. Further, in embodiments, when the reference picture width is different from the current picture width, the wrap-around motion compensation process may be disabled during the interpolation process for motion compensation.
[0098] Embodiments may be used separately or combined in any order. Further, each of the method (or embodiment), encoder, and decoder may be implemented by a processing circuit (e.g., one or more processors, or one or more integrated circuits). In one example, the one or more processors execute a program stored in a non-transitory computer-readable medium.
[0099] FIG. 7 shows an exemplary syntax table according to embodiments. In embodiments, that sps_ref_wraparound_enabled_flag (701) is equal to 1 may specify that horizontal wraparound motion compensation is applied to inter prediction. That sps_ref_wraparound_enabled_flag (701) is equal to 0 may specify that horizontal wraparound motion compensation is not applied. When the value of (CtbSizeY / MinCbSizeY + 1) is less than or equal to (pic_width_in_luma_samples / MinCbSizeY - 1), where pic_width_in_luma_samples is the value of pic_width_in_luma_samples in any PPS that refers to the SPS, sps_ref_wraparound_enabled_flag (701) may be equal to 0. That the value of sps_ref_wraparound_enabled_flag (701) is equal to 0 when vps_independent_layer_flag[GeneralLayerIdx[nuh_layer_id]] is equal to 0 may be a bitstream compliance requirement. If not present, the value of sps_ref_wraparound_enabled_flag (701) may be presumed to be equal to 0.
[0100] In embodiments, refPicWidthInLumaSamples may be the pic_width_in_luma_samples of the current reference picture of the current picture. In embodiments, when refPicWidthInLumaSamples is equal to the pic_width_in_luma_samples of the current picture, refWraparoundEnabledFlag may be set equal to sps_ref_wraparound_enabled_flag. Otherwise, refWraparoundEnabledFlag may be set equal to 0.
[0101] The luma position (xInti, yInti) in the complete sample unit may be derived as follows for i = 0..1.
[0102] When subpic_treated_as_pic_flag[SubPicIdx] is equal to 1, the following may apply:
Number
[0103] Otherwise (when subpic_treated_as_pic_flag[subPicIdx] is equal to 0), the following may apply:
Number
[0104] The luma position (xInt, yInt) in the complete sample unit may be derived as follows.
[0105] When subpic_treated_as_pic_flag[SubPicIdx] is equal to 1, the following may apply:
Number
[0106] Otherwise, the following may apply:
Number
[0107] The predicted luma sample value preSampleLXL may be derived as follows: preSampleLXL = refPicLXL[xInt][yInt] << shift3.
[0108] The chroma positions (xInti, yInti) in the complete sample unit may be derived as follows for i = 0..3.
[0109] If subpic_treated_as_pic_flag[SubPicIdx] is equal to 1, the following may apply:
Number
[0110] Otherwise (subpic_treated_as_pic_flag[subPicIdx] is equal to 0), the following may apply:
Number
[0111] The chroma positions (xInti, yInti) in the complete sample unit may be further modified as follows for i = 0..3:
Number
[0112] Figures 8A - 8C are flowcharts of exemplary processes 800A, 800B, and 800C for generating an encoded video bitstream according to various embodiments. In various embodiments, any of processes 800A, 800B, and 800C, or any part of processes 800A, 800B, and 800C, may be combined in any desired order in any combination or permutation. In some implementations, one or more process blocks of Figures 8A - 8C may be executed by decoder 210. In some implementations, one or more process blocks of Figures 8A - 8C may be executed by another device or group of devices separate from or including decoder 210, such as encoder 203.
[0113] As shown in FIG. 8A, process 800A may include making a first determination as to whether the current layer of the current picture is an independent layer (block 811).
[0114] As further shown in FIG. 8A, process 800A may include making a second determination as to whether reference picture resampling is enabled for the current layer (block 812).
[0115] As further shown in FIG. 8A, process 800A may include disabling wrap-around compensation for the current picture based on the first determination and the second determination (block 813).
[0116] As further shown in FIG. 8A, process 800A may include encoding the current picture without wrap-around compensation (block 814).
[0117] In some embodiments, the first determination may be made based on a first flag signaled in a first syntax structure, and the second determination may be made based on a second flag signaled in a second syntax structure lower than the first syntax structure.
[0118] In some embodiments, the first flag may be signaled in a video parameter set, and the second flag may be signaled in a sequence parameter set.
[0119] In some embodiments, wrap-around compensation may be disabled based on the absence of the second flag in the sequence parameter set.
[0120] As shown in FIG. 8B, process 800B may include determining whether the current layer of the current picture is an independent layer (block 821).
[0121] As further shown in FIG. 8B, if it is determined that the current layer is not an independent layer (NO in block 821), the process 800B may proceed to block 822, where the wrap-around motion compensation may be disabled.
[0122] As further shown in FIG. 8B, if it is determined that the current layer is an independent layer (YES in block 821), the process 800B may proceed to block 823.
[0123] As further shown in FIG. 8B, the process 800B may include determining whether reference picture resampling is enabled (block 823).
[0124] As further shown in FIG. 8B, if it is determined that reference picture resampling is enabled (YES in block 823), the process 800B may proceed to block 822, where the wrap-around motion compensation may be disabled.
[0125] As further shown in FIG. 8B, if it is determined that reference picture resampling is not enabled (NO in block 823), the process 800B may proceed to block 824, where the wrap-around motion compensation may be enabled.
[0126] As shown in FIG. 8C, the process 800C may include determining whether the current layer of the current picture is an independent layer (block 831).
[0127] As further shown in FIG. 8C, the process 800C may include determining whether reference picture resampling is enabled (block 832).
[0128] As further shown in FIG. 8C, the process 800C may include determining whether the width of the current picture is different from the width of the current reference picture (block 833).
[0129] As further shown in FIG. 8C, if it is determined that the width of the current picture is different from the width of the current reference picture (YES in block 833), process 800C may proceed to block 834, where wrap-around motion compensation may be disabled.
[0130] As further shown in FIG. 8C, if it is determined that the width of the current picture is the same as the width of the current reference picture (NO in block 821), process 800C may proceed to block 835, where wrap-around motion compensation may be enabled.
[0131] In some embodiments, block 833 may be executed during an interpolation process for motion compensation.
[0132] FIGS. 8A - 8C illustrate exemplary blocks of processes 800A, 800B, and 800C. However, in some implementations, process 800 may include additional blocks, fewer blocks, different blocks, or blocks in a different arrangement than those shown in FIGS. 8A - 8C. Additionally or alternatively, two or more of the blocks of processes 800A, 800B, and 800C may be executed in parallel.
[0133] Furthermore, the proposed method may be implemented by a processing circuit (e.g., one or more processors or one or more integrated circuits). In one example, the one or more processors execute a program stored on a non - transitory computer - readable medium to perform one or more of the proposed methods.
[0134] The techniques described above can be implemented as computer software using computer - readable instructions and can be physically stored on one or more computer - readable media. For example, FIG. 9 shows a computer system 900 suitable for implementing certain embodiments of the disclosed subject matter.
[0135] Computer software can be coded using any suitable machine code or computer language and be the subject of assembly, compilation, linking, or similar mechanisms to create code containing instructions that can be executed directly by a computer central processing unit (CPU), a graphics processing unit (GPU), etc., or through interpretation, microcode execution, etc.
[0136] Instructions can be executed on various types of computers or their components, including, for example, personal computers, tablet computers, servers, smartphones, gaming devices, Internet of Things devices, etc.
[0137] The components shown in FIG. 9 for computer system 900 are illustrative in nature and are not intended to imply any limitations regarding the use or functionality scope of the computer software implementing embodiments of the present disclosure. Nor should the component configuration be construed as having any dependency or requirement regarding any one or combination of the components shown in the exemplary embodiments of computer system 900.
[0138] Computer system 900 can include certain human interface input devices. Such human interface input devices can respond to input by one or more human users through, for example, tactile input (e.g., keystrokes, swipes, movement of a data glove), voice input (e.g., voice, clapping), visual input (e.g., gestures), olfactory input (not shown). Also, the human interface device can be used to capture certain media that are not necessarily directly related to conscious human input, such as voice (e.g., speech, music, ambient sound), images (e.g., scanned images, photographic images obtained from a still image camera), video (e.g., 2D video, 3D video including stereoscopic video).
[0139] The input human interface device may include one or more of a keyboard 901, a mouse 902, a trackpad 903, a touch screen 910 and an accompanying graphics adapter 950, a data glove, a joystick 905, a microphone 906, a scanner 907, a camera 908 (only one of each is depicted).
[0140] The computer system 900 may also include certain human interface output devices. Such human interface output devices may, for example, stimulate the senses of one or more human users through tactile output, sound, light, and smell / taste. Such human interface output devices include tactile output devices (e.g., tactile feedback by the touch screen 910, data glove or joystick 905; however, there may also be tactile feedback devices that do not function as input devices), audio output devices (e.g., speakers 909, headphones (not shown)), visual output devices (e.g., a screen 910 including a cathode ray tube (CRT) screen, a liquid crystal display (LCD) screen, a plasma screen, an organic light emitting diode (OLED) screen; each may or may not have a touch screen input function, each may or may not have a tactile feedback function, and some of them may be able to output higher than three-dimensional output through means such as two-dimensional visual output or stereoscopic output; virtual reality glasses (not shown), holographic displays and smoke tanks (not shown)), and a printer (not shown).
[0141] Computer system 900 can also include an optical medium including a CD / DVD ROM / RW 920 along with a human-accessible memory device and associated media, such as a CD / DVD or similar media 921, a thumb drive 922, a removable hard drive or solid state drive 923, legacy magnetic media such as tapes and floppy disks (not shown), specialized ROM / ASIC / PLD-based devices such as security dongles (not shown), and the like.
[0142] One of ordinary skill in the art should also understand that the term "computer-readable medium" as used in connection with the presently disclosed subject matter does not encompass a transmission medium, a carrier wave, or other transient signals.
[0143] Computer system 900 can also include an interface to one or more communication networks (955). The network can be, for example, wireless, wired, or optical. The network can further be local, wide area, metropolitan area, in-vehicle and industrial, real-time, delay tolerant, etc. Examples of networks include Ethernet®, wireless LAN, Global System for Mobile Communications (GSM), third generation (3G), fourth generation (4G), fifth generation (5G), Long Term Evolution (LTE), etc., cellular networks, cable TV, satellite TV, TV wired or wireless wide area digital networks including terrestrial broadcast TV, in-vehicle and industrial including CANBus, etc. Some types of networks typically require an external network interface adapter (954) attached to a certain general-purpose data port or peripheral bus (949) (such as the Universal Serial Bus (USB) port of computer system 900). Others are typically integrated into the core of computer system 900 by attachment to a system bus as described later (such as an Ethernet interface to a PC computer system or a cellular network interface to a smartphone computer system). Using any of these networks, computer system 900 can communicate with other entities. Such communication can be unidirectional, receive-only (such as broadcast TV), transmit-only unidirectional (such as CANbus to certain CANbus devices), or bidirectional to other computer systems using, for example, local or wide area digital networks. For each of the networks and network interfaces (1154) as described above, certain protocols and protocol stacks can be used.
[0144] The aforementioned human interface device, human-accessible storage device, and network interface can be attached to the core 940 of computer system 900.
[0145] The core 940 can include one or more central processing units (CPUs) 941, a graphics processing unit (GPU) 942, a specialized programmable processing device in the form of a field programmable gate array (FPGA) 943, a hardware accelerator 944 for certain tasks, etc. These devices can be connected through a system bus 948 together with built-in mass storage devices such as read-only memory (ROM) 945, random access memory (RAM) 946, an internal hard drive that is not user-accessible, a solid state device (SSD) 947, etc. In some computer systems, the system bus 948 may be accessible in the form of one or more physical plugs to enable expansion by additional CPUs, GPUs, etc. Peripheral devices can be attached directly to the core system bus 948 or through a peripheral bus 949. Architectures for peripheral buses include Peripheral Component Interconnect (PCI), USB, etc.
[0146] The CPU 941, GPU 942, FPGA 943, and accelerator 944 can execute certain instructions that can together constitute the above-mentioned computer code. That computer code can be stored in the ROM 945 or RAM 946. Temporary data can also be stored in the RAM 946, while persistent data can be stored, for example, in the internal mass storage device 947. By using cache memory that can be closely associated with one or more CPUs 941, GPUs 942, mass storage device 947, ROM 945, RAM 946, etc., fast storage and retrieval to any of the memory devices can be enabled.
[0147] A computer-readable medium can have computer code thereon for performing various computer-implemented operations. The medium and the computer code may be specially designed and constructed for the purposes of this disclosure, or they may be of the kind well-known and available to those having skill in the art of computer software.
[0148] By way of example and not limitation, a computer system having an architecture 900, specifically a core 940, can provide functionality as a result of a processor (including a CPU, GPU, FPGA, accelerator, etc.) executing software embodied on one or more tangible computer-readable media. Such computer-readable media can be media associated with user-accessible mass storage as introduced above, as well as certain storage of the core 940 of a non-transitory nature such as the mass storage device 947 or ROM 945 inside the core. The software implementing various embodiments of the present disclosure can be stored on such a device and executed by the core 940. The computer-readable media can include one or more memory devices or chips depending on specific needs. The software can include defining data structures stored in the RAM 946 and modifying such data structures according to processes defined by the software, causing specific processes or specific portions of specific processes described herein to be executed by the core 940 and specifically the processors (including CPU, GPU, FPGA, etc.) therein. Additionally or alternatively, the computer system can provide functionality as a result of logic wired within a circuit (e.g., accelerator 944) or otherwise embodied, which can operate instead of or in conjunction with software for executing specific processes or specific portions of specific processes described herein. References to software include logic and vice versa as appropriate. References to computer-readable media can, as appropriate, include circuits (e.g., integrated circuits (ICs)) storing software for execution, circuits embodying logic for execution, or both. The present disclosure encompasses any suitable combination of hardware and software.
[0149] Although the present disclosure has described several exemplary embodiments, there are changes, substitutions, and various alternative equivalents that fall within the scope of the present disclosure. Thus, it will be understood by those skilled in the art that many systems and methods can be devised that embody the principles of the present disclosure and thus are within the spirit and scope of the present disclosure, even though not explicitly shown or described herein.
Claims
1. A method for decoding encoded video data by at least one processor, the method comprising: decoding a picture of the current layer of the encoded video data with or without wrap-around compensation; wherein wrap-around compensation is disabled when the width of the current picture is different from the width of the current reference picture; and wrap-around compensation is enabled when the width of the current picture is the same as the width of the current reference picture. A method.
2. The method according to claim 1, wherein the determination as to whether the width of the current picture is different from the width of the current reference picture is made when reference picture resampling is enabled, and wrap-around compensation is enabled when the reference picture resampling is not enabled.
3. The method according to claim 2, wherein the determination as to whether the reference picture resampling is enabled is based on a first flag signaled in a first syntax structure. The method according to claim 2.
4. The method according to claim 3, wherein the first flag is signaled in a sequence parameter set. The method according to claim 3.
5. The method according to any one of claims 1 to 4, wherein the determination as to whether the width of the current picture is different from the width of the current reference picture is made during an interpolation process for motion compensation.
6. An apparatus for decoding encoded video data, the apparatus comprising: at least one memory configured to store program code; and at least one processor configured to read the program code and operate as instructed by the program code, wherein the program code is for causing the at least one processor to execute the method according to any one of claims 1 to 5. An apparatus.
7. A computer program for causing one or more processors to execute the method according to any one of claims 1 to 5.
8. A method for encoding a picture by at least one processor, the method comprising: encoding a picture of the current layer of video data with or without wrap-around compensation. The wrap-around compensation is disabled when the width of the current picture is different from the width of the current reference picture, and the wrap-around compensation is enabled when the width of the current picture is the same as the width of the current reference picture, method. **Claim 9** A method for providing an encoded video bitstream by at least one processor, comprising: encoding the current layer of the current picture with or without wrap-around compensation, wherein the wrap-around compensation is disabled when the width of the current picture is different from the width of the current reference picture, and the wrap-around compensation is enabled when the width of the current picture is the same as the width of the current reference picture; including the encoded current picture in the video bitstream; and outputting the video bitstream. Method.