Signaling reference picture resampling in video bitstream along with resampling picture size indication
Patent Information
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2025-04-10
- Publication Date
- 2026-03-16
AI Technical Summary
Existing video coding technologies struggle with efficiently handling changes in picture size within a coded video sequence, particularly in modern codecs like VVC, where reference picture resampling is required but current signaling methods are inefficient and inflexible, leading to issues in applications like 360-degree video and surveillance where different parts of a scene may require different adaptive resolution settings.
The proposed method involves signaling a compliance window and resampling ratio through flags and parameters in the video bitstream, allowing for accurate resampling of reference pictures based on the compliance window size and resampling picture size, enabling flexible and efficient adaptation of resolution changes within a video sequence.
This approach allows for more precise and efficient reference picture resampling, improving inter-prediction quality and adaptability in video coding, particularly in applications with diverse resolution requirements such as 360-degree video and surveillance, by ensuring accurate scaling factors are calculated for different parts of the scene.
Smart Images

Figure 00000000_0000_ABST
Abstract
Description
Background Art
[0001] Cross - reference to related applications This application claims priority to U.S. Provisional Application No. 62 / 903,639, filed September 20, 2019, U.S. Provisional Application No. 62 / 905,319, filed September 24, 2019, and U.S. Patent Application No. 17 / 010,163, filed September 2, 2020, the entireties of which are incorporated herein by reference.
[0002] Technical field The disclosed subject matter relates to video coding and decoding, and more particularly, to signaling of reference picture resampling with resampling picture size indication.
[0003] Background art Video coding and decoding that utilize inter-picture prediction with motion compensation are known. Uncompressed digital video can be composed of a series of pictures, each picture having a spatial size of, for example, 1920×1080 luminance samples and associated chrominance samples. A series of pictures can have a constant or variable picture rate, for example, 60 pictures per second, i.e., 60 Hz (informally also known as the frame rate). Uncompressed video has significant bitrate requirements. For example, 1080p60 4:2:0 video at 8 bits per sample (1920×1080 luminance sample resolution at 60 Hz frame rate) requires a bandwidth approaching 1.5 Gbit / s. One hour of such video requires a storage space of over 600 GB.
[0004] One purpose of video coding and decoding may be to reduce the redundancy of the input video signal by compression. Compression may help to reduce the aforementioned bandwidth or memory space requirements, in some cases by more than two orders of magnitude. Both lossless and lossy compression, as well as combinations thereof, can be used. Lossless compression refers to a technique in which an exact copy of the original signal can be reconstructed from the compressed original signal. When using lossy compression, the reconstructed signal may not be identical to the original signal, but the distortion between the original signal and the reconstructed signal is small enough to make the reconstructed signal useful for the intended application. In the case of video, lossy compression is widely used. The amount of tolerable distortion depends on the application. For example, users of certain consumer streaming applications may tolerate higher distortion than users of applications contributed by television. The achievable compression ratio can reflect that higher acceptable / tolerable distortion can result in a higher compression ratio.
[0005] Video encoders and decoders can utilize a number of broad and categorical techniques, including, for example, motion compensation, transformation, quantization, and entropy coding, some of which are introduced below.
[0006] Historically, video encoders and decoders have tended to operate at a given picture size, which in most cases was defined and remained constant for a coded video sequence (CVS), group of pictures (GOP), or similar multi-picture time frame. For example, in MPEG-2, it is known that the system design can vary the horizontal resolution (and thus the picture size) depending on factors such as scene activity, but only in I pictures, and thus typically for a GOP. Resampling of reference pictures for use with different resolutions within a CVS is known, for example, from ITU-T Rec. H.263 Annex P. However, in this case, the picture size remains unchanged and only the reference pictures are resampled, resulting in only a portion of the picture canvas being used (in the case of downsampling) or only a portion of the scene being used (in the case of upsampling). Further, H.263 Annex Q allows individual macroblocks (in each dimension) to be resampled up or down by a factor of two. Again, the picture size remains the same. The size of the macroblocks is fixed in H.263 and thus does not need to be signaled.
[0007] In modern video coding, the change of the picture size of the predicted picture is becoming even more mainstream. For example, VP9 allows for the change of the resolution of the entire picture and reference picture resampling. Similarly, a certain proposal for VVC (including, for example, Hendry, et. al, “On adaptive resolution change (ARC) for VVC”, Joint Video Team document JVET-M0135-v1, Jan 9-19, 2019, which is incorporated herein by reference in its entirety) allows for resampling of the entire reference picture to a different - higher or lower - resolution. In that document, it is suggested that candidates for different resolutions are coded within the sequence parameter set and are referenced by the picture-by-picture syntax elements within the picture parameter set. SUMMARY OF THE INVENTION
[0008] In an embodiment, a method for decoding an encoded video bitstream using at least one processor is provided, the method comprising obtaining a first flag indicating that a compliance window exists for a current picture; obtaining a second flag indicating whether the compliance window is used for reference picture resampling based on the first flag indicating that the compliance window exists; determining a resampling ratio between the current picture and a reference picture based on the second flag indicating that the compliance window is used for the reference picture resampling and based on a compliance window size of the compliance window; determining the resampling ratio based on a resampling picture size based on the second flag indicating that the compliance window is not used for reference picture resampling; and performing the reference picture resampling for the current picture using the resampling ratio.
[0009] In an embodiment, an apparatus for decoding an encoded video bitstream is provided. The apparatus includes at least one memory configured to store program code, and at least one processor configured to read the program code and operate as instructed by the program code. The program code includes a first acquisition code configured to cause the at least one processor to acquire a first flag indicating the presence of a compliance window in the current picture, a second acquisition code configured to cause the at least one processor to acquire a second flag indicating whether the compliance window is used for reference picture resampling based on the first flag indicating the presence of the compliance window, a first determination code configured to cause the at least one processor to determine a resampling ratio between the current picture and the reference picture based on the second flag indicating that the compliance window is used for reference picture resampling and based on the compliance window size of the compliance window, a second determination code configured to cause the at least one processor to determine the resampling ratio based on the resampling picture size based on the second flag indicating that the compliance window is not used for reference picture resampling, and an execution code configured to cause the at least one processor to perform reference picture resampling on the current picture using the resampling ratio.
[0010] In an embodiment, a non-transitory computer-readable medium storing instructions is provided, and when the instructions are executed by one or more processors of a device to decode an encoded video bitstream, the one or more processors are caused to: obtain a first flag indicating that a compliance window exists in a current picture; obtain a second flag indicating whether the compliance window is used for reference picture resampling based on the first flag indicating that the compliance window exists; determine a resampling ratio between the current picture and a reference picture based on the second flag indicating that the compliance window is used for reference picture resampling, based on a compliance window size of the compliance window; determine the resampling ratio based on a resampling picture size based on the second flag indicating that the compliance window is not used for reference picture resampling; and perform reference picture resampling on the current picture using the resampling ratio.
Brief Description of the Drawings
[0011] Further features, properties, and various advantages of the disclosed subject matter will become more apparent from the following detailed description and the accompanying drawings.
[0012]
Figure 1
[0013]
Figure 2
[0014]
Figure 3
[0015]
Figure 4
[0016]
Figure 5
[0017]
Figure 6A
Figure 6B
[0018]
Figure 7
[0019]
Figure 8A
Figure 8B
[0020]
Figure 9
[0021]
Figure 10
[0022]
Figure 11
Mode for Carrying Out the Invention
[0023] FIG. 1 shows a simplified block diagram of a communication system (100) according to an embodiment of the present disclosure. The system (100) may include at least two terminals (110-120) interconnected via a network (150). For unidirectional data transmission, the first terminal (110) can code video data at a local location for transmission to the other terminal (120) via the network (150). The second terminal (120) can receive the coded video data of the other terminal from the network (150), decode the coded data, and display the restored video data. Unidirectional data transmission is common in media serving applications and the like.
[0024] FIG. 1 shows a second pair of terminals (130, 140) provided to support bidirectional transmission of coded video that may occur, for example, during a video conference. For bidirectional data transmission, each terminal (130, 140) can code video data captured at a local location for transmission to the other terminal via the network (150). Each terminal (130, 140) can also receive the coded video data transmitted by the other terminal, decode the coded data, and display the restored video data on a local display device.
[0025] In FIG. 1, the terminals (110-140) may be illustrated as servers, personal computers, or smartphones, but the principles of the present disclosure are not limited thereto. Embodiments of the present disclosure find use in laptop computers, tablet computers, media players, and / or dedicated video conferencing devices. The network (150) represents any number of networks that transmit coded video data between the terminals (110-140), including, for example, wired and / or wireless communication networks. The communication network (150) can exchange data over circuit-switched and / or packet-switched channels. Representative networks include communication networks, local area networks, wide area networks, and / or the Internet. For the purposes of this description, the architecture and topology of the network (150) may not be important for the operation of the present disclosure, unless otherwise described below.
[0026] FIG. 2 shows the arrangement of video encoders and decoders in a streaming environment as an example application of the disclosed subject matter. The disclosed subject matter may be similarly applicable to other video-enabled applications, including, for example, video conferencing, digital TV, storage of compressed video on digital media (including CDs, DVDs, memory sticks, etc.).
[0027] A streaming system may include a capture subsystem (213) that can include, for example, a video source (201), such as a digital camera, that generates, for example, an uncompressed video sample stream (202). This sample stream (202) is depicted as a thick line to emphasize a higher data volume when compared to an encoded video bitstream and can be processed by an encoder (203) coupled to the camera (201). The encoder (203) can include hardware, software, or a combination thereof and can enable or implement aspects of the disclosed subject matter as detailed below. The encoded video bitstream (204) is depicted as a thin line to emphasize a lower data volume when compared to the sample stream and can be stored at a streaming server (205) for future use. One or more streaming clients (206, 208) can access the streaming server (205) to retrieve a copy (207, 209) of the encoded video bitstream (204). The client (206) can include a video decoder (210) that decodes an incoming copy of the encoded video bitstream (207) and generates a progressive video sample stream (211) that can be rendered at a display (212) or other rendering device (not shown). In some streaming systems, the video bitstreams (204, 207, 209) can be encoded according to a particular video coding / compression standard. Specific examples of these standards include ITU-T Recommendation H.265. A video coding standard, informally known as Versatile Video Coding (VVC), is under development.
[0028] FIG. 3 can be a functional block diagram of a video decoder (210) according to an embodiment of the present disclosure.
[0029] The receiver (310) can receive one or more codec video sequences to be decoded by the decoder (210), or in the same or another embodiment, one coded video sequence at a time, in which case the decoding of each coded video sequence is independent of other coded video sequences. The coded video sequence can be received from a channel (312), which may be a hardware / software link to a storage device storing the encoded video data. The receiver (310) can receive the encoded video data together with other data, such as encoded audio data and / or auxiliary data streams, and these data can be transferred respectively using an entity (not shown). The receiver (310) can separate the coded video sequence from other data. To handle network jitter, a buffer memory (315) may be coupled between the receiver (310) and the entropy decoder / parser (320) (hereinafter referred to as the "parser"). If the receiver (310) is receiving data from a store-and-forward device with sufficient bandwidth and controllability, or from an isochronous network, the buffer (315) may not be necessary or may be small. For use in a best-effort packet network such as the Internet, the buffer (315) may be required, can be relatively large, and advantageously can be of an adaptable size.
[0030] Video decoder (210) may include a parser (320) for reconstructing symbols (321) from an entropy-coded video sequence. The categories of these symbols include information used to manage the operation of the decoder (210) and information that, while not an essential part of the decoder, may control a rendering device such as a display (212) that can be coupled to the decoder, as shown in FIG. 3. The control information for the rendering device may be in the form of supplementary enhancement information (SEI (Supplementary Enhancement Information) message) or a video usability information (VUI) parameter set fragment (not shown). The parser (320) can analyze / entropy-decode the received coded video sequence. The coding of the coded video sequence can follow video coding techniques or standards and can follow principles well-known to those skilled in the art, including variable length coding, Huffman coding, arithmetic coding with or without context influence, etc. The parser (320) can extract a set of subgroup parameters for at least one of the subgroups of pixels within the video decoder based on at least one parameter corresponding to the group. The subgroups can include group of pictures (GOP), picture, sub-picture, tile, slice, brick, macroblock, coding tree unit (CTU), coding unit (CU), block, transform unit (TU), prediction unit (PU), etc. A tile may indicate a rectangular region of CUs / CTUs within a specific tile column and row within a picture. A brick may indicate a rectangular region among the CU / CTU rows (multiple rows) within a specific tile. A slice may indicate one or more bricks of a picture included in an NAL unit. A sub-picture may indicate a rectangular region of one or more slices within one picture.The entropy decoder / parser can also extract information such as conversion coefficients, quantization parameter values, motion vectors, etc. from the coded video sequence.
[0031] The parser (320) can perform an entropy decoding / analysis operation on the video sequence received from the buffer (315) to create a symbol (321).
[0032] The reconstruction of the symbol (321) can include multiple different units depending on the type of the coded video picture or a part thereof (e.g., inter-picture, intra-picture, inter-block, intra-block) and other factors. How each unit is involved can be controlled by subgroup control information analyzed by the parser (320) from the coded video sequence. Such a flow of subgroup control information between the parser (320) and multiple subsequent units is not depicted for clarity.
[0033] The decoder 210 can conceptually be subdivided into multiple functional units as described below, beyond the function blocks already mentioned. In a practical implementation operating under commercial constraints, many of these units can interact closely with each other and can be at least partially integrated with each other. However, for the purpose of explaining the disclosed subject matter, the following conceptual subdivision into functional units is appropriate.
[0034] The first unit is the scaler / inverse transform unit (351). The scaler / inverse transform unit (351) receives the quantized conversion coefficients as well as control information, which includes, as symbols (321) from the parser (320), which conversion to use, block size, quantization coefficients, quantization scaling matrix, etc. The unit can output a block containing sample values that can be input to the aggregator (355).
[0035] In some cases, the output samples of the scaler / inverse transform (351) can be associated with intra-coded blocks, i.e., blocks that do not use prediction information from previously reconstructed pictures but can use prediction information from previously reconstructed parts of the current picture. Such prediction information can be provided by the intra-picture prediction unit (352). In some cases, the intra-picture prediction unit (352) uses the surrounding already reconstructed information taken from the current (partially reconstructed) picture (358) to generate blocks of the same size and shape as the block being reconstructed. The aggregator (355) adds, in some cases sample by sample, the prediction information generated by the intra prediction unit (352) to the output sample information as provided by the scaler / inverse transform unit (351).
[0036] In other cases, the output samples of the scaler / inverse transform unit (351) can be associated with inter-coded, potentially motion-compensated blocks. In such cases, the motion compensation prediction unit (353) can access the reference picture memory (357) to retrieve the samples used for prediction. After motion-compensating the retrieved samples according to the symbols (321) associated with the block, these samples can be added by the aggregator (355) to the output of the scaler / inverse transform unit (in this case called residual samples or residual signal) to generate output sample information. When the motion compensation unit retrieves prediction samples, the address within the reference picture memory form can be controlled by, for example, motion vectors available to the motion compensation unit in the form of symbols (321) that can have X, Y, and reference picture components. Also, motion compensation can include interpolation of sample values fetched from the reference picture memory, a motion vector prediction mechanism, etc. when sub-sample exact motion vectors are used.
[0037] The output samples of the aggregator (355) may be affected by various loop filtering techniques within the loop filter unit (356). Video compression techniques can include in-loop filtering techniques, which are controlled by parameters included in the coded video bitstream and made available as symbols (321) from the parser (320) to the loop filter unit (356), but can respond to meta information obtained during the decoding of previous parts (in decoding order) of the coded picture or coded video sequence, and can also respond to previously reconstructed and loop filtered sample values.
[0038] The output of the loop filter unit (356) can be output to the rendering device (212) and can be a sample stream that can be stored in the reference picture memory for use in future inter-picture prediction.
[0039] Some coded pictures, once fully reconstructed, can be used as reference pictures for future prediction. When a coded picture is fully reconstructed and the coded picture is identified as a reference picture (e.g., by the parser (320)), the current reference picture (358) can become part of the reference picture buffer (357), and a new current picture memory can be reallocated before starting the reconstruction of subsequent coded pictures.
[0040] The video decoder 210 can perform a decoding operation according to a predetermined video compression technique that may be documented in a standard such as ITU-T Rec. H.265. The coded video sequence can conform to the syntax defined by the video compression technique or standard in the sense that it conforms to the syntax of the video compression technique or standard as defined in the video compression technique document or standard, particularly in the profile document therein. Also, for compliance, it may be necessary that the complexity of the coded video sequence is within the range determined by the level of the video compression technique or standard. In some cases, the level restricts the maximum picture size, maximum frame rate, maximum reconstruction sample rate (e.g., measured in megasamples per second), maximum reference picture size, etc. The restrictions set by the level may, in some cases, be further restricted by the metadata for HRD buffer management and the Hypothetical Reference Decoder (HRD) specification signaled in the coded video sequence.
[0041] In an embodiment, the receiver (310) can receive additional (redundant) data along with the encoded video. The additional data may be included as part of the coded video sequence. The additional data may be used by the video decoder (210) to properly decode the data and / or to more accurately reconstruct the original video data. The additional data can be in the form of, for example, temporal, spatial, or SNR enhancement layers, redundant slices, redundant pictures, forward error correction codes, etc.
[0042] FIG. 4 can be a functional block diagram of a video encoder (203) according to an embodiment of the present disclosure.
[0043] The encoder (203) can receive video samples from a video source (201) (which is not part of the encoder) that is capable of capturing the video image to be coded by the encoder (203).
[0044] The video source (201) can provide a source video sequence to be coded by the encoder (203) in the form of a digital video sample stream that can be of any suitable bit depth (e.g., 8 bits, 10 bits, 12 bits, ···), any color space (e.g., BT.601 YCrCb, RGB, ···), and any suitable sampling structure (e.g., YCrCb 4:2:0, YCrCb 4:4:4). In a media serving system, the video source (201) may be a storage device that stores pre-prepared video. In a video conferencing system, the video source (203) may be a camera that captures local image information as a video sequence. The video data may be provided as a plurality of individual pictures that convey motion when viewed in sequence. Each picture itself may be organized as a spatial array of pixels, and each pixel can contain one or more samples depending on the sampling structure, color space, etc. in use. A person skilled in the art can easily understand the relationship between pixels and samples. The following description focuses on samples.
[0045] According to an embodiment, the encoder (203) can code and compress pictures of the source video sequence into the coded video sequence (443) in real time or under any other arbitrary time constraints required by the application. Enforcing an appropriate coding speed is one function of the controller (450). The controller controls other functional units as described below and is functionally coupled to these units. The coupling is not depicted for clarity. The parameters set by the controller can include rate control related parameters (picture skip, quantizer, lambda value of rate distortion optimization techniques, ···), picture size, layout of the group of pictures (GOP), maximum motion vector search range, etc. Those skilled in the art can easily identify other functions of the controller (450) because they may be related to the video encoder (203) optimized for a specific system design.
[0046] Some video encoders operate in what those skilled in the art would readily recognize as a "coding loop". As an extremely simplified explanation, the coding loop can be composed of an encoder (430) (hereinafter referred to as the "source coder") (which is responsible for generating symbols based on the input picture and reference pictures to be coded), and an encoder (203) incorporated (local) decoder (433). The (local) decoder (433) reconstructs the symbols to create sample data. (In the video compression technology considered in the subject matter disclosed, any compression between the symbols and the coded video bitstream is lossless, so) the (remote) decoder will also create its sample data. The reconstructed sample stream is input into the reference picture memory (434). Since the decoding of the symbol stream results in a bit-exact result that is independent of the decoder's location (local or remote), the content of the reference picture buffer is also bit-exact between the local encoder and the remote encoder. In other words, the prediction part of the encoder "sees" the same sample values as the reference picture samples as what the decoder would "see" when using prediction during decoding. This basic principle of reference picture synchronization (for example, if synchronization cannot be maintained due to channel errors, resulting in drift) is well known to those skilled in the art.
[0047] The operation of the "local" decoder (433) can be assumed to be the same as that of the "remote" decoder (210), which has already been described in detail in connection with FIG. 3. However, referring briefly to FIG. 4, since symbols are available and the encoding / decoding of symbols by the entropy coder (445) and the parser (320) for the coded video sequence is lossless, the entropy decoding part of the decoder (210) including the channel (312), the receiver (310), the buffer (315) and the parser (320) may not be fully implemented in the local decoder (433).
[0048] The insight that can be made at this point is that any decoder technology that exists in the decoder, except for analysis / entropy decoding, must necessarily exist in the corresponding encoder in a substantially identical functional form. For this reason, the disclosed subject matter focuses on the operation of the decoder. The description of the encoder technology can be omitted because it is the reverse of the decoder technology described comprehensively. More detailed descriptions are only required in certain fields and are provided below.
[0049] The source coder (430) can perform motion-compensated predictive coding as part of its operation, which predicts the input frame by referring to one or more previously coded frames from the video sequence designated as the "reference frame". In this way, the coding engine (432) codes the difference between the pixel block of the input frame and the pixel block of the reference frame that can be selected as the prediction reference for the input frame.
[0050] The local video decoder (433) can decode the encoded video data of a frame that can be designated as a reference frame based on the symbols generated by the source coder (430). The operation of the coding engine (432) can advantageously be a lossless process. When the coded video data is decoded by a video decoder (not shown in FIG. 4), the reconstructed video sequence may typically be a replica of the source video sequence with some errors. The local video decoder (433) can reproduce the decoding process that can be executed by the video decoder for the reference frame, causing the reconstructed reference frame to be stored in the reference picture cache (434). In this way, the encoder (203) can store a copy of the locally reconstructed reference frame having the same content as the reconstructed reference frame obtained by the video decoder at the remote end (when there are no transmission errors).
[0051] The predictor (435) can perform a prediction search for the coding engine (432). That is, for a new frame to be coded, the predictor (435) can search the reference picture memory (434) for sample data (as candidate reference pixel blocks) or specific metadata (e.g., reference picture motion vectors, block shapes, etc.), which may serve as appropriate prediction references for the new picture. The predictor (435) can operate on a per-block basis for the samples to find an appropriate prediction reference. In some cases, the input picture may have a prediction reference drawn from a plurality of reference pictures stored in the reference picture memory (434) as determined by the search results obtained by the predictor (435).
[0052] The controller (450) can manage the coding operation of the video coder (430), including, for example, the setting of parameters and subgroup parameters used to encode video data.
[0053] The outputs of all the aforementioned functional units may be affected by entropy coding in the entropy coder (445). The entropy coder converts the symbols generated by the various functional units into a coded video sequence by losslessly compressing the symbols according to techniques known to those skilled in the art, such as, for example, Huffman coding, variable length coding, arithmetic coding, etc.
[0054] The transmitter (440) can buffer the coded video sequence as created by the entropy coder (445) and prepare it for transmission via the communication channel (460), which may be a hardware / software link to a storage device that stores the encoded video data. The transmitter (440) can merge the coded video data from the video coder (430) with other data to be transmitted, such as, for example, coded audio data and / or auxiliary data streams (sources not shown).
[0055] The controller (450) can manage the operation of the encoder (203). During coding, the controller (450) can assign a specific coded picture type to each coded picture, which may affect the coding technique applicable to the individual picture. For example, a picture can often be designated as one of the following frame types:
[0056] An Intra Picture (I Picture) can be encoded and decoded without using any other frames in the sequence as a source of prediction. Some video codecs allow different types of Intra Pictures, including, for example, independent decoder refresh pictures. Those skilled in the art are familiar with these variations of I Pictures, as well as their respective uses and characteristics.
[0057] A Predicted Picture (P Picture) can be encoded and decoded using intra prediction or inter prediction with at most one motion vector and a reference index to predict the sample values of each block.
[0058] A Bi - Directional Predicted Picture (B Picture) can be encoded and decoded using intra prediction or inter prediction with at most two motion vectors and a reference index to predict the sample values of each block. Similarly, multiple predicted pictures can use more than two reference pictures and associated metadata for the reconstruction of one block.
[0059] The source picture is typically spatially divided into a plurality of sample blocks (e.g., blocks of 4×4, 8×8, 4×8, or 16×16 samples each) and can be coded block by block. The blocks can be predictive coded with reference to other (already coded) blocks as determined by the coding assignment applied to each picture of the block. For example, blocks of an I picture may be coded non-predictively, or they may be predictive coded (spatial prediction or intra prediction) with reference to already coded blocks of the same picture. Pixel blocks of a P picture may be coded non-predictively by spatial prediction or by temporal prediction with reference to one previously coded reference picture. Blocks of a B picture may be coded non-predictively by spatial prediction or by temporal prediction with reference to one or two previously coded reference pictures.
[0060] The video coder (203) can perform coding operations according to a predetermined video coding technology or standard such as ITU-T Rec. H.265. During the operation, the video coder (203) can perform various compression operations including predictive coding operations that utilize the temporal and spatial redundancies in the input video sequence. Accordingly, the coded video data can conform to the syntax specified by the video coding technology or standard being used.
[0061] In an embodiment, the transmitter (440) can transmit additional data along with the encoded video. The video coder (430) can include such data as part of the coded video sequence. The additional data may include other forms of redundant data such as temporal / spatial / SNR enhancement layers, redundant pictures and slices, supplementary enhancement information (SEI) messages, visual user utility information (VUI) parameter set fragments, and the like.
[0062] Recently, compression domain aggregation or extraction of a single video picture from multiple semantically independent picture parts has attracted attention. In particular, for example, in the context of 360 coding or certain surveillance applications, multiple semantically independent source pictures (e.g., the six cube surfaces of a cube projection 360 scene, or the individual camera inputs in the case of a multi-camera surveillance setup) may require different adaptive resolution settings to handle the activities for each of the various scenes at a given point in time. In other words, the encoder may choose to use different resampling factors for different semantically independent pictures that make up the 360 or surveillance scene as a whole at a given point in time. When combined into a single picture, it requires that reference picture resampling be performed and that adaptive resolution coding signaling be available for the coded picture parts.
[0063] Some terms that will be referred to in the remainder of this description are introduced below.
[0064] A sub - picture can, in some cases, refer to a rectangular array of samples for an entity such as a sample, block, macro - block, coding unit, or a similar entity that is semantically grouped and that can potentially be coded independently at a modified resolution. One or more sub - pictures can form a picture. One or more coded sub - pictures can form a coded picture. One or more sub - pictures can be assembled into one picture, and one or more sub - pictures can be extracted from one picture. In certain environments, one or more coded sub - pictures can be assembled in the compressed domain without transcoding the sample level to the coded picture, and in the same or other cases, one or more coded sub - pictures can be extracted from the coded picture in the compressed domain.
[0065] Reference picture resampling (RPR) or adaptive resolution change (ARC) can refer to a mechanism that enables the change of the resolution of a picture or sub - picture within a coded video sequence, for example, by reference picture resampling. Hereinafter, the RPR / ARC parameters refer to the control information required to perform an adaptive resolution change, which can include, for example, filter parameters, scaling factors, the resolution of the output and / or reference pictures, various control flags, etc.
[0066] In an embodiment, coding and decoding can be performed on a single semantically independent coded video picture. Before explaining the additional complexity of the meaning and implications of coding / decoding multiple sub - pictures with independent RPR / ARC parameters, the options for signaling the RPR / ARC parameters will be explained.
[0067] Regarding FIGS. 5A - 5E, several embodiments for signaling RPR / ARC parameters are shown. As mentioned in each of the embodiments, they may have certain advantages and certain disadvantages from the viewpoints of coding efficiency, complexity, and architecture. A video coding standard or technology may select one or more of these embodiments, or options known from related technologies, for signaling RPR / ARC parameters. The embodiments may not be mutually exclusive and may be interchangeable with each other as long as they are considered based on the needs of the application, the related standardization technology, or the choice of the encoder.
[0068] The classes of RPR / ARC parameters may include the following:
[0069] - Up / downsampling factors. Separate or combined in the X and Y dimensions.
[0070] - Up / downsampling factors. Adding the time dimension and indicating a constant speed zoom - in / out for a given number of pictures.
[0071] - Either of the above two may include the coding of one or more perhaps short syntax elements that can point to a table containing factors.
[0072] - Resolution. In the X or Y dimension in units of samples, blocks, macroblocks, coding units (CUs), or any other suitable granularity for each of the input picture, output picture, reference picture, coded picture, or combinations thereof. If there are more than one resolution (e.g., one for the input picture and one for the reference picture), in some cases, one set of values may be inferred from another set of values. These may be gated, for example, by using flags. For more detailed examples, see below.
[0073] - "Warping" coordinates. Similar to those used in H.263 Annex P and, as described above, at an appropriate granularity here. H.263 Annex P defines one efficient method for coding such warping coordinates, but other potentially more efficient methods are also conceivable. The variable-length reversible "Huffman"-style coding of the Annex P warping coordinates can be replaced with a binary coding of appropriate length, in which case the length of the binary code word is derived, for example, from the maximum picture size, perhaps by multiplying by a specific factor and offsetting by a specific value, allowing "warping" outside the boundaries of the maximum picture size.
[0074] - Up or downsampling filter parameters. In an embodiment, there may be only a single filter for upsampling and / or downsampling. However, in an embodiment, it may be desirable to allow more flexibility in filter design, which may require signaling of filter parameters. Such parameters can be selected by an index within a list of possible filter designs, and the filter can be fully specified (e.g., by a list of filter coefficients using appropriate entropy coding techniques), and the filter can be implicitly selected by the up / downsample ratio signaled according to any of the mechanisms described above.
[0075] The following description assumes the coding of a finite set of up / down sample factors (the same factors used in both the X and Y dimensions) specified by codewords. These codewords can be variable length coded using, for example, the Ext - Golomb code common to certain syntax elements in video coding specifications such as H.264 and H.265. One suitable mapping of values to up / down sample factors can be, for example, according to Table 1: Table 1 [Table 1]
[0076] Depending on the needs of the application and the capabilities of the up and downscale mechanisms available in video compression technologies and standards, many similar mappings can be devised. The table can be extended to more values. The values can also be represented by entropy coding mechanisms other than the Ext - Golomb code, for example, using binary coding. This can have certain advantages outside of the video processing engine itself (where the encoder and decoder are most important), for example, by MANE, when the resampling factor is of concern. In situations where a change in resolution is not required, it is possible to select a short Ext - Golomb code; note that in the above table, it is only 1 bit. This can potentially have an advantage in coding efficiency over using binary coding for the most common cases.
[0077] The number of entries in the table and their semantics may be fully or partially configurable. For example, the basic outline of the table may be conveyed in a "high" parameter set such as a sequence or decoder parameter set. In an embodiment, one or more such tables may be defined in a video coding technology or standard and may be selected, for example, through a decoder or sequence parameter set.
[0078] Next, it will be described how the upsampling / downsampling factors (ARC information) coded as described above can be included in a video coding technology or standard syntax. Similar considerations may apply to one or more codewords that control up / downsampling filters. If a relatively large amount of data is required for the filter or other data structure, refer to the following description.
[0079] As shown in FIG. 5, H.263 Annex P includes ARC information (502) in the picture header (501) in the form of four warping coordinates, specifically in the H.263 PLUS PTYPE (503) header extension. This may be a wise design choice when a) there is an available picture header and b) frequent changes to the ARC information are expected. However, the overhead when using H.263-style signaling can be extremely large, and the picture header may be of a temporary nature, so the scaling factor may not be involved between picture boundaries.
[0080] In the same or another embodiment, the signaling of ARC parameters can follow the detailed example outlined in FIGS. 6A-6B. FIGS. 6A-6B show a syntax diagram in a kind of representation that generally follows C-style programming such as that used in video coding standards since at least 1993. The thick lines indicate syntax elements present in the bitstream, and the non-bold lines often indicate control flow and variable settings.
[0081] As shown in FIG. 6A, a tile group header (601), as an exemplary syntax structure of a header applicable to a (presumably rectangular) portion of a picture, can conditionally include a variable-length Exp-Golomb coded syntax element dec_pic_size_idx (602) (bold). The presence of this syntax element in the tile group header can be gate-controlled by the use of adaptive resolution (603) - here by the value of a flag not drawn in bold, which means that the flag is present in the bitstream at the point where it occurs within the syntax diagram. Whether adaptive resolution is being used for this picture or a part thereof can be signaled in any high-level syntax structure inside or outside the bitstream. In the illustrated example, it is signaled in the sequence parameter set as outlined below.
[0082] Referring to FIG. 6B, an excerpt of the sequence parameter set (610) is shown. The first syntax element shown is the adaptive_pic_resolution_change_flag (611). When true, this flag can specify the use of adaptive resolution, which may require certain control information. In this example, such control information is conditionally present based on the value of a flag based on the if() statements in the parameter set (612) and the tile group header (601).
[0083] When adaptive resolution is in use, in this example, the output resolution is coded in sample units (613). The number 613 refers to both output_pic_width_in_luma_samples and output_pic_height_in_luma_samples, and together they can define the resolution of the output picture. It is possible to define specific constraints for any value somewhere in the video coding technology or standard. For example, the definition of a level can limit the total number of output samples that can be, for example, the product of the values of these two syntax elements. Also, a specific video coding technology or standard, or an external technology or standard, such as a system standard, may limit the range of numbers (e.g., one or both dimensions must be divisible by a power of two) or may limit the aspect ratio (e.g., the width and height must be in a relationship such as 4:3 or 16:9). Such limitations may be introduced to facilitate hardware implementation or for other reasons and are well known in the art.
[0084] In certain applications, it may be desirable for the encoder to instruct the decoder to use a specific reference picture size rather than implicitly assuming that its size is the output picture size. In this example, the syntax element reference_pic_size_present_flag (614) gates the conditional presence of the reference picture dimensions (615) (where again the numbers refer to both width and height).
[0085] Finally, a table of possible decoded picture widths and heights is presented. Such a table can be represented, for example, by table indication (num_dec_pic_size_in_luma_samples_minus1) (616). "minus1" indicates the interpretation of the value of the syntax element. For example, if the coded value is zero, there is one table entry. If the value is 5, there are six table entries. For each "line" in the table, the width and height of the decoded picture are then included in the syntax (617).
[0086] The table entries (617) to be presented can be indexed using the syntax element dec_pic_size_idx (602) in the tile group header, thereby allowing various decoded sizes (effective zoom factors) for each tile group.
[0087] Certain video coding techniques or standards, such as VP9, support spatial scalability and enable spatial scalability by implementing a specific form of reference picture resampling (signaled quite differently from the subject matter being disclosed) along with temporal scalability. In particular, certain reference pictures can be upsampled to higher resolutions using ARC-style techniques and form the basis of the spatial enhancement layer. These upsampled pictures can be refined using the normal prediction mechanism at the higher resolution, so details can be added.
[0088] The embodiments described in the present application can be used in such an environment. In certain cases, in the same or different embodiments, the value of the NAL unit header, for example, the Temporal ID (Temporal ID) field, can be used to indicate not only the temporal layer but also the spatial layer. Doing so may have certain advantages for a particular system design; for example, an existing selective forwarding unit (SFU) created and optimized for the selected transfer of the temporal layer based on the Temporal ID value of the NAL unit header can be used without modification for a scalable environment. To enable this, there may be conditions on the mapping between the coded picture size and the temporal layer, specified by the temporal ID field of the NAL unit header.
[0089] Recently, compression domain aggregation or extraction of a single video picture from multiple semantically independent picture parts has attracted attention. In particular, for example, in the context of 360 coding or certain surveillance applications, multiple semantically independent source pictures (e.g., the six cube surfaces of a cube projection 360 scene, or the individual camera inputs in the case of a multi-camera surveillance setup) may require different adaptive resolution settings to handle the activities for each of the various scenes at a given point in time. In other words, the encoder may choose to use different resampling factors for different semantically independent pictures that make up the 360 or surveillance scene as a whole at a given point in time. When combined into a single picture, it requires that reference picture resampling be performed and that adaptive resolution coding signaling be available for the parts of the coded picture.
[0090] In an embodiment, the encoder specifies that the compliance window may be a rectangular sub-part of the picture, and samples within the sub-part are output by the decoder, while samples outside the window may be omitted from the decoder output.
[0091] One of the issues addressed here is that the compliance window can serve two purposes. One purpose is to define the output size as described above. The output size can be application-driven and, for example, in a stereoscopic or 360 cube map application, can be quite different from the coded picture size. The second purpose is that the compliance window may also be input to the reference picture resampling described above.
[0092] For both the coded picture and the resampled picture, when having a specific combination of compliance window size and picture size, samples outside the padding area may be referenced. This problem may be solvable by bitstream constraints regarding the compliance window size, but the solution would have the drawback that the compliance window can no longer be freely selected by the application. Separating the two functionalities may be advantageous in some applications.
[0093] In an embodiment, the compliance window size may be signaled in the PPS. The compliance window parameter can be used to calculate the resampling ratio when the compliance window size of the reference picture differs from that of the current picture. The decoder may need to recognize the compliance window size of each picture to determine whether the resampling process is necessary.
[0094] In an embodiment, the scale factor for reference picture resampling (RPR) can be calculated based on the output width and output height that can be derived from the compliance window parameters between the current picture and the reference picture. This may allow the scaling factor to be calculated more accurately compared to using the decoded picture size. This may work well for most video sequences, where the output picture size is approximately the same as the decoded picture size and has a small padded area.
[0095] However, this can also cause various problems. For example, for immersive media applications (e.g., 360 cube maps, stereoscopy, point clouds), if the compliance window size is very different from the decoded picture size due to a large offset value, the calculation of the scaling factor based on the compliance window size may not guarantee the quality of inter-prediction at different resolutions. In extreme cases, there may be no collocated region in the reference picture at the equivalent position of the current CU. When RPR is used for multi-layer scalability, the compliance window offset may not be used in the calculation of the region referenced between layers. Note that in the HEVC Scalability Extension (SHVC), the region referenced in each directly dependent layer is explicitly signaled in the PPS-extension. When a sub-bitstream targeting a specific region (sub-picture) is extracted from the entire bitstream, the compliance window size does not exactly match the picture size. Note that once the bitstream is encoded, the compliance window parameters cannot be updated as long as the parameters are used for scaling calculations.
[0096] Due to the above potential problems, the calculation of the scaling factor based on the conformance window size may have corner cases, in which alternative parameters are required. As a fallback, it is proposed to signal reference region parameters that can be used to calculate the scaling of RPR and scalability if the conformance window parameters are not used in the calculation of the scaling factor.
[0097] Referring to FIG. 7, in an embodiment, conformance_window_flag can be signaled in the PPS. A conformance_window_flag equal to 1 can specify that the conformance cropping window offset parameter follows next within the PPS. A conformance_window_flag equal to 0 can specify that there is no conformance cropping window offset parameter.
[0098] Continuing to refer to FIG. 7, in an embodiment, conf_win_left_offset, conf_win_right_offset, conf_win_top_offset, and conf_win_bottom_offset specify samples of the picture that refer to the PPS output by the decoding process, from the perspective of a rectangular region specified in the output picture coordinates. When conformance_window_flag is equal to 0, the values of conf_win_left_offset, conf_win_right_offset, conf_win_top_offset, and conf_win_bottom_offset may be presumed to be equal to 0.
[0099] In an embodiment, the compliance cropping window may include luma samples with horizontal picture coordinates from SubWidthC * conf_win_left_offset to pic_width_in_luma_samples - (SubWidthC * conf_win_right_offset + 1) and vertical coordinates from SubHeightC * conf_win_top_offset to pic_height_in_luma_samples - (SubHeightC * conf_win_bottom_offset + 1), including both ends.
[0100] Assume that the value of SubWidthC * (conf_win_left_offset + conf_win_right_offset) may be smaller than pic_width_in_luma_samples, and the value of SubHeightC * (conf_win_top_offset + conf_win_bottom_offset) may be smaller than pic_height_in_luma_samples.
[0101] The variables PicOutputWidthL and PicOutputHeightL may be derived as shown in the following equations (Equation 1) and (Equation 2): PicOutputWidthL = pic_width_in_luma_samples - SubWidthC * (conf_win_right_offset + conf_win_left_offset) (Equation 1) PicOutputHeightL = pic_height_in_pic_size_units - SubHeightC * (conf_win_bottom_offset + conf_win_top_offset) (Equation 2)
[0102] If ChromaArrayType is not equal to 0, the corresponding specified samples of the two chroma arrays can be considered samples with picture coordinates (x / SubWidthC, y / SubHeightC), where (x,y) are the picture coordinates of the specified luma sample.
[0103] In an embodiment, the flag can be present within the PPS or other parameter set, and can indicate whether the resampled picture size (width and height) is explicitly signaled within the PPS or other parameter set. If the resampled picture size parameter is explicitly signaled, the resampling ratio between the current picture and the reference picture may be calculated based on the resampled picture size parameter.
[0104] Referring to FIG. 8A, in an embodiment, a use_conf_win_for_rpr_flag equal to 0 can specify that resampled_pic_width_in_luma_samples and resampled_pic_height_in_luma_samples follow next in an appropriate location, for example within the PPS.
[0105] In an embodiment, a use_conf_win_for_rpr_flag equal to 1 can specify that resampling_pic_width_in_luma_samples and resampling_pic_height_in_luma_samples do not exist.
[0106] In an embodiment, resampling_pic_width_in_luma_samples can specify the width of each reference picture that refers to the PPS for resampling in units of luma samples. resampling_pic_width_in_luma_samples may not be equal to 0, may be an integer multiple of Max(8, MinCbSizeY), and may be less than or equal to pic_width_max_in_luma_samples.
[0107] In an embodiment, resampling_pic_height_in_luma_samples can specify the height of each reference picture that refers to the PPS for resampling in units of luma samples. resampling_pic_height_in_luma_samples may not be equal to 0, may be an integer multiple of Max(8, MinCbSizeY), and may be less than or equal to pic_height_max_in_luma_samples.
[0108] If the syntax element resampling_pic_width_in_luma_samples does not exist, the value of resampling_pic_width_in_luma_samples may be assumed to be equal to PicOutputWidthL.
[0109] If the syntax element resampling_pic_height_in_luma_samples does not exist, the value of resampling_pic_height_in_luma_samples may be assumed to be equal to PicOutputHeightL.
[0110] In an embodiment, a partial interpolation process example involving reference picture resampling can be processed as follows:
[0111] The inputs to this process may be: the luma position (xSb, ySb) specifying the top-left sample of the current coding sub-block relative to the top-left luma sample of the current picture, the variable sbWidth specifying the width of the current coding sub-block, the variable sbHeight specifying the height of the current coding sub-block, the motion vector offset mvOffset, the refined motion vector refMvLX, the selected reference picture sample array refPicLX, the half-sample interpolation filter index hpelIfIdx, the bi-directional optical flow flag bdofFlag, and the variable cIdx specifying the color component index of the current block.
[0112] The output of this process may be an array predSamplesLX of (sbWidth + brdExtSize) x (sbHeight + brdExtSize) of predicted sample values.
[0113] The prediction block border extension size brdExtSize may be derived according to Equation (Equation 3): brdExtSize = ( bdofFlag || ( inter_affine_flag[ xSb ][ ySb ] && sps_affine_prof_enabled_flag ) )? 2 : 0 (Equation 3)
[0114] The variable fRefWidth may be set equal to resampling_pic_width_in_luma_samples of the reference picture in luma samples.
[0115] The variable fRefHeight may be set equal to resampling_pic_height_in_luma_samples of the reference picture in luma samples.
[0116] The motion vector mvLX may be set equal to ( refMvLX ― mvOffset ).
[0117] If cIdx is equal to 0, the scaling factors and their fixed-point representations may be determined according to the following equations (Equation 4) and (Equation 5): hori_scale_fp = ((fRefWidth << 14 ) + ( resampling_pic_width_in_luma_samples >> 1)) / resampling_pic_width_in_luma_samples (Equation 4) vert_scale_fp = ((fRefHeight << 14 ) + ( resampling_pic_height_in_luma_samples >> 1)) / resampling_pic_height_in_luma_samples (Equation 5)
[0118] Let (xIntL, yIntL) be the luma position given in full-sample units and (xFracL, yFracL) be the offset given in 1 / 16-sample units. These variables can be used to specify the fractional-sample position within the reference sample array refPicLX.
[0119] The top-left coordinates (xSbIntL, ySbIntL) of the boundary block of the reference sample padding may be set equal to (xSb + (mvLX
[0000] >> 4), ySb + (mvLX
[0001] >> 4)).
[0120] For each luma sample position (xL = 0..sbWidth ― 1 + brdExtSize, yL = 0..sbHeight ― 1 + brdExtSize) in the predicted luma sample array predSamplesLX, the corresponding predicted luma sample value predSamplesLX[xL][yL] may be derived as follows: - (refxSb L , refySbL ), and (refx L , refy L ) are assumed to be the luma positions indicated by the motion vectors (refMvLX, refMvLX) given in 1 / 16 sample units. The variable refxSb L , refx L , refySb L , and refy L may be derived as shown in the following equations (Equation 6) to (Equation 9): refxSb L = ((xSb << 4) + refMvLX
[0000] ) * hori_scale_fp (Equation 6) refx L = ((Sign(refxSb) * ((Abs(refxSb) + 128) >> 8) + x L * ((hori_scale_fp + 8) >> 4)) + 32) >> 6 (Equation 7) refySb L = ((ySb << 4) + refMvLX
[0001] ) * vert_scale_fp (Equation 8) refyL = ((Sign(refySb) * ((Abs(refySb) + 128) >> 8) + yL * ((vert_scale_fp + 8) >> 4)) + 32) >> 6 (Equation 9) - The variables xInt L , yInt L , xFrac L and yFrac L may be derived as shown in the following equations (Equation 10) to (Equation 13): xInt L = refx L >> 4 (Equation 10) yInt L= refy L >> 4 (Equation 11) xFrac L = refx L & 15 (Equation 12) yFrac L = refy L & 15 (Equation 13)
[0121] If bdofFlag is equal to TRUE, or (sps_affine_prof_enabled_flag is equal to TRUE and inter_affine_flag[xSb][ySb] is equal to TRUE), and one or more of the following conditions are TRUE, the predicted luma sample value predSamplesLX[xL][yL] may be derived by initiating the luma integer sample acquisition process with (xIntL + (xFracL >> 3) - 1), yIntL + (yFracL >> 3) - 1) as the input: - xL is equal to 0. - xL is equal to sbWidth + 1. - yL is equal to 0. - yL is equal to sbHeight + 1.
[0122] Otherwise, the predicted luma sample value predSamplesLX[xL][yL] is derived by initiating the luma sample 8-tap interpolation filtering process with (xIntL - (brdExtSize > 0? 1 : 0), yIntL - (brdExtSize > 0? 1 : 0)), (xFracL, yFracL), (xSbIntL, ySbIntL), refPicLX, hpelIfIdx, sbWidth, sbHeight and (xSb, ySb) as the input.
[0123] Otherwise (when cIdx is not equal to 0), the following may be applied:
[0124] Let (xIntC, yIntC) be the chroma position given in full - sample units and (xFracC, yFracC) be the offset given in 1 / 32 - sample units. These variables can be used to specify a general fractional - sample position within the reference sample array refPicLX.
[0125] The upper - left coordinate (xSbIntC, ySbIntC) of the bounding box for reference sample padding may be set equal to ((xSb / SubWidthC)+(mvLX
[0000] >>5), (ySb / SubHeightC)+(mvLX
[0001] >>5)).
[0126] For each chroma - sample position (xC = 0..sbWidth - 1, yC = 0..sbHeight - 1) in the predicted chroma - sample array preSamplesLX, the corresponding predicted chroma - sample value preSamplesLX[xC][yC] may be derived as follows: - Let (refxSb C , refySb C ) and (refx C , refy C ) be the chroma positions indicated by the motion vector (mvLX
[0000] , mvLX
[0001] ) given in 1 / 32 - sample units. The variables refxSb C , refySb C, refx C and refy C may be derived as shown by the following equations (Equation 14) through (Equation 17): refxSb C = ((xSb / SubWidthC << 5)+mvLX
[0000] )*hori_scale_fp (Equation 14) refx C = ((Sign(refxSbC ) * ( ( Abs( refxSb C ) + 256 ) >> 9 ) + xC * ( ( hori_scale_fp + 8 ) >> 4 ) ) + 16 ) >> 5 (Equation 15) refySb C = ( ( ySb / SubHeightC << 5 ) + mvLX
[0001] ) * vert_scale_fp (Equation 16) refy C = ( ( Sign( refySb C ) * ( ( Abs( refySb C ) + 256 ) >> 9 ) + yC* ( ( vert_scale_fp + 8 ) >> 4 ) ) + 16 ) >> 5 (Equation 17) - variable xInt C , yInt C , xFrac C and yFrac C may be derived as shown in the following equations (Equation 18) to (Equation 21): xInt C = refx C >> 5 (Equation 18) yInt C = refy C >> 5 (Equation 19) xFrac C = refy C & 31 (Equation 20) yFrac C = refy C & 31 (Equation 21)
[0127] The predicted sample value predSamplesLX[xC][yC] may be derived by starting a process with (xIntC, yIntC), (xFracC, yFracC), (xSbIntC, ySbIntC), sbWidth, sbHeight, and refPicLX as inputs.
[0128] Referring to FIG. 8B, in an embodiment, use_conf_win_for_rpr_flag equal to 0 can specify that resampled_pic_width_in_luma_samples and resampled_pic_height_in_luma_samples follow next within the PPS. use_conf_wid_for_rpr_flag equal to 1 specifies that resampling_pic_width_in_luma_samples and resampling_pic_height_in_luma_samples do not exist.
[0129] In an embodiment, ref_region_left_offset can specify the horizontal offset between the top-left luma samples of the reference region within the decoded picture. The value of ref_region_left_offset shall be within the range including both ends of -2 14 ~2 14 -1. If it does not exist, the value of ref_region_left_offset may be presumed to be equal to conf_win_left_offset.
[0130] In an embodiment, ref_region_top_offset can specify the vertical offset between the top-left luma samples of the reference region within the decoded picture. The value of ref_region_top_offset shall be within the range including both ends of -2 14 ~2 14 -1. If it does not exist, the value of ref_region_top_offset may be presumed to be equal to conf_win_right_offset.
[0131] In an embodiment, ref_region_right_offset can specify the horizontal offset between the bottom-right luma samples of the reference region within the decoded picture. The value of ref_layer_right_offset is -2 14 ~2 14 and shall be within the range including both ends of -1. If it does not exist, the value of ref_region_right_offset may be presumed to be equal to conf_win_top_offset.
[0132] In an embodiment, ref_region_bottom_offset can specify the vertical offset between the bottom-right luma samples of the reference region within the decoded picture. The value of ref_layer_bottom_offset is -2 14 ~2 14 and shall be within the range including both ends of -1. If it does not exist, the value of ref_region_bottom_offset[ ref_loc_offset_layer_id[ i ] ] may be presumed to be equal to conf_win_bottom_offset.
[0133] The variables PicRefWidthL and PicRefHeightL may be derived as shown in the following equations (Equation 22) to (Equation 23): PicRefWidthL = pic_width_in_luma_samples ― SubWidthC * ( ref_region_right_offset + ref_region_left_offset ) (Equation 22) PicRefHeightL = pic_height_in_pic_size_units ― SubHeightC * ( ref_region_bottom_offset + ref_region_top_offset) (Equation 23)
[0134] The variable fRefWidth may be set equal to PicRefWidthL of the reference picture of the luma samples.
[0135] The variable fRefHeight may be set equal to PicRefHeightL of the reference picture of the luma samples.
[0136] The motion vector mvLX may be set equal to (refMvLX ― mvOffset).
[0137] When cIdx is equal to 0, the scaling factors and their fixed-point representations may be determined as shown in Equations (Equation 24) and (Equation 25) as follows: hori_scale_fp = ( ( fRefWidth << 14 ) + ( PicRefWidthL >> 1 ) ) / PicRefWidthL (Equation 24) vert_scale_fp = ( ( fRefHeight << 14 ) + ( PicRefHeightL >> 1 ) ) / PicRefHeight (Equation 25)
[0138] The upper-left coordinates (xSbInt L , ySbInt L ) of the bounding box for reference sample padding may be set equal to (xSb + (mvLX
[0000] >> 4), ySb + (mvLX
[0001] >> 4)).
[0139] For each luma sample position (x L = 0..sbWidth ― 1 + brdExtSize, y L = 0..sbHeight ― 1 + brdExtSize) in the predicted luma sample array predSamplesLX, the corresponding predicted luma sample value predSamplesLX[x L [y Lmay be derived as follows: - (refxSb L , refySb L ) and (refx L , refy L ) are assumed to be the luma positions indicated by the motion vectors (refMvLX, refMvLX) given in 1 / 16 sample units. Variables refxSb L , refx L , refySb L , and refy L may be derived as shown by the following equations (Equation 26) to (Equation 29): refxSb L = (( (xSb + ref_region_left_offset) << 4) + refMvLX
[0000] ) * hori_scale_fp (Equation 26) refx L = ((Sign(refxSb) * ((Abs(refxSb) + 128) >> 8) + x L * ((hori_scale_fp + 8) >> 4)) + 32) >> 6 (Equation 27) refySb L = (( (ySb + ref_region_top_offset) << 4) + refMvLX
[0001] ) * vert_scale_fp (Equation 28) refyL = ((Sign(refySb) * ((Abs(refySb) + 128) >> 8) + yL * ((vert_scale_fp + 8) >> 4)) + 32) >> 6 (Equation 29)
[0140] Referring to FIG. 9, in an embodiment, when a resampling ratio for reference picture resampling is calculated based on a compliance window size, a decoder that executes a current block decoding process with inter-picture motion compensation prediction may access one or more pixels outside the decoded pixels in the reference picture. In some examples, the decoding process may use a decoded picture that is unavailable in the DPB, and as a result, the decoder may crash or undesirable decoding behavior may occur.
[0141] For example, as shown in FIG. 9, based on a decoded picture having a decoded picture size 903 and a compliance window size 904 and a reference picture having a decoded picture size 901 and a compliance window size 902, the decoding process for the current block 906 may use a reference block 905 that is unavailable.
[0142] To address such problems, in some cases, it is possible to use explicit signaling of the resampling picture size in the calculation of the resampling ratio for reference picture resampling.
[0143] In an embodiment, if the compliance window size does not match the actual resampling ratio, the compliance window size may not be appropriate for use in calculating the resampling ratio for reference picture resampling. For example, if the current picture is composed of stereoscopic sub-pictures (left and right), or multiple faces of a 360 picture (e.g., 6 faces of a cube map projection), the compliance window size may be different from the desired resampling ratio. In some cases, the compliance window size may cover a sub-region of the decoded picture. In such cases, alternatively, the resampling picture size may be used for the calculation of the resampling ratio.
[0144] FIG. 10 is an exemplary process 1000 for decoding an encoded video bitstream. In some implementations, one or more process blocks of FIG. 10 may be performed by decoder 210. In some implementations, one or more process blocks of FIG. 10 may be performed by another device or group of devices separate from decoder 210, such as decoder 203, or including decoder 210.
[0145] As shown in FIG. 10, process 1000 can include obtaining a first flag indicating that a compliance window exists for the current picture (block 1010).
[0146] As further shown in FIG. 10, process 1000 can include obtaining a second flag indicating whether the compliance window is used for reference picture resampling based on the first flag indicating that the compliance window exists (block 1020).
[0147] As further shown in FIG. 10, process 1000 may include determining whether a second flag indicates that the compliance window is used for reference picture resampling (block 1030). If it is determined that the second flag indicates that the compliance window is used for reference picture resampling (YES in block 1030), process 1000 may proceed to block 1040 and then to block 1060. In block 1040, process 1000 may include determining a resampling ratio between the current picture and the reference picture based on the compliance window size of the compliance window.
[0148] If it is determined that the second flag does not indicate that the compliance window is used for reference picture resampling (NO in block 1030), process 1000 may proceed to block 1050 and then to block 1060. In block 1050, process 1000 may include determining a resampling ratio based on the resampling picture size.
[0149] As further shown in FIG. 10, process 1000 may include performing reference picture resampling on the current picture using the resampling ratio (block 1060).
[0150] In an embodiment, the first flag and the second flag may be signaled in a picture parameter set.
[0151] In an embodiment, the compliance window size may be determined based on at least one offset distance from the boundary of the current picture.
[0152] In an embodiment, the at least one offset distance may be signaled in a picture parameter set.
[0153] In an embodiment, the resampled picture size may be specified in an encoded video bitstream as at least one of the width of the resampled picture size and the height of the resampled picture size.
[0154] In an embodiment, at least one of the width and the height may be signaled in a picture parameter set.
[0155] In an embodiment, at least one of the width and the height may be expressed as the number of a plurality of luma samples included in at least one of the width and the height.
[0156] In an embodiment, the resampled picture size may be determined based on at least one offset distance from the top left luma sample of a reference region of the current picture.
[0157] In an embodiment, at least one offset distance may be signaled in a picture parameter set.
[0158] FIG. 10 shows exemplary blocks of process 1000, but in some implementations, process 1000 may include additional blocks, fewer blocks, different blocks, or blocks arranged otherwise than those shown in FIG. 10. Additionally or alternatively, two or more of the blocks of process 1000 may be executed in parallel.
[0159] Furthermore, the proposed method may be implemented by a processing circuit (e.g., one or more processors or one or more integrated circuits). In one example, one or more processors execute a program stored in a non-transitory computer-readable medium to execute one or more of the proposed methods.
[0160] The above-described technology can be implemented as computer software using computer-readable instructions and can be physically stored on one or more computer-readable media. For example, FIG. 11 shows a computer system 1100 suitable for implementing a particular embodiment of the disclosed subject matter.
[0161] The computer software can be coded using any suitable machine code or computer language that can be affected by mechanisms such as assembly, compilation, and linking to generate code that includes instructions that can be executed by a computer central processing unit (CPU), a graphics processing unit (GPU), etc., either directly or through interpretation, microcode execution, etc.
[0162] The instructions can be executed on various types of computers or their components, including, for example, personal computers, tablet computers, servers, smartphones, gaming devices, the Internet of Things (IoT), etc.
[0163] The components shown in FIG. 11 with respect to the computer system 1100 are exemplary in nature and are not intended to suggest any limitation as to the scope of the use or functionality of the computer software for implementing embodiments of the present disclosure. Also, the configuration of the components should not be construed as having any dependency or requirement with respect to any one or combination of the components shown in the exemplary embodiment of the computer system 1100.
[0164] Computer system 1100 may include a specific human interface input device. Such a human interface input device can respond to input by one or more human users, for example, via tactile input (e.g., keystrokes, swipes, movement of a data glove), audio input (e.g., voice, clapping), visual input (e.g., gesture), and olfactory input (not shown). Further, the human interface device can also be used to capture specific media that is not necessarily directly related to conscious human input, such as audio (e.g., conversation, music, ambient sound), images (e.g., scanned images, photographic images obtained from a still image camera), and video (e.g., two-dimensional video, three-dimensional video including stereoscopic video).
[0165] The input human interface device may include one or more of a keyboard 1101, a mouse 1102, a trackpad 1103, a touch screen 1110 and associated graphics adapter 1150, a data glove, a joystick 1105, a microphone 1106, a scanner 1107, and a camera 1108.
[0166] Computer system 1100 can also include a specific human interface output device. Such a human interface output device can stimulate the senses of one or more human users, for example, through tactile output, sound, light, and smell / taste. Such a human interface output device can be a tactile output device (e.g., tactile feedback by a touch screen 1110, a data glove, or a joystick 1105, although there can also be a tactile feedback device that does not function as an input device), an audio output device (e.g., a speaker 1109, headphones (not shown)), a visual output device (e.g., a screen 1110 including a cathode ray tube (CRT) screen, a liquid crystal display (LCD) screen, a plasma screen, an organic light emitting diode (OLED) screen (each may or may not have a touch screen input function, each may or may not have tactile feedback capability, and some of them can output an output of three or more dimensions by means such as two-dimensional visual output or stereoscopic output), virtual reality glasses (not shown), a holographic display, and a smoke tank (not shown)), and a printer (not shown).
[0167] Computer system 1100 can also include an optical medium such as a CD / DVD ROM / RW 1120 having a CD / DVD or similar medium 1121, a thumb drive 1122, a removable hard drive or solid state drive 1123, legacy magnetic media such as tapes and floppy disks (not shown), specialized ROM / ASIC / PLD-based devices such as security dongles (not shown), and other human-accessible storage devices and their associated media.
[0168] Those skilled in the art should also understand that the term "computer-readable medium" as used in connection with the subject matter of this disclosure does not include a transmission medium, a carrier wave, or other transient signals.
[0169] Computer system 1100 may also include an interface to one or more communication networks (1155). The network can be, for example, wireless, wired, or optical. The network can further be local, wide area, metropolitan, vehicular and industrial, real-time, delay-tolerant, etc. Examples of networks include Ethernet, wireless LAN, cellular networks (including Global System for Mobile Communications (GSM), 3rd generation (3G), 4th generation (4G), 5th generation (5G), Long Term Evolution, etc. for mobile communications), TV wired or wireless wide area digital networks (including cable TV, satellite TV, terrestrial broadcast TV), vehicular and industrial including CANBus, etc. A particular network generally requires an external network interface adapter (1154) connected to a particular general-purpose data port or peripheral bus (1149) (for example, the Universal Serial Bus (USB) port of computer system 1100); other networks are generally incorporated into the core of computer system 1100 by connection to a system bus as described later (for example, an Ethernet interface is incorporated into a PC computer system and a cellular network interface is incorporated into a smartphone computer system). As an example, network 1155 may be connected to peripheral bus 1149 using network interface 1154. Using any of these networks, computer system 1100 can communicate with other entities. Such communication can be unidirectional, receive-only (e.g., broadcast TV), unidirectional transmit-only (e.g., CANbus for a particular CANbus device), or bidirectional, for example, for other computer systems using local or wide area digital networks. Specific protocols and protocol stacks can be used for each of those networks and network interfaces (1154) as described above.
[0170] The foregoing human interface device, the human-accessible memory device, and the network interface can be attached to the core 1140 of the computer system 1100.
[0171] The core 1140 can include one or more central processing units (CPUs) 1141, a graphics processing unit (GPU) 1142, a specific programmable processing unit in the form of a field programmable gate array (FPGA) 1143, a hardware accelerator 1144 for a specific task, and the like. These devices may be connected via a system bus 1148 together with a read-only memory (ROM) 1145, a random access memory (RAM) 1146, an internal mass storage, such as an internal hard drive inaccessible to users, a solid state drive (SSD), and the like 1147. In some computer systems, the system bus 1148 can be made accessible in the form of one or more physical plugs to enable expansion by additional CPUs, GPUs, and the like. Peripheral devices can be attached directly to the core system bus 1148 or via a peripheral bus 1149. The architecture of the peripheral bus includes peripheral component interconnect (PCI), USB, and the like.
[0172] The CPU 1141, GPU 1142, FPGA 1143, and accelerator 1144 can be combined to execute specific instructions capable of constructing the above-described computer code. The computer code can be stored in the ROM 1145 or the RAM 1146. Temporary data can also be stored in the RAM 1146, while persistent data can be stored, for example, in the internal mass storage 1147. By using a cache memory that can be closely associated with one or more CPUs 1141, GPUs 1142, mass storage 1147, ROM 1145, RAM 1146, etc., fast storage and retrieval for any memory device can be enabled.
[0173] A computer-readable medium can have thereon computer code for performing various computer-implemented operations. The medium and the computer code can be specially designed and constructed for the purposes of this disclosure, or they can be of the kind well-known and available to those of ordinary skill in the field of computer software.
[0174] By way of example and not limitation, a computer system having an architecture 1100, specifically a core 1140, can provide functions derived from a processor (including a CPU, GPU, FPGA, accelerator, etc.) that executes software embodied on one or more tangible computer-readable media. Such computer-readable media can be, for example, media associated with user-accessible mass storage as described above, as well as specific storage of the core 1140 of a non-transitory nature such as core internal mass storage 1147 or ROM 1145. The software implementing various embodiments of the present disclosure can be stored on such a device and executed by the core 1140. The computer-readable media can include one or more memory devices or chips, depending on specific needs. The software causes the core 1140, particularly the processor (including a CPU, GPU, FPGA, etc.) therein, to define a data structure stored in the RAM 1146 and modify such a data structure according to a process defined by the software, thereby causing the execution of a specific process or a portion of a specific process described herein. Further or alternatively, the computer system can provide functions as a result of logic wired within a circuit (e.g., accelerator 1144) or otherwise incorporated, and the circuit can operate instead of or in conjunction with software to execute a specific process or a specific portion of a specific process described herein. References to software can include logic, and vice versa if necessary. References to computer-readable media can include a circuit (e.g., an integrated circuit (IC)) that stores software for execution, a circuit that embodies logic for execution, or both if appropriate. The present disclosure encompasses any suitable combination of hardware and software.
[0175] Although several exemplary embodiments have been described, there are changes, substitutions, and various alternative equivalents that fall within the scope of the present disclosure. Therefore, although not explicitly illustrated or described in this application, it will be recognized that those skilled in the art will be able to devise many systems and methods that embody the principles of the present disclosure and are thus within its spirit and scope.
[0176] (Appendix 1) A method of decoding an encoded video bitstream using at least one processor, comprising: obtaining a first flag indicating that a compliance window exists in the current picture; obtaining a second flag indicating whether the compliance window is used for reference picture resampling based on the first flag indicating that the compliance window exists; determining a resampling ratio between the current picture and a reference picture based on the second flag indicating that the compliance window is used for reference picture resampling and based on a compliance window size of the compliance window; determining the resampling ratio based on a resampling picture size based on the second flag indicating that the compliance window is not used for reference picture resampling; performing reference picture resampling on the current picture using the resampling ratio; and a method comprising: (Appendix 2) The method according to Appendix 1, wherein the first flag and the second flag are signaled in a picture parameter set. (Appendix 3) The method according to Appendix 1, wherein the compliance window size is determined based on at least one offset distance from a border of the current picture. (Appendix 4) The at least one offset distance is signaled in a picture parameter set, according to the method described in Appendix 3. (Appendix 5) The resampled picture size is specified in the encoded video bitstream as at least one of the width of the resampled picture size and the height of the resampled picture size, according to the method described in Appendix 1. (Appendix 6) At least one of the width and the height is signaled in a picture parameter set, according to the method described in Appendix 5. (Appendix 7) At least one of the width and the height is expressed as the number of luma samples included in at least one of the width and the height, according to the method described in Appendix 5. (Appendix 8) The resampled picture size is determined based on at least one offset distance from the top-left luma sample of the reference region of the current picture, according to the method described in Appendix 1. (Appendix 9) The at least one offset distance is signaled in a picture parameter set, according to the method described in Appendix 8. (Appendix 10) An apparatus for decoding an encoded video bitstream, at least one memory configured to store program code, at least one processor configured to read the program code and operate as instructed by the program code comprising, the program code being a first acquisition code configured to cause the at least one processor to obtain a first flag indicating the presence of a compliance window in the current picture, Causing the at least one processor to obtain a second flag indicating whether the compliance window is used for reference picture resampling based on the first flag indicating the presence of the compliance window. Causing the at least one processor to determine a resampling ratio between the current picture and the reference picture based on the second flag indicating that the compliance window is used for the reference picture resampling, based on the compliance window size of the compliance window. Causing the at least one processor to determine the resampling ratio based on the resampling picture size based on the second flag indicating that the compliance window is not used for the reference picture resampling. Execution code configured to cause the at least one processor to perform the reference picture resampling on the current picture using the resampling ratio. An apparatus including. (Appendix 11) The apparatus according to Appendix 10, wherein the first flag and the second flag are signaled in a picture parameter set. (Appendix 12) The apparatus according to Appendix 10, wherein the compliance window size is determined based on at least one offset distance from the border of the current picture. (Appendix 13) The apparatus according to Appendix 12, wherein the at least one offset distance is signaled in a picture parameter set. (Appendix 14) The resampled picture size is the apparatus according to Appendix 10, where at least one of the width of the resampled picture size and the height of the resampled picture size is specified in the encoded video bitstream. (Appendix 15) The apparatus according to Appendix 14, where at least one of the width and the height is signaled in a picture parameter set. (Appendix 16) The apparatus according to Appendix 14, where at least one of the width and the height is expressed as the number of luma samples included in at least one of the width and the height. (Appendix 17) The resampled picture size is determined based on at least one offset distance from the top-left luma sample of the reference region of the current picture, for the apparatus according to Appendix 10. (Appendix 18) The apparatus according to Appendix 17, where the at least one offset distance is signaled in a picture parameter set. (Appendix 19) A non-transitory computer-readable medium storing instructions, which, when executed by one or more processors of an apparatus to decode an encoded video bitstream, cause the one or more processors to obtain a first flag indicating the presence of a compliance window in the current picture; obtain a second flag indicating whether the compliance window is used for reference picture resampling, based on the first flag indicating the presence of the compliance window; determine a resampling ratio between the current picture and a reference picture based on the second flag indicating that the compliance window is used for the reference picture resampling, based on the compliance window size of the compliance window; Based on the second flag indicating that the conformity window is not used for the reference picture resampling, determining the resampling ratio based on the resampling picture size; Executing the reference picture resampling for the current picture using the resampling ratio; A non - transitory computer - readable medium for causing the above to be executed. (Appendix 20) The non - transitory computer - readable medium according to Appendix 19, wherein the first flag and the second flag are signaled in a picture parameter set.
Claims
1. A method for decoding an encoded video bitstream using at least one processor, The steps include obtaining a first syntax element in PPS that indicates the presence of a conformance window in the current picture, If the first syntax element indicates the existence of the conformance window, the steps include obtaining the conformance window parameter in PPS, The steps include obtaining a second syntax element indicating that the conformance window parameter is used for reference picture resampling (RPR), and analyzing its value, When the second syntax element indicates that the conformance window is used for the RPR, the steps include determining the resampling ratio between the current picture and the reference picture based on the conformance window size of the conformance window of the current picture and the reference picture, If the second syntax element indicates that the conformance window is not used for the RPR, the steps include determining the resampling ratio based on the size of the reference picture and the size of the resampled picture of the current picture, The steps include: performing the RPR on the current picture using the resampling ratio; A method comprising the following: the size of the resampled picture is determined based on the size of the resampled reference picture using a resampled picture size parameter, the resampled picture size parameter, when explicitly signaled, indicates the top, bottom, left, and right offsets of the resampled picture relative to the resampled reference picture.
2. The method according to claim 1, wherein the conformance window size is determined based on at least one offset distance from the border of the current picture.
3. The method according to claim 1, wherein if the resampling picture size parameter is not explicitly signaled, the top, bottom, left, and right offsets are set to be equal to the top, bottom, left, and right offsets of the conformance window relative to the current picture, respectively.
4. A method for encoding a video bitstream using at least one processor, The steps include generating a first syntax element in PPS that indicates the presence of a conformance window in the current picture, The steps include generating conformance window parameters in PPS when the first syntax element indicates the existence of the conformance window, The steps include generating a second syntax element indicating that the conformance window parameter is used for reference picture resampling (RPR), and setting its value, When the second syntax element indicates that the conformance window is used for the RPR, the steps include determining the resampling ratio between the current picture and the reference picture based on the conformance window size of the conformance window of the current picture and the reference picture, If the second syntax element indicates that the conformance window is not used for the RPR, the steps include determining the resampling ratio based on the size of the reference picture and the size of the resampled picture of the current picture, The steps include: performing the RPR on the current picture using the resampling ratio; A method comprising the following: the size of the resampled picture is determined based on the size of the resampled reference picture using a resampled picture size parameter, the resampled picture size parameter, when explicitly signaled, indicates the top, bottom, left, and right offsets of the resampled picture relative to the resampled reference picture.
5. A method for encoding a video bitstream using at least one processor, and for transmitting or storing the encoded bitstream in a storage medium, The steps include generating a first syntax element in PPS that indicates the presence of a conformance window in the current picture, The steps include generating conformance window parameters in PPS when the first syntax element indicates the existence of the conformance window, The steps include generating a second syntax element indicating that the conformance window parameter is used for reference picture resampling (RPR), and setting its value, When the second syntax element indicates that the conformance window is used for the RPR, the steps include determining the resampling ratio between the current picture and the reference picture based on the conformance window size of the conformance window of the current picture and the reference picture, If the second syntax element indicates that the conformance window is not used for the RPR, the steps include determining the resampling ratio based on the size of the reference picture and the size of the resampled picture of the current picture, The steps include: performing the RPR on the current picture using the resampling ratio; A method comprising the following: the size of the resampled picture is determined based on the size of the resampled reference picture using a resampled picture size parameter, the resampled picture size parameter, when explicitly signaled, indicates the top, bottom, left, and right offsets of the resampled picture relative to the resampled reference picture.
6. At least one memory configured to store program code, A processor that reads the program code and operates as instructed by the program code. An apparatus comprising, wherein the program code causes the at least one processor to execute the method according to any one of claims 1 to 5.
7. A computer program that causes one or more processors to perform the method according to any one of claims 1 to 5.