Signaling of Inter-layer Prediction in Video Bitstreams
The introduction of a new syntax for signaling scaling in video bitstreams, utilizing RPR or ARC, addresses the inefficiencies in inter-layer prediction, thereby improving coding and decoding efficiency and supporting scalability.
Patent Information
- Application Number
- JP2023215273
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Priority Date
- 2020-09-14
- Filing Date
- 2023-12-20
- Publication Date
- 2025-06-03
- Estimated Expiration
- 2040-09-18
AI Technical Summary
Existing video coding and decoding technologies face challenges in efficiently signaling inter-layer prediction in a scalable video bitstream, which affects coding and decoding efficiency.
A new syntax is designed for signaling scaling in a video bitstream, utilizing reference picture resampling (RPR) or adaptive resolution change (ARC) to support scalability, thereby improving coding efficiency by employing inter-layer prediction at the block level.
The proposed solution enhances coding and decoding efficiency by eliminating the need for additional resampling processes, while supporting scalability through effective inter-layer prediction mechanisms.
Smart Images

Figure 0007688106000002 
Figure 0007688106000003 
Figure 0007688106000004
Abstract
Description
Technical Field
[0001] This application claims the benefit of U.S. Provisional Patent Application No. 62 / 903,652, filed Sep. 20, 2019, and U.S. Patent Application No. 17 / 019,713, filed Sep. 14, 2020, the entireties of which are incorporated herein by reference.
[0002] The disclosed subject matter relates to video coding and decoding, and more particularly, to signaling of inter-layer prediction in a video bitstream.
Background Art
[0003] Video coding and decoding using inter-picture prediction with motion compensation has been known for decades. Uncompressed digital video can be composed of a series of pictures, each picture having, for example, the spatial dimensions of 1920×1080 luminance samples and associated chrominance samples. A series of pictures can have, for example, a fixed or variable (also known informally as the frame rate) picture rate of 60 pictures per second or 60 Hz. Uncompressed video has significant bitrate requirements. For example, 1080p60 4:2:0 video (1920×1080 luminance sample resolution at 60 Hz frame rate) at 8 bits per sample requires a bandwidth close to 1.5 Gbit / s. One hour of such video requires storage space exceeding 600 GB.
[0004] One purpose of video coding and decoding can be to reduce the redundancy of the input video signal through compression. Compression can help reduce the aforementioned bandwidth or storage space requirements, in some cases by more than an order of magnitude. Both reversible compression and irreversible compression, as well as combinations thereof, can be employed. Reversible compression refers to techniques that can restore an exact copy of the original signal from the compressed original signal. When using irreversible compression, the restored signal may not be identical to the original signal, but the distortion between the original signal and the restored signal is small enough to make the restored signal useful for the intended application. In the case of video, irreversible compression is widely adopted. The amount of distortion tolerated depends on the application. For example, users of certain consumer streaming applications can tolerate higher distortion than users of television contribution applications. It can be shown that the achievable compression ratio can be higher as the tolerated / tolerable distortion is larger.
[0005] Video encoders and video decoders can utilize techniques from several broad categories, including, for example, motion compensation, transformation, quantization, and entropy coding, some of which are introduced below.
[0006] Historically, video encoders and video decoders have tended to operate at a given picture size that was defined for, and remained constant for, most cases, a coded video sequence (CVS), a group of pictures (GOP), or a similar multi-picture time frame. For example, in MPEG-2, the system design is known to change the horizontal resolution (and thereby the picture size) according to factors such as the activity of the scene, but only in I pictures, and thus usually for GOPs. Resampling of reference pictures for use with different resolutions within a CVS is known, for example, from ITU-T Rec. H.263 Annex P. However, here the picture size does not change, only the reference pictures are resampled, potentially resulting in only a portion of the picture canvas being used (in the case of downsampling) or only a portion of the scene being captured (in the case of upsampling). Further, H.263 Annex Q enables resampling of individual macroblocks by a factor of two up or down (in each dimension). Again, the picture size remains the same. Since the size of the macroblocks is fixed in H.263, it does not need to be signaled.
[0007] Changing the picture size of predicted pictures has become more mainstream in recent video coding. For example, VP9 enables reference picture resampling (RPR) and changing the resolution of the entire picture. Similarly, some proposals made for VVC (e.g., Hendry et al., "On adaptive resolution change (ARC) for VVC", Joint Video Team document JVET-M0135-v1, January 9 - 19, 2019, which is hereby incorporated by reference in its entirety) enable resampling of entire reference pictures to different, higher or lower resolutions. In that document, various candidate resolutions are proposed that are coded within the sequence parameter set and referenced by per-picture syntax elements within the picture parameter set. SUMMARY OF THE INVENTION
Means for Solving the Problem
[0008] To address one or more different technical problems, the present disclosure describes a new syntax designed for signaling scaling in a video bitstream and its use. In this way, an improvement in coding (decoding) efficiency can be achieved.
[0009] According to embodiments herein, for scalability support, an additional burden may be achieved by modifying the high-level syntax (HLS) using reference picture resampling (RPR) or adaptive resolution change (ARC). In technical terms, inter-layer prediction is employed in a scalable system to improve the coding efficiency of the enhancement layer. In addition to the spatial and temporal motion compensation prediction available in a single-layer codec, inter-layer prediction uses the resampled video data of the restored reference picture from the reference layer to predict the current enhancement layer. Then, by modifying the existing interpolation process for motion compensation, the resampling process for inter-layer prediction is executed at the block level. This means that no additional resampling process is required to support scalability. In the present disclosure, high-level syntax elements that support spatial / quality scalability using RPR are disclosed.
[0010] A method and apparatus are included that comprise a memory configured to store computer program code, and one or more processors configured to access the computer program code and operate as instructed by the computer program code. The computer program code comprises syntax analysis code configured to cause at least one processor to perform syntax analysis on at least one video parameter set (VPS) that includes at least one syntax element indicating whether at least one layer in a scalable bitstream is one of a dependent layer of the scalable bitstream and an independent layer of the scalable bitstream; determination code configured to cause at least one processor to determine, based on a plurality of flags included in the VPS, a number of dependent layers of the scalable bitstream including its dependent layer; first decoding code configured to cause at least one processor to decode pictures in a dependent layer by performing syntax analysis and interpretation on an inter-layer reference picture (ILRP) list; and second decoding code configured to cause at least one processor to decode pictures in an independent layer without performing syntax analysis and interpretation on the ILRP list.
[0011] According to an embodiment, the second decoding code is further configured to cause at least one processor to decode pictures in an independent layer by performing syntax analysis and interpretation on a reference picture list that does not include any decoded pictures of other layers.
[0012] According to an embodiment, the inter-layer reference picture list includes decoded pictures of other layers.
[0013] According to an embodiment, the syntax analysis code is further configured to cause at least one processor to perform syntax analysis on at least one VPS by determining whether other syntax elements indicate a maximum number of layers.
[0014] According to an embodiment, the syntax analysis code is further configured to cause at least one VPS to be syntax-analyzed by at least one processor by determining whether the VPS includes a flag indicating whether another layer in the scalable bitstream is a reference layer for at least one layer.
[0015] According to an embodiment, the syntax analysis code is further configured to cause at least one VPS to be syntax-analyzed by at least one processor by determining whether a flag indicates that another layer is a reference layer for at least one layer by specifying the index of the other layer and the index of at least one layer, and the syntax analysis code is further configured to cause at least one VPS to be syntax-analyzed by at least one processor by determining whether it includes other syntax elements indicating a value less than the number of dependent layers determined for the VPS.
[0016] According to an embodiment, the syntax analysis code is further configured to cause at least one VPS to be syntax-analyzed by at least one processor by determining whether a flag indicates that another layer is not a reference layer for at least one layer by specifying the index of the other layer and the index of at least one layer, and the syntax analysis code is further configured to cause at least one VPS to be syntax-analyzed by at least one processor by determining whether it includes other syntax elements indicating a value less than the number of dependent layers determined for the VPS.
[0017] According to an embodiment, the syntax analysis code is further configured to cause at least one VPS to be syntax-analyzed by at least one processor by determining whether the VPS includes a flag indicating whether a plurality of layers including at least one layer should be decoded by interpreting an ILRP list.
[0018] According to an embodiment, the syntax analysis code is further configured to cause at least one VPS to perform syntax analysis on at least one processor by determining whether the VPS includes a flag indicating whether a plurality of layers including at least one layer should be decoded without interpreting the ILRP list.
[0019] According to an embodiment, the syntax analysis code is further configured to cause at least one VPS to perform syntax analysis on at least one processor, further including determining whether the VPS includes a flag indicating whether a plurality of layers including at least one layer should be decoded by interpreting the ILRP list.
[0020] Further features, properties, and various advantages of the disclosed subject matter will be more apparent from the following forms for implementing the invention and the accompanying drawings.
Brief Description of the Drawings
[0021]
Figure 1
Figure 2
Figure 3
Figure 4
Figure 5A
Figure 5B
Figure 5C
[0022]
Figure 5D
Figure 5E
Figure 6
Figure 7
Figure 8
Figure 9
Figure 10
[0023] The proposed features described below may be used separately or combined in any order. Further, embodiments may be implemented by a processing circuit (e.g., one or more processors or one or more integrated circuits). In one example, one or more processors execute a program stored in a non-transitory computer-readable medium.
[0024] FIG. 1 shows a simplified block diagram of a communication system (100) according to an embodiment of the present disclosure. The communication system (100) may include at least two terminals (110 and 120) interconnected via a network (150). For unidirectional data transmission, the first terminal (110) can encode video data at a local location for transmission to another terminal (120) via the network (150). The second terminal (120) can receive the encoded video data of another terminal from the network (150), decode the encoded data, and display the restored video data. Unidirectional data transmission may be common in media serving applications and the like.
[0025] FIG. 1 shows a second pair of terminals (130, 140) provided to support bidirectional transmission of encoded video that may occur, for example, during a video conference. For bidirectional data transmission, each terminal (130, 140) can encode video data captured at a local location for transmission to another terminal via the network (150). Each terminal (130, 140) can also receive the encoded video data transmitted by another terminal, decode the encoded data, and display the restored video data on a local display device.
[0026] In the example of FIG. 1, the terminals (110, 120, 130, 140) may be shown as servers, personal computers, and smartphones, but the principles of the present disclosure need not be so limited. Embodiments of the present disclosure find applications using laptop computers, tablet computers, media players, and / or dedicated video conferencing equipment. The network (150) represents any number of networks that transfer encoded video data among the terminals (110, 120, 130, 140), including, for example, wired and / or wireless communication networks. The communication network (150) can exchange data in circuit-switched channels and / or packet-switched channels. Representative networks include telecommunications networks, local area networks, wide area networks, and / or the Internet. For the purposes of this description, the architecture and topology of the network (150) may not be important for the operation of the present disclosure, unless described below herein.
[0027] FIG. 2 shows the arrangement of a video encoder and a video decoder in a streaming environment as an example for an application of the disclosed subject matter. The disclosed subject matter may be equally applicable to other video-enabled applications, including, for example, video conferencing, digital television, storage of compressed video on digital media such as CDs, DVDs, memory sticks, etc.
[0028] A streaming system may include, for example, a video source (201) that creates an uncompressed video sample stream (202), such as a digital camera, and a capture subsystem (213). That sample stream (202), depicted as a thick line to emphasize its high data volume compared to the encoded video bitstream, can be processed by an encoder (203) coupled to the camera (201). The encoder (203) can include hardware, software, or a combination thereof to enable or implement aspects of the disclosed subject matter, as described in more detail below. The encoded video bitstream (204), depicted as a thin line to emphasize its lower data volume compared to the sample stream, can be stored in a streaming server (205) for future use. One or more streaming clients (206, 208) can access the streaming server (205) to retrieve a copy (207, 209) of the encoded video bitstream (204). The client (206) can include a video decoder (210) that decodes an input copy of the encoded video bitstream (207) and creates an output video sample stream (211) that can be rendered on a display (212) or other rendering device (not depicted). In some streaming systems, the video bitstreams (204, 207, 209) can be encoded according to a particular video coding / compression standard. Examples of those standards include ITU-T Recommendation H.265. A video coding standard, informally known as Multipurpose Video Coding or VVC, is under development. The disclosed subject matter may be used in connection with VVC.
[0029] FIG. 3 can be a functional block diagram of a video decoder (210) according to an embodiment of the present disclosure.
[0030] The receiver (310) receives one or more codec video sequences to be decoded by the decoder (210), and in the same or another embodiment, can receive one encoded video sequence at a time, and the decoding of each encoded video sequence is independent of other encoded video sequences. The encoded video sequence may be received from a channel (312), and the channel (312) may be a hardware / software link to a storage device storing the encoded video data. The receiver (310) can receive the encoded video data together with other data, such as encoded audio data, and / or an auxiliary data stream, which can be transferred to their respective using entities (not depicted). The receiver (310) can separate the encoded video sequence from other data. To counter network jitter, a buffer memory (315) may be coupled between the receiver (310) and the entropy decoder / parser (320) (hereinafter, "parser"). When the receiver (310) is receiving data from a store-and-forward device with sufficient bandwidth and controllability, or from a synchronous network, the buffer (315) may not be required or may have a low probability. For use in a best-effort packet network such as the Internet, the buffer (315) may be required and may have a relatively high probability and advantageously may be of an adaptive size.
[0031] Video decoder (210) may include a parser (320) to recover symbols (321) from an entropy-coded video sequence. The categories of those symbols include information used to manage the operation of decoder (210), and potentially, information for controlling a rendering device such as a display (212) that can be coupled to the decoder, although not an essential part of the decoder, as shown in FIG. 2. The control information for the rendering device may be in the form of supplementary enhancement information (SEI message) or a parameter set fragment of video user data information (VUI) (not depicted). Parser (320) can perform syntax analysis / entropy decoding on the received coded video sequence. The coding of the coded video sequence can follow video coding techniques or standards and can follow principles well known to those skilled in the art, including variable length coding, Huffman coding, arithmetic coding, etc., regardless of whether it is context-sensitive. Parser (320) can extract a set of subgroup parameters for at least one of the subgroups of pixels in the video decoder from the coded video sequence based on at least one parameter corresponding to a group. The subgroups can include picture groups (GOPs), pictures, tiles, slices, macroblocks, coding units (CUs), blocks, transform units (TUs), prediction units (PUs), etc. The entropy decoder / parser can also extract information such as transform coefficients, quantizer parameter values, motion vectors, etc. from the coded video sequence.
[0032] Parser (320) can perform entropy decoding / syntax analysis operations on the video sequence received from buffer (315) to create symbols (321).
[0033] The restoration of symbol (321) can include multiple different units, depending on the type of the coded video picture or a part thereof (such as inter-picture and intra-picture, inter-block and intra-block), and other factors. How each unit is involved can be controlled by subgroup control information parsed from the coded video sequence by the parser (320). Such a flow of subgroup control information between the parser (320) and the multiple units below is not depicted for clarity.
[0034] In addition to the functional blocks already described, decoder 210 can be conceptually subdivided into several functional units as described below. In an actual implementation operating under commercial constraints, many of these units interact closely with each other and can be at least partially integrated with each other. However, for describing the disclosed subject matter, a conceptual subdivision into the following functional units is appropriate.
[0035] The first unit is the scaler / inverse transform unit (351). The scaler / inverse transform unit (351) receives, as symbol (321), from the parser (320) the quantization transform coefficients and control information including which transform to use, block size, quantization coefficients, quantization scaling matrix, etc. It can output a block containing sample values that can be input to the aggregator (355).
[0036] In some cases, the output samples of the scaler / inverse transform (351) may be related to blocks that are intra-coded, i.e., do not use prediction information from previously reconstructed pictures, but can use prediction information from previously reconstructed parts of the current picture. Such prediction information can be provided by the intra-picture prediction unit (352). In some cases, the intra-picture prediction unit (352) uses the surrounding already reconstructed information fetched from the current (partially reconstructed) picture (356) to generate a block of the same size and shape as the block being reconstructed. The aggregator (355) may, in some cases, add, for each sample, the prediction information generated by the intra-prediction unit (352) to the output sample information provided by the scaler / inverse transform unit (351).
[0037] In other cases, the output samples of the scaler / inverse transform unit (351) may be related to inter-coded and potentially motion-compensated blocks. In such cases, the motion-compensation prediction unit (353) can access the reference picture memory (357) to fetch the samples used for prediction. After motion-compensating the samples fetched according to the symbols (321) related to the block, these samples can be added by the aggregator (355) to the output of the scaler / inverse transform unit to generate the output sample information (in this case, called residual samples or residual signal). The address in the reference picture memory where the motion-compensation prediction unit fetches the prediction samples can be controlled, for example, by the motion vectors available to the motion-compensation prediction unit in the form of symbols (321) that can have X, Y, and reference picture components. Motion compensation can also include interpolation of the sample values fetched from the reference picture memory when exact sub-sample motion vectors are used, a motion vector prediction mechanism, etc.
[0038] The output samples of the aggregator (355) can undergo various loop filtering techniques in the loop filter unit (356). The video compression technology is controlled by parameters included in the encoded video bitstream and can include in-loop filter techniques made available to the loop filter unit (356) as symbols (321) from the parser (320), but it can respond not only to meta information obtained during the decoding of the previous part (in decoding order) of the encoded picture or encoded video sequence, but also to previously restored and loop-filtered sample values.
[0039] The output of the loop filter unit (356) can be a sample stream that is not only output to the rendering device (212) but can also be stored in the reference picture memory (356) for use in future inter-picture prediction.
[0040] When a particular encoded picture is fully restored, it can be used as a reference picture for future prediction. When the encoded picture is fully restored and the encoded picture is identified as a reference picture (e.g., by the parser (320)), the current reference picture (356) can become part of the reference picture buffer (357), and the unused current picture memory can be reallocated before starting the restoration of the next encoded picture.
[0041] The video decoder 320 can perform a decoding operation according to a predetermined video compression technique that can be documented in a standard such as ITU-T Rec.H.265. The encoded video sequence can conform to the syntax of the video compression technique or standard specified in the document or standard of the video compression technique, specifically the profile document therein, in the sense that the encoded video sequence complies with the syntax specified by the video compression technique or standard being used. Also, for compliance, it is necessary that the complexity of the encoded video sequence is within the range defined by the level of the video compression technique or standard. In some cases, depending on the level, the maximum picture size, the maximum frame rate, the maximum restoration sample rate (measured, for example, in megasamples per second), the maximum reference picture size, etc. are restricted. The restrictions set by the level can, in some cases, be further restricted by the specifications of the hypothetical reference decoder (HRD) and the metadata for HRD buffer management signaled within the encoded video sequence.
[0042] In one embodiment, the receiver (310) can receive additional (redundant) data together with the encoded video. The additional data may be included as part of the encoded video sequence. The additional data may be used by the video decoder (320) to properly decode the data and / or to more accurately restore the original video data. The additional data can be in the form of, for example, temporal, spatial, or SNR extension layers, redundant slices, redundant pictures, forward error correction codes, etc.
[0043] FIG. 4 can be a functional block diagram of a video encoder (203) according to an embodiment of the present disclosure.
[0044] The encoder (203) can receive video samples from a video source (201) (not part of the encoder) that can capture the video images encoded by the encoder (203).
[0045] The video source (201) can provide a source video sequence to be encoded by an encoder (203) in the form of a digital video sample stream that can be of any suitable bit depth (e.g., 8-bit, 10-bit, 12-bit, …), any color space (e.g., BT.601 Y CrCB, RGB, …), and any suitable sampling structure (e.g., Y CrCb 4:2:0, Y CrCb 4:4:4). In a media serving system, the video source (201) may be a storage device that stores previously prepared video. In a video conferencing system, the video source (203) may be a camera that captures local image information as a video sequence. The video data may be provided as a plurality of individual pictures that convey motion when viewed in sequence. Each picture itself may be organized as a spatial array of pixels, and each pixel can include one or more samples depending on the sampling structure, color space, etc. in use. One skilled in the art can easily understand the relationship between pixels and samples. The following description focuses on samples.
[0046] According to one embodiment, the encoder (203) can encode and compress the pictures of the source video sequence into an encoded video sequence (443) in real time or under any other time constraints required by the application. Enforcing an appropriate coding rate is one function of the controller (450). The controller controls and is functionally coupled to the other functional units described below. For clarity, the couplings are not depicted. The parameters set by the controller can include rate control related parameters (picture skip, quantizer, lambda value of rate distortion optimization techniques, …), picture size, layout of picture groups (GOP), maximum motion vector search range, etc. One skilled in the art can easily identify the other functions of the controller (450) as may be relevant for a video encoder (203) optimized for a certain system design.
[0047] Some video encoders operate in what those skilled in the art would readily recognize as a “coding loop”. As an overly simplified explanation, the coding loop can consist of an encoding portion of an encoder (430) (hereinafter “source encoder”) that is involved in creating symbols based on input pictures and reference pictures to be coded, and a (local) decoder (433) incorporated in an encoder (203) that restores the symbols and creates sample data that a (remote) decoder also creates, since in the video compression techniques contemplated in the disclosed subject matter any compression between the symbols and the coded video bit stream is reversible. The restored sample stream is input into a reference picture memory (434). Since the decoding of the symbol stream leads to bit-exact results regardless of the location of the decoder (local or remote), the contents of the reference picture buffer are also bit-exact between the local encoder and the remote encoder. In other words, the prediction portion of the encoder “sees” the same sample values as the decoder “sees” as reference picture samples when using prediction during decoding. This basic principle of reference picture synchronization (and the resulting drift if synchronization cannot be maintained, for example, due to channel errors) is well known to those skilled in the art.
[0048] The operation of the “local” decoder (433) can be the same as the operation of the “remote” decoder (210) already described in detail above in conjunction with FIG. 3. However, also referring briefly to FIG. 3, since symbols are available and the encoding / decoding of the symbols into the coded video sequence by the entropy encoder (445) and the parser (320) can be reversible, the entropy decoding portion of the decoder (210) including the channel (312), the receiver (310), the buffer (315), and the parser (320) may not be fully implemented in the local decoder (433).
[0049] The observation that can be made at this point is that any decoder technology other than the syntax analysis / entropy decoding present in the decoder must necessarily exist in substantially the same functional form within the corresponding encoder. For this reason, the disclosed subject matter focuses on the operation of the decoder. The description of the encoder technology can be omitted since it is the reverse of the decoder technology described comprehensively. A more detailed description is necessary only for specific areas and is provided below.
[0050] As part of its operation, the source coder (430) can perform motion-compensated predictive coding that predictively codes an input frame by referring to one or more previously coded frames from a video sequence designated as a "reference frame". In this way, the coding engine (432) codes the difference between a pixel block of the input frame and a pixel block of a reference frame that can be selected as a prediction reference for the input frame.
[0051] The local video decoder (433) can decode the coded video data of a frame that can be designated as a reference frame based on the symbols created by the source coder (430). The operation of the coding engine (432) can advantageously be a non-invertible process. When the coded video data can be decoded by a video decoder (not shown in FIG. 4), the restored video sequence can typically be a replica of the source video sequence with some errors. The local video decoder (433) can replicate the decoding process that can be performed by the video decoder for the reference frame and cause the restored reference frame to be stored in the reference picture cache (434). In this way, the encoder (203) can locally store a copy of the restored reference frame having common content as the restored reference frame obtained by a remote video decoder (in the absence of transmission errors).
[0052] Predictor (435) can perform predictive search for the coding engine (432). That is, for a new frame to be coded, predictor (435) can search the reference picture memory (434) for sample data (as candidate reference pixel blocks), or specific metadata such as reference picture motion vectors, block shapes, etc., which can serve as appropriate prediction criteria for the new picture. Predictor (435) can operate on a sample block - to - pixel block basis to find appropriate prediction criteria. In some cases, the input picture can have prediction criteria drawn from a plurality of reference pictures stored in the reference picture memory (434), as determined by the search results obtained by predictor (435).
[0053] Controller (450) can manage the coding operations of video coder (430), including, for example, setting parameters and subgroup parameters used for encoding video data.
[0054] The outputs of all the aforementioned functional units can undergo entropy coding within entropy coder (445). The entropy coder converts the symbols generated by various functional units into a coded video sequence by reversibly compressing the symbols according to techniques known to those skilled in the art, such as Huffman coding, variable - length coding, arithmetic coding, etc.
[0055] Transmitter (440) can buffer the coded video sequence created by entropy coder (445) in preparation for transmission via communication channel (460), which may be a hardware / software link to a storage device storing the encoded video data. Transmitter (440) can merge the coded video data from video coder (430) with other data to be transmitted, such as coded audio data and / or an auxiliary data stream (from a source not shown).
[0056] The controller (450) can manage the operation of the video encoder (203). During coding, the controller (450) can assign a specific coded picture type to each coded picture, which may affect the coding technique applicable to each picture. For example, a picture may often be assigned as one of the following frame types.
[0057] An intra picture (I picture) can be a picture that can be coded and decoded without using any other frame in the sequence as a prediction source. Some video codecs allow various types of intra pictures, including, for example, independent decoder refresh pictures. Those skilled in the art know those variants of I pictures, as well as their respective uses and characteristics.
[0058] A predicted picture (P picture) can be a picture that can be coded and decoded using intra prediction or inter prediction that uses at most one motion vector and a reference index to predict the sample values of each block.
[0059] A bi-directionally predicted picture (B picture) can be a picture that can be coded and decoded using intra prediction or inter prediction that uses at most two motion vectors and reference indices to predict the sample values of each block. Similarly, a multi-predicted picture can use three or more reference pictures and related metadata for the restoration of a single block.
[0060] The source picture is typically spatially subdivided into a plurality of sample blocks (e.g., blocks of 4×4, 8×8, 4×8, or 16×16 samples each) and can be coded block by block. The blocks can be coded predictively with reference to other (already coded) blocks as determined by the coding assignment applied to each picture of the block. For example, blocks of an I picture can be coded non-predictively or they can be coded predictively with reference to already coded blocks of the same picture (spatial prediction or intra prediction). Pixel blocks of a P picture can be coded non-predictively via spatial prediction or via temporal prediction with reference to one previously coded reference picture. Blocks of a B picture can be coded non-predictively via spatial prediction or via temporal prediction with reference to one or two previously coded reference pictures.
[0061] The video encoder (203) can perform coding operations according to a given video coding technology or standard such as ITU-T Rec.H.265. In its operation, the video coder (203) can perform various compression operations including predictive coding operations that utilize the temporal and spatial redundancy in the input video sequence. Thus, the coded video data can conform to the syntax specified by the video coding technology or standard being used.
[0062] In one embodiment, the transmitter (440) can transmit additional data along with the encoded video. The video coder (430) may include such data as part of the encoded video sequence. The additional data may include temporal / spatial / SNR enhancement layers, other forms of redundant data such as redundant pictures and slices, supplementary enhancement information (SEI) messages, visual user utility information (VUI) parameter set fragments, and the like.
[0063] Before describing some aspects of the disclosed subject matter in more detail, some terms that are referred to throughout the remainder of this description need to be introduced.
[0064] Hereinafter, a sub-picture refers to a rectangular arrangement of samples, blocks, macroblocks, coding units, or similar entities that can be semantically grouped and independently coded at a modified resolution in some cases. One or more sub-pictures can form one picture. One or more coded sub-pictures can form one coded picture. One or more sub-pictures may be assembled into a picture, and one or more sub-pictures may be extracted from a picture. In certain environments, one or more coded sub-pictures may be assembled within a compressed region without transcoding the coded picture to the sample level, and in the same or certain other cases, one or more coded sub-pictures may be extracted from the coded picture within the compressed region.
[0065] Hereinafter, reference picture resampling (RPR) or adaptive resolution change (ARC) refers to a mechanism that enables, for example, changing the resolution of a picture or sub-picture within a coded video sequence by reference picture resampling. Hereinafter, RPR / ARC parameters refer to the control information necessary to perform adaptive resolution change, which may include, for example, filter parameters, scaling factors, the resolution of the output and / or reference pictures, various control flags, and the like.
[0066] The above description focuses on the coding and decoding of a single semantically independent coded video picture. Options for signaling RPR / ARC parameters are described before describing the implications of coding / decoding multiple sub-pictures with independent RPR / ARC parameters and the additional complexity thereby implied.
[0067] Referring to FIG. 5, several new options for signaling RPR / ARC parameters are shown. As described for each option, they have certain advantages and certain disadvantages from the perspectives of coding efficiency, complexity, and architecture. A video coding standard or technology can select one or more of these options, or options known from the prior art, for signaling RPR / ARC parameters. The options may not be mutually exclusive and may be exchanged depending on, among other things, the needs of the application, the associated standard technology, or the choice of encoder.
[0068] The classes of RPR / ARC parameters may include the following. - Upsample / downsample coefficients separated or combined in the X and Y dimensions, - Upsample / downsample coefficients with an added temporal dimension indicating a constant rate of zoom in / zoom out for a given number of pictures, - Either of the above two may include the coding of one or more possibly short syntax elements that can refer to a table containing coefficients, - Resolution in the X or Y dimension of input pictures, output pictures, reference pictures, samples of coded pictures, blocks, macroblocks, CUs, or any other suitable granularity unit that is combined or separated (if there are two or more resolutions, such as for the input picture resolution, reference picture resolution, etc., in some cases, a set of values may be inferred from another set of values. It can be gated, for example, by using a flag. For more detailed examples, see below.). - Similarly, at an appropriate granularity as described above, "warping" coordinates similar to those used in H.263 Annex P (H.263 Annex P defines one efficient way to code such warping coordinates, but other potentially more efficient ways may also be devised. For example, according to an embodiment, the variable-length reversible "Huffman"-style coding of the warping coordinates of Annex P is replaced by a binary coding of appropriate length, and the length of the binary codeword is derived, for example, from the maximum picture size, optionally multiplied by a specific factor, and offset by a specific value to enable "warping" outside the boundaries of the maximum picture size.), and / or - Upsampling or downsampling filter parameters (in the simplest case, there may be only a single filter for upsampling and / or downsampling. However, in some cases, it may be advantageous to allow more flexibility in filter design, which may require signaling of filter parameters. Such parameters may be selected via an index within a list of possible filter designs, the filter may be fully specified (e.g., via a list of filter coefficients using appropriate entropy coding techniques), and the filter may be implicitly selected via the upsampling / downsampling ratio signaled according to any of the mechanisms described above).
[0069] Subsequently, the description assumes the coding of a finite set of upsampling / downsampling coefficients (the same coefficients used in both the X and Y dimensions) indicated via codewords. The codewords can advantageously be variable-length coded using, for example, the Exp-Golomb code common to certain syntax elements in video coding specifications such as H.264 and H.265. One appropriate mapping of values to upsampling / downsampling coefficients can follow, for example, Table 1 below.
[0070]
Table 1
[0071] According to the needs of the application and the capabilities of the upscaling mechanism and downscaling mechanism available in the video compression technology or standard, many similar mappings can be devised. The table can be extended to more values. The values may also be represented by an entropy coding mechanism other than the Ext-Golomb code, for example, using binary coding. This can have certain advantages, for example, by the MANE, when the resampling factor is an object external to the video processing engine (initially the encoder and decoder) itself. Note that in the most common case where resolution change is not required (presumably), a short Ext-Golomb code of only 1 bit in the above table can be selected. It can have the advantage of coding efficiency over using binary coding in the most common cases.
[0072] The number of items in the table, as well as their semantics, may be fully or partially configurable. For example, the basic outline of the table may be transmitted in a "high" parameter set such as a sequence parameter set or a decoder parameter set. Alternatively or additionally, one or more such tables may be defined in the video coding technology or standard and may be selected, for example, via a decoder parameter set or a sequence parameter set.
[0073] Subsequently, how the upsampling / downsampling factor (ARC information) coded as described above can be included in the syntax of the video coding technology or convention is described. Similar considerations may apply to one or several codewords that control the upsampling / downsampling filter. For an explanation when a relatively large amount of data is required for the filter or other data structure, see below.
[0074] As shown in the example of FIG. 5A, the illustration (500A) shows that H.263 Annex P includes ARC information (502) in the form of four warping coordinates in the picture header (501), specifically in the H.263 PLUS PTYPE (503) header extension. This can be a wise design choice when a) there is an available picture header and b) frequent changes to the ARC information are expected. However, the overhead when using H.263-style signaling can be very high, and the picture header can be of a temporary nature, so the scaling factor may not be related between picture boundaries. Further, as shown in the example of FIG. 5B, the illustration (500B) shows that JVET-M0135 includes PPS information (504), ARC reference information (505), SPS information (507), and target resolution table information (506).
[0075] According to an exemplary embodiment, FIG. 5C shows an example (500C) in which tile group header information (508) and ARC information (509) are shown, FIG. 5D shows an example (500D) in which tile group header information (514), ARC reference information (513), SPS information (516), and ARC information (515) are shown, and FIG. 5E shows an example (500E) in which adaptive parameter set (APS) information (511) and ARC information (512) are shown.
[0076] FIG. 6 shows an example (600) of a table in which the adaptive resolution is in use, in which example the output resolution is coded in sample units (613). The number 613 refers to both output_pic_width_in_luma_samples and output_pic_height_in_luma_samples, which can together define the resolution of the output picture. In other places in the video coding technology or standard, specific restrictions for any value can be defined. For example, the level definition can limit the total number of output samples that can be the product of the values of those two syntax elements. Also, a particular video coding technology or standard, or an external technology or standard such as a system standard, for example, can limit the numbering range (for example, one or both dimensions must be divisible by a power of two), or the aspect ratio (for example, the width and height must be in a relationship such as 4:3 or 16:9). Such restrictions may be introduced to facilitate hardware implementation or for other reasons, as will be understood by those skilled in the art in view of the present disclosure.
[0077] For a particular application, it may be desirable for the encoder to instruct the decoder to use a particular reference picture size instead of implicitly assuming that the size is the output picture size. In this example, the syntax element reference_pic_size_present_flag (614) gates the conditional presence of the reference picture dimensions (615) (similarly, the number refers to both the width and height).
[0078] Certain video coding techniques or standards, such as VP9, support spatial scalability by performing a specific form of reference picture resampling (signaled in a manner completely different from the disclosed subject matter) along with temporal scalability to enable spatial scalability. In particular, certain reference pictures can be upsampled to a higher resolution using ARC-style techniques to form the basis of a spatial enhancement layer. Those upsampled pictures can be improved using normal prediction mechanisms at high resolution to add details.
[0079] The disclosed subject matter can be used and is used in such an environment according to embodiments. In some cases, in the same or different embodiments, values in the NAL unit header, such as the temporal ID field, can be used to indicate not only the temporal layer but also the spatial layer. By doing so, there are specific advantages in certain system designs. For example, existing selective forwarding units (SFUs) created and optimized for temporal layer selection forwarding based on the value of the NAL unit header temporal ID can be used without modification for a scalable environment. To enable that, there may be requirements for the mapping between the coded picture size and the temporal layer, which is indicated by the temporal ID field in the NAL unit header.
[0080] In embodiments, information regarding inter-layer dependency relationships may be signaled within the VPS (or DPS, SPS, or SEI message). The inter-layer dependency information may be used to identify which layer can be used as a reference layer for decoding the current layer. The decoded picture picA within the direct dependency layer where nuh_layer_id is equal to m may be used as a reference picture for the picture picB where nuh_layer_id is equal to n when n is greater than m and the two pictures picA and picB belong to the same access unit.
[0081] In the same or other embodiments, the inter-layer reference picture (ILRP) list may be the inter-prediction reference picture (IPRP) list within the slice header (or parameter set). together It may be explicitly signaled. Both the ILRP list and the IPRP list may be used in constructing the forward and backward prediction reference picture lists.
[0082] In the same or other embodiments, the syntax elements within the VPS (or other parameter set) can indicate whether each layer is dependent or independent. Referring to the example (700) of FIG. 7, the syntax element vps_max_layers_minus1 (703) plus 1 can specify the maximum number of layers allowed in one or potentially all CVSs that refer to the VPS (701). The vps_all_independent_layers_flag (704) equal to 1 can specify that all layers within the CVS are coded independently, i.e., without using inter-layer prediction. The vps_all_independent_layers_flag (704) equal to 0 can specify that one or more of the layers within the CVS can use inter-layer prediction. When not present, the value of the vps_all_independent_layers_flag may be presumed to be equal to 1. When the vps_all_independent_layers_flag is equal to 1, the value of the vps_independent_layer_flag[i] (706) may be presumed to be equal to 1. When the vps_all_independent_layers_flag is equal to 0, the value of the vps_independent_layer_flag[0] is presumed to be equal to 1.
[0083] Referring to FIG. 7, vps_independent_layer_flag[i] (706) equal to 1 can specify that the layer with index i does not use inter-layer prediction. vps_independent_layer_flag[i] equal to 0 can specify that the layer with index i can use inter-layer prediction and that vps_layer_dependency_flag[i] exists within the VPS. vps_direct_dependency_flag[i][j] (707) equal to 0 can specify that the layer with index j is not a direct reference layer for the layer with index i. vps_direct_dependency_flag[i][j] equal to 1 can specify that the layer with index j is a direct reference layer for the layer with index i. When vps_direct_dependency_flag[i][j] does not exist for i and j in the range of 0 or more and vps_max_layers_minus1 or less, it may be assumed to be equal to 0.
[0084] The variable DirectDependentLayerIdx[i][j] that specifies the j-th direct dependent layer of the i-th layer, and the variable NumDependentLayers[i] that specifies the number of dependent layers of the i-th layer are derived as follows. for(i = 1; i < vps_max_layers_minus1; i--) if(!vps_independent_layer_flag[ i ]){ for(j = i, k = 0; j >= 0; j--) if(vps_direct_dependency_flag[ i ][ j ]) DirectDependentLayerIdx[ i ][ k++]=j NumDependentLayers[ i ]=k }
[0085] In the same or another embodiment, referring to FIG. 7, when vps_max_layers_minus1 is greater than 0 and the value of vps_all_independent_layers_flag is equal to 0, vps_output_layers_mode and vps_output_layer_flags[i] may be signaled. A vps_output_layers_mode (708) equal to 0 can specify that only the topmost layer is output. A vps_output_layer_mode equal to 1 specifies that all layers can be output. A vps_output_layer_mode equal to 2 can specify that the layers to be output are the layers where vps_output_layer_flag[i] (709) is equal to 1. The value of vps_output_layers_mode shall be in the range from 0 to 2 inclusive. A value of 3 for vps_output_layer_mode may be reserved for future use. When absent, the value of vps_output_layers_mode may be assumed to be equal to 1. A vps_output_layer_flag[i] equal to 1 can specify that the i-th layer is output. A vps_output_layer_flag[i] equal to 0 can specify that the i-th layer is not output. The list OutputLayerFlag[i], where a value of 1 can specify that the i-th layer is output and a value of 0 can specify that the i-th layer is not output, is derived as follows. OutputLayerFlag[ vps_max_layers_minus1 ]=1 for(i = 0; i < vps_max_layers_minus1; i++) if(vps_output_layer_mode == 0) OutputLayerFlag[ i ]=0 else if(vps_output_layer_mode == 1) OutputLayerFlag[ i ]=1 else if (vps_output_layer_mode == 2) OutputLayerFlag[i] = vps_output_layer_flag[i]
[0086] In the same or another embodiment, the output of the current picture may be specified as follows. - If PictureOutputFlag is equal to 1 and DpbOutputTime[n] is equal to CpbRemovalTime[n], the current picture is output. - Otherwise, if PictureOutputFlag is equal to 0, the current picture is not output but is stored in the DPB as specified in the slice. - Otherwise (PictureOutputFlag is equal to 1 and DpbOutputTime[n] is greater than CpbRemovalTime[n]), the current picture is output later, stored in the DPB (as specified in the slice), and is output at time DpbOutputTime[n] unless it is indicated by decoding or inferring no_output_of_prior_pics_flag equal to 1 at a time preceding DpbOutputTime[n] that it is not to be output. When output, the picture is cropped using the compliance cropping window specified in the PPS for the picture.
[0087] In the same or another embodiment, PictureOutputFlag may be set as follows. - If any one of the following conditions is true, PictureOutputFlag is set equal to 0, - The current picture is a RASL picture and NoIncorrectPicOutputFlag of the associated IRAP picture is equal to 1. - gdr_enabled_flag is equal to 1 and the current picture is a GDR picture with NoIncorrectPicOutputFlag equal to 1. - The gdr_enabled_flag is equal to 1, the current picture is associated with a GDR picture where NoIncorrectPicOutputFlag is equal to 1, and the PicOrderCntVal of the current picture is less than the RpPicOrderCntVal of the associated GDR picture. - The vps_output_layer_mode is equal to 0 or 2, and OutputLayerFlag[GeneralLayerIdx[nuh_layer_id]] is equal to 0. - Otherwise, PictureOutputFlag is set equal to pic_output_flag.
[0088] Alternatively, in the same or other embodiments, PictureOutputFlag may be set as follows. - If one of the following conditions is true, PictureOutputFlag is set equal to 0, - The current picture is a RASL picture and the NoIncorrectPicOutputFlag of the associated IRAP picture is equal to 1. - The gdr_enabled_flag is equal to 1 and the current picture is a GDR picture where NoIncorrectPicOutputFlag is equal to 1. - The gdr_enabled_flag is equal to 1, the current picture is associated with a GDR picture where NoIncorrectPicOutputFlag is equal to 1, and the PicOrderCntVal of the current picture is less than the RpPicOrderCntVal of the associated GDR picture. - The vps_output_layer_mode is equal to 0, the current access unit has a PictureOutputFlag equal to 1, has a nuh_layer_id nuhLid greater than the current picture, and includes a picture belonging to the output layer (i.e., OutputLayerFlag[GeneralLayerIdx[nuhLid]] is equal to 1). - The -vps_output_layer_mode is equal to 2 and OutputLayerFlag[GeneralLayerIdx[nuh_layer_id]] is equal to 0. - Otherwise, PictureOutputFlag is set equal to pic_output_flag.
[0089] In the same or other embodiments, a flag within the VPS (or another parameter set) can indicate whether an ILRP list is signaled for the current slice (or picture). For example, referring to the example (800) of FIG. 8, an inter_layer_ref_pics_present_flag equal to 0 can specify that no ILRP is used for inter prediction of any coded picture within the CVS. An inter_layer_ref_pics_flag equal to 1 can specify that an ILRP can be used for inter prediction of one or more coded pictures within the CVS.
[0090] In the same or other embodiments, when the k-th layer is a dependent layer, the inter-layer reference picture (ILRP) list for pictures within the k-th layer may or may not be signaled. However, when the k-th layer is an independent layer, the ILRP list for pictures within the k-th layer is not signaled and no ILRP is included in the reference picture list.
[0091] The value of inter_layer_ref_pics_present_flag may be set equal to 0 when sps_video_parameter_set_id is equal to 0, when nuh_layer_id is equal to 0, or when vps_independent_layer_flag[GeneralLayerIdx[nuh_layer_id]] is equal to 1.
[0092] In the same or other embodiments, referring to the example (900) of FIG. 9, a set of syntax elements that explicitly indicate an ILRP list may be signaled within an SPS, PPS, APS, or slice header. The ILRP list may be used to construct the reference picture list of the current picture.
[0093] In the same or other embodiments, the ILRP list may be used to identify active or non - active reference pictures within a decoded picture buffer (DPB). An active reference picture may be used as a reference picture for decoding the current picture, and a non - active reference picture may not be used for decoding the current picture and may be used for decoding subsequent pictures in decoding order.
[0094] In the same or another embodiment, the ILRP list may be used to identify which reference pictures can be stored in the DPB, or can be output from and removed from the DPB. That information may be used to operate the decoder based on a hypothetical reference decoder (HRD) model and parameters.
[0095] In the same or another embodiment, the syntax element ilrp_idc[listIdx][rplsIdx][i] may be signaled within a VPS, SPS, PPS, APS, or slice header. The syntax element ilrp_idc[listIdx][rplsIdx][i] specifies the index of the ILRP of the i - th item in the ref_pic_list_struct(listIdx,rplsIdx) syntax structure with respect to the list of direct - dependency layers. The value of ilrp_idc[listIdx][rplsIdx][i] shall be in the range from 0 to GeneralLayerIdx[nuh_layer_id] - 1.
[0096] In the same embodiment, the syntax element ilrp_idc[listIdx][rplsIdx][i] may be an index indicating an ILRP picture in the direct dependency layer, which is identified by vps_direct_dependency_flag[i][j] signaled within the VPS. In this case, the value of ilrp_idc[listIdx][rplsIdx][i] shall be in the range of 0 or more and NumDependentLayers[GeneralLayerIdx[nuh_layer_id]] - 1 or less.
[0097] In the same embodiment, when the nuh_layer_id of the current layer is equal to k, signaling an index indicating an ILRP among all layers where nuh_layer_id is less than k, signaling an index indicating an ILRP in the direct dependency layer can be bit-efficient as compared.
[0098] In the same or another embodiment, further referring to FIG. 9, the reference picture lists RefPicList[0] and RefPicList[1] may be constructed as follows. for(i = 0; i < 2; i++){ for(j = 0, k = 0, pocBase = PicOrderCntVal; j < num_ref_entries[ i ][ RplsIdx[ i ] ]; j++){ if(!(inter_layer_ref_pic_flag[ i ][ RplsIdx[ i ] ][ j ] && GeneralLayerIdx[ nuh_layer_id ])) { if(st_ref_pic_flag[ i ][ RplsIdx[ i ] ][ j ]){ RefPicPocList[ i ][ j ] = pocBase - DeltaPocValSt[ i ][ RplsIdx[ i ] ][ j ] if(a reference picture picA having the same nuh_layer_id as the current picture exists in the DPB, and PicOrderCntVal is equal to RefPicPocList[i][j) RefPicList[ i ][ j ]=picA else RefPicList[ i ][ j ]=“no reference picture” (-) pocBase=RefPicPocList[ i ][ j ] } else { if(!delta_poc_msb_cycle_lt[ i ][ k ]){ if(a reference picA with the same nuh_layer_id as the current picture exists in the DPB, and PicOrderCntVal & (MaxPicOrderCntLsb - 1) is equal to PocLsbLt[i][k]) RefPicList[ i ][ j ]=picA else RefPicList[ i ][ j ]=“no reference picture” RefPicLtPocList[ i ][ j ]=PocLsbLt[ i ][ k ] } else { if(a reference picA with the same nuh_layer_id as the current picture exists in the DPB, and PicOrderCntVal is equal to FullPocLt[i][k]) RefPicList[ i ][ j ]=picA else RefPicList[ i ][ j ]=“no reference picture” RefPicLtPocList[ i ][ j ]=FullPocLt[ i ][ k ] } k++ } } else { layerIdx = DirectDependentLayerIdx[GeneralLayerIdx[nuh_layer_id]][ilrp_idc[i][RplsIdx[i]][j]] refPicLayerId = vps_layer_id[layerIdx] if (a reference picture picA where nuh_layer_id is equal to refPicLayerId exists in the DPB, and there exists a PicOrderCntVal that is the same as the current picture) RefPicList[i][j] = picA else RefPicList[i][j] = “no reference picture” } } }
[0099] The techniques for signaling the adaptation resolution parameters described above are implemented as computer software using computer-readable instructions and can be physically stored on one or more computer-readable media. For example, FIG. 10 shows a computer system (1000) suitable for implementing some embodiments of the disclosed subject matter.
[0100] The computer software can be coded using any suitable machine code or computer language that can receive an assembly, compilation, linking, or similar mechanism to create code that includes instructions that can be executed directly, or via interpretation, microcode execution, etc., by a computer central processing unit (CPU), a graphics processing unit (GPU), etc.
[0101] The instructions can be executed on various types of computers or their components, including, for example, personal computers, tablet computers, servers, smartphones, game devices, Internet of Things devices, etc.
[0102] The components shown in FIG. 10 for the computer system (1000) are essentially exemplary and do not imply any limitation on the scope of use or functions of the computer software implementing the embodiments of the present disclosure. The configuration of the components should not be construed as having any dependency or requirement regarding any one or combination of the components shown in the exemplary embodiments of the computer system (1000).
[0103] The computer system (1000) may include a specific human interface input device. Such a human interface input device can respond to input by one or more human users via, for example, tactile input (such as keystrokes, swipes, movement of a data glove), audio input (such as voice, clapping), visual input (such as gestures), and olfactory input (not depicted). The human interface device can also be used to capture specific media that is not necessarily directly related to conscious input by humans, such as audio (such as voice, music, ambient sound), images (such as scanned images, photographic images obtained from a still camera), and video (such as 2D video, 3D video including stereoscopic video).
[0104] The input human interface device may include one or more of a keyboard (1001), a mouse (1002), a trackpad (1003), a touch screen (1010), a joystick (1005), a microphone (1006), a scanner (1007), and a camera (1008) (only one of each is depicted).
[0105] The computer system (1000) may also include certain human interface output devices. Such human interface output devices may, for example, stimulate the senses of one or more human users via tactile output, sound, light, and smell / taste. Such human interface output devices include tactile output devices (e.g., a touch screen (1010), or tactile feedback by a joystick (1005), although there may also be tactile feedback devices that do not function as input devices), audio output devices (such as a speaker (1009), headphones (not depicted)), visual output devices (such as a screen (1010) including a CRT screen, an LCD screen, a plasma screen, an OLED screen, regardless of whether each has a touch screen input function and regardless of whether each has a tactile feedback function, some of which may be capable of outputting two-dimensional visual output or three-dimensional or higher output via means such as stereographic output, virtual reality glasses (not depicted), holographic displays, and smoke tanks (not depicted)), and a printer (not depicted).
[0106] The computer system (1000) can also include human-accessible storage devices and associated media such as an optical medium including a CD / DVD ROM / RW (1020) having a CD / DVD or similar medium (1021), a thumb drive (1022), a removable hard drive or solid state drive (1023), legacy magnetic media such as tapes and floppy disks (not depicted), and special ROM / ASIC / PLD-based devices such as security dongles (not depicted).
[0107] One of ordinary skill in the art should also understand that the term "computer-readable medium" as used in connection with the presently disclosed subject matter does not include a transmission medium, a carrier wave, or other transient signals.
[0108] The computer system (1000) can also include an interface to one or more communication networks. The network can be, for example, wireless, wired, optical. The network can further be local, wide area, metropolitan, vehicle and industrial, real-time, delay tolerant, etc. Examples of networks include local area networks such as Ethernet, wireless LAN, cellular networks including GSM, 3G, 4G, 5G, LTE, etc., wired or wireless wide area digital networks for TV including cable TV, satellite TV, and terrestrial broadcast TV, vehicle and industrial including CANBus. A particular network typically requires an external network interface adapter attached to a particular general-purpose data port (such as a USB port of the computer system (1000)) or a peripheral bus (1049), and other networks are typically integrated into the core of the computer system (1000) by attaching to the system bus described below (such as an Ethernet interface to a PC computer system or a cellular network interface to a smartphone computer system). Using any of these networks, the computer system (1000) can communicate with other entities. Such communication can be only unidirectional reception (such as broadcast TV), only unidirectional transmission (such as CANbus to a particular CANbus device), or bidirectional with other computer systems using, for example, local or wide area digital networks. Particular protocols and protocol stacks can be used with each of those networks and network interfaces described above.
[0109] The aforementioned human interface device, human-accessible storage device, and network interface can be attached to the core (1040) of the computer system (1000).
[0110] The core (1040) can include one or more central processing units (CPUs) (1041), a graphics processing unit (GPU) (1042), a special programmable processing device in the form of a field programmable gate array (FPGA) (1043), a hardware accelerator (1044) for specific tasks, etc. These devices may be connected via a system bus (1048) together with a read-only memory (ROM) (1045), a random access memory (1046), an internal mass storage such as a hard drive or SSD that is not accessible to internal users (1047). In some computer systems, the system bus (1048) may be accessible in the form of one or more physical plugs to enable expansion by additional CPUs, GPUs, etc. Peripheral devices can be attached directly to the system bus (1048) of the core or via a peripheral bus (1049). Architectures for peripheral buses include PCI, USB, etc.
[0111] The CPU (1041), GPU (1042), FPGA (1043), and accelerator (1044) can execute specific instructions that can, in combination, constitute the aforementioned computer code. The computer code can be stored in the ROM (1045) or RAM (1046). Migration data can also be stored in the RAM (1046), while persistent data can be stored, for example, in the internal mass storage (1047). Fast storage and retrieval for any of the memory devices can be enabled using a cache memory that can be closely associated with one or more CPUs (1041), GPUs (1042), mass storage (1047), ROM (1045), RAM (1046), etc.
[0112] A computer-readable medium can have computer code thereon for performing various computer-implemented operations. The medium and the computer code can be specially designed and constructed for the purposes of this disclosure, or they can be of the kind well-known and available to those having skill in the computer software arts.
[0113] As an example, and not by way of limitation, a computer system (1000) having an architecture, specifically a core (1040), can provide functionality as a result of software embodied within one or more tangible computer-readable media being executed by a processor (including a CPU, GPU, FPGA, accelerator, etc.). Such computer-readable media can be the user-accessible mass storage introduced above, as well as media associated with specific storage of the core (1040) of a non-transitory nature such as core internal mass storage (1047) or ROM (1045). The software implementing various embodiments of the present disclosure can be stored in such devices and executed by the core (1040). The computer-readable media can include one or more memory devices or chips depending on specific needs. The software can cause the core (1040), and specifically the processor (including a CPU, GPU, FPGA, etc.) therein, to define data structures stored in RAM (1046) and modify such data structures according to processes defined by the software, thereby executing specific processes or specific portions of specific processes described herein. Additionally or alternatively, the computer system can provide functionality as a result of logic wired or otherwise embodied within a circuit (e.g., an accelerator (1044)) that operates instead of or in conjunction with software to execute specific processes or specific portions of specific processes described herein. Optionally, references to software can include logic and vice versa. Optionally, references to computer-readable media can include a circuit (such as an integrated circuit (IC)) that stores software for execution, a circuit that embodies logic for execution, or both. The present disclosure encompasses any suitable combination of hardware and software.
[0114] Although this disclosure describes some exemplary embodiments, there are changes, substitutions, and various alternative equivalents within the scope of this disclosure. Therefore, it will be understood that those skilled in the art can devise numerous systems and methods that embody the principles of this disclosure and are thus within its spirit and scope, even though not explicitly illustrated or described herein.
Explanation of Reference Numerals
[0115] 100 Communication system 110 Terminal 120 Terminal 130 Terminal 140 Terminal 150 Network 201 Video source 202 Uncompressed video sample stream 203 Video encoder 204 Encoded video bitstream 205 Streaming server 206 Streaming client 207 Copy of encoded video bitstream 208 Streaming client 209 Copy of encoded video bitstream 210 Video decoder 211 Output video sample stream 212 Display 213 Capture subsystem 310 Receiver 312 Channel 315 Buffer memory 320 Entropy decoder / parser 321 Symbol 351 Scaler / inverse transform unit 352 Intra-picture prediction unit 353 Motion compensation prediction unit 355 Aggregator 356 Current picture / loop filter unit 357 Reference picture memory 430 Source Coder 432 Coding Engine 433 (Local) Decoder 434 Reference Picture Memory 435 Predictor 440 Transmitter 443 Encoded Video Sequence 445 Entropy Coder 450 Controller 460 Communication Channel 500A Exemplification 500B Exemplification 501 Picture Header 502 ARC Information 503 H.263 PLUS PTYPE 504 PPS Information 505 ARC Reference Information 506 Target Resolution Table Information 507 SPS Information 508 Tile Group Header Information 509 ARC Information 511 Adaptive Parameter Set (APS) Information 512 ARC Information 513 ARC Reference Information 514 Tile Group Header Information 515 ARC Information 516 SPS Information 600 Example of Table 615 Reference Picture Dimension 700 Example 701 VPS 800 Example 900 Example 1000 Computer System 1001 Keyboard 1002 Mouse 1003 Track Pad 1005 Joystick 1006 Microphone 1007 Scanner 1008 Camera 1009 Speaker 1010 Touch screen 1020 CD / DVD ROM / RW 1021 CD / DVD or similar media 1022 Thumb drive 1023 Removable hard drive or solid state drive 1040 Core 1041 Central Processing Unit (CPU) 1042 Graphics Processing Unit (GPU) 1043 Field Programmable Gate Array (FPGA) 1044 Hardware accelerator 1045 Read Only Memory (ROM) 1046 Random Access Memory (RAM) 1047 Internal mass storage 1048 System bus 1049 Peripheral bus 1050 Graphics adapter 1054 Network interface
Claims
1. A method for video encoding of a scalable bitstream executed by at least one processor, comprising: encoding at least one video parameter set (VPS) including at least one syntax element indicating whether at least one layer in the scalable bitstream is one of a dependent layer and an independent layer of the scalable bitstream into the scalable bitstream; signaling an inter-layer reference picture (ILRP) list used for decoding pictures in the dependent layer together with an inter-prediction reference picture (IPRP) list in the VPS, wherein the ILRP list includes an index of the ILRP, and when the index of the IPRP indicates an ILRP picture in the direct dependent layer, the value of the index of the ILRP is in a range from 0 to a value obtained by subtracting 1 from the number of the dependent layers; A method comprising the above steps.
2. The method according to claim 1, wherein the VPS includes a plurality of flags, the plurality of flags including a first flag specifying that each of the at least one layer is the independent layer and a second flag specifying that each of the at least one layer is a direct dependent layer, and the number of dependent layers including the dependent layer of the scalable bitstream is determined for each of the at least one layer based on the plurality of flags by determining whether the first flag is affirmative or negative, determining whether the second flag is affirmative or negative when the first flag is negative, determining that the layer for which the second flag is affirmative is the direct dependent layer when the second flag is affirmative, and determining that the number of the direct dependent layers is the number of the dependent layers.
3. The method according to claim 1 or 2, wherein the inter-layer reference picture list includes decoded pictures of other layers.
4. The method according to any one of claims 1 to 3, wherein the at least one VPS includes other syntax elements indicating the maximum number of layers. **Claim 5**: The method according to claim 2, wherein the second flag indicates whether another layer within the scalable bitstream is a reference layer for the at least one layer. **Claim 6** The at least one VPS indicates that the second flag indicates the other layer as the reference layer for the at least one layer by specifying the index of the other layer and the index of the at least one layer, and the at least one VPS is used to determine whether the at least one VPS includes another syntax element indicating a value less than the determined number of dependent layers. The method according to claim 5. **Claim 7** The at least one VPS indicates that the second flag does not indicate the other layer as the reference layer for the at least one layer by specifying the index of the other layer and the index of the at least one layer, and the at least one VPS is used to determine whether the at least one VPS includes another syntax element indicating a value less than the determined number of dependent layers. The method according to claim 5. **Claim 8** The method according to claim 2, wherein the at least one VPS is used to determine whether the at least one VPS includes the first flag indicating whether a plurality of layers including the at least one layer should be decoded by interpreting the ILRP list. **Claim 9** The method according to claim 2, wherein the at least one VPS is used to determine whether the at least one VPS includes the first flag indicating whether a plurality of layers including the at least one layer should be decoded without interpreting the ILRP list. **Claim 10** An apparatus configured to execute the method according to any one of claims 1 to 9. **Claim 11** A program for causing a computer to execute the method according to any one of claims 1 to 9.
Citation Information
Patent Citations
Image decoding apparatus and computer-executable program
JP2010226376A
Video decoder error handling
US20090213938A1
High level syntax for HEVC extensions
US20150103886A1
Image decoding device and image coding device
US20160191933A1
Method and apparatus for encoding multilayer video, and method and apparatus for decoding multilayer video
US20160227232A1