Decoded picture buffer management for video encoding - Patent Application 20070122997
By identifying and removing sub-layer non-reference pictures from the decoded picture buffer, the method optimizes video coding efficiency by reducing unnecessary storage and improving resource utilization in decoded picture buffer management.
Patent Information
- Application Number
- JP2023045837
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Priority Date
- 2020-03-11
- Filing Date
- 2023-03-22
- Publication Date
- 2025-09-09
- Estimated Expiration
- 2040-03-12
Smart Images

Figure 0007736381000005 
Figure 0007736381000006 
Figure 0007736381000007
Abstract
Description
[Technical Field]
[0001] CROSS-REFERENCE TO RELATED APPLICATIONS This application claims priority to U.S. Provisional Patent Application No. 62 / 819,460, filed March 15, 2019, and U.S. Patent Application No. 16 / 815,710, filed March 11, 2020, the disclosures of which are incorporated herein by reference in their entireties.
[0002] This disclosure relates to a set of advanced video coding techniques, and more specifically to decoded picture buffer management. [Background technology]
[0003] The system for decoding includes a decoded picture buffer for storing pictures used as references in the decoding. Summary of the Invention [Means for solving the problem]
[0004] Some embodiments of the present disclosure improve decoded picture buffer management, for example, by accommodating temporal sub-layer adaptation.
[0005] In some embodiments, a method is provided that includes storing previously decoded pictures of a video stream, including a first plurality of pictures of a same temporal sublayer, in a picture buffer, the first plurality of pictures including at least one sublayer reference picture for predicting a current picture of the video stream, determining whether a picture of the first plurality of pictures is a sublayer non-reference (“SLNR”) picture based on an indicator presented in at least one of a slice header and a picture header, removing the SLNR picture from the picture buffer based on determining that the picture is an SLNR picture, and predicting the current picture using one or more of the at least one sublayer reference picture stored in the picture buffer after removing the SLNR picture from the picture buffer.
[0006] In one embodiment, determining whether a picture of the first plurality of pictures is an SLNR picture includes identifying a network abstraction layer (NAL) unit type of the picture and determining whether the picture is an SLNR picture based on the identified NAL unit type.
[0007] In one embodiment, the method further includes providing an identifier for the picture determined to be an SLNR picture based on the picture being determined to be an SLNR picture, and the removing step includes removing the picture from the fpicture buffer based on the identifier. In one embodiment, the method further includes forming a reference picture list including an entry for each of the first plurality of pictures, and the providing step includes providing an identifier for an entry in the reference picture list corresponding to the picture determined to be an SLNR picture.
[0008] In one embodiment, the previously decoded pictures stored in the picture buffer include a second picture that is a reference picture, and the method further includes determining whether a temporal sub-layer value of the second picture is greater than a predetermined value and removing the second picture from the picture buffer based on determining that the temporal sub-layer value of the second picture is greater than the predetermined value. In one embodiment, the method further includes providing an identifier of the second picture based on determining that the temporal sub-layer value of the second picture is greater than the predetermined value, and removing the second picture includes removing the second picture from the picture buffer based on the identifier. In one embodiment, the method further includes comparing the predetermined value with a value corresponding to the highest temporal sub-layer identification number, and determining whether the temporal sub-layer value of the second picture is greater than the predetermined value occurs based on determining that the predetermined value is not equal to the value corresponding to the highest temporal sub-layer identification number. In one embodiment, the method further includes the steps of determining whether the current picture is an Intra Random Access Point (IRAP) picture; determining whether a flag indicates that there will be no output of a Random Access Skip Leading ("RASL") picture; and determining whether to set a respective identifier for each reference picture stored in the picture buffer based on whether the current picture is determined to be an IRAP picture and whether the flag is determined to indicate that there will be no output of a RASL picture, wherein the respective identifier for each reference picture indicates whether the reference picture should be removed from the picture buffer.
[0009] In one embodiment, the temporal sub-layer value of the second picture is greater than the temporal sub-layer value of the first plurality of pictures stored in the picture buffer.
[0010] In one embodiment, the method further comprises removing pictures not referenced by the reference picture list from the picture buffer based on the pictures not referenced by the reference picture list.
[0011] In some embodiments, a decoder for decoding a video stream is provided, the decoder comprising: a memory configured to store computer program code; and at least one processor configured to access and operate as instructed by the computer program code, the computer program code configured to cause the at least one processor to store previously decoded pictures of the video stream, including a first plurality of pictures of a same temporal sub-layer, in a picture buffer, the first plurality of pictures including at least one sub-layer reference picture for predicting a current picture of the video stream; and to cause the at least one processor to store previously decoded pictures of the video stream, including a first plurality of pictures of a same temporal sub-layer, in a picture buffer, the first plurality of pictures including at least one sub-layer reference picture for predicting a current picture of the video stream. and the picture header; removal code configured to cause the at least one processor to remove the SLNR picture from the picture buffer based on the determination that the picture is an SLNR picture; and prediction code configured to cause the at least one processor to predict the current picture using one or more of the at least one sub-layer reference pictures stored in the picture buffer after removing the SLNR picture from the picture buffer.
[0012] In one embodiment, the decision code is configured to cause the at least one processor to identify a network abstraction layer (NAL) unit type of the picture and determine whether the picture is an SLNR picture based on the identified NAL unit type.
[0013] In one embodiment, the computer program code further includes providing code configured to cause the at least one processor to provide an identifier of the picture determined to be an SLNR picture based on the picture being determined to be an SLNR picture, and the removal code is configured to cause the at least one processor to remove the picture from the picture buffer based on the identifier. In one embodiment, the computer program code further includes forming code configured to cause the at least one processor to form a reference picture list including an entry for each of the first plurality of pictures, and the providing code is configured to cause the at least one processor to provide an identifier of an entry in the reference picture list corresponding to the picture determined to be an SLNR picture.
[0014] In one embodiment, the previously decoded pictures stored in the picture buffer include a second picture that is a reference picture, the decision code is configured to cause the at least one processor to determine whether a value of a temporal sub-layer of the second picture is greater than a predetermined value, and the removal code is configured to cause the at least one processor to remove the second picture from the picture buffer based on determining that the value of the temporal sub-layer of the second picture is greater than the predetermined value.
[0015] In one embodiment, the providing code is configured to cause the at least one processor to provide an identifier of the second picture based on determining that the temporal sub-layer value of the second picture is greater than a predetermined value, and the removal code is configured to cause the at least one processor to remove the second picture from the picture buffer based on the identifier. In one embodiment, the determining code is configured to cause the at least one processor to compare the predetermined value with a value corresponding to the highest temporal sub-layer identification number and determine whether the temporal sub-layer value of the second picture is greater than the predetermined value based on determining that the predetermined value is not equal to the value corresponding to the highest temporal sub-layer identification number. In one embodiment, the decision code is configured to cause the at least one processor to determine whether the current picture is an Intra Random Access Point (IRAP) picture and to determine whether a flag indicates that there will be no output of a Random Access Skip Leading ("RASL") picture, and the computer program code further includes providing code configured to cause the at least one processor to set a respective identifier for each reference picture stored in the picture buffer if the current picture is determined to be an IRAP picture and the flag is determined to indicate that there will be no output of a RASL picture, wherein the respective identifier for each reference picture indicates whether the each reference picture should be removed from the picture buffer.
[0016] In one embodiment, the temporal sub-layer value of the second picture is greater than the temporal sub-layer value of the first plurality of pictures stored in the picture buffer.
[0017] In some embodiments, a non-transitory computer-readable medium storing computer instructions is provided that, when executed by at least one processor, cause the at least one processor to store previously decoded pictures of a video stream, including a first plurality of pictures of a same temporal sublayer, in a picture buffer, the first plurality of pictures including at least one sub-layer reference picture for predicting a current picture of the video stream, determine whether a picture of the first plurality of pictures is a sub-layer non-reference (“SLNR”) picture based on an indicator presented in at least one of a slice header and a picture header, remove the SLNR picture from the picture buffer based on determining that the picture is an SLNR picture, and predict the current picture using one or more of the at least one sub-layer reference pictures stored in the picture buffer after removing the SLNR picture from the picture buffer.
[0018] Further features, nature and various advantages of the disclosed subject matter will become more apparent from the following detailed description and accompanying drawings. [Brief explanation of the drawings]
[0019] [Figure 1] FIG. 1 is a schematic diagram of a simplified block diagram of a communication system according to one embodiment. [Figure 2] FIG. 1 is a schematic diagram of a simplified block diagram of a streaming system according to one embodiment. [Figure 3] FIG. 2 is a schematic diagram of a simplified block diagram of a video decoder and display according to one embodiment. [Figure 4] FIG. 2 is a schematic diagram of a simplified block diagram of a video encoder and a video source according to one embodiment. [Figure 5] 1 is a flow diagram illustrating a process performed by one embodiment. [Figure 6] 1 is a flow diagram illustrating a process performed by one embodiment. [Figure 7] FIG. 1 illustrates a device according to one embodiment. [Figure 8] FIG. 1 illustrates a computer system suitable for implementing embodiments. DETAILED DESCRIPTION OF THE INVENTION
[0020] 1 shows a simplified block diagram of a communication system 100 according to one embodiment of the present disclosure. The system 100 may include at least two terminals 110, 120 interconnected via a network 150. In the case of unidirectional data transmission, a first terminal 110 may encode video data at a local location for transmission to another terminal 120 via the network 150. A second terminal 120 may receive the other terminal's encoded video data from the network 150, decode the encoded data, and display the recovered video data. Unidirectional data transmission may be common in media serving applications, for example.
[0021] 1 shows a second pair of terminals 130, 140 provided to support bidirectional transmission of encoded video, such as may occur during a video conference. For bidirectional transmission of data, each terminal 130, 140 may encode video data captured at a local location for transmission to the other terminal over network 150. Each terminal 130, 140 may also receive encoded video data transmitted by the other terminal, decode the encoded data, and display the recovered video data on a local display device.
[0022] In FIG. 1 , terminals 110-140 may be, for example, servers, personal computers, and smartphones, and / or any other type of terminal. For example, terminals 110-140 may be laptop computers, tablet computers, media players, and / or dedicated video conferencing equipment. Network 150 represents any number of networks that convey encoded video data between terminals 110-140, including, for example, wired and / or wireless communication networks. Communication network 150 may exchange data over circuit-switched and / or packet-switched channels. Exemplary networks include telecommunications networks, local area networks, wide area networks, and / or the Internet. For purposes of this discussion, the architecture and topology of network 150 may not be important to the operation of the present disclosure unless otherwise described herein.
[0023] 2 illustrates the placement of a video encoder and decoder in a streaming environment as an example of an application of the disclosed subject matter. The disclosed subject matter may be used in other video-enabled applications, including, for example, video conferencing, digital TV, and storage of compressed video on digital media including CDs, DVDs, memory sticks, etc.
[0024] 2, the streaming system 200 may include a capture subsystem 213 that includes a video source 201 and an encoder 203. The streaming system 200 may further include at least one streaming server 205 and / or at least one streaming client 206.
[0025] A video source 201 may, for example, create an uncompressed video sample stream 202. The video source 201 may, for example, be a digital camera. The sample stream 202 may be processed by an encoder 203, depicted as a thick line to emphasize the large amount of data when compared to an encoded video bitstream, coupled to the camera 201. The encoder 203 may include hardware, software, or a combination thereof to enable or implement aspects of the disclosed subject matter, as described in more detail below. The encoder 203 may also generate an encoded video bitstream 204. The encoded video bitstream 204, depicted as a thin line to emphasize the smaller amount of data when compared to the uncompressed video sample stream 202, may be stored on a streaming server 205 for future use. One or more streaming clients 206 may access the streaming server 205 to obtain a video bitstream 209, which may be a copy of the encoded video bitstream 204.
[0026] The streaming client 206 may include a video decoder 210 and a display 212. The video decoder 210 may, for example, decode a video bitstream 209, which is an incoming copy of the encoded video bitstream 204, and generate an outgoing video sample stream 211 that may be rendered on a display 212 or another rendering device (not shown). In some streaming systems, the video bitstreams 204, 209 may be encoded according to a particular video encoding / compression standard. Examples of such standards include, but are not limited to, ITU-T Recommendation H.265. A video encoding standard informally known as Versatile Video Coding (VVC) is under development. Embodiments of the present disclosure may be used in the context of VVC.
[0027] FIG. 3 illustrates an example functional block diagram of a video decoder 210 attached to a display 212, according to one embodiment of the present disclosure.
[0028] The video decoder 210 may include a channel 312, a receiver 310, a buffer memory 315, an entropy decoder / parser 320, a scaler / inverse transform unit 351, an intra prediction unit 352, a motion compensated prediction unit 353, an aggregator 355, a loop filter unit 356, a reference picture memory 357, and a current picture memory 358. In at least one embodiment, the video decoder 210 may include an integrated circuit, a series of integrated circuits, and / or other electronic circuitry. The video decoder 210 may also be embodied partially or entirely in software running on one or more CPUs with associated memory.
[0029] In this and other embodiments, the receiver 310 can receive one or more coded video sequences, one coded video sequence at a time, decoded by the decoder 210, with the decoding of each coded video sequence being independent of the other coded video sequences. The coded video sequences can be received from a channel 312, which can be a hardware / software link to a storage device that stores the coded video data. The receiver 310 can receive the coded video data along with other data, such as coded audio data and / or ancillary data streams, which can be forwarded to a respective using entity (not shown). The receiver 310 can separate the coded video sequences from the other data. To suppress network jitter, a buffer memory 315 can be coupled between the receiver 310 and the entropy decoder / parser 320 (hereinafter, "parser"). When the receiver 310 is receiving data from a store-and-forward device with sufficient bandwidth and controllability, or from a synchronous network, the buffer 315 may not be used or may be small. For use in a best effort packet network such as the Internet, buffer 315 may be necessary and may be relatively large and adaptively sized.
[0030] The video decoder 210 may include a parser 320 for reconstructing symbols 321 from the entropy-encoded video sequence. These symbol categories include, for example, information used to manage the operation of the decoder 210 and potential information for controlling a rendering device, such as the display 212, which may be coupled to the decoder as shown in FIG. 2. The rendering device control information may be in the form of, for example, a Supplemental Enhancement Information (SEI) message or a Video Usability Information (VUI) parameter set fragment (not shown). The parser 320 may parse / entropy decode the received coded video sequence. The coding of the coded video sequence may follow a video coding technique or standard and may follow principles well known to those skilled in the art, including variable-length coding, Huffman coding, arithmetic coding with or without context-dependent coding, etc. The parser 320 may extract at least one set of subgroup parameters for a subgroup of pixels for the video decoder from the coded video sequence based on at least one parameter corresponding to the group. The subgroups may include groups of pictures (GOPs), pictures, tiles, slices, macroblocks, coding units (CUs), blocks, transform units (TUs), prediction units (PUs), etc. The parser 320 can also extract from the coded video sequence information such as transform coefficients, quantization parameter values, motion vectors, etc.
[0031] Parser 320 may perform entropy decoding / parsing operations on the video sequence received from buffer 315 to create symbols 321 .
[0032] The reconstruction of symbols 321 may involve multiple different units, depending on the type of coded video picture or portion thereof (e.g., inter-picture and intra-picture, inter-block and intra-block), and other factors. Which units are involved and how they are involved may be controlled by subgroup control information parsed from the coded video sequence by parser 320. The flow of such subgroup control information between parser 320 and multiple units described below is not shown for clarity.
[0033] Beyond the functional blocks already mentioned, the decoder 210 may be conceptually subdivided into several functional units, as described below. In an actual implementation operating under commercial constraints, many of these units will interact closely with each other and may be at least partially integrated with each other. However, for purposes of describing the disclosed subject matter, the following conceptual subdivision into functional units is appropriate.
[0034] One unit may be a scaler / inverse transform unit 351. The scaler / inverse transform unit 351 may receive quantized transform coefficients as well as control information including the transform to use, block size, quantization coefficients, quantization scaling matrix, etc. as symbols 321 from the parser 320. The scaler / inverse transform unit 351 may output blocks containing sample values that may be input to an aggregator 355.
[0035] In some cases, the output samples of the scaler / inverse transform unit 351 may relate to intra-coded blocks, i.e., blocks that do not use prediction information from a previously reconstructed picture but can use prediction information from a previously reconstructed portion of the current picture. Such prediction information may be provided by the intra-picture prediction unit 352. In some cases, the intra-picture prediction unit 352 generates blocks of the same size and shape as the block being reconstructed using surrounding already reconstructed information fetched from the current (partially reconstructed) picture in the current picture memory 358. The aggregator 355 may add, on a sample-by-sample basis, the prediction information generated by the intra-prediction unit 352 to the output sample information provided by the scaler / inverse transform unit 351.
[0036] In other cases, the output samples of the scaler / inverse transform unit 351 may relate to an inter-coded, potentially motion-compensated block. In such cases, the motion-compensated prediction unit 353 may access the reference picture memory 357 to fetch samples used for prediction. After motion-compensating the samples fetched with the symbols 321 associated with the block, these samples may be added to the output of the scaler / inverse transform unit 351 by the aggregator 355 (in this case, referred to as residual samples or residual signals) to generate output sample information. The addresses in the reference picture memory 357 from which the motion-compensated prediction unit 353 fetches the prediction samples may be controlled by motion vectors. The motion vectors may be available to the motion-compensated prediction unit 353, for example, in the form of symbols 321 that may have x, y, and reference picture components. Motion compensation may also include interpolation of sample values fetched from the reference picture memory 357 when sub-sample accurate motion vectors are used, motion vector prediction mechanisms, etc.
[0037] The output samples of aggregator 355 may be subject to various loop filtering techniques in loop filter unit 356. Video compression techniques may include in-loop filter techniques controlled by parameters contained in the coded video bitstream and made available to loop filter unit 356 as symbols 321 from parser 320, but may also be responsive to meta-information obtained during decoding of a previous (in decoding order) part of the coded picture or coded video sequence, as well as to previously reconstructed and loop-filtered sample values.
[0038] The output of the loop filter unit 356 may be a sample stream that may be output to a rendering device such as the display 212, as well as stored in a reference picture memory 357 for use in future inter-picture prediction.
[0039] Once a particular coded picture is fully reconstructed, it can be used as a reference picture for future prediction. Once a coded picture is fully reconstructed and the coded picture is identified as a reference picture (e.g., by parser 320), the current reference picture stored in current picture memory 358 can become part of reference picture buffer 357, and a new current picture memory can be reallocated before starting reconstruction of the next coded picture.
[0040] The video decoder 210 can perform decoding operations according to a predetermined video compression technique, which may be documented in a standard such as ITU-T Rec. H.265. The encoded video sequence may conform to the syntax specified in the video compression technique document or standard being used, in the sense of conforming to the syntax of the video compression technique or standard as specified in the video compression technique document or standard and particularly the profile documents therein. To comply with some video compression techniques or standards, the complexity of the encoded video sequence may also be within a range defined by the level of the video compression technique or standard. In some cases, the level limits the maximum picture size, maximum frame rate, maximum reconstruction sample rate (e.g., measured in megasamples per second), maximum reference picture size, etc. The limits set by the level may, in some cases, be further limited by the specification of a hypothetical reference decoder (HRD) and HRD buffer management metadata signaled in the encoded video sequence.
[0041] In one embodiment, the receiver 310 can receive additional (redundant) data along with the encoded video. The additional data may be included as part of the encoded video sequence. The additional data may be used by the video decoder 210 to properly decode the data and / or to more accurately reconstruct the original video data. The additional data may be in the form of, for example, temporal, spatial, or SNR enhancement layers, redundant slices, redundant pictures, forward error correction codes, etc.
[0042] FIG. 4 illustrates an example functional block diagram of a video encoder 203 associated with a video source 201 according to one embodiment of the present disclosure.
[0043] The video encoder 203 may include, for example, an encoder that is a source coder 430, a coding engine 432, a (local) decoder 433, a reference picture memory 434, a predictor 435, a transmitter 440, an entropy coder 445, a controller 450, and a channel 460.
[0044] The encoder 203 may receive video samples from a video source 201 (not part of the encoder) that may capture video images to be encoded by the encoder 203 .
[0045] The video source 201 may provide a source video sequence to be encoded by the encoder 203 in the form of a digital video sample stream, which may be of any suitable bit depth (e.g., x-bit, 10-bit, 12-bit, ...), any color space (e.g., BT.601 Y CrCB, RGB, ...), and any suitable sampling structure (e.g., Y CrCb 4:2:0, Y CrCb 4:4:4). In a media serving system, the video source 201 may be a storage device that stores previously prepared video. In a video conferencing system, the video source 203 may be a camera that captures local image information as a video sequence. The video data may be provided as multiple individual pictures that, when viewed in sequence, impart motion. The pictures themselves may be organized as a spatial array of pixels, each of which may contain one or more samples, depending on the sampling structure, color space, etc., in use. Those skilled in the art will readily understand the relationship between pixels and samples. The following description focuses on samples.
[0046] According to one embodiment, the encoder 203 can encode and compress pictures of a source video sequence into an encoded video sequence 443 in real time or under any other time constraint required by the application. Performing an appropriate encoding rate may be a function of the controller 450. The controller 450 may also control and be operatively coupled to other functional units, as described below. For clarity, couplings are not shown. Parameters set by the controller 450 may include rate control-related parameters (e.g., picture skip, quantizer, lambda value for rate-distortion optimization techniques), picture size, group-of-picture (GOP) layout, maximum motion vector search range, etc. Those skilled in the art can readily identify other functions of the controller 450 as they may relate to a video encoder 203 optimized for a particular system design.
[0047] Some video encoders operate in what a skilled person would easily recognize as a "coding loop." As a simplified explanation, when the compression of symbols into the encoded video bitstream is lossless for a particular video compression technique, the encoding loop may consist of a source coder 430 encoding portion (responsible for creating symbols based on the input picture to be coded and the reference picture), and a (local) decoder 433 embedded in the encoder 203, which reconstructs the symbols and creates sample data that also creates a (remote) decoder. That reconstructed sample stream may be input to a reference picture memory 434. Because decoding the symbol stream produces bit-accurate results regardless of the decoder's location (local or remote), the contents of the reference picture memory are also bit-accurate between the local and remote encoders. In other words, the prediction portion of the encoder "sees" the exact same sample values as the decoder "sees" when using prediction during decoding. This basic principle of reference picture synchrony (and the resulting drift when synchrony cannot be maintained, for example, due to channel errors) is known to those skilled in the art.
[0048] The operation of the "local" decoder 433 may be substantially the same as the operation of the "remote" decoder 210, which has already been described in detail above in connection with Figure 3. However, because symbols are available and the encoding / decoding of symbols into an encoded video sequence by the entropy coder 445 and parser 320 may be lossless, the entropy decoding portion of the decoder 210, including the channel 312, receiver 310, buffer 315, and parser 320, may not be fully implemented in the local decoder 433.
[0049] An observation that can be made at this time is that decoder techniques other than analysis / entropy decoding present in a decoder may also need to be present in a corresponding encoder in substantially identical functional form. For this reason, the disclosed subject matter focuses on decoder operation. Descriptions of encoder techniques may be omitted, as they may be the inverse of the decoder techniques described generically. Only in certain areas are more detailed descriptions necessary and are provided below.
[0050] As part of its operation, the source coder 430 may perform motion-compensated predictive coding, which predictively codes an input frame with reference to one or more previously coded frames from the video sequence designated as “reference frames.” In this manner, the coding engine 432 codes the differences between pixel blocks of the input frame and pixel blocks of reference frames that may be selected as predictive references for the input frame.
[0051] The local video decoder 433 may decode the encoded video data of a frame that may be designated as a reference frame based on the symbols created by the source coder 430. The operation of the encoding engine 432 may advantageously be a lossy process. When the encoded video data is decoded by a video decoder (not shown in FIG. 4), the reconstructed video sequence may be a replica of the source video sequence, typically with some errors. The local video decoder 433 may replicate the decoding process that may be performed by the video decoder on the reference frame and store the reconstructed reference frame in the reference picture memory 434. In this way, the encoder 203 may locally store a copy of the reconstructed reference frame that has common content as the reconstructed reference frame (without transmission errors) obtained by the far-end video decoder.
[0052] The predictor 435 may perform the prediction search of the coding engine 432. That is, for a new frame to be encoded, the predictor 435 may search the reference picture memory 434 for sample data (as candidate reference pixel blocks) or specific metadata, such as reference picture motion vectors, block shapes, etc., which may serve as suitable prediction references for the new picture. The predictor 435 may operate on sample block by pixel block to find suitable prediction references. In some cases, as determined by the search results obtained by the predictor 435, the input picture may have prediction references drawn from multiple reference pictures stored in the reference picture memory 434.
[0053] Controller 450 may manage the encoding operations of video coder 430, including, for example, setting parameters and subgroup parameters used to encode the video data.
[0054] The output of all the aforementioned functional units may be subjected to entropy coding in entropy coder 445. The entropy coder converts the symbols produced by the various functional units into an encoded video sequence by losslessly compressing the symbols by techniques known to those skilled in the art, such as Huffman coding, variable length coding, arithmetic coding, etc.
[0055] The transmitter 440 may buffer the encoded video sequence created by the entropy coder 445 and prepare it for transmission over a communication channel 460, which may be a hardware / software link to a storage device that stores the encoded video data. The transmitter 440 may merge the encoded video data from the video coder 430 with other data to be transmitted, such as encoded audio data and / or auxiliary data streams (sources not shown).
[0056] The controller 450 may manage the operation of the encoder 203. During encoding, the controller 450 may assign a particular coded picture type to each coded picture, which may affect the coding technique that may be applied to the respective picture. For example, a picture may typically be assigned as an intra-picture (I-picture), a predicted picture (P-picture), or a bidirectionally predicted picture (B-picture).
[0057] An intra picture (I-picture) may be one that can be coded and decoded without using other frames of the sequence as a source of prediction. Some video codecs can use various types of intra pictures, such as independent decoder refresh (IDR) pictures. Those skilled in the art are aware of these variations of I-pictures and their respective uses and characteristics.
[0058] A predictive picture (P picture) may be one that can be coded and decoded using intra- or inter-prediction, which uses at most one motion vector and reference index to predict the sample values of each block.
[0059] Bidirectionally predicted pictures (B-pictures) may be coded and decoded using intra- or inter-prediction, which uses up to two motion vectors and reference indices to predict the sample values of each block. Similarly, multiple predicted pictures may use more than two reference pictures and associated metadata for the reconstruction of a single block.
[0060] A source picture is typically spatially subdivided into multiple sample blocks (e.g., blocks of 4x4, 8x8, 4x8, or 16x16 samples each) and may be coded block by block. Blocks may be predictively coded with reference to other (already coded) blocks, as determined by the coding assignment applied to the block's respective picture. For example, blocks of an I-picture may be nonpredictively coded, or they may be predictively coded with reference to already coded blocks of the same picture (spatial prediction or intra-prediction). Pixel blocks of a P-picture may be nonpredictively coded via spatial prediction or via temporal prediction with reference to one previously coded reference picture. Pixel blocks of a B-picture may be nonpredictively coded via spatial prediction or via temporal prediction with reference to one or two previously coded reference pictures.
[0061] Video coder 203 may perform encoding operations according to a predetermined video encoding technique or standard, such as ITU-T Rec. H.265. In its operation, video coder 203 may perform various compression operations, including predictive encoding operations that exploit temporal and spatial redundancy in the input video sequence. Thus, the encoded video data may conform to a syntax specified by the video encoding technique or standard being used.
[0062] In one embodiment, the transmitter 440 can transmit additional data along with the coded video. The video coder 430 may include such data as part of the coded video sequence. The additional data may include temporal / spatial / SNR enhancement layers, other forms of redundant data such as redundant pictures and slices, supplemental enhancement information (SEI) messages, visual usability information (VUI) parameter set fragments, etc.
[0063] Encoders and decoders of this disclosure may implement the decoded picture buffer management of this disclosure with respect to a decoded picture buffer (DPB), such as reference picture memory 357 and reference picture memory 434, for example.
[0064] The decoded picture buffer may store decoded pictures that are available for reference to reconstruct subsequent pictures in the decoding process. For example, pictures stored in the decoded picture buffer may be available for use as references in the prediction process of one or more subsequent pictures.
[0065] Encoders and decoders of this disclosure may construct and / or use one or more reference picture lists (e.g., syntax element "RefPicList[ i ]"), each listing a respective picture stored in the decoded picture buffer. For example, each index in the reference picture list may correspond to a respective picture in the decoded picture buffer. The reference picture list may, for example, refer to a list of reference pictures that may be used for inter-prediction.
[0066] Several aspects of the decoded picture buffer management of this disclosure are described below.
[0067] Some embodiments of the present disclosure improve decoded picture buffer management by accommodating temporal sub-layer adaptation. The term "sub-layer" may refer to a temporal scalable layer of a temporal scalable bitstream, including VCL NAL units and associated non-VCL NAL units that have a particular value of the TemporalId variable.
[0068] For example, in one embodiment, the network abstraction layer (NAL) units "TRAIL_NUT", "STSA_NUT", "RASL_NUT", and "RADL_NUT" are redesignated as ("TRAIL_N", "TRAIL_R"), ("STSA_N", "STSA_R"), ("RASL_N, RASL_R"), and ("RADL_N, RASL_R"), respectively, to indicate whether pictures of the same temporal sublayer are reference or non-reference pictures. RefPicList[ i ] may include non-reference pictures that have the same temporal identifier as the current picture being decoded.
[0069] In one embodiment, "sps_max_dec_pic_buffering_minus1" is signaled for each highest temporal identifier in the sequence parameter set ("SPS").
[0070] In one embodiment, a list of unused reference pictures for each highest temporal identifier is signaled in the tile group header.
[0071] In one embodiment, when the value of the specified highest temporal identifier (e.g., syntax element "HighestTid") is not equal to "sps_max_sub_layers_minus1", all reference pictures with temporal identifiers (e.g., syntax element "TemporalId") greater than the specified highest temporal identifier are marked as "unused for reference".
[0072] According to some embodiments of the present disclosure, NAL units that are not used to predict and reconstruct other subsequent NAL units in the same temporal sublayer may or may not be discarded from the decoded picture buffer depending on the target or available bitrate of the network.
[0073] For example, FIG. 5 is a flow diagram illustrating how encoders and decoders of this disclosure can process corresponding NAL units by parsing and interpreting NAL unit types. As shown in FIG. 5, a decoder (or encoder) can perform process 500. Process 500 may include parsing a NAL unit header of a NAL unit (501) and identifying the NAL unit type of a current NAL unit (502). Subsequently, the decoder (or encoder) can determine whether the current NAL unit will be used to predict and reconstruct a subsequent NAL unit of the same temporal sublayer (503). Based on this determination, the decoder (or encoder) may reconstruct / forward the subsequent NAL unit using the current NAL unit (504), or alternatively, discard the current NAL unit from the decoded picture buffer without using the NAL unit to predict and reconstruct a subsequent NAL unit (505). For example, if it is determined that the current NAL unit will be used to predict and reconstruct a subsequent NAL unit of the same temporal sublayer, the decoder (or encoder) may reconstruct / transfer the subsequent NAL unit using the current NAL unit stored in the decoded picture buffer (504). If the NAL unit will not be used to predict and reconstruct a subsequent NAL unit, the decoder (or encoder) may discard the current NAL unit from the decoded picture buffer without using the NAL unit to predict and reconstruct a subsequent NAL unit (505). Predicting and reconstructing a subsequent NAL unit may refer to decoding the current picture by predicting and reconstructing the current picture using the decoded picture buffer.
[0074] The embodiments of the present disclosure may be used separately or combined in any order. Furthermore, each of the methods, encoders, and decoders of the present disclosure may be implemented by processing circuitry (e.g., one or more processors or one or more integrated circuits). In one example, one or more processors execute a program stored on a non-transitory computer-readable medium to perform the functions of the methods, encoders, and decoders described in the present disclosure.
[0075] As mentioned above, the NAL unit types "TRAIL_NUT", "STSA_NUT", "RASL_NUT", and "RADL_NUT" are split and defined as ("TRAIL_N", "TRAIL_R"), ("STSA_N", "STSA_R"), ("RASL_N", "RASL_R"), and ("RADL_N", "RASL_R") to indicate non-reference pictures of the same sublayer. Thus, encoders and decoders of this disclosure may use, for example, the NAL units listed in Table 1 below.
[0076] [Table 1]
[0077] A picture of a sub-layer may have one of the above NAL unit types. If a picture has a NAL unit type (e.g., syntax element "nal_unit_type") equal to "TRAIL_N", "TSA_N", "STSA_N", "RADL_N", or "RASL_N", the picture is a sub-layer non-reference (SLNR) picture. Otherwise, the picture is a sub-layer reference picture. An SLNR picture may be a picture that contains, in decoding order, samples that cannot be used for inter prediction in the decoding process of a subsequent picture of the same sub-layer. A sub-layer reference picture may be a picture that contains, in decoding order, samples that can be used for inter prediction in the decoding process of a subsequent picture of the same sub-layer. A sub-layer reference picture may also be used for inter prediction in the decoding process of a subsequent picture of a higher sub-layer in decoding order.
[0078] By providing NAL units (e.g., VCL NAL units, etc.) that indicate non-reference pictures, unnecessary NAL units can be discarded for bitrate adaptation. RefPicList[ i ] may include non-reference pictures that have the same temporal ID (indicating the temporal sub-layer to which the picture belongs) as the current picture. In this regard, in one embodiment, non-reference pictures may be marked as "unused reference pictures" and can be quickly removed from the decoded picture buffer.
[0079] For example, in one embodiment, a decoder (or encoder) may determine whether a picture is an SLNR picture based on the NAL unit associated with the picture, and if the picture is an SLNR picture, mark the picture as an "unused reference picture." A picture that may be stored in the decoded picture buffer may be marked by entering an identifier in the picture's entry in a reference picture list, the identifier being, for example, "no reference picture" or "unused reference picture." The decoder (or encoder) may perform such an aspect as part of step 503 of process 500, as shown in FIG. 5. The decoder (or encoder) may then remove the picture from the decoded picture buffer based on the picture being marked. The decoder (or encoder) may perform such an aspect as part of step 505 of process 500, as shown in FIG. 5.
[0080] In one embodiment, the reference picture lists "RefPicList
[0000] " and "RefPicList
[0001] " may be constructed as follows: for(i=0;i < 2;i++){ for(j=0,k=0,pocBase=PicOrderCntVal;j < num_ref_entries[ i ][ RplsIdx[ i ] ];j++){ if(st_ref_pic_flag[ i ][ RplsIdx[ i ] ][ j ]){ RefPicPocList[ i ][ j ]=pocBase-DeltaPocSt[ i ][ RplsIdx[ i ] ][ j ] if(RefPicPocList[ i ][ j ] has a PicOrderCntVal equal to reference picture picA in DPB) && Reference picA is not an SLNR picture with TemporalId equal to that of the current picture) RefPicList[ i ][ j ]=picA else RefPicList[ i ][ j ] = “No reference picture” (8-5) pocBase=RefPicPocList[ i ][ j ] }else{ if(!delta_poc_msb_cycle_lt[ i ][ k ]){ if(poc_lsb_lt[ i ][ RplsIdx[ i ] ][ j ]) there exists a reference picA in DPB with PicOrderCntVal&(MaxPicOrderCntLsb-1) equal to && Reference picA is not an SLNR picture with TemporalId equal to that of the current picture) RefPicList[ i ][ j ]=picA else RefPicList[ i ][ j ] = “No reference pictures” }else{ if(FullPocLt[ i ][ RplsIdx[ i ] ][ j ] There is a reference picA in the DPB with a PicOrderCntVal equal to && Reference picA is not an SLNR picture with TemporalId equal to that of the current picture) RefPicList[ i ][ j ]=picA else RefPicList[ i ][ j ] = “No reference pictures” } k+++ } } }
[0081] In one embodiment, constraints can be applied to bitstream conformance. For example, an encoder or decoder can be constrained to ensure that there are no active entries in RefPicList
[0000] or RefPicList
[0001] for which one or more of the following are true: (1) the entry is equal to "No Reference Picture", or (2) this entry is an SLNR picture and has the same "TemporalId" as the current picture.
[0082] As mentioned above, in one embodiment, the syntax element "sps_max_dec_pic_buffering_minus1" may be signaled for each highest temporal identifier of the SPS (eg, the syntax element "HighestTid").
[0083] The value of the variable "HighestTid" may be determined by external means if such means are available. Otherwise, "HighestTid" may be set equal to the syntax element "sps_max_sub_layers_minus1". The decoder can then estimate the maximum required size of the decoded picture buffer for a given "HighestTid" value.
[0084] In an embodiment, the SPS may include the following example syntax shown in Table 2.
[0085] [Table 2]
[0086] "sps_max_dec_pic_buffering_minus1[ i ]" plus 1 specifies the maximum required size of the decoded picture buffer for a coded video sequence ("CVS") in a unit of the picture storage buffer when "HighestTid" is equal to i. The value of "sps_max_dec_pic_buffering_minus1[ i ]" may range from 0 to 'MaxDpbSize' - 1, inclusive, where 'MaxDpbSize' is specified elsewhere.
[0087] As mentioned above, in one embodiment, the list of unused reference pictures for each highest temporal ID may be signaled in the tile group header.
[0088] Depending on the value of "HighestTid", some reference pictures of each temporal sub-layer may not be used as references for subsequent pictures. In one embodiment, unused reference pictures corresponding to each "HighestTid" value of the tile group header may be explicitly signaled. By explicitly signaling unused reference pictures corresponding to each "HighestTid" value of the tile group header, unused decoded reference pictures may be quickly removed from the DPB.
[0089] In an embodiment, the SPS may include the following example syntax shown in Table 3:
[0090] [Table 3]
[0091] "unused_ref_pic_signaling_enabled_flag" equal to 0 specifies that "num_unused_ref_pic" and "delta_poc_unused_ref_pic[ i ]" are not present in the tile group header and the timing of removal of decoded pictures from the DPB is determined implicitly. "unused_ref_pic_signaling_enabled_flag" equal to 1 specifies that "num_unused_ref_pic" and "delta_poc_unused_ref_pic[ i ]" are present in the tile group header and the timing of removal of decoded pictures from the DPB is explicitly determined by parsing "delta_poc_unused_ref_pic[ i ]".
[0092] In an embodiment, the tile group header may include the following example syntax shown in Table 4.
[0093] [Table 4]
[0094] "num_unused_ref_pic" specifies the number of unused reference picture entries. If not present, the value of this field may be set equal to 0.
[0095] "delta_poc_unused_ref_pic[i]" specifies the absolute difference in picture order count value between the current picture and the unused decoded picture referenced by the i-th entry. The value of "delta_poc_unused_ref_pic[i]" must be between 0 and 2. 15 It can be less than or equal to -1.
[0096] If "unused_ref_pic_signaling_enabled_flag" is equal to 1, the following applies: for(i=0;i < num_unused_ref_pic[ HighestTid ];i++) if (a reference picture picX with PicOrderCntVal equal to (current picture PicOrderCntVal-delta_poc_unused_ref_pic [ HighestTid ][ i ]) exists in the DPB picX is marked as "unused for reference".
[0097] In one embodiment, the decoder (or encoder) may determine whether the picture should be marked as an "unused reference picture" based on the above determination. The decoder (or encoder) may perform such an aspect as part of step 503 of process 500 shown in FIG. 5. The decoder (or encoder) may then remove the picture from the decoded picture buffer based on the picture being marked. The decoder (or encoder) may perform such an aspect as part of step 505 of process 500 shown in FIG. 5.
[0098] According to one aspect of an embodiment, when the value of "HighestTid" is not equal to "sps_max_sub_layers_minus1", all reference pictures whose "TemporalId" is greater than HighestTid may be marked as "unused for reference".
[0099] The value of "HighestTid" can be changed instantly by external means. The "HighestTid" may be used as an input for the sub-bitstream extraction process.
[0100] For example, the process may be invoked once per picture after the decoding of the tile group header and the decoding processes for building the reference picture list for the tile group, but before the decoding of the tile group data. This process may cause one or more reference pictures in the DPB to be marked as "unused for reference" or "used for long-term reference."
[0101] In one embodiment, a decoded picture in the DPB may be marked as "unused for reference," "used for short-term reference," or "used for long-term reference," but only one of these three at any given moment during the operation of the decoding process. Assigning one of these markings to a picture may implicitly remove another of these markings, when applicable. When a picture is referred to as being marked "used for reference," this refers collectively to a picture that is marked as "used for short-term reference" or "used for long-term reference" (but not both).
[0102] Decoded pictures of a DPB may be identified (e.g., indexed) or stored differently within the DPB based on their markings. For example, short-term reference pictures ("STRPs") may be identified by their "PicOrderCntVal" values. Long-term reference pictures ("LTRPs") may be identified by the Log2(MaxLtPicOrderCntLsb) LSBs of their "PicOrderCntVal" values.
[0103] If the current picture is an IRAP picture with "NoRaslOutputFlag" equal to 1, all reference pictures currently in the DPB (if any) are marked as "unused for reference." "NoRaslOutputFlag" equal to 1 may indicate no output of IRAP pictures by the decoder.
[0104] When the value of "HighestTid" is not equal to "sps_max_sub_layers_minus1", all reference pictures whose "TemporalId" is greater than "HighestTid" are marked as "unused for reference".
[0105] As an example, referring to FIG. 6, decoders and encoders of this disclosure may perform process 600. Process 600 may be performed based on a determination that the value of “HighestTid” is not equal to “sps_max_sub_layers_minus1.” As shown in FIG. 6, the decoder (or encoder) may determine a temporal ID value of a reference picture (601), e.g., the first reference picture listed in the DPB or reference picture list. Subsequently, the decoder (or encoder) may compare the temporal ID value of the reference picture with a predetermined value (e.g., the value of “HighestTid”) (602). If the temporal ID value is greater than the predetermined value, the decoder (or encoder) may mark the reference picture as “unused for reference” (603). In one embodiment, the decoder (or encoder) may provide the mark in the DPB or reference picture list.
[0106] Regardless of whether the temporal ID value is greater than the predetermined value, the decoder (or encoder) may then determine in step 602 (604) whether another reference picture exists in the DPB (or reference picture list) that does not have a temporal ID value compared to the predetermined value. If the decoder (or encoder) determines in step 602 that there is another reference picture in the DPB (or reference picture list) that does not have a temporal ID value compared to the predetermined value, the decoder (or encoder) may repeat steps 601-604 for all reference pictures in the DPB (or reference picture list). If the decoder (or encoder) determines in step 602 that all reference pictures in the DPB (or reference picture list) have their respective temporal ID values compared to the predetermined value, the decoder (or encoder) may remove reference pictures marked as "unused for reference" from the DPB (605). The decoder (or encoder) may decode the current picture using the DPB with any number of pictures removed from the DPB (606).
[0107] In an embodiment, the decoder (and encoder) may also perform other functions to decode the current picture using the DPB. For example, the decoder (and encoder) may alternatively or additionally apply the following: (1) for each LTRP entry in RefPicList<0000> or RefPicList<0001>, if the referenced picture is an STRP, the decoder (or encoder) may mark the picture as "used for long-term reference." (2) The decoder (or encoder) may mark each reference picture in the DPB that is not referenced by any entry in RefPicList<0000> or RefPicList<0001> as "unused for reference."
[0108] In one embodiment, the decoder (or encoder) can either remove all reference pictures in the DPB that are marked as "unused for reference" before using the DPB to decode the current picture, or can keep such reference pictures in the DPB and ignore the reference pictures when using the DPB to decode the current picture.
[0109] In an embodiment, the device 800 may comprise a memory storing computer program code that, when executed by at least one processor, may cause the at least one processor to perform the functions of the decoder and encoder described above.
[0110] For example, referring to FIG. 7, the computer program code for device 800 may include storage code 810 , determination code 820 , removal code 830 , and decode code 840 .
[0111] The storage code 810 may be configured to cause the at least one processor to store, in a decoded picture buffer, previously decoded pictures of the video stream, including multiple first pictures of the same temporal sub-layer, where the multiple first pictures include at least one sub-layer reference picture for predicting a current picture of the video stream.
[0112] The decision code 820 may be configured to cause at least one processor to make a decision using one or more of the techniques described above. For example, the decision code 820 may be configured to cause at least one processor to determine whether a picture among the plurality of first pictures is a sub-layer non-reference (“SLNR”) picture. Alternatively or additionally, the decision code 820 may be configured to cause at least one processor to identify a network abstraction layer (NAL) unit type of the picture and determine whether the picture is an SLNR picture based on the identified NAL unit type. Alternatively or additionally, the decision code 820 may be configured to cause at least one processor to determine whether a temporal sub-layer value of the picture is greater than a predetermined value (e.g., the value of “HighestTid”). Alternatively or additionally, the decision code 820 may be configured to cause at least one processor to compare a predetermined value (e.g., the value of “HighestTid”) with a value corresponding to the highest temporal sub-layer identification number. Alternatively or additionally, decision code 820 may be configured to cause at least one processor to determine whether a temporal sub-layer value of the picture is greater than a predetermined value (e.g., the value of "HighestTid") when it is determined that the predetermined value is not equal to the value corresponding to the highest temporal sub-layer identification number. Alternatively or additionally, decision code 820 may be configured to cause at least one processor to determine whether the current picture is an Intra Random Access Point (IRAP) picture and whether a flag indicates that a Random Access Skip Leading ("RASL") picture should not be output.
[0113] The removal code 830 may be configured to cause at least one processor to remove one or more pictures from the decoded picture buffer via one or more of the techniques described above. For example, the removal code 830 may be configured to cause at least one processor to remove an SLNR picture from the decoded picture buffer based on determining that the picture is an SLNR picture. Alternatively or additionally, the removal code 830 may be configured to cause at least one processor to remove a picture from the decoded picture buffer based on determining that the temporal sub-layer value of the picture is greater than a predetermined value (e.g., the value of “HighestTid”). In an embodiment, the removal code 830 may be configured to cause at least one processor to remove a picture from the decoded picture buffer based on an identifier (e.g., a marking such as “unused for reference” or “no reference”).
[0114] The decoding code 840 may be configured to cause the at least one processor to decode a current picture using the decoded picture buffer according to one or more of the techniques described above. For example, in one embodiment, the decoding code 840 includes prediction code configured to cause the at least one processor to predict the current picture using one or more of the at least one sub-layer reference pictures stored in the decoded picture buffer after removing a picture from the decoded picture buffer (e.g., an SLNR picture or a picture marked with an identifier such as "unused for reference" or "no reference").
[0115] In one embodiment, the computer program code may further include providing code 850 and forming code 860 .
[0116] The providing code 850 may be configured to cause at least one processor to provide an identifier via one or more of the techniques described above. The identifier may indicate, for example, that a specified picture is “unused for reference,” “used for short-term reference,” or “used for long-term reference.” For example, the providing code 850 may be configured to cause the at least one processor to provide an identifier of a picture determined to be an SLNR picture (e.g., a marking such as “unused for reference” or “no reference”) based on the picture being determined to be an SLNR picture. Alternatively or additionally, the providing code 850 may be configured to cause the at least one processor to provide an identifier of an entry in a reference picture list corresponding to the picture determined to be an SLNR picture. Alternatively or additionally, the providing code 850 may be configured to cause the at least one processor to provide an identifier of a picture based on determining that a value of a temporal sub-layer of the picture is greater than a predetermined value (e.g., the value of “HighestTid”). Alternatively or additionally, the providing code 850 may be configured to, if the current picture is determined to be an IRAP picture and the flag is determined to indicate that there will be no output of an IRAP picture, cause at least one processor to set an identifier for each reference picture currently stored in the DPB indicating that each currently stored reference picture should be removed from the DPB.
[0117] The formation code 860 may be configured to cause the at least one processor to form one or more reference picture lists according to one or more of the techniques described above. For example, the formation code 860 may be configured to cause the at least one processor to form a reference picture list that includes entries for one or more pictures in the DPB.
[0118] The techniques described above may be implemented using computer-readable instructions and as computer software physically stored on one or more computer-readable media. For example, Figure 8 illustrates a computer system 900 suitable for implementing certain embodiments of the disclosure.
[0119] Computer software can be encoded using any suitable machine code or computer language that may rely on assembly, compiling, linking, or similar mechanisms to create code containing instructions that can be executed by a computer central processing unit (CPU), graphics processing unit (GPU), etc., directly, or through interpretation, microcode execution, etc.
[0120] The instructions may be executed on various types of computers or components thereof, including, for example, personal computers, tablet computers, servers, smartphones, gaming devices, Internet of Things devices, and the like.
[0121] 8 for computer system 900 are examples and are not intended to suggest any limitation on the scope of use or functionality of the computer software implementing embodiments of the present disclosure. Neither the arrangement of components should be interpreted as having any dependency or requirement relating to any one or combination of components illustrated in the non-limiting embodiment of computer system 900.
[0122] The computer system 900 may include certain human interface input devices. Such human interface input devices may respond to input by one or more human users, for example, via tactile input (e.g., keystrokes, swipes, data glove movements), audio input (e.g., voice, clapping), visual input (e.g., gestures), or olfactory input (not shown). The human interface devices may also be used to capture certain media not necessarily directly associated with conscious human input, such as audio (e.g., voice, music, ambient sounds), images (e.g., scanned images, photographic images obtained from a still image camera), and video (e.g., two-dimensional video, three-dimensional video including stereoscopic video).
[0123] The input human interface devices may include one or more of a keyboard 901, a mouse 902, a trackpad 903, a touchscreen 910, a data glove, a joystick 905, a microphone 906, a scanner 907, and a camera 908 (only one of each is shown).
[0124] The computer system 900 may also include certain human interface output devices. Such human interface output devices may stimulate one or more of the human user's senses, for example, through tactile output, sound, light, and smell / taste. Such human interface output devices may include haptic output devices (e.g., haptic feedback via a touchscreen 910, data gloves, or joystick 905, although haptic feedback devices that do not function as input devices may also be present). For example, such devices may include audio output devices (such as speakers 909, headphones (not shown)), visual output devices (such as screens 910, including CRT screens, LCD screens, plasma screens, and OLED screens, each with or without touchscreen input capabilities, each with or without haptic feedback capabilities, some of which may output two-dimensional visual output or output in more than three dimensions through means such as stereographic output, virtual reality glasses (not shown), holographic displays, and smoke tanks (not shown)), and printers (not shown).
[0125] The computer system 900 may include human-accessible storage devices and their associated media, such as optical media including CD / DVD ROM / RW 920 with media such as CD / DVD 921, thumb drives 922, removable hard drives or solid state drives 923, legacy magnetic media such as tape or floppy disks (not shown), specialized ROM / ASIC / PLD-based devices such as security dongles (not shown), etc.
[0126] Those skilled in the art will also understand that the term "computer-readable medium" as used in connection with the presently disclosed subject matter does not include transmission media, carrier waves, or other transitory signals.
[0127] The computer system 900 may also include interfaces to one or more communication networks. The networks may be, for example, wireless, wired, or optical. The networks may further be local, wide area, metropolitan, vehicular and industrial, real-time, delay-tolerant, etc. Examples of networks include local area networks such as Ethernet, cellular networks including WLAN, GSM, 3G, 4G, 5G, LTE, etc., TV wired or wireless wide area digital networks including cable TV, satellite TV, terrestrial broadcast TV, and vehicular and industrial networks including CANBus, etc. Particular networks generally require an external network interface adapter connected to a particular general-purpose data port or peripheral bus 949 (e.g., a USB port on computer system 900), while others are generally integrated into the core of computer system 900 by connection to a system bus as described below (e.g., an Ethernet interface to a PC computer system or a cellular network interface to a smartphone computer system). Using any of these networks, computer system 900 can communicate with other entities. Such communication can be unidirectional, receive only (e.g., broadcast TV), unidirectional transmit only (e.g., a CANbus to a particular CANbus device), or bidirectional, e.g., to other computer systems using local or wide-area digital networks. Such communication can include communication to a cloud computing environment 955. As discussed above, particular protocols and protocol stacks can be used with each of these networks and network interfaces.
[0128] The aforementioned human interface devices, human-accessible storage devices, and network interface 954 may be connected to core 940 of computer system 900 .
[0129] The core 940 may include one or more central processing units (CPUs) 941, graphics processing units (GPUs) 942, specialized programmable processing units in the form of field programmable gate arrays (FPGAs) 943, hardware accelerators 944 for specific tasks, etc. These devices may be connected through a system bus 948, along with read-only memory (ROM) 945, random access memory 946, and internal mass storage 947, such as a non-user-accessible internal hard drive or SSD. In some computer systems, the system bus 948 is accessible in the form of one or more physical plugs, allowing expansion with additional CPUs, GPUs, etc. Peripheral devices may be connected directly to the core's system bus 948 or via a peripheral bus 949. Peripheral bus architectures include PCI, USB, etc. A graphics adapter 950 may also be included in the core 940.
[0130] The CPU 941, GPU 942, FPGA 943, and accelerator 944 may combine to execute specific instructions that may constitute the aforementioned computer code, which may be stored in ROM 945 or RAM 946. Transient data may also be stored in RAM 946, while permanent data may be stored in, for example, internal mass storage 947. Fast storage and retrieval to any memory device may be enabled through the use of cache memory, which may be closely associated with one or more of the CPU 941, GPU 942, mass storage 947, ROM 945, RAM 946, etc.
[0131] The computer-readable medium may have computer code thereon for performing various computer-implemented operations. The medium and computer code may be those specially designed and constructed for the purposes of the present disclosure, or they may be of the kind well known and available to those skilled in the computer software arts.
[0132] By way of example, and not limitation, a computer system having architecture 900, and particularly core 940, can provide functionality as a result of a processor (including a CPU, GPU, FPGA, accelerator, etc.) executing software embodied in one or more tangible computer-readable media. Such computer-readable media may be media associated with user-accessible mass storage, as introduced above, as well as specific storage of the core 940 that is non-transitory in nature, such as core internal mass storage 947 or ROM 945. Software implementing various embodiments of the present disclosure may be stored on such devices and executed by the core 940. The computer-readable media may include one or more memory devices or chips, depending on particular needs. The software may cause the core 940, and particularly the processor therein (including a CPU, GPU, FPGA, etc.), to perform particular processes or particular portions of particular processes described herein, including the definition of data structures stored in RAM 946 and the modification of such data structures by software-defined processes. Additionally, or alternatively, a computer system may provide functionality as a result of logic hardwired or otherwise embodied in circuitry (e.g., accelerator 944) that can operate in place of or together with software to perform particular processes or portions of particular processes described herein. References to software can include logic, and vice versa, as appropriate. References to computer-readable media can encompass circuitry (such as an integrated circuit (IC)) that stores software for execution, circuitry that embodies logic for execution, or both, as appropriate. The present disclosure encompasses any suitable combination of hardware and software.
[0133] While this disclosure describes several non-limiting embodiments, there are modifications, permutations, and various substitute equivalents that are within the scope of this disclosure. It will thus be appreciated that those skilled in the art will be able to devise numerous systems and methods that, although not explicitly shown or described herein, embody the principles of the present disclosure and are therefore within its spirit and scope. [Explanation of symbols]
[0134] 100 Communication Systems 110 Terminal 120 terminals 130 terminals 140 terminals 150 Network 200 Streaming System 201 Video Sources 202 uncompressed video sample streams 203 Encoder 204 encoded video bitstream 205 Streaming Server 206 Streaming Client 209 Encoded Video Bitstream 210 Video Decoder 211 video sample streams 212 Display 213 Capture Subsystem 310 Receiver 312 channels 315 Buffer Memory 320 Parser 321 Symbol 351 Scaler / Inverse Conversion Unit 352 Intra Prediction Units 353 Motion Compensation Prediction Unit 355 Aggregator 356 Loop Filter Unit 357 Reference Picture Memory 358 Current Screen 430 Source Coder 432 encoding engine 433 decoder 434 Reference Picture Memory 435 Predictor 440 Transmitter 443 coded video sequence 445 Entropy Coder 450 Controller 460 channels 500 processes 600 processes 800 devices 810 Memory Code 820 Decision Code 830 Removal Code 840 Decryption Code 850 Offer Code 860 Formation Code 900 Computer Systems 901 Keyboard 902 Mouse 903 Trackpad 905 Joystick 906 Mike 907 Scanner 908 Camera 909 Speaker 910 Touchscreen 920 CD / DVD ROM / RW 921 CD / DVD and other media 922 thumb drive 923 Removable Hard Drive or Solid State Drive 940 cores 941 Central Processing Unit (CPU) 942 Graphics Processing Unit (GPU) 943 Field Programmable Gate Area (FPGA) 944 Accelerator 945 Read-Only Memory (ROM) 946 Random Access Memory (RAM) 947 Mass Storage 948 System Bus 949 Peripheral Bus 950 graphics adapter 954 network interface 955 Cloud Computing Environment
Claims
1. 1. A method for decoding a video stream, comprising: storing previously decoded pictures of the video stream, including a first plurality of pictures of a same temporal sub-layer, in a decoded picture buffer, the first plurality of pictures including at least one sub-layer reference picture for predicting a current picture of the video stream, and the previously decoded pictures stored in the picture buffer including a second picture that is a reference picture; identifying a Network Abstraction Layer (NAL) unit type of a picture of the first plurality of pictures, The NAL unit type of the picture Non-Stepwise Temporal Sub-Layer Access (STSA) coded tile groups of subsequent pictures, coded tile groups of STSA pictures, A coded tile group of a random-access skip-reading (RASL) picture, or Coded Tile Groups for Random Access Decodable Reading (RADL) Pictures identifying the object as including when it is determined that the predetermined value is not equal to the value corresponding to the highest temporal sub-layer identification number, determining whether the value of the temporal sub-layer of the second picture is greater than the predetermined value; comparing the predetermined value with a value corresponding to the highest temporal sub-layer identification number; removing the identified picture from the decoded picture buffer based on the NAL unit type of the identified picture indicating a non-reference picture; removing the second picture from the decoded picture buffer based on determining that the value of the temporal sub-layer of the second picture is greater than the predetermined value; decoding the current picture using the decoded picture buffer, predicting the current picture using one or more of the at least one sub-layer reference pictures stored in the decoded picture buffer after removing the picture from the decoded picture buffer; A method comprising:
2. 2. The method of claim 1, wherein identifying the NAL unit type comprises identifying the NAL unit type of the picture as containing the coded tile group of the non-STSA subsequent picture.
3. The method of claim 1 , wherein identifying the NAL unit type comprises identifying the NAL unit type of the picture as containing the coded tile group of the STSA picture.
4. The method of claim 1 , wherein identifying the NAL unit type comprises identifying the NAL unit type of the picture as containing the coded tile group of the RASL picture.
5. The method of claim 1 , wherein identifying the NAL unit type comprises identifying the NAL unit type of the picture as containing the coded tile group of the RADL picture.
6. providing an identifier for the picture based on the step of identifying the NAL unit type of the picture. further comprising The method of claim 1 , wherein the removing step comprises removing the picture from the decoded picture buffer based on the identifier.
7. forming a reference picture list including an entry for each of the first plurality of pictures; further comprising The method of claim 6 , wherein the step of providing the identifier comprises providing the identifier to the entry in the reference picture list that corresponds to the picture.
8. providing an identifier for the second picture based on determining that the value of the temporal sub-layer of the second picture is greater than the predetermined value. further comprising The method of claim 1 , wherein removing the second picture comprises removing the second picture from the decoded picture buffer based on the identifier.
9. 1. A decoder for decoding a video stream, comprising: a memory configured to store computer program code; at least one processor configured to access said computer program code and to operate as instructed by said computer program code; the computer program code comprising: storage code configured to cause the at least one processor to store previously decoded pictures of the video stream, including a first plurality of pictures of a same temporal sub-layer, in a decoded picture buffer, the first plurality of pictures including at least one sub-layer reference picture for predicting a current picture of the video stream, and the previously decoded pictures stored in the picture buffer including a second picture that is a reference picture; and and determining whether the temporal sub-layer value of the second picture is greater than a predetermined value and comparing the predetermined value with the value corresponding to the highest temporal sub-layer identification number when it is determined that the predetermined value is not equal to a value corresponding to a highest temporal sub-layer identification number. removal code configured to cause the at least one processor to remove the identified picture from the decoded picture buffer based on the NAL unit type of the picture indicating a non-reference picture, and to remove the second picture from the decoded picture buffer based on determining that the temporal sub-layer value of the second picture is greater than the predetermined value; decoding code configured to cause the at least one processor to decode the current picture using the decoded picture buffer, the decoding code including prediction code configured to cause the at least one processor to predict the current picture using one or more of the at least one sub-layer reference pictures stored in the decoded picture buffer after removing the picture from the decoded picture buffer; and Including, a decoder.
10. 10. The decoder of claim 9, wherein the decision code is configured to cause the at least one processor to identify the NAL unit type of the picture as containing the coded tile group of the non-STSA subsequent picture.
11. 10. The decoder of claim 9, wherein the decision code is configured to cause the at least one processor to identify the NAL unit type of the picture as containing the coded tile group of the STSA picture.
12. 10. The decoder of claim 9, wherein the decision code is configured to cause the at least one processor to identify the NAL unit type of the picture as containing the coded tile group of the RASL picture.
13. 10. The decoder of claim 9, wherein the decision code is configured to cause the at least one processor to identify the NAL unit type of the picture as containing the coded tile group of the RADL picture.
14. the computer program code further includes providing code configured to cause the at least one processor to provide an identifier of the picture based on identifying the NAL unit type of the picture; 10. The decoder of claim 9, wherein the removal code is configured to cause the at least one processor to remove the picture from the decoded picture buffer based on the identifier.
15. the computer program code further includes forming code configured to cause the at least one processor to form a reference picture list including an entry for each of the first plurality of pictures; The decoder of claim 14 , wherein the providing code is configured to cause the at least one processor to provide the identifier to the entry in the reference picture list that corresponds to the picture.
16. When executed by at least one processor, the method causes the at least one processor to: storing previously decoded pictures of a video stream, including a first plurality of pictures of a same temporal sub-layer, in a decoded picture buffer, the first plurality of pictures including at least one sub-layer reference picture for predicting a current picture of the video stream, the previously decoded pictures stored in the picture buffer including a second picture that is a reference picture; identifying a network abstraction layer (NAL) unit type of a picture of the first plurality of pictures, including identifying the NAL unit type of the picture as comprising a coded tile group of a non-stepwise temporal sub-layer access (STSA) subsequent picture, a coded tile group of an STSA picture, a coded tile group of a random access skip reading (RASL) picture, or a coded tile group of a random access decodable reading (RADL) picture; determining whether the value of the temporal sub-layer of the second picture is greater than the predetermined value when it is determined that the predetermined value is not equal to the value corresponding to the highest temporal sub-layer identification number; comparing the predetermined value with a value corresponding to the highest temporal sub-layer identification number; removing the picture from the decoded picture buffer based on the NAL unit type of the identified picture indicating a non-reference picture; removing the second picture from the decoded picture buffer based on determining that the value of the temporal sub-layer of the second picture is greater than the predetermined value; using the decoded picture buffer to decode the current picture by predicting the current picture using one or more of the at least one sub-layer reference pictures stored in the decoded picture buffer after removing the picture from the decoded picture buffer. A non-transitory computer-readable medium that stores computer instructions.
Citation Information
Patent Citations
Image decoding device and image encoding device
JP2015035642A
Signalization change in output layer set
JP2016518763A
Providing a common set of parameters for sub-layers of coded video
WO2014059051A1
Image decoding device and image decoding method
WO2015137432A1