Video encoding using encoded picture buffer
By using interpolation technology in video codecs to manage encoded picture buffer (CPB) parameters, the problem of difficulty in managing CPB size and bit rate in the prior art is solved, and the effect of safe and correct CPB operation and reducing the overhead of HRD parameter transmission is achieved.
Patent Information
- Application Number
- CN202380072527.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2022-10-13
- Filing Date
- 2023-10-12
- Publication Date
- 2025-05-30
AI Technical Summary
The prior art is difficult to effectively manage the encoded picture buffer (CPB) size and bit rate in video encoding, resulting in the possibility of buffer overflow or underflow, especially when multiple bit rate scenarios are required, the overhead of transmitting a large number of HRD parameters is high.
The CPB parameters are determined by using interpolation techniques in video codecs, especially in different bit rate scenarios, and the selected bit rate is used to interpolate between the time offset and time removal delay of different CPB parameters to determine the time offset and time removal delay after interpolation.
It realizes that the buffer underflow and overflow is avoided while ensuring safe and correct CPB operations, and reduces the overhead of HRD parameter transmission, providing a better trade-off effect.
Smart Images

Figure CN120077657A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to video coding and the use of coded picture buffers in video coding. Background Art
[0002] The Hypothetical Reference Decoder and its use for checking bitstream and decoder consistency are fundamental components of every video coding standard (e.g., VVC).
[0003] For performing such consistency checks, an HRD buffer model is defined, which includes a Hypothetical Stream Scheduler (HSS), a Coded Picture Buffer (CPB), a decoding process (which is considered instantaneous), a Decoded Picture Buffer (DBP), and an output clipping process, as Figure 17 shown.
[0004] This model defines the timing and bitrate at which the bitstream is fed into the coded picture buffer, the time at which its decoding units (access units or VCL NAL units in low-latency operation mode) are removed from the CPB and immediately decoded, and the output time at which pictures are output from the DPB.
[0005] Only by doing so is it possible to define the CPB size required by the decoder to avoid buffer overflows (more data is sent to the decoder than can be retained in the CPB) or underflows (less data is sent to the decoder at a lower bitrate than required), and the necessary data from the AU does not appear at the decoder at the correct decoding time.
[0006] The latest video coding standards define different parameters to describe the bitstream and HRD requirements, as well as the buffer model.
[0007] For example, in HEVC, hrd_parameters are defined for each sublayer and describe one or more tuples of Bitrate(i) and CPBsize(i), which indicate that no overflow or underflow will occur if the HSS feeds a CPB of size CPBsize(i) at a bitrate of Bitrate(i). In other words, when these bitrate and CPB size tuples are adhered to, continuous decoding can be guaranteed.
[0008] In combination with the hrd_parameter syntax element, additional timing information exists in the bitstream, which specifies the time at which each picture is removed from the CPB, i.e., this information indicates the time at which the VCL NAL units belonging to each picture are sent for decoding.
[0009] The relevant information is present in the cached picture SEI messages with the syntax elements or variables InitialCPBRemovalDelay(i), InitialCPBRemovalDelayOffset(i), and AuCPBRemovalDelay, and in the picture timing SEI messages with AuCPBRemovalDelay.
[0010] However, depending on the application and the transmission channel, information about HRD parameters for many bitrates will be needed in order to be able to fine-tune according to the bitrate. However, this will require the transmission of a large number of bits consuming HRD parameters for the intensive selection of bitrate(i). For a large number of bitrates with reasonable overhead for sending HRD information, it would be advantageous to have a concept that allows correct HRD parameterization, i.e., a concept that does not result in CPB underflow or overflow. Summary of the Invention
[0011] Accordingly, an object of the present invention is to provide a video codec that uses coded picture buffering operations that result in a better trade-off between the bit consumption for HRD signaling on the one hand and an efficient way of determining HRD parameters for many bitrate scenarios on the other hand.
[0012] An embodiment may have an apparatus for video decoding, the apparatus having a coded picture buffer (CPB) and a decoded picture buffer (DPB), configured to receive a data stream having pictures of a video encoded therein in coded order as a sequence of access units (AUs), sequentially feed the sequence of access units into the CPB at a selected bitrate, wherein feeding of access units not yet having reached a virtual availability time for raster removal according to a time frame is paused until the virtual availability time is reached, wherein the time frame for raster removal advances a selected time removal delay for a first access unit in the coded order, and for subsequent access units in the coded order advances a sum of the selected time removal delay and a selected time offset; remove AUs from the CPB on a per-AU basis using a time raster [RemovalTime], extract a first CPB parameter related to a first operating point and a second CPB parameter related to a second operating point from the data stream, each of the first CPB parameter and the second CPB parameter indicating a CPB size, a predetermined time offset, a predetermined time removal delay, and a predetermined bitrate, wherein the first CPB parameter is different from the second CPB parameter at least in terms of the predetermined bitrate, determine the selected time offset by interpolating at the selected bitrate between the predetermined time offset indicated by the first CPB parameter and the predetermined time offset indicated by the second CPB parameter, and determine the selected time removal delay by interpolating at the selected bitrate between the predetermined time removal delay indicated by the first CPB parameter and the predetermined time removal delay indicated by the second CPB parameter, decode a current AU removed from the CPB using inter-picture prediction with reference pictures stored in the DPB to obtain a decoded picture, and insert the decoded picture into the DPB, assign to each reference picture stored in the DPB a classification as one of a short-term reference picture, a long-term reference picture, and a picture not for reference, read DPB mode information from the current AU, if the DPB mode information indicates a first mode, remove one or more reference pictures classified as short-term pictures from the DPB according to a first-in first-out (FIFO) policy, if the DPB mode information indicates a second mode, read memory management control information having at least one command from the current AU, and execute the at least one command to change the classification assigned to at least one of the reference pictures stored in the DPB, and use the classification of the reference pictures in the DPB to manage removal of the reference pictures from the DPB.
[0013] Another embodiment may have an apparatus for encoding video into a data stream, wherein the data stream should be decoded by feeding the data stream into a decoder including a coded picture buffer (CPB). The apparatus is configured to: encode pictures of the video encoded in coding order into the data stream as an access unit (AU) sequence, determine a first CPB parameter associated with a first operating point and a second CPB parameter associated with a second operating point, each of the first CPB parameter and the second CPB parameter indicating a CPB size, a predetermined time offset, a predetermined time removal delay, and a predetermined bit rate, wherein the first CPB parameter is different from the second CPB parameter at least in terms of the predetermined bit rate, and perform the determination such that interpolation between the predetermined time offset of the first CPB parameter and the predetermined time offset of the second CPB parameter at each of a plurality of selected bit rates results in an interpolated time offset and an interpolated time removal delay. Thus, the data stream is fed into the decoder via the CPB by: sequentially feeding the AU sequence into the CPB using the corresponding selected bit rate, wherein feeding of access units that have not reached the virtual available time for raster removal according to the time frame is paused until the virtual available time is reached, wherein the time frame raster removal advances the interpolated time removal delay for the first access unit in the coding order, and advances the sum of the interpolated time removal delay and the interpolated time offset for subsequent access units in the coding order; and removing AUs from the CPB on a per-AU basis using a time raster without causing any underflow and any overflow, and encoding the CPB parameters into the data stream. The apparatus is configured to: when encoding the AU, use inter-picture prediction to encode the current picture into the current AU based on the referenced reference pictures stored in the decoded picture buffer (DPB), and insert the decoded version of the current picture in the DPB into the DPB, assign a classification of one of a short-term reference picture, a long-term reference picture, and a picture not for reference to each reference picture stored in the DPB, write DPB mode information into the current AU, if the DPB mode information indicates a first mode, remove one or more reference pictures classified as short-term pictures from the DPB according to a first-in first-out (FIFO) policy, if the DPB mode information indicates a second mode, write memory management control information with at least one command into the current AU, the command indicating a change in the classification assigned to at least one of the reference pictures stored in the DPB, wherein the classification of the reference pictures in the DPB is used to manage the removal of the reference pictures from the DPB.
[0014] According to another embodiment, a method for video decoding by using a coded picture buffer (CPB) and a decoded picture buffer (DPB) may have the following steps: receiving a data stream that has pictures of a video encoded therein in coded order as an access unit (AU) sequence, sequentially feeding the access unit sequence into the CPB at a selected bitrate, wherein feeding of access units that have not reached the virtual available time for removing the raster according to a time frame is paused until the virtual available time is reached, wherein the time frame removes the delay by a selected time in advance for a first access unit in the coded order, and for subsequent access units in the coded order, the time for removing the delay in advance is the sum of the selected time and a selected time offset; removing AUs from the CPB on a per-AU basis using a time raster [RemovalTime], extracting a first CPB parameter related to a first operating point and a second CPB parameter related to a second operating point from the data stream, each of the first CPB parameter and the second CPB parameter indicating a CPB size, a predetermined time offset, a predetermined time removal delay, and a predetermined bitrate, wherein the first CPB parameter is different from the second CPB parameter at least in terms of the predetermined bitrate, determining the selected time offset by interpolating at the selected bitrate between the predetermined time offset indicated by the first CPB parameter and the predetermined time offset indicated by the second CPB parameter, and determining the selected time removal delay by interpolating at the selected bitrate between the predetermined time removal delay indicated by the first CPB parameter and the predetermined time removal delay indicated by the second CPB parameter, decoding a current AU removed from the CPB using inter-picture prediction based on reference pictures stored in the DPB to obtain a decoded picture, and inserting the decoded picture into the DPB, assigning a classification of one of a short-term reference picture, a long-term reference picture, and a picture not for reference to each reference picture stored in the DPB, reading DPB mode information from the current AU, if the DPB mode information indicates a first mode, removing one or more reference pictures classified as short-term pictures from the DPB according to a first-in first-out (FIFO) policy, if the DPB mode information indicates a second mode, reading memory management control information having at least one command from the current AU, and executing the at least one command to change the classification assigned to at least one of the reference pictures stored in the DPB, and using the classification of the reference pictures in the DPB to manage removal of the reference pictures from the DPB.
[0015] Another embodiment may have a data stream into which a video is encoded and having a first CPB parameter and a second CPB parameter such that the above inventive method does not cause CPB overflow and underflow.
[0016] The basic idea of the present invention is that a good compromise between the CPB parameter transmission capacity and the CPB parameterization effectiveness can be achieved by using interpolation between CPB (or HRD) parameters signaled explicitly at a selected bitrate, and can be specifically achieved in an efficient manner, i.e., in a manner that results in safe and correct CPB operation without underflow and overflow, and in the following manner: according to this manner, for example, the CPB size indicated by the explicitly signaled CPB parameters does not have to be provided with a safety offset to account for interpolation-related contingencies, even if the explicitly signaled CPB parameters indicate, in addition to the CPB size and bitrate for the explicitly signaled operating points, a predetermined time offset and a predetermined time removal delay for these operating points. Specifically, according to this idea, on the decoding side, both the time offset and the time removal delay for the selected bitrate can be determined by interpolating between the corresponding values of the offset and the delay at the selected bitrate based on the signaled CPB parameters. Then, this interpolated / selected time offset can be used to sequentially feed the sequence of access units of the video data stream into the coded picture buffer using the selected bitrate (i.e., by pausing the feeding of access units that have not yet reached the virtual available time for removal from the time frame until the said virtual available time is reached, where the time frame removal raster is advanced by the selected / interpolated time removal delay for the first access unit in coding order, and by the sum of the selected / interpolated time removal delay and the selected time offset for subsequent access units in coding order). Using the time raster, access units can then be removed from the coded picture buffer. Although only interpolation has to be performed on the decoding side to determine the selected time offset and the selected time removal delay, the encoder sets the explicitly signaled CPB parameters related to the operating points of the explicitly prepared video data stream in a manner that takes interpolation into account (i.e., in a manner such that the corresponding selected / interpolated values for the time offset and the time removal delay do not result in underflow or overflow). Description of the Drawings
[0017] Embodiments of the present application will be described below with reference to the drawings, in which:
[0018] Figure 1 A block diagram showing a possible implementation of an encoder according to which embodiments of the present application can be implemented is shown;
[0019] Figure 2 A block diagram showing a possible implementation of a decoder according to which embodiments of the present application can be implemented is shown, and the decoder is suitable for Figure 1 the encoder;
[0020] Figure 3 A block diagram showing a device for decoding according to an embodiment of the present application is shown;
[0021] Figure 4 A schematic diagram showing the CPB parameter encoding and interpolation therebetween;
[0022] Figure 5 On the left, a filling state diagram of the CPB feeder is shown, and on the right, a CPB filling state diagram is shown, where the upper half shows an example where subsequent non-first access units do not use time offset, and the lower half shows an example where subsequent non-first access units use time offset;
[0023] Figures 6 to 14 Shows the CPB buffer filling state for different examples of CPB parameter bit rate, the selected bit rate therebetween, and time offset settings;
[0024] Figure 15 Shows Figure 3 A schematic diagram of an example of the operating mode of a device for managing the DPB;
[0025] Figure 16 Shows a block diagram of a device for encoding according to an embodiment of the present application; and
[0026] Figure 17 Shows a block diagram of a known HRD buffer model. Detailed Description
[0027] Before continuing with the introductory part of this specification and explaining the problems related to the desire to provide a high degree of flexibility in terms of operating points for HRD operations, a preliminary example of a video codec is provided, in which the subsequent described embodiments can be established. However, it should be noted that these examples of the video codec should not be considered as limiting the embodiments described subsequently in this application.
[0028] Figure 1 Shows an encoder 10 configured to encode video 12 into a bitstream 14. The encoder 10 encodes pictures 16 of the video 12 into the bitstream 14 using a picture coding order, which may be different from the presentation time order 18 in which the pictures 16 are presented or output sequentially when presenting the video 12. Figure 1 Also shows a possible implementation of the encoder 10, but again note that the details Figure 1 Elaborated do not limit the embodiments of the present application described in more detail below. Although applied subsequently, the encoding of the encoder 10 may not include intra prediction, may not include inter prediction, may not operate in blocks, may not use transform residual coding, and may operate lossy or lossless or in a combination thereof.
[0029] Figure 1The encoder 10 performs encoding by using prediction. In a block-by-block manner, the encoder 10 predicts the current picture, or more precisely, the current encoded part of the picture, and forms a prediction residual 20 by subtracting the prediction signal 24 from the original version of the current picture 26 at the subtractor 22. Then, the residual encoder 28 encodes the prediction residual 20 into the bitstream 14, where the residual encoding can be lossy and can include, for example, transforming the residual signal 20 into a transform domain and performing entropy encoding on the transform coefficients obtained from the transform. To obtain the prediction signal 24 based on a reconstructable version of the encoded part of the video 12, the residual decoder 30 reverses the residual encoding and generates a residual signal 32 from the transform coefficients by inverse transform, and the residual signal 32 differs from the residual signal 20 by the loss introduced by the residual encoder 28. To reconstruct the current picture, or more precisely, the current encoded block of the current picture, the residual signal 32 is added to the prediction signal 24 by the adder 34 to generate a reconstruction signal 36. Optionally, the loop filter 38 subjects the reconstruction signal 38 to some loop filtering, and the filtered signal 40 is input into the loop buffer 42. Thus, the loop buffer 42 caches the reconstructable versions of the encoded pictures and the reconstructable parts of the current picture respectively. Based on these reconstructable versions 44, and optionally, based on the unfiltered reconstructable version 36 of the encoded part of the current picture, the prediction stage 46 determines the prediction signal 24.
[0030] The encoder 10 uses rate-distortion optimization to make many encoding decisions. For example, the predictor 46 selects one of several encoding modes, which include, for example, one or more inter-prediction modes and one or more intra-prediction modes, and optionally, combinations thereof at the granularity of the encoded blocks. At the granularity of these encoded blocks, or alternatively, at the granularity of the prediction blocks into which these encoded blocks are further subdivided, the predictor 46 determines prediction parameters suitable for the selected prediction mode, such as one or more motion vectors for an inter-prediction block or an intra-prediction mode for an intra-prediction block. The residual encoder 28 performs residual encoding at the granularity of residual blocks, which may optionally coincide with any one of the encoded blocks or prediction blocks, or may be a further subdivision of any one of these blocks or may result from another independent subdivision of the current picture into residual blocks. Even the aforementioned subdivisions are determined by the encoder 10. These encoding decisions (i.e., subdivision information, prediction mode, prediction parameters, and residual data) are encoded by the encoder 10 into the bitstream 14 using, for example, entropy encoding.
[0031] Each picture 16 is encoded by the encoder 10 into a consecutive part 48 (referred to as an access unit) of the bitstream 14. Thus, the sequence of access units 48 in the bitstream 14 has pictures 16 encoded into it sequentially (i.e., in the aforementioned picture encoding order).
[0032] Figure 2Shown suitable for Figure 1 The decoder 100 of the encoder 10 is configured to decode the video 12' from the bitstream 14 by decoding the corresponding picture 16' of the video 12' from each access unit 48. Figure 1 To this end, the decoder 100 is internally interpreted as follows Figure 1 The decoder 100 is a reconstruction part of the prediction loop of the encoder 10. That is, the decoder 100 includes a residual decoder 130 that reconstructs the residual signal 32 from the bitstream 14. The prediction signal 124 is added to the residual signal 132 at the adder 134 to produce a reconstructed signal 136. The optional loop filter 138 filters the reconstructed signal 136 to produce a filtered reconstructed signal 140, which is then buffered in a loop buffer 142. The buffered and filtered reconstructed signal 144 is output from the buffer by the decoder 100, that is, the buffered and reconstructed signal contains the reconstructed pictures 16', and these pictures 16' are output from the buffer 142 in the presentation time order. In addition, a predictor or prediction unit 146 performs prediction based on the signal 144 and, optionally, the reconstructed signal 136 to produce the prediction signal 124. The decoder obtains all necessary information for decoding, such as subdivision information, prediction mode decisions, prediction parameters and residual data from the bitstream 14, for example using entropy decoding and determined by the encoder 10 using rate / distortion optimization. As mentioned above, the residual data may include transform coefficients.
[0033] The encoder 10 can perform its encoding task in such a way that, on average, the video 12 is encoded in the bitstream 14 at a specific bit rate (i.e., so that, on average, the pictures 16 are encoded into the bitstream 14 using a specific number of bits). However, due to different picture content complexities, changing scene content, and differently encoded pictures (e.g., I-frames, P-frames, and B-frames), the number of bits spent in the bitstream 14 for each picture 16 may vary. That is, the size or number of bits of each access unit 48 may vary. In order to ensure uninterrupted playback of the video 12' at the decoder 100, the encoder 10 provides CPB parameters for the bitstream 14. If fed to the decoder 100 via the decoded picture buffer 200 in some predefined manner, these CPB parameters ensure that the decoder 100 performs such uninterrupted or problem-free decoding. In other words, the CPB parameters refer to Figure 3 The device shown, wherein the feeder 202 feeds the decoder 100 via the encoded picture buffer 200, the feeder 202 receives the bit stream 14 and feeds the bit stream 14 to the decoder 100 via the encoded picture buffer 200, so that the decoder 100 then timely accesses the access unit 48 of the bit stream 14, so that the picture 16' of the video 12' can be output without interruption in the presentation time order.
[0034] For a number of so-called operating points OP i , the CPB parameters are written into the bitstream 14 by the encoder 10. Each operating point OP i refers to a different bitrate(i) at which the feeder 202 feeds the bitstream 14 (i.e., the sequence of access units 48) into the coded picture buffer 200. That is, for each operating point OP i , the CPB parameter 300 indicates the bitrate at which they apply. In addition, the CPB parameter 300 indicates the coded picture buffer size of the coded picture buffer 200, which is sufficient to contain the most complete state when feeding to the decoder 100 at the corresponding bitrate. In addition, the information indicated by the CPB parameter 300 i indicates the time delay at which, relative to the time point at which the first bit of the bitstream 14 is input into the coded picture buffer 200, respectively, the first access unit is removed from the coded picture buffer 200 and passed to the decoder 100. The term "first" may refer to the picture coding order and a certain buffer cycle, i.e., a subsequence of pictures. In addition, the CPB parameter 300 i indicates the time offset at which the feeding of subsequent access units following the aforementioned first access unit is allowed to be fed into the decoded picture buffer 200 before their regular feeding, which is delayed by the aforementioned time delay and determined by the regular time raster. Figure 4 is not shown, but optionally, the CPB parameter 300 i indicates another piece of information, such as information that reveals or indicates or allows the derivation of the following: the time raster at which the just-mentioned bitstream 14 is removed from the coded picture buffer 200 access unit by access unit for decoding by the decoder 100, as just mentioned, the time raster being delayed by the time delay. Thus, the time raster is related to the frame rate of the video 12 in order to allow the decoder 100 to recover the pictures 16' at a rate sufficient to output these pictures 16' at that frame rate. The optional indication of the time raster may be common to all operating points and indicated commonly in the bitstream for all operating points. In addition, instead of signaling any information about the time raster, the information may be fixed and pre-known to the encoder and decoder.
[0035] Figure 4 shows the CPB parameter 300 i for two operating points OP i , i.e., two operating points OP i-1 and OP iRefers to two different bitrates and CPB sizes. For the bitrates between bitrate(i - 1) and bitrate(i), there are no other instances of such CPB parameters in bitstream 14. As already indicated above, embodiments of the present application will fill this gap by the possibility of deriving such missing instances of CPB parameters for a selected bitrate between bitrate(i - 1) and bitrate(i) through interpolation.
[0036] Note that due to the fact that the aforementioned time raster is related to the frame rate, encoder 10 can indicate this time raster or information about it only once, commonly for all CPB parameters or all instances of CPB parameters (or even in other words, commonly for all operating points). Additionally, there may even be no information about the time raster transmitted in the data stream, in which case the time raster is pre-known between the encoder and the decoder. For example, due to the pre-known predetermined frame rates for videos 12 and 12' respectively between the encoder and the decoder, and the relationship between a certain group of pictures (GOP) structure and picture coding order (on one hand) and the presentation time order 18 (on the other hand).
[0037] Now continue with the description of the introduction part of this specification. As described above, CPB parameters can be transmitted through SEI messages. InitialCPBRemovalDelay corresponds to Figure 4 the time delay, and InitialCPBRemovalDelayOffset corresponds to Figure 4 the time offset. AuCPBRemovalDelay indicates the time raster, that is, the time distance between consecutive AUs removed from the DPB. As illustrated, Figure 4 the CPB parameters of can be transmitted in the cache cycle SEI message, which indicates the correct scheduling for feeding to decoder 100 via encoded picture buffer 200 for a so-called cache cycle (i.e., a picture sequence of the video corresponding to a certain sequence of access units including the first access unit and subsequent access units of this cache cycle).
[0038] As described in the introduction part of this specification, it is known that CPB parameters need to be transmitted in the bitstream, but these CPB parameters refer to certain specific bitrates.
[0039] For the most basic operations, only InitialCPBRemovalDelay(i) and AuCPBRemovalDelay are used.
[0040] In that case, the first decoded access unit is a random access point with its corresponding cached period SEI message, and time 0 is defined as the time when the first bit of the random access point enters the CPB. Then, at time InitialCPBRemovalDelay(i), the picture corresponding to this random access point is removed from the CPB. For another non-RAP picture, the removal of the CPB occurs at InitialCPBRemovalDelay(i) + AuCPBRemovalDelay (legacy codecs can define some additional parameters to convert the indicated delay into a time increment, i.e., ClockTick, but this is ignored here for simplicity).
[0041] When the next RAP arrives, the removal time is calculated for non-RAP pictures as before, i.e., InitialCPBRemovalDelay(i) + AuCPBRemovalDelay, and this new value is used as an anchor for other increments to another RAP, i.e., anchorTime = InitialCPBRemovalDelay(i) + AuCPBRemovalDelay. Then, the removal of the picture becomes anchorTime + AuCPBRemovalDelay, and anchorTime is updated at the next RAP using the cached SEI message, anchorTime = anchorTime + AuCPBRemovalDelay, etc.
[0042] In other words, the calculation of RemovalTime for the first access unit (AU with cached period SEI) that initializes the decoder is as follows:
[0043] RemovalTime[0] = InitialCPBRemovalDelay(i)
[0044] Note that InitialCPBRemovalDelay can be derived from the bitstream as initial_cpb_removal_delay[i] ÷ 90000.
[0045] The RemovalTime of an AU (not the first access unit that initializes the decoder, but the first AU of another cached period, i.e., the AU with a cached period SEI message that is not the first AU of the initializing decoder) is calculated as:
[0046] RemovalTime[n] = RemovalTime[n b + AuCPBRemovalDelay
[0047] where nb refers to the index of the first AU in the previous buffer cycle (the AU before the current AU that also has a buffer cycle SEI message), and AuCPBRemovalDelay can be derived from the bitstream as t c ×cpb_removal_delay(n), and t c is clockTicks (the unit that gives the cpb_removal_delay syntax to convert a given value to time).
[0048] The RemovalTime of an AU (which is neither the first access unit that initializes the decoder nor the first AU of another buffer cycle (i.e., the AU that has a buffer cycle SEI message and is not the first AU that initializes the decoder)) is calculated as:
[0049] RemovalTime[n]=RemovalTime[n b +AuCPBRemovalDelay
[0050] where n b refers to the index of the first AU in the current buffer cycle (the AU before the current AU that has a buffer cycle SEI message), and AuCPBRemovalDelay can be derived from the bitstream as t c ×cpb_removal_delay(n), and t c is clockTicks (the unit that gives the cpb_removal_delay syntax to convert a given value to time).
[0051] The drawback of the described model is that the defined InitialCPBRemovalDelay implicitly sets a limit on the available / usable CPB size. Therefore, in order to use the CPB buffer, a large time delay (InitialCPBRemovalDelay) will be required to remove the first access unit. In fact, assuming that the encoded pictures at the decoder are sent immediately after they are encoded, the time when each picture arrives at the decoder will not be earlier than the following time:
[0052] initArrivalEarliestTime[n]=RemovalTime[n]-InitCpbRemovalDelay(i)
[0053] That is, its removal time minus the InitialCPBRemovalDelay, which is the time the decoder waits to remove the first AU since it received the corresponding first bit of that AU in the CPB.
[0054] Alternatively, in the case where the picture before the current picture is so large that its last bit arrives (AuFinalArrivalTime[n - 1]) later than RemovalTime[n] - InitCpbRemovalDelay(i), the initial arrival time (the time when the first bit of the current picture is fed into the CPB) is equal to:
[0055] initArrivalTime[n]=Max(AuFinalArrivalTime[n - 1],initArrivalEarliestTime[n])
[0056] This means that, for example, if an AU with a new cache cycle SEI message cannot enter the CPB earlier than InitialCPBRemovalDelay(i) of its removal time, it is impossible to achieve a CPB A greater than the CPB B , because feeding the CPB at Bitrate(i) during InitialCPBRemovalDelay(i) only achieves a CPB A fill level of the CPB.
[0057] To solve this problem, the idea is to assume that the transmitter (or Figure 17 the HSS in Figure 4 or the feeder 202 in Figure 5 ) schedules the first RAP with a cached SEI message with a given time offset InitialCPBRemovalDelayOffset(i) as shown in
[0058] That is, Figure 5 the upper part of Figure 17 shows the feeding of the virtual buffer fill level of the HSS in Figure 3 or the feeder 202 in ai1 to the coded picture buffer 200, assuming that the "transmitter" instantaneously obtains the access unit at the aforementioned removal raster, where the transmitter sequentially sends the access unit sequence to the coded picture buffer at a specific bitrate derivable from the slope in the figure. The right side shows the receiver side, and more precisely, shows the CPB fill level, which reveals the feeding of the AU into the CPB and the removal of the access unit from the coded picture buffer. The removal occurs instantaneously and is fed at a specific bitrate. Again, the bitrate is derivable from the slope in the right side figure. The trailing edges in this figure indicate the instances where the access unit sequence is removed from the CPB access unit by access unit and fed into the decoder. However, they appear at a time raster that is delayed by the time delay at the start between the arrival of the first bit of the first access unit T
[0059] Figure 5 The lower part of [figure] shows the effect of having a time offset: the available or usable CPB size is not limited by the amount determined by the time delay used to remove the first access unit. Instead, the time offset enables feeding a subsequent access unit after the first access unit at a time instance before the time raster with the time delay removed in advance, i.e., at the maximum of the time advance of the time offset. Here, in the example of [figure], correspondingly, for feeding, feeding the fifth access unit can be resumed immediately after the fourth access unit without having to wait or stop feeding until the time for the fifth access unit according to the time raster with only that time delay advanced is reached. Figure 5 In the example of [figure], correspondingly, for feeding, feeding the fifth access unit can be resumed immediately after the fourth access unit without having to wait or stop feeding until the time for the fifth access unit according to the time raster with only that time delay advanced is reached.
[0060] By doing so, the scheduling changes as follows:
[0061] initArrivalEarliestTime[n]=RemovalTime[n] - InitCpbRemovalDelay(i) -InitialCPBRemovalDelayOffset(i)
[0062] This means that the CPB size greater than the CPB A of the CPB B can correspond to the size achieved by feeding the CPB at Bitrate(i) for InitCpbRemovalDelay(i)+InitialCPBRemovalDelayOffset(i).
[0063] To summarize the working principle described above, in terms of how initArrivalEarliestTime is calculated, there are two types of frames. For the first picture (or access unit) in the caching period (i.e., the caching period is defined as the period from the AU with the caching period SEI message until the next AU carrying the caching period SEI message), initArrivalEarliestTime is calculated as its RemovalTime minus InitCpbRemovalDelay. For any other AU within the caching period that is not the first AU (i.e., an AU not carrying the caching period SEI message), InitArrivalEarliestTime is calculated as its RemovalTime minus InitCpbRemovalDelay minus InitialCPBRemovalDelayOffset.
[0064] An encoder typically sends a bitstream having a single value or several values (referred to as scheduling options) with parameters related to HRD operation. For example, different rate values used by a decoder to operate its CPB, i.e., the rates at which the CPB can be fed.
[0065] However, there may be scenarios where different rates are desired. For example, when the channel sends data to the decoder in a bursty manner, data is sent at a high bitrate (whenever any data is available), and the channel is used to send other data when there is no video data to send.
[0066] To solve this problem, it is necessary to calculate the HRD or CPB parameters through some fitting, which can be piecewise linear fitting / regression.
[0067] For example, assume there are two parameter sets corresponding to two rates R 0 and R 1 , where R 0 < R 1 , as shown at Figure 4 300 in i-1 . And assume R sel is selected such that it lies within R 0 and R 1 and corresponds to a value equal to the sum of 90%R 0 and 10%R 1 . Values (such as the CPB size) for such a rate can be calculated similarly by using the same formula (i.e., in the case of this example, the new CPB size is equal to the linear fit of the sum of 90%CPB 0 and 10%CPB 1 ).
[0068] However, when it comes to determining the offset of the earliest arrival time of each picture (i.e., InitialCPBRemovalDelayOffset), the same calculation cannot be performed. For this, there are the following different reasons:
[0069] 1) The buffer fullness (for which the CPB size is the ultimate limit) is affected jointly by the feeding rate used and InitialCPBRemovalDelayOffset, so the linear fit of the InitialCPBRemovalDelayOffset value does not work.
[0070] 2) The actual initial arrival time is not the earliest possible arrival time, but the maximum between the final arrival time of the previous access unit and the earliest possible arrival time, i.e., initArrivalTime[n] = Max(AuFinalArrivalTime[n - 1], initArrivalEarliestTime[n])
[0071] Therefore, at each access unit, the additional data fed into the buffer due to InitialCPBRemovalDelayOffset is proportional to Max(initArrivalTime[n] - AuFinalArrivalTime[n - 1], 0).
[0072] When calculating the fitting of the operating points not signaled by the discrete parameters does not follow a piecewise linear fitting but a more complex fitting (e.g., cubic or polynomial fitting), the calculation of this problem may become more complex.
[0073] Therefore, according to an embodiment of the present application, when using interpolation between operating points (including the interpolated version between the time offsets indicated by one operating point and the time offsets indicated by another operating point) to manage the feeding and emptying of the coded picture buffer 200, the receiving side can rely on an operating mode without problems (i.e., an operating mode without underflow and overflow). The encoder notes that the resulting interpolation does not cause buffer problems.
[0074] That is, according to an embodiment, Figure 3 the device operates as follows. The device will receive the bitstream 14 as the access unit sequence 48. The feeder 202 will sequentially feed the access unit sequence into the coded buffer 200 using the selected bitrate 302 somewhere between the bitrate(i - 1) at one operating point OP i-1 and the bitrate(i) at another operating point OP i for which the CPB parameters 300 i-1 at one operating point OP i and the CPB parameters 300 i-1 and 300 i are Figure 1Encoder 10 writes the bitstream 14. However, feeder 200 feeds the coded picture buffer 200 discontinuously at the selected bitrate, but can pause or stop feeding certain access units. Specifically, feeder 202 stops or pauses feeding any access unit that has not reached the virtual available time (referred to as InitArrivalEarliestTime in the previous description) defined by the time removal raster until that virtual available time is reached, where the time frame removal raster is the time removal delay interpolated in advance for the first access unit in the picture coding order, and for subsequent access units in the picture coding order, it is the sum of the interpolated time removal delay and the interpolated time offset. Similarly, due to the difference in calculating these times, access units that are not the first access unit can stay in the coded picture buffer 200 for a longer time compared to the first access unit. For example, if for a certain access unit, the time removal raster indicates a removal time t removal , but t removal minus the interpolated time removal delay (in the case where the access unit is the first access unit) or t removal minus the sum of the interpolated time removal delay and (plus) the interpolated time offset (in the case where the access unit is a subsequent access unit but not the first access unit) has not arrived, then the feeding of that access unit to the coded picture buffer 200 is delayed until that time is reached. In the above description, InitArrivalEarliestTime in the previous description, the pause or stop is reflected by the max function discussed above, InitialCPBRemovalDelayOffset is used to indicate the time offset, and the time raster is defined by the array RemovalTime. newInitialCPBRemovalDelayOffset is used hereinafter to represent the interpolated time offset. As described above, the time raster can be derived from the data stream 14 in the form of the time difference between the removal of the first access unit and the removal of each of the subsequent access units (e.g., the difference for each subsequent AU), measuring the time between the removal of that AU and the removal of the first AU (of the current buffer cycle).
[0075] However, Figure 3 the device uses the time raster to remove access units from the coded picture buffer 200 access unit by access unit. Due to the fact that the feeding has been advanced by the time removal delay, the removal of the first access unit actually occurs at that time removal delay. Decoder 100 receives the removed access units, decodes them, and outputs the decoded / reconstructed pictures.
[0076] To obtain the interpolated value, Figure 3 the device performs the following operations. It extracts from the bitstream 14 the data related to the first operating point OPi-1 The first CPB parameter 300 related thereto i-1 and the second operating point OP i The second CPB parameter 300 related thereto i , wherein each of the first CPB parameter and the second CPB parameter indicates a CPB size, a predetermined time offset, a predetermined time removal delay, and a predetermined bit rate, and wherein the CPB parameter 300 i-1 and 300 i are for different predetermined bit rates. Then, the apparatus determines an interpolated time offset by interpolating between the predetermined time offsets indicated by the first CPB parameter 300 i-1 at the selected bit rate 302 and the predetermined time offsets indicated by the second CPB parameter 300 i , and the apparatus determines an interpolated time removal delay by interpolating between the predetermined time removal delays indicated by the first CPB parameter 300 i-1 at the selected bit rate 302 and the predetermined time removal delays indicated by the second CPB parameter 300 i . Specific embodiments will be described later, which indicate for example that the interpolation can be linear interpolation. That is, Figure 3 The apparatus can linearly interpolate between the predetermined time offsets indicated by the CPB parameters 300 i-1 and 300 i , and linearly interpolate between the predetermined time removal offsets indicated by these adjacent CPB parameters. However, different kinds of interpolation can alternatively be used. It is even possible that the encoder determines one or more interpolation parameters to parameterize the interpolation and sends one or more interpolation parameters in the data stream, Figure 3 The apparatus derives this information from the data stream 14 and performs interpolation accordingly. The interpolation methods for the time offset and the time removal delay can be the same. Additionally, Figure 3 The apparatus can also interpolate between the CPB sizes indicated by the first CPB parameter 300 i-1 at the selected bit rate 302 and the CPB sizes indicated by the second CPB parameter 300 i to obtain an interpolated CPB size, and Figure 3The apparatus can rely on the fact that the interpolated CPB size for the CPB 200 is sufficient to accommodate any full state that occurs when feeding and emptying the coded picture buffer 200 using the interpolated values for the time removal delay and the time offset. As previously mentioned, the encoder ensures that this commitment holds. For example, for a discrete set of intermediate bitrates between bitrate(i - 1) and bitrate(i) (e.g., in units of one - tenth or another fraction of the interval between these bitrates), the encoder can limit this commitment on the decoding side. Then, the encoder can test whether any overflow or underflow situations occur, and if so, adjust any values accordingly. For example, if an overflow occurs, the encoder can increase one of the CPB sizes indicated by the operating points OP i and OP i-1 Alternatively, the encoder can resume the entire encoding of the bitstream and the determination of the CPB parameters and the conflict - free check of the interpolated values.
[0077] In other words, according to an embodiment, the encoder can ensure that a weighted linear combination of two values among the indicated discrete InitialCPBRemovalDelayOffset values can be calculated and used as newInitialCPBRemovalDelayOffset for calculating the earliest arrival time, such that the HRD constraints (for the CPB size and the bitrate) calculated when fitting the CPB and bitrate curves from the corresponding indicated discrete values result in valid decoder operation. As discussed previously, the curve fitting of the CPB size and bitrate curves can be:
[0078] · Linear
[0079] · Cubic
[0080] · Polynomial
[0081] That is to say, the encoder and decoder can use interpolation other than the piece - wise linear interpolation between the OP i values of the operating points (and specifically, between their InitCpbRemovalDelayOffset i values for the time offset).
[0082] According to an embodiment, the encoder can indicate the weights (α 0 and α 1 ) to be used for interpolation, i.e., newInitialCPBRemovalDelayOffset = α 0 ×InitCpbRemovalDelayOffset 0 +α 1×InitCpbRemovalDelayOffset 1
[0083] As another alternative, instead of signaling α as in the previous equation 0 and α 1 , two other weights (ß 0 and ß 1 ) are provided, which together with the selected rate (R sel ) allow the calculation of the actually used (α 0 and α 1 ).
[0084] α 0 is equal to ß 0 / R sel , and α 1 is equal to ß 1 / R sel
[0085] As another alternative, the provided weights (ß 0 and ß1) can be equal to the bit rate for discrete operating points, and the calculated α 0 and α1 are scaled respectively by the normalized distance between the selected rate R sel and the bit rates provided as HRD parameters R 0 and R 1 . That is, the interpolation can be newInitialCPBRemovalDelayOffset=
[0086] (R sel - R 0 ) / (R 1 - R 0 ) × (R 0 / R sel ) × InitCpbRemovalDelayOffset 0 +
[0087] ( 1 - (R sel - R 0 ) / (R 1 - R 0 )) × (R 1 / R sel ) ×InitCpbRemovalDelayOffset 1
[0088] where R 0 and R 1 are the operating points OP0 and OP 1 for the bit rate, for said operating point OP 0 and OP 1 of the CPB parameter 300 0 and 300 1 are respectively indicated as the time offset InitCpbRemovalDelayOffset 0 and InitCpbRemovalDelayOffset 1 and the selected rate is R sel . That is to say, here, the interpolation is performed by using (R sel -R 0 ) / (R 1 –R 0 )×(R 0 / R sel ) to weight InitCpbRemovalDelayOffset 0 and using (1-(R sel -R 0 ) / (R 1 -R 0 ))×(R 1 / R sel ) to weight InitCpbRemovalDelayOffset 1 , that is, using the product of two ratios.
[0089] The encoder again ensures that the provided bitstream and indication are applied not only to a given discrete operating point, but also to the values that can be calculated therebetween. To do this, the encoder can sample different values between the given discrete values and ensure that there is no buffer underflow or overflow for such values.
[0090] An embodiment of the bitstream is at bit rates R 0 and R 1The described weights are transmitted in a form to calculate the intermediate initial removal offset. A bitstream processing device (e.g., or including a video decoder on the client side) that receives the bitstream generated by the encoder as described above will have to rely on the fact that the encoder generates the bitstream in a manner that meets the above time limit. The device can check the bitstream to ensure that it is actually a legal bitstream, and such a check can be part of the normal decoding process. The check can include: parsing the corresponding bitstream syntax that conveys the CPB and timing information (e.g., the initial removal offset), deriving directly related variables and other variables through their respective equations, and monitoring the values of the variables to comply with the temporal level limits indicated in other syntax of the bitstream (e.g., the level indicator). In addition, a system consisting of an encoder and a decoder can include steps to check the characteristics of the bitstream.
[0091] Figure 6 and Figure 7 shows the CPB buffer filling levels for two operating points where R 0 and R 1 are equal to 13.5 and 16 Mbps respectively. The CPB sizes CPB 0 and CPB 1 are 1400 and 1000 megabytes respectively. And the above offsets, namely InitCpbRemovalDelayOffset 0 and InitCpbRemovalDelayOffset 1 are 30 and 0 milliseconds respectively.
[0092] The following figure shows the CPB fullness for different values of R sel with a value equal to 14.75 Mbps (i.e., exactly between 13.5 and 16 Mbps) and newInitialCPBRemovalDelayOffset. In Figure 8 the latter is equal to InitCpbRemovalDelayOffset 0 (30 ms), resulting in an overflow, while in Figure 9 the latter is equal to InitCpbRemovalDelayOffset 1 (0 ms), resulting in an underflow. When only using the linear fit of the given operating point InitCpbRemovalDelayOffset 0 and InitCpbRemovalDelayOffset 1 (15 ms), the result is Figure 10 's case. Using the calculated value (16.3 ms) discussed above, the result is Figure 11 's case. Figure 12Shows the use of InitCpbRemovalDelayOffset 0 with any value (20 ms) between InitCpbRemovalDelayOffset 1 also results in an overflow. It can be seen that equal to InitCpbRemovalDelayOffset 0 and InitCpbRemovalDelayOffset 1 Any number of offsets may cause problems, namely buffer underflow or overflow. In the above example, the linear fit results in a valid operating value and the explicit weighted pre-configuration provided as discussed above. However, in other cases, the linear fit may result in an overflow (if the encoder doesn't care). For example, Figure 13 depicts the linear fit for the case where R 0 and R 1 are equal to 14.5 and 18 Mbps respectively, and CPB 0 and CPB 1 are 1250 and 1300 megabytes respectively, and InitCpbRemovalDelayOffset 0 and InitCpbRemovalDelayOffset 1 are equal to 15 and 40 respectively. Here, it results in an overflow. In Figure 14 an interpolation using interpolation with a two-factor weight is shown.
[0093] Figure 3 The decoder 100 reconstructs the corresponding picture 16' from each incoming AU, where the implementation details that are possible but optional and can be applied individually or in combination have been described above with respect to Figure 2 . Some options for implementing the decoder and encoder are described in more detail below.
[0094] The first option relates to the processing of the decoded pictures 16' and their caching in the decoded picture buffer. Figure 2 The loop buffer 142 of Figure 3 may include such a DPB. In
[0095] According to an embodiment, Figure 3The apparatus distinguishes between two types of reference pictures, namely, short-term and long-term. The encoder does the same when simulating the DPB filling state of the decoder 100 at each time point during decoding. A reference picture can be marked as "not for reference" when it is no longer needed for prediction reference. The transitions between these three states (short-term, long-term, and not for reference) are controlled by the decoded reference picture marking process. There are two alternative decoded reference picture marking mechanisms, the implicit sliding window process and the explicit memory management control operation (MMCO) process. For each current decoded picture or each current decoded AU, it is signaled in the data stream 14 which process will be used for DPB management. When the number of reference frames is equal to a given maximum number (max-num-ref-frames in the SPS), the sliding window process marks short-term reference pictures as "not for reference". The short-term reference pictures are stored in a first-in-first-out manner such that the most recently decoded short-term picture remains in the DPB. The explicit MMCO process is controlled via a number of MMCO commands. If this mode is selected for the current AU or current decoded picture, the bitstream contains one or more of these commands for that AU or within that AU. The MMCO commands can be any of the following: 1) mark one or more short-term or long-term reference pictures as "not for reference", 2) mark all pictures as "not for reference", or 3) mark the current reference picture or an existing short-term reference picture as long-term and assign a long-term picture index to that long-term picture. The reference picture marking operation and any output (for presentation) and removal of pictures from the DPB can be performed after the pictures have been decoded.
[0096] That is to say, returning to Figure 3, according to an embodiment, the decoder 100 may decode 400 the current AU 402 removed from the CPB 200 and thus received from the CPB 200 to obtain a decoded picture 16'. The decoder 100 may use inter-picture prediction based on reference pictures 404 stored in the aforementioned DPB 406 and included in the aforementioned loop buffer 142. The decoded picture 16' may be inserted 408 into the DPB 406. The insertion 408 may be performed for each current decoded picture 16', that is, each decoded picture may be placed in the DPB 406 after its decoding 400. The insertion 408 may occur instantaneously (i.e., at the CPB removal time of the AU 402 when ignoring the decoding time, or the CPB removal time plus the required decoding time). However, the insertion 408 may also be additionally dependent on specific facts, as indicated at 410. For example, each newly decoded picture is inserted into the DPB 406 unless it is to be output 412 at its CPB removal time, that is, unless it is an immediately output picture, that is, a picture to be immediately output for presentation, and it is a non-reference picture, which can be indicated in the corresponding AU 402 by corresponding syntax elements. The apparatus may assign a classification of one of a short-term reference picture, a long-term reference picture, and a picture not for reference to each reference picture 414 stored in the DPB 406. The apparatus further reads DPB mode information 416 from the current AU 402, and if the mode information 416 indicates an intrinsic mode or a first mode, the intrinsic DPB management process 418 is activated and one or more reference pictures 414 classified as short-term pictures are removed 424 from the DPB 406 according to the FIFO policy. If the mode information 416 indicates an explicit mode or a second mode, the explicit DPB management process 420 is activated, and one or more commands included in the memory management control information included in the current AU 402 are executed to change the classification assigned to at least one of the reference pictures 414 stored in the DPB 406, and the classification of the reference pictures 414 in the DPB 406 is used to manage the removal 424 of the reference pictures from the DPB 406. Regardless of whether process 418 or 420 is selected for the current picture 16', any picture 414 in the DPB 406 whose picture output time has reached is output 422 for presentation. The pictures 414 that are no longer output and are classified as pictures not for reference are removed 424 from the DPB 406.
[0097] Some possible details of the reference picture marking mechanism discussed below Figure 15 are described. 1) The first aspect relates to gaps in frame numbers and non-existent pictures. Although not described above, each reference picture 414 in the DPB 406 may be associated with a frame number, which may be Figure 3The device is derived from the frame number syntax element in AU 402, which indicates the AU ordering in decoding order. Typically, for each reference picture 414, the frame number is incremented by 1, but gaps in the frame number can be allowed by setting the corresponding high-level (e.g., sequence-level) flag (which can be referred to as the parameter-gaps-in-frame-num-allowed-flag) to 1. For example, to allow the encoder or MANE (Media-Aware Network Element) to transmit a bitstream in which the frame number for a reference picture is incremented by more than 1 relative to the previous reference picture in decoding order. This can be advantageous for supporting temporal scalability. Figure 3 A device that receives an AU sequence with gaps in the frame number can be configured to create non-existent pictures to fill the gaps. The non-existent pictures are assigned frame number values in the gaps and are treated as reference pictures during decoding reference picture marking, but will not be used for output (and thus not displayed). The non-existent pictures ensure that the state of the DPB regarding the frame numbers of the pictures residing therein is the same for a decoder that has received the pictures as for a decoder that has not received the pictures.
[0098] Another possible aspect relates to the loss of a reference picture when using a sliding window. When a reference picture is lost, Figure 3 the device can attempt to conceal the picture and, if the feedback channel is available when the loss is detected, may report the loss to the encoder. If gaps in the frame number are not allowed, a discontinuity in the frame number value indicates an unexpected loss of a reference picture. If gaps in the frame number are allowed, the discontinuity in the frame number value may be due to the intentional removal of a temporal layer or subsequence or an unexpected picture loss, and the decoder (e.g., Figure 3 the device) should only infer a picture loss when a non-existent picture is referenced during the inter-picture prediction process. The picture order count of the concealed picture may be unknown, which can cause the decoder (e.g., Figure 3 the device) to use an incorrect reference picture when decoding a B picture without detecting any errors.
[0099] Another possible aspect relates to the loss of a reference picture that contains an MMCO that marks a short-term reference picture as "not for reference". When the reference picture containing the MMCO command is lost, the reference picture state in the DPB becomes incorrect, and thus the reference picture list for several pictures after the lost picture may become incorrect. If a picture containing an MMCO command related to a long-term reference picture is lost, there is a risk that the number of long-term reference pictures in the DPB is different from the number that would be long-term reference pictures when the picture is received, resulting in an "incorrect" sliding window process for all subsequent pictures. That is, the encoder and decoder (i.e., Figure 3The device) will contain different numbers of short-term reference pictures, resulting in an asynchronous behavior of the sliding window process. To make matters worse, the decoder will not necessarily know that the sliding window process is asynchronous.
[0100] The following diagrams illustrate the possible MMCO commands described above. One or more or all of these commands can be applied to produce different embodiments:
[0101] memory_management_control_operation Memory management control operation 0 End of memory_management_control_operation syntax element loop 1 Mark short-term reference pictures as "not for reference" 2 Mark long-term reference pictures as "not for reference" 3 Mark short-term reference pictures as "for long-term reference" and assign a long-term frame index to them 4 Specify the maximum long-term frame index and mark long-term reference pictures with all long-term frame indices greater than the maximum value as "not for reference" 5 Mark all reference pictures as "not for reference" and set the MaxLongTermFrameldx variable to "no long-term frame index" 6 Mark the current picture as "for long-term reference" and assign a long-term frame index to it
[0102] Another option for the implementation of the decoder and encoder is now described, which can optionally be combined with the options described previously regarding DPB management and which relates to entropy decoding a certain syntax element (e.g., residual data in the form of transform coefficients) into the bitstream 14. Lossless entropy coding of lossy quantized transform coefficients is a key part of an efficient video codec. One such method is called context-adaptive variable-length coding (CAVLC), where the encoder switches between different variable-length code (VLC) tables for various syntax elements, depending on the value of the previously transmitted syntax element in the same slice in a context-adaptive manner. The encoder and decoder can use CAVLC. Due to the fact that each syntax element is encoded into the bitstream 14 by writing the corresponding codeword (which has been selected from the context-adaptively chosen code table for that syntax element) into the bitstream, each CAVLC-encoded bit in the bitstream can be associated with a single syntax element. Thus, when using CAVLC, the relevant information regarding the level of transform coefficients that appear in the bitstream 14 in scan order can be obtained in a directly accessible form as a syntax element. The encoder and decoder can use CAVLC to signal the transform coefficients in the bitstream 14. The following syntax elements can be used, i.e., syntax elements with the following semantics:
[0103] - A syntax element (indicated by CoeffToken) that indicates the total number of non-zero transform coefficient levels in a transform block
[0104] - One or more syntax elements that indicate the signs of several trailing-ones transform coefficient levels (i.e., a series of syntax elements (all 1s) that appear when scanning the syntax elements in scan order until the end of the last non-zero syntax element)
[0105] - One or more syntax elements for each non-zero transform coefficient other than the trailing-ones transform coefficient, which indicate the transform coefficient level value
[0106] - A syntax element that indicates the total number of zero-valued transform coefficient levels
[0107] - A syntax element that indicates the number of consecutive transform coefficient levels with zero values before a non-zero transform coefficient level is encountered in scan order from the current scan position forward.
[0108] Alternatively or additionally, the encoder may select between the use of CABAC (thus Context-Adaptive Binary Arithmetic Coding) and the use of CAVLC and signal this selection in the bitstream 14, and the decoder reads this signal and uses the indicated manner of decoding the residual data.
[0109] Another option for the implementation of the decoder and encoder is now described, which may optionally be combined with any of the previously described options regarding DPB management and options regarding CAVLC, and which relates to the quarter-pixel interpolation filter. To allow for inter prediction at a finer granularity than the regular full-pixel sample grid, a sample interpolation process is used to derive sample values at sub-pixel sample positions, the range of which can be from half-pixel positions to quarter-pixel positions. The encoder and decoder may use a method of performing quarter-pixel interpolation, and the method is as follows. First, 6-tap FIR filters are used to generate sample values at half-pixel positions, and then the generated sample values at half-pixel positions are averaged by interpolation to generate sample values at quarter-pixel positions for the luminance component.
[0110] For completeness, Figure 16 An encoder of a device suitable Figure 3 for encoding video 12 into a bitstream 14 is shown, which encoder may or may not be in accordance with Figure 1 and provides CPB parameters 300 for the bitstream 14 in a manner such that the interpolation to be performed by the Figure 3 device results in an operation without underflow and overflow.
[0111] Accordingly, the following embodiments or aspects may be derived from the above description, and the following embodiments or aspects may in turn be further extended individually or in combination by any of the above details and facts.
[0112] According to a first aspect, an apparatus for video decoding may include an encoded picture buffer 200 and be configured to receive a data stream 14 having pictures 16 of video 12 encoded therein in coded order as a sequence of access units 48, sequentially feed the sequence of access units 48 into the CPB using a selected bitrate 302, wherein feeding of an access unit that has not reached the virtual available time for raster removal according to a time frame is paused until the virtual available time is reached, wherein the time frame raster removal advances the selected time removal delay for a first access unit in the coded order, and for subsequent access units in the coded order advances the selected time removal delay by the sum of the selected time offset; remove AUs from the CPB on a per-AU basis using a time raster [RemovalTime], extract a first CPB parameter 300 related to a first operating point from the data stream i-1 and a second CPB parameter 300 related to a second operating point i , each of the first CPB parameter and the second CPB parameter indicating a CPB size, a predetermined time offset, a predetermined time removal delay, and a predetermined bitrate, wherein the first CPB parameter 300 i-1 differs from the second CPB parameter 300 at least in terms of the predetermined bitrate i , determine the selected time offset by interpolating between the predetermined time offsets indicated by the first CPB parameter 300 i-1 at the selected bitrate and indicated by the second CPB parameter 300 i , and determine the selected time removal delay by interpolating between the predetermined time removal delays indicated by the first CPB parameter 300 i-1 at the selected bitrate and indicated by the second CPB parameter 300 i .
[0113] According to a second aspect, when referring back to the first aspect, the apparatus may be configured to: derive one or more interpolation parameters from the data stream and parameterize the interpolation using the one or more interpolation parameters.
[0114] According to a third aspect, when referring back to the first aspect or the second aspect, the apparatus may be configured to: perform the interpolation using a weighted sum of the predetermined time offset indicated by the first CPB parameter weighted by a first weight and the predetermined time offset indicated by the second CPB parameter weighted by a second weight.
[0115] According to a fourth aspect, when referring back to the third aspect, the apparatus may be configured to determine the first weight and the second weight based on the selected bit rate, the predetermined bit rate indicated by the first CPB parameter, and the predetermined bit rate indicated by the second CPB parameter.
[0116] According to a fifth aspect, when referring back to the third aspect, the apparatus may be configured to calculate a linear interpolation weight by dividing a difference between the selected bit rate and the predetermined bit rate indicated by the first CPB parameter by a difference between the predetermined bit rate indicated by the first CPB parameter and the predetermined bit rate indicated by the second CPB parameter; and use the linear interpolation weight to determine the first weight and the second weight.
[0117] According to a sixth aspect, when referring back to the fifth aspect, the apparatus may be configured to determine the first weight such that the first weight is the linear interpolation weight or a product where one of the factors is the linear interpolation weight; and determine the second weight such that the second weight is a difference between the linear interpolation weight and 1 or a product where one of the factors is a difference between the linear interpolation weight and 1.
[0118] According to a seventh aspect, when referring back to the fifth aspect, the apparatus may be configured to determine the first weight such that the first weight is a product where a first factor of the product is the linear interpolation weight and a second factor of the product is the predetermined bit rate indicated by the first CPB parameter divided by the selected bit rate; and determine the second weight such that the second weight is a product where one of the factors of the product is a difference between the linear interpolation weight and 1 and a second factor of the product is the predetermined bit rate indicated by the second CPB parameter divided by the selected bit rate.
[0119] According to an eighth aspect, when referring back to any one of the first to seventh aspects, the apparatus may further include a decoded picture buffer (DPB) and be configured to: decode the current AU 402 removed from the CPB 200 using inter-picture prediction based on a reference picture 404 stored in the DPB that is being referred to, to obtain a decoded picture 16', and insert 408 the decoded picture into the DPB; assign to each reference picture 414 stored in the DPB one of a classification as a short-term reference picture, a long-term reference picture, and a picture not for reference; read DPB mode information 416 from the current AU; if the DPB mode information indicates a first mode, remove 424 from the DPB one or more reference pictures classified as short-term pictures according to a FIFO policy; if the DPB mode information indicates a second mode, read memory management control information including at least one command from the current AU and execute the at least one command to change the classification assigned to at least one of the reference pictures stored in the DPB; and manage the removal 424 of reference pictures from the DPB using the classification of the reference pictures in the DPB.
[0120] According to a ninth aspect, when referring back to the eighth aspect, the apparatus may be configured to: read from the current AU an indication as to whether the decoded picture is not for inter-picture prediction; if the decoded picture is not indicated as not for inter-picture prediction or not directly output, perform inserting the decoded picture into the DPB; and if the decoded picture is indicated as not for inter-picture prediction and directly output, directly output the decoded picture without caching the decoded picture in the DPB.
[0121] According to a tenth aspect, when referring back to the eighth or ninth aspect, the apparatus may be configured to: assign a frame index to each reference picture classified as a long-term picture in the DPB; and if the frame index assigned to a predetermined reference picture classified as a long-term picture in the DPB is referred to in the current AU, use the predetermined reference picture as the reference picture being referred to in the DPB.
[0122] According to the eleventh aspect, when referring back to the tenth aspect, the apparatus may be configured to perform one or more of the following: if at least one command in the current AU is the first command, reclassify a reference picture classified as a short-term reference picture in the DPB as a picture not for reference; if at least one command in the current AU is the second command, reclassify a reference picture classified as a long-term reference picture in the DPB as a picture not for reference; if at least one command in the current AU is the third command, reclassify a reference picture classified as a short-term picture in the DPB as a long-term reference picture and assign a frame index to the reclassified reference picture; if at least one command in the current AU is the fourth command, set an upper limit of the frame index according to the fourth command, and reclassify all reference pictures classified as long-term pictures in the DPB and to which a frame index exceeding the upper limit of the frame index has been assigned as pictures not for reference; if at least one command in the current AU is the fifth command, classify the current picture as a long-term picture and assign a frame index to the current picture.
[0123] According to the twelfth aspect, when referring back to any one of the eighth to eleventh aspects, the apparatus may be configured to: remove from the DPB any reference picture that is classified as a picture not for reference and is no longer output.
[0124] According to the thirteenth aspect, when referring back to any one of the first to twelfth aspects, the apparatus may be configured to: read an entropy coding mode indicator from a data stream, and if the entropy coding mode indicator indicates a context adaptive variable length coding mode, use the context adaptive variable length coding mode to decode prediction residual data from the current AU, and if the entropy coding mode indicator indicates a context adaptive binary arithmetic coding mode, use the context adaptive binary arithmetic coding mode to decode prediction residual data from the current AU. Optionally but not mandatorily, decoding prediction residual data from the current AU using the context adaptive variable length coding mode may involve using a first syntax element indicating the total number of non-zero transform coefficient levels in a transform block, a second syntax element indicating the total number of zero-valued transform coefficient levels in the transform block, a third syntax element indicating the number of consecutive zero-valued transform coefficient levels from the current scan position forward before encountering a non-zero transform coefficient level in scan order, one or more fourth syntax elements for each non-zero transform coefficient other than the trailing 1 transform coefficient, the one or more fourth syntax elements indicating the transform coefficient level value of the corresponding non-zero value transform coefficient, and one or more fifth syntax elements indicating the sign of the transform coefficient level of the trailing 1 transform coefficient.
[0125] According to the fourteenth aspect, when referring back to any one of the first to thirteenth aspects, the apparatus may be configured to: derive quarter-pixel values in a reference picture being referred to, based on motion vectors in a current AU and using a 6-tap FIR filter to derive half-pixel values and average adjacent half-pixel values.
[0126] According to the fifteenth aspect, when referring back to any one of the first to fourteenth aspects, the apparatus may be configured to: derive information about a temporal raster from a data stream by a time difference between removal of a first access unit and removal of each of subsequent access units.
[0127] According to the sixteenth aspect, when referring back to any one of the first to fifteenth aspects, the apparatus may be configured to: interpolate between a CPB size indicated by a first CPB parameter 300 and a CPB size indicated by a second CPB parameter 300 at a selected bitrate 302 to obtain an interpolated CPB size, thereby determining a minimum CPB size of an encoded picture buffer 200. i-1 i
[0128] According to the seventeenth aspect, when referring back to any one of the first to sixteenth aspects, in the apparatus, the selected bitrate may be between a predetermined bitrate indicated by a first CPB parameter 300 i-1 i and a predetermined bitrate indicated by a second CPB parameter 300.
[0129] According to the eighteenth aspect, when referring back to any one of the first to seventeenth aspects, the apparatus may be configured to: operate in units of a caching period, and a first access unit in coding order is the first access unit of a current caching period.
[0130] According to a nineteenth aspect, an apparatus for encoding video into a data stream, wherein the data stream is to be decoded by feeding the data stream to a decoder including a coded picture buffer (CPB). The apparatus may be configured to: encode pictures of the video encoded in coding order into the data stream as an access unit (AU) sequence; determine a first CPB parameter associated with a first operating point and a second CPB parameter associated with a second operating point, each of the first CPB parameter and the second CPB parameter indicating a CPB size, a predetermined time offset, a predetermined time removal delay, and a predetermined bit rate, wherein the first CPB parameter is different from the second CPB parameter at least in terms of the predetermined bit rate; and perform the determination such that interpolation between the predetermined time offset of the first CPB parameter and the predetermined time offset of the second CPB parameter at each of a plurality of selected bit rates results in an interpolated time offset and an interpolated time removal delay. Thus, the data stream is fed to the decoder via the CPB by: sequentially feeding the AU sequence to the CPB using the respective selected bit rate, wherein feeding of access units not having reached the virtual available time according to the time frame removal raster is paused until the virtual available time is reached, wherein the time frame removal raster advances the interpolated time removal delay for a first access unit in the coding order, and advances the interpolated time removal delay plus the interpolated time offset for subsequent access units in the coding order; and removing AUs from the CPB on a per-AU basis using a time raster without causing any underflow and any overflow, i.e., no underflow and no overflow occurs, and encode the CPB parameter into the data stream.
[0131] According to a twentieth aspect, when referring back to the nineteenth aspect, in the apparatus, the interpolation may be parameterized using an interpolation parameter, and the apparatus may be configured to encode the interpolation parameter into the data stream.
[0132] According to a twenty-first aspect, when referring back to the nineteenth or twentieth aspect, in the apparatus, the interpolation is performed using a weighted sum of the predetermined time offset indicated by the first CPB parameter weighted by a first weight and the predetermined time offset indicated by the second CPB parameter weighted by a second weight.
[0133] According to a twenty-second aspect, when referring back to the twenty-first aspect, in the apparatus, the first weight and the second weight are determined based on the selected bit rate, the predetermined bit rate indicated by the first CPB parameter, and the predetermined bit rate indicated by the second CPB parameter.
[0134] According to the twenty-third aspect, when referring back to the twenty-first aspect, in the apparatus, the linear interpolation weight determined by dividing the difference between the selected bitrate and the predetermined bitrate indicated by the first CPB parameter by the difference between the predetermined bitrate indicated by the first CPB parameter and the predetermined bitrate indicated by the second CPB parameter can be used to determine the first weight and the second weight.
[0135] According to the twenty-fourth aspect, when referring back to the twenty-third aspect, in the apparatus, the first weight can be determined such that the first weight is the linear interpolation weight or a product where one of the factors is the linear interpolation weight; and the second weight can be determined such that the second weight is the difference between the linear interpolation weight and 1 or a product where one of the factors is the difference between the linear interpolation weight and 1.
[0136] According to the twenty-fifth aspect, when referring back to the twenty-third aspect, in the apparatus, the first weight is determined such that the first weight is a product: the first factor of the product is the linear interpolation weight, and the second factor of the product is the predetermined bitrate indicated by the first CPB parameter divided by the selected bitrate; and the second weight can be determined such that the second weight is a product: one of the factors of the product is the difference between the linear interpolation weight and 1, and the second factor of the product is the predetermined bitrate indicated by the second CPB parameter divided by the selected bitrate.
[0137] According to the twenty-sixth aspect, when referring back to any one of the nineteenth aspect to the twenty-fifth aspect, the apparatus can be configured to: when encoding the AU, use inter-picture prediction based on the reference picture referred to stored in the decoded picture buffer (DPB) to encode the current picture into the current AU, and insert the decoded version of the current picture in the DPB into the DPB, assign to each reference picture stored in the DPB a classification as one of a short-term reference picture, a long-term reference picture, and a picture not for reference, write the DPB mode information into the current AU, if the DPB mode information indicates the first mode, remove one or more reference pictures classified as short-term pictures from the DPB according to the FIFO policy, if the DPB mode information indicates the second mode, write memory management control information including at least one command into the current AU, the command indicating a change in the classification assigned to at least one of the reference pictures stored in the DPB, wherein the classification of the reference pictures in the DPB is used to manage the removal of the reference pictures from the DPB.
[0138] According to the twenty-seventh aspect, when referring back to the twenty-sixth aspect, the apparatus may be configured to: write an indication of whether the decoded picture is not used for inter-picture prediction into the current AU; wherein, if the decoded picture is not indicated as not used for inter-picture prediction or not directly output, the decoded picture will be inserted into the DPB, and if the decoded picture is indicated as not used for inter-picture prediction and directly output, the decoded picture will be directly output without caching the decoded picture in the DPB.
[0139] According to the twenty-eighth aspect, when referring back to the twenty-sixth or twenty-seventh aspect, in the apparatus, a frame index will be assigned to each reference picture classified as a long-term picture in the DPB, and if the frame index of a predetermined reference picture classified as a long-term picture in the DPB is referred to in the current AU, the predetermined reference picture will be used as the referenced reference picture in the DPB.
[0140] According to the twenty-ninth aspect, when referring back to the twenty-eighth aspect, in the apparatus, one or more of the following are performed: if at least one command in the current AU is a first command, the reference pictures classified as short-term reference pictures in the DPB will be re-classified as pictures not for reference; if at least one command in the current AU is a second command, the reference pictures classified as long-term reference pictures in the DPB will be re-classified as pictures not for reference; if at least one command in the current AU is a third command, the reference pictures classified as short-term pictures in the DPB will be re-classified as long-term reference pictures, and a frame index will be assigned to the re-classified reference pictures; if at least one command in the current AU is a fourth command, the frame index upper limit will be set according to the fourth command, and all reference pictures classified as long-term pictures in the DPB to which a frame index exceeding the frame index upper limit has been assigned will be re-classified as pictures not for reference; if at least one command in the current AU is a fifth command, the current picture will be classified as a long-term picture, and a frame index will be assigned to the current picture.
[0141] According to the thirtieth aspect, when referring back to any one of the twenty-sixth aspect to the twenty-ninth aspect, in the apparatus, pictures classified as not for reference and from which no reference pictures are output any longer will be removed from the DPB.
[0142] According to the thirty - first aspect, when referring back to any one of the nineteenth aspect to the thirty - first aspect, the apparatus may be configured to: write an entropy coding mode indicator into the data stream, and if the entropy coding mode indicator indicates the context - adaptive variable - length coding mode, encode the prediction residual data into the current AU using the context - adaptive variable - length coding mode, and if the entropy coding mode indicator indicates the context - adaptive binary arithmetic coding mode, encode the prediction residual data into the current AU using the context - adaptive binary arithmetic coding mode.
[0143] According to the thirty - second aspect, when referring back to any one of the nineteenth aspect to the thirty - first aspect, the apparatus may be configured to: derive quarter - pixel values in a reference picture being referred to based on motion vectors in the current AU and using a 6 - tap FIR filter to derive half - pixel values and average adjacent half - pixel values.
[0144] According to the thirty - third aspect, when referring back to any one of the nineteenth aspect to the thirty - second aspect, the apparatus may be configured to: provide information about the temporal raster for the data stream by the time difference between the removal of the first access unit and the removal of each of the subsequent access units.
[0145] According to the thirty - fourth aspect, when referring back to any one of the nineteenth aspect to the thirty - third aspect, in the apparatus, interpolation will be performed between the CPB size indicated by the first CPB parameter 300 i-1 and the CPB size indicated by the second CPB parameter 300 i at the selected bitrate 302 to obtain an interpolated CPB size, thereby determining the minimum CPB size of the coded picture buffer 200.
[0146] According to the thirty - fifth aspect, when referring back to any one of the nineteenth aspect to the thirty - fourth aspect, in the apparatus, the selected bitrate may be between the predetermined time offset indicated by the first CPB parameter 300 i-1 and the predetermined time offset indicated by the second CPB parameter 300i.
[0147] According to the thirty - sixth aspect, when referring back to any one of the nineteenth aspect to the thirty - fifth aspect, the apparatus may be configured to operate in units of cache cycles, and the first access unit in the coding order is the first access unit of the current cache cycle.
[0148] According to the thirty-seventh aspect, when referring back to any one of the nineteenth to thirty-sixth aspects, the apparatus may be configured to: perform the determination by determining a preliminary version of a first CPB parameter and a second CPB parameter; perform interpolation at a plurality of selected bitrates to obtain an interpolated time offset and an interpolated time removal delay for each of the plurality of selected bitrates; and for each of the plurality of selected bitrates, check whether feeding the data stream to the decoder via the CPB using the interpolated time offset and the interpolated time removal delay obtained for each selected bitrate results in underflow and overflow, and if so, recover the encoding in a different way, modify the preliminary version of the first CPB parameter and the second CPB parameter, or recover the interpolation in a different way, and if not, determine that the first CPB parameter and the second CPB parameter are equal to the preliminary version.
[0149] According to the thirty-eighth aspect, a method for video decoding using an encoded picture buffer 200 may have the following steps: receiving a data stream 14 having pictures 16 of a video 12 encoded therein in coded order as a sequence of access units 48, sequentially feeding the sequence of the access units 48 into the CPB using a selected bitrate 302, where feeding of access units that have not reached the virtual available time according to the time frame removal raster is paused until the virtual available time is reached, where the time frame removal raster advances a selected time removal delay for a first access unit in the coded order, and advances a sum of the selected time removal delay and a selected time offset for subsequent access units in the coded order; removing AUs from the CPB on a per-AU basis using a time raster [RemovalTime], extracting a first CPB parameter 300 related to a first operating point from the data stream i-1 and a second CPB parameter 300 related to a second operating point i , each of the first CPB parameter and the second CPB parameter indicating a CPB size, a predetermined time offset, a predetermined time removal delay, and a predetermined bitrate, where the first CPB parameter 300 i-1 differs from the second CPB parameter 300 at least in terms of the predetermined bitrate i , determining the selected time offset by interpolating between the predetermined time offsets indicated by the first CPB parameter 300 i-1 at the selected bitrate and the second CPB parameter 300 i , and determining the selected time removal delay by interpolating between the predetermined time removal delays indicated by the first CPB parameter 300 i-1 at the selected bitrate and the second CPB parameter 300 i .
[0150] The thirty-ninth aspect may have a data stream into which video can be encoded and which may include a first CPB parameter and a second CPB parameter, such that the method according to the thirty-eighth aspect does not result in CPB overflow and underflow.
[0151] According to the fortieth aspect, a method for encoding video into a data stream (wherein the data stream is to be decoded by feeding the data stream into a decoder including a coded picture buffer (CPB)) may have the following steps: encoding pictures of the video encoded in coding order into the data stream as an access unit (AU) sequence, determining a first CPB parameter associated with a first operating point and a second CPB parameter associated with a second operating point, each of the first CPB parameter and the second CPB parameter indicating a CPB size, a predetermined time offset, a predetermined time removal delay, and a predetermined bit rate, wherein the first CPB parameter is different from the second CPB parameter at least in terms of the predetermined bit rate, and performing the determination such that interpolation between the predetermined time offset of the first CPB parameter and the predetermined time offset of the second CPB parameter at each of a plurality of selected bit rates results in an interpolated time offset and an interpolated time removal delay, and thereby feeding the data stream into the decoder via the CPB by: sequentially feeding the AU sequence into the CPB using the respective selected bit rate, wherein feeding of access units not having reached the virtual available time according to the time frame removal raster is paused until the virtual available time is reached, wherein the time frame removal raster advances the interpolated time removal delay for a first access unit in the coding order, and advances the sum of the interpolated time removal delay and the interpolated time offset for subsequent access units in the coding order; and removing AUs from the CPB on a per-AU basis using a time raster without causing any underflow and any overflow.
[0152] The forty-first aspect may have a data stream generated by the method according to the fortieth aspect.
[0153] It should be understood that in this specification, a signal on a line is sometimes named by the reference numeral of the line, or sometimes represented by the reference numeral itself attributed to the line. Thus, this marking manner enables the line having a certain signal to indicate the signal itself. The line may be a physical line in a hardwired implementation. However, in a computerized implementation, there is no physical line, but the signal represented by the line is sent from one computing module to another computing module.
[0154] Although the present invention has been described in the context of a block diagram where blocks represent physical or logical hardware components, the present invention can also be implemented by a computer-implemented method. In the latter case, the blocks represent corresponding method steps, where these steps represent functions performed by corresponding logical or physical hardware blocks.
[0155] Although some aspects have been described in the context of an apparatus, it will be clear that these aspects also represent a description of a corresponding method, where the blocks or apparatus correspond to method steps or features of method steps. Similarly, aspects described in the context of method steps also represent a description of corresponding items or features of corresponding blocks or corresponding apparatuses. Some or all of the method steps can be performed by (or using) a hardware device such as a microprocessor, a programmable computer, or an electronic circuit. In some embodiments, one or more of the most important method steps can be performed by such a device.
[0156] The transmitted or encoded signals of the present invention can be stored on a digital storage medium or can be transmitted on a transmission medium such as a wireless transmission medium or a wired transmission medium (such as the Internet).
[0157] Depending on certain implementation requirements, embodiments of the present invention can be implemented in hardware or in software. The implementation can be performed by using a digital storage medium (e.g., a floppy disk, a DVD, a Blu-Ray, a CD, a ROM, a PROM, an EPROM, an EEPROM, or a flash memory) storing electronically readable control signals that cooperate (or are capable of cooperating) with a programmable computer system to perform the corresponding method. Thus, the digital storage medium can be computer-readable.
[0158] Some embodiments according to the present invention include a data carrier having electronically readable control signals that are capable of cooperating with a programmable computer system to perform one of the methods described herein.
[0159] Generally, embodiments of the present invention can be implemented as a computer program product having program code that is operable to perform one of the methods when the computer program product runs on a computer. The program code can be stored, for example, on a machine-readable carrier.
[0160] Other embodiments include a computer program stored on a machine-readable carrier for performing one of the methods described herein.
[0161] In other words, embodiments of the method of the present invention are thus a computer program having program code for performing one of the methods described herein when the computer program runs on a computer.
[0162] Accordingly, another embodiment of the method of the present invention is a data carrier (or a non-transitory storage medium such as a digital storage medium or a computer-readable medium) comprising a computer program recorded thereon for performing one of the methods described herein. The data carrier, digital storage medium or recording medium is generally tangible and / or non-transitory.
[0163] Accordingly, another embodiment of the method of the present invention is a data stream or signal sequence representing a computer program for performing one of the methods described herein. The data stream or signal sequence can be configured, for example, to be transmitted via a data communication connection (e.g., via the Internet).
[0164] Another embodiment includes a processing device, e.g., a computer or a programmable logic device, configured or adapted to perform one of the methods described herein.
[0165] Another embodiment includes a computer having installed thereon a computer program for performing one of the methods described herein.
[0166] Another embodiment according to the present invention includes a device or system configured to transmit (e.g., electronically or optically) to a receiver a computer program for performing one of the methods described herein. The receiver can be, for example, a computer, a mobile device, a storage device, etc. The device or system can include, for example, a file server for transmitting the computer program to the receiver.
[0167] In some embodiments, a programmable logic device (e.g., a field programmable gate array) can be used to perform some or all of the functions of the methods described herein. In some embodiments, a field programmable gate array can cooperate with a microprocessor to perform one of the methods described herein. Generally, these methods can be performed by any hardware device.
[0168] The above embodiments are merely illustrative of the principles of the present invention. It should be understood that modifications and variations of the arrangements and details described herein will be apparent to other those skilled in the art. Accordingly, it is intended to be limited only by the scope of the appended patent claims rather than by the specific details given by the description and explanation of the embodiments herein.
[0169] References
[0170] [1] From Sjoberg, Rickard, et al. "Overview of HEVC high-level syntaxand reference picture management." IEEE transactions on Circuits and Systemsfor Video Technology 22.12 (2012): 1858-1870。
Claims
1. An apparatus for video decoding, the apparatus comprising a coded picture buffer (CPB) and a decoded picture buffer (DPB), configured to: Receive a data stream having pictures of a video encoded therein in coded order as an access unit (AU) sequence, Feed the access unit sequence sequentially into the CPB at a selected bitrate, wherein feeding of access units not having reached the virtual available time for removing the raster according to a time frame is paused until the virtual available time is reached, wherein the time frame removes the delay in advance by a selected time for a first access unit in the coded order, and for subsequent access units in the coded order, the delay removed in advance by a selected time is the sum of the selected time and a selected time offset; Remove AUs from the CPB on a per-AU basis using the time raster, Extract a first CPB parameter related to a first operating point and a second CPB parameter related to a second operating point from the data stream, each of the first CPB parameter and the second CPB parameter indicating a CPB size, a predetermined time offset, a predetermined time removal delay, and a predetermined bitrate, wherein, the first CPB parameter is different from the second CPB parameter at least in terms of the predetermined bitrate, determine the selected time offset by interpolating between the predetermined time offset indicated by the first CPB parameter and the predetermined time offset indicated by the second CPB parameter at the selected bitrate, and determine the selected time removal delay by interpolating between the predetermined time removal delay indicated by the first CPB parameter and the predetermined time removal delay indicated by the second CPB parameter at the selected bitrate, Decode a current AU removed from the CPB using inter-picture prediction according to a referenced reference picture stored in the DPB to obtain a decoded picture, and Insert the decoded picture into the DPB, Assign to each reference picture stored in the DPB one of a classification as a short-term reference picture, a long-term reference picture, and a picture not for reference, Read DPB mode information from the current AU, If the DPB mode information indicates a first mode, remove one or more reference pictures classified as short-term pictures from the DPB according to a first-in first-out (FIFO) policy, If the DPB mode information indicates a second mode, then Read memory management control information in the current AU including at least one command and execute the at least one command to change the classification assigned to at least one of the reference pictures stored in the DPB, and use the classification of the reference pictures in the DPB to manage removal of the reference pictures from the DPB.
2. The apparatus according to claim 1, configured to: Derive one or more interpolation parameters from the data stream, and Parameterize the interpolation using the one or more interpolation parameters.
3. The apparatus according to claim 1, configured to: The interpolation is performed using a weighted sum of a predetermined time offset indicated by the first CPB parameter weighted by a first weight and a predetermined time offset indicated by the second CPB parameter weighted by a second weight.
4. The apparatus according to claim 3, configured to: Determine the first weight and the second weight based on the selected bitrate, a predetermined bitrate indicated by the first CPB parameter, and a predetermined bitrate indicated by the second CPB parameter.
5. The apparatus according to claim 3, configured to: Calculate a linear interpolation weight by dividing a difference between the selected bitrate and a predetermined bitrate indicated by the first CPB parameter by a difference between the predetermined bitrate indicated by the first CPB parameter and a predetermined bitrate indicated by the second CPB parameter, and Use the linear interpolation weight to determine the first weight and the second weight.
6. The apparatus according to claim 5, configured to: Determine the first weight such that the first weight is the linear interpolation weight or a product where one of the factors is the linear interpolation weight, and Determine the second weight such that the second weight is a difference between the linear interpolation weight and 1 or a product where one of the factors is a difference between the linear interpolation weight and 1.
7. The apparatus according to claim 5, configured to: Determine the first weight such that the first weight is a product: a first factor of the product is the linear interpolation weight, and a second factor of the product is a predetermined bitrate indicated by the first CPB parameter divided by the selected bitrate, and Determine the second weight such that the second weight is a product: one factor of the product is a difference between the linear interpolation weight and 1, and a second factor of the product is a predetermined bitrate indicated by the second CPB parameter divided by the selected bitrate.
8. The apparatus according to claim 1, configured to: Read an indication from the current AU as to whether the decoded picture is not used for inter-picture prediction; If the decoded picture is not indicated as not used for inter-picture prediction or not directly output, perform inserting the decoded picture into the DPB, and if the decoded picture is indicated as not used for inter-picture prediction and directly output, directly output the decoded picture without caching the decoded picture in the DPB.
9. The apparatus according to claim 1, configured to: Assign a frame index to each reference picture classified as a long-term picture in the DPB, and If the frame index assigned to a predetermined reference picture classified as a long-term picture in the DPB is referenced in the current AU, use the predetermined reference picture as the referenced reference picture in the DPB.
10. The apparatus according to claim 9, configured to perform one or more of the following: If at least one command in the current AU is a first command, reclassify a reference picture classified as a short-term reference picture in the DPB as a picture not for reference, If at least one command in the current AU is a second command, reclassify the reference pictures in the DPB that are classified as long-term reference pictures as pictures not for reference. If at least one command in the current AU is a third command, reclassify the reference pictures in the DPB that are classified as short-term pictures as long-term reference pictures and assign a frame index to the reclassified reference pictures. If at least one command in the current AU is a fourth command, set an upper limit of the frame index according to the fourth command, and reclassify all the reference pictures in the DPB that are classified as long-term pictures and to which a frame index exceeding the upper limit of the frame index has been assigned as pictures not for reference. If at least one command in the current AU is a fifth command, classify the current picture as a long-term picture and assign a frame index to the current picture.
11. The apparatus according to claim 1, configured to: Remove from the DPB any reference pictures that are classified as pictures not for reference and are no longer output.
12. The apparatus according to claim 1, configured to: Read an entropy coding mode indicator from the data stream, and If the entropy coding mode indicator indicates a context-adaptive variable length coding mode, decode prediction residual data from the current AU using the context-adaptive variable length coding mode, and if the entropy coding mode indicator indicates a context-adaptive binary arithmetic coding mode, decode prediction residual data from the current AU using the context-adaptive binary arithmetic coding mode.
13. The apparatus according to claim 1, configured to: Derive quarter-pixel values in a reference picture being referred to based on motion vectors in the current AU and using a 6-tap FIR filter to derive half-pixel values and average adjacent half-pixel values.
14. The apparatus according to claim 1, configured to: Derive information about the time raster from the data stream by the time difference between the removal of the first access unit and the removal of each subsequent access unit in the subsequent access units.
15. The apparatus according to claim 1, configured to: Interpolate between the CPB size indicated by the first CPB parameter and the CPB size indicated by the second CPB parameter at a selected bit rate to obtain an interpolated CPB size, thereby determining the minimum CPB size of the coded picture buffer.
16. The apparatus according to claim 1, wherein, the selected bit rate is between the bit rate indicated by the first CPB parameter and the bit rate indicated by the second CPB parameter.
17. The apparatus according to claim 1, configured to operate in units of cache cycles, where the first access unit in the coding order is the first access unit of the current cache cycle.
18. An apparatus for encoding video into a data stream, wherein, the data stream should be decoded by feeding the data stream to a decoder including a coded picture buffer CPB, and the apparatus is configured to: Encode pictures of the video into a data stream in coding order as an access unit (AU) sequence, determine a first CPB parameter associated with a first operation point and a second CPB parameter associated with a second operation point, each of the first CPB parameter and the second CPB parameter indicating a CPB size, a predetermined time offset, a predetermined time removal delay, and a predetermined bit rate, wherein the first CPB parameter is different from the second CPB parameter at least in terms of the predetermined bit rate, and perform the determination such that interpolation between the predetermined time offset of the first CPB parameter and the predetermined time offset of the second CPB parameter at each of a plurality of selected bit rates results in an interpolated time offset and an interpolated time removal delay, and thereby feed the data stream to the decoder via the CPB by: sequentially feed the AU sequence into the CPB using a respective selected bit rate, wherein feeding of access units not having reached the virtual available time for removing the raster according to the time frame is paused until the virtual available time is reached, wherein the time frame removal raster advances the interpolated time removal delay for a first access unit in the coding order, and advances the interpolated time removal delay plus the interpolated time offset for subsequent access units in the coding order; and remove AUs from the CPB on a per-AU basis using the time raster, without causing any underflow and any overflow, and encode the CPB parameter into the data stream, wherein the apparatus is configured to, when encoding the AU, encode a current picture into the current AU using inter-picture prediction based on a referenced reference picture stored in a decoded picture buffer (DPB), and insert a decoded version of the current picture in the DPB into the DPB, assign to each reference picture stored in the DPB one of classifications as a short-term reference picture, a long-term reference picture, and a picture not for reference, write DPB mode information into the current AU, if the DPB mode information indicates a first mode, remove one or more reference pictures classified as short-term pictures from the DPB according to a first-in first-out (FIFO) policy, if the DPB mode information indicates a second mode, write memory management control information including at least one command into the current AU, the command indicating a change in the classification assigned to at least one of the reference pictures stored in the DPB, wherein the classification of the reference pictures in the DPB is used to manage removal of the reference pictures from the DPB.
19. The apparatus according to claim 18, wherein, parameterize the interpolation using interpolation parameters, and the apparatus is configured to encode the interpolation parameters into the data stream.
20. The apparatus according to claim 18, wherein, The interpolation will be performed using a weighted sum of a predetermined time offset indicated by the first CPB parameter weighted by a first weight and a predetermined time offset indicated by the second CPB parameter weighted by a second weight.
21. The apparatus according to claim 20, wherein, the first weight and the second weight are determined based on the selected bit rate, a predetermined bit rate indicated by the first CPB parameter, and a predetermined bit rate indicated by the second CPB parameter.
22. The apparatus according to claim 20, wherein, a linear interpolation weight determined by dividing a difference between the selected bit rate and a predetermined bit rate indicated by the first CPB parameter by a difference between the predetermined bit rate indicated by the first CPB parameter and a predetermined bit rate indicated by the second CPB parameter is used to determine the first weight and the second weight.
23. The apparatus according to claim 22, wherein, the first weight is determined such that the first weight is the linear interpolation weight or a product where one of the factors is the linear interpolation weight, and the second weight is determined such that the second weight is a difference between the linear interpolation weight and 1 or a product where one of the factors is a difference between the linear interpolation weight and 1.
24. The apparatus according to claim 22, wherein, the first weight is determined such that the first weight is a product: a first factor of the product is the linear interpolation weight, and a second factor of the product is a predetermined bit rate indicated by the first CPB parameter divided by the selected bit rate, and the second weight is determined such that the second weight is a product: one factor of the product is a difference between the linear interpolation weight and 1, and a second factor of the product is a predetermined bit rate indicated by the second CPB parameter divided by the selected bit rate.
25. The apparatus according to claim 18, configured to: write an indication of whether the decoded picture is not used for inter - picture prediction into the current AU; wherein, if the decoded picture is not indicated as not used for inter - picture prediction or not directly output, the decoded picture will be inserted into the DPB, and if the decoded picture is indicated as not used for inter - picture prediction and directly output, the decoded picture will be directly output without caching the decoded picture in the DPB.
26. The apparatus according to claim 18, wherein, a frame index will be assigned to each reference picture classified as a long - term picture in the DPB, and if the frame index of a predetermined reference picture classified as a long - term picture in the DPB is referenced in the current AU, the predetermined reference picture will be used as the referenced reference picture in the DPB.
27. The apparatus according to claim 26, wherein one or more of the following are performed: if at least one command in the current AU is a first command, a reference picture classified as a short - term reference picture in the DPB will be re - classified as a picture not for reference, If at least one command in the current AU is a second command, reference pictures in the DPB that are classified as long-term reference pictures will be re-classified as pictures not for reference. If at least one command in the current AU is a third command, reference pictures in the DPB that are classified as short-term pictures will be re-classified as long-term reference pictures, and a frame index will be assigned to the re-classified reference pictures. If at least one command in the current AU is a fourth command, a frame index upper limit is set according to the fourth command, and all reference pictures in the DPB that are classified as long-term pictures and to which a frame index exceeding the frame index upper limit has been assigned will be re-classified as pictures not for reference. If at least one command in the current AU is a fifth command, the current picture will be classified as a long-term picture, and a frame index will be assigned to the current picture.
28. The apparatus according to claim 18, wherein, Any reference picture that is classified as a picture not for reference and is no longer output will be removed from the DPB.
29. A method for video decoding by using an encoded picture buffer CPB and a decoded picture buffer DPB, the method comprising: receiving a data stream having pictures of a video encoded therein in coded order as an access unit AU sequence, sequentially feeding the access unit sequence into the CPB at a selected bitrate, where feeding of an access unit that has not reached the virtual available time for removing the raster according to the time frame is paused until the virtual available time is reached, where the time frame removes the delay in advance by a selected time for the first access unit in the coded order, and the delay removed in advance by a selected time for subsequent access units in the coded order is the sum of the selected time offset; removing AUs from the CPB on a per-AU basis using the time raster, extracting a first CPB parameter related to a first operating point and a second CPB parameter related to a second operating point from the data stream, each of the first CPB parameter and the second CPB parameter indicating a CPB size, a predetermined time offset, a predetermined time removal delay, and a predetermined bitrate, wherein the first CPB parameter is different from the second CPB parameter at least in terms of the predetermined bitrate, determining the selected time offset by interpolating between the predetermined time offsets indicated by the first CPB parameter and the second CPB parameter at the selected bitrate, and determining the selected time removal delay by interpolating between the predetermined time removal delays indicated by the first CPB parameter and the second CPB parameter at the selected bitrate, decoding the current AU removed from the CPB using inter-picture prediction based on the reference pictures stored in the DPB to obtain a decoded picture, and inserting the decoded picture into the DPB. Assign a classification of one of a short-term reference picture, a long-term reference picture, and a picture not for reference to each reference picture stored in the DPB. Read DPB mode information from the current AU. If the DPB mode information indicates a first mode, remove one or more reference pictures classified as short-term pictures from the DPB according to a first-in first-out (FIFO) policy. If the DPB mode information indicates a second mode, read memory management control information including at least one command in the current AU, and execute the at least one command to change the classification assigned to at least one of the reference pictures stored in the DPB, and use the classification of the reference pictures in the DPB to manage removal of the reference pictures from the DPB.
30. A method for encoding video into a data stream wherein the data stream is to be decoded by feeding the data stream to a decoder including a coded picture buffer (CPB), the method comprising: encoding pictures of the video into the data stream in coded order as a sequence of access units (AUs); determining a first CPB parameter related to a first operating point and a second CPB parameter related to a second operating point, each of the first CPB parameter and the second CPB parameter indicating a CPB size, a predetermined time offset, a predetermined time removal delay, and a predetermined bit rate, wherein the first CPB parameter is different from the second CPB parameter at least in terms of the predetermined bit rate, and performing the determination such that interpolation between the predetermined time offset of the first CPB parameter and the predetermined time offset of the second CPB parameter at each of a plurality of selected bit rates results in an interpolated time offset and an interpolated time removal delay, and feeding the data stream to the decoder via the CPB by: sequentially feeding the AU sequence to the CPB using a respective selected bit rate, wherein feeding of an access unit not yet reaching a virtual available time according to a time frame removal raster is paused until the virtual available time is reached, wherein the time frame removal raster advances the interpolated time removal delay for a first access unit in the coded order, and advances the interpolated time removal delay plus the interpolated time offset for subsequent access units in the coded order; and removing AUs from the CPB on a per-AU basis using the time raster without causing any underflow and any overflow, and encoding the CPB parameters into the data stream wherein the method comprises, when encoding the AU, encoding a current picture into the current AU using inter-picture prediction according to reference pictures to be referenced and stored in a decoded picture buffer (DPB) at the decoder side, wherein the decoder is to insert a decoded version of the current picture into the DPB, and assign a classification of one of a short-term reference picture, a long-term reference picture, and a picture not for reference to each reference picture stored in the DPB. Write the DPB mode information into the current AU, where if the DPB mode information indicates the first mode, the decoder removes one or more reference pictures classified as short-term pictures from the DPB according to the first-in-first-out (FIFO) strategy, if the DPB mode information indicates the second mode, write memory management control information including at least one command into the current AU, the command indicating a change in the classification assigned to at least one of the reference pictures stored in the DPB, where the classification of the reference pictures in the DPB is used to manage the removal of the reference pictures from the DPB.
31. A data stream into which video is encoded and includes first CPB parameters and second CPB parameters such that the method according to claim 29 does not cause CPB overflow and underflow.