Video Data Encoding and Decoding Using a Symbolic Picture Buffer
By adjusting the maximum encoding picture buffer size and MinCR scaling factor, the system ensures efficient storage and decoding of high-resolution video data, addressing the challenge of insufficient buffer sizes at minimum compression ratios.
Patent Information
- Application Number
- JP2022558087
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Priority Date
- 2020-04-03
- Filing Date
- 2021-03-08
- Publication Date
- 2025-06-18
- Estimated Expiration
- 2041-03-08
AI Technical Summary
Existing video data compression and decompression systems face challenges in efficiently storing and decoding high-resolution video data, particularly at minimum compression ratios, due to insufficient encoding picture buffer sizes.
The proposed solution involves adjusting the maximum encoding picture buffer size and the minimum compression ratio (MinCR) scaling factor to ensure that the buffer can store a full picture even at minimum compression, thereby preventing buffer overflow and underrun issues.
This adjustment allows for efficient storage and decoding of high-resolution video data, ensuring that the encoding picture buffer size is sufficient to handle the maximum luminance picture size even at minimum compression ratios, thus preventing data loss and improving system reliability.
Smart Images

Figure 0007694578000007 
Figure 0007694578000008 
Figure 0007694578000009
Abstract
Description
Technical Field
[0001] The present disclosure relates to the encoding and decoding of video data.
Background Art
[0002] The description of "Background Art" provided herein is for the purpose of generally presenting the context of the present disclosure. The research of the present inventors is not admitted, explicitly or implicitly, as prior art to the present invention to the extent that it is described in this section of the background art and cannot otherwise be considered prior art at the time of filing.
[0003] There are several systems, such as video / image data encoding and decoding systems, that involve converting video data into a frequency domain representation, quantizing the frequency domain coefficients, and then applying some form of entropy encoding to the quantized coefficients. This allows the video data to be compressed. To restore the reconstructed original video data, corresponding decoding or decompression techniques are applied.
Prior Art Documents
Non-Patent Documents
[0004]
Non-Patent Document 1
Summary of the Invention
Problems to be Solved by the Invention
[0005] The present disclosure addresses or mitigates the problems arising from this process.
Means for Solving the Problems
[0006] Each aspect and feature of the present disclosure is defined in the appended claims.
[0007] It should be understood that both the foregoing general description and the following detailed description are exemplary of the technology and not restrictive thereof.
[0008] A more complete understanding of the present disclosure and many of its attendant advantages will be readily obtained as the same becomes better understood by reference to the following detailed description, when considered in connection with the accompanying drawings.
Brief Description of the Drawings
[0009]
Figure 1
Figure 2
Figure 3
Figure 4
Figure 5
Figure 6
Figure 7
Figure 8
Figure 9
Figure 10
Figure 11
Figure 12
Figure 13
Figure 14
Figure 15
Figure 16
Figure 17
Figure 18
Figure 19
Figure 20
Figure 21
Figure 22
Figure 23
Figure 24
Figure 25
BEST MODE FOR CARRYING OUT THE INVENTION
[0010] Referring now to the drawings, FIGS. 1 to 4 are provided to give a schematic view of an apparatus or system that utilizes a compression and / or decompression device to be described below in connection with various embodiments of the present technology.
[0011] All of the data compression and / or decompression devices described below may be implemented in hardware, in software executed on a general-purpose data processing device such as a general-purpose computer, in programmable hardware such as an application specific integrated circuit (ASIC) or a field programmable gate array (FPGA), or as a combination thereof. When embodiments are implemented by software and / or firmware, such software and / or firmware, and the non-transitory data storage medium in which such software and / or firmware is stored or otherwise provided, will be understood to be considered embodiments of the present technology.
[0012] FIG. 1 schematically shows an audio / video data transmission / reception system using video data compression and decompression. In this example, the data values to be encoded or decoded represent image data.
[0013] The input audio / video signal 10 is supplied to a video data compression device 20, which compresses at least the video component of the audio / video signal 10 for transmission along a transmission path 30 such as a cable, an optical fiber, a wireless link, etc. The compressed signal is processed by a decompression device 40 to provide an output audio / video signal 50. In the case of the return path, a compression device 60 compresses the audio / video signal for transmission to a decompression device 70 along the transmission path 30.
[0014] Accordingly, the compression device 20 and the decompression device 70 can form one node of the transmission link. The decompression device 40 and the compression device 60 can form the other node of the transmission link. Of course, if the transmission link is unidirectional, only one of these nodes will require a compression device and the other node will only require a decompression device.
[0015] Figure 2 schematically shows a video display system using video data decompression. In particular, the compressed audio / video signal 100 is processed by the decompression device 110 to provide a decompressed signal that can be displayed on the display 120. The decompression device 110 can be implemented as an integral part of the display 120, for example, provided within the same housing as the display device. Alternatively, the decompression device 110 may be provided as a so-called set-top box (STB), for example. Note that the expression "set-top" does not imply that the box is placed in any particular orientation or position with respect to the display 120, but is simply a term used in the art to indicate a device that can be connected to the display as a peripheral device.
[0016] Figure 3 schematically shows an audio / video storage system using video data compression and decompression. The input audio / video signal 130 is supplied to the compression device 140, which generates a compressed signal for storage in a storage device 150 such as a magnetic disk device, an optical disk device, a magnetic tape device, a solid-state storage device such as a semiconductor memory, or other storage devices. For playback, the compressed data is read from the storage device 150 and passed to the decompression device 160 for decompression to provide the output audio / video signal 170.
[0017] It will be understood that compressed or encoded signals, and storage media such as computer-readable non-transitory storage media storing such signals, are considered embodiments of the present technology.
[0018] Figure 4 schematically shows a video camera using video data compression. In Figure 4, an imaging device 180 such as a charge-coupled device (CCD) image sensor and associated control and readout electronics generates a video signal that is passed to the compression device 190. A microphone (or microphones) 200 generates an audio signal that is passed to the compression device 190. The compression device 190 generates a compressed audio / video signal 210 that is stored and / or transmitted (generally shown as stage 220).
[0019] The techniques described below relate primarily to video data compression and decompression. It will be appreciated that, in conjunction with the video data compression techniques described, many existing techniques for audio data compression may be used to generate a compressed audio / video signal. Accordingly, no separate explanation of audio data compression will be given. It will also be appreciated that the data rate associated with video data, particularly broadcast quality video data, is generally very much higher than that associated with audio data (whether compressed or uncompressed). Accordingly, it will be appreciated that uncompressed audio data may form part of a compressed audio / video signal, along with compressed video data. This embodiment (shown in FIGS. 1 to 4) is related to audio / video data, but it will be further appreciated that the techniques described below are applicable in a system for simply processing (i.e., compressing, decompressing, storing, displaying, and / or transmitting) video data. That is, the embodiments can be applied to video data compression without processing any associated audio data at all.
[0020] Accordingly, FIG. 4 provides an example of a video capture device comprising an image sensor and an encoding device of the type described below. Accordingly, FIG. 2 provides an example of a decoding device of the type described below and a display on which a decoded image is output.
[0021] The combination of FIGS. 2 and 4 can provide a video capture device comprising an image sensor 180, an encoding device 190, a decoding device 110, and a display 120 on which a decoded image is output.
[0022] FIG. 5 and FIG. 6 schematically show a storage medium that stores compressed data generated by, for example, devices 20 and 60, or compressed data input to devices 110, a storage medium, or stages 150 and 220. FIG. 5 schematically shows a disk storage medium such as a magnetic disk or an optical disk, and FIG. 6 schematically shows a solid-state storage medium such as a flash memory. Note that FIGS. 5 and 6 can also provide an example of a computer-readable non-transitory storage medium that stores computer software that causes a computer to execute one or more of the methods described below when executed by the computer.
[0023] Therefore, the above configuration provides an example of a video storage device, a capture device, a transmission device, or a reception device that implements any of the present techniques.
[0024] FIG. 7 is a schematic diagram showing an overview of a video / image data compression (encoding) and decompression (decoding) device for encoding and / or decoding video / image data representing one or more images.
[0025] The controller 343 controls the operation of the entire device. In particular, when referring to the compression mode, the controller 343 controls the trial encoding process by acting as a selector that selects various operation modes such as block size and shape, and whether to encode video data reversibly or in another way. The controller is considered to form (in some cases) part of an image encoder or an image decoder. Successive images of the input video signal 300 are supplied to the adder 310 and the image predictor 320. The image predictor 320 will be described in more detail below with reference to FIG. 8. The features of the device in FIG. 7 can be utilized by adding the intra-image predictor of FIG. 8 to the image encoder / decoder (in some cases). However, this does not mean that the image encoder / decoder necessarily requires all the features of FIG. 7.
[0026] The adder 310 actually receives the input video signal 300 at the “+” input, receives the output of the image predictor 320 at the “−” input, and performs a subtraction (negative addition) operation that subtracts the predicted image from the input image. As a result, a so-called residual image signal 330 representing the difference between the actual image and the predicted image is generated.
[0027] One reason for generating the residual image signal is as follows. The data encoding technique to be described, i.e., the technique that will be applied to the residual image signal, tends to work more efficiently when there is less “energy” in the image to be encoded. Here, the term “efficiently” refers to the generation of a small amount of encoded data, and at a specific image quality level, it is desirable (and considered “efficient”) to generate as little data as possible. The reference to “energy” in the residual image relates to the amount of information contained in the residual image. If the predicted image is identical to the actual image, the difference between the two images (i.e., the residual image) contains zero information (zero energy) and is very easy to encode into a small amount of encoded data. Generally, when the prediction process can be reasonably and well-performed such that the content of the predicted image is similar to the content of the image to be encoded, it is expected that the residual image data will contain less information (less energy) than the input image, making it easier to encode into a small amount of encoded data.
[0028] Therefore, encoding (using the adder 310) includes predicting the image region of the image to be encoded and generating a residual image region that depends on the difference between the predicted image region and the corresponding region of the image to be encoded. In the context of the techniques described below, an ordered array of data values includes the data values representing the residual image region. Decoding includes predicting the image region of the image to be decoded, generating a residual image region indicating the difference between the predicted image region and the corresponding region of the image to be decoded, an ordered array of data values including the data values representing the residual image region, and combining the predicted image region and the residual image region.
[0029] Here, the remainder of the apparatus that acts as an encoder (for encoding a residual image or a difference image) will be described.
[0030] The residual image data 330 is supplied to a conversion unit / circuit 340 that generates a discrete cosine transform (DCT) representation of a block or region of the residual image data. The DCT technique itself is well-known and will not be described in detail here. Also, note that the use of DCT is merely an illustration of one exemplary configuration. Examples of other transforms that can be used include, for example, discrete sine transform (DST). The transform can also include a sequence or cascade of individual transforms, such as a configuration where one transform is followed by another (whether directly or not). The choice of transform may be explicitly determined and / or may depend on side information used to configure the encoder and decoder. In other examples, a so-called "transform skip" mode where no transform is applied can be selectively used.
[0031] Accordingly, in various embodiments, the encoding and / or decoding method includes predicting an image region of the image to be encoded and generating a residual image region that depends on the difference between the predicted image region and the corresponding region of the image to be encoded, and the ordered array of data values (described below) includes the data values of the representation of the residual image region.
[0032] The output of the conversion unit 340, that is, (in one example) the set of DCT coefficients for each transformed block of the image data, is supplied to a quantizer 350. In the field of video data compression, various quantization techniques are known, ranging from simple multiplication by a quantization scaling factor to the application of a complex look-up table under the control of quantization parameters. There are two general objectives. First, the quantization process reduces the number of possible values of the transformed data. Second, the quantization process can increase the likelihood that the values of the transformed data are zero. Both of these enable the entropy encoding process described below to function more efficiently in generating a small amount of compressed video data.
[0033] The data scanning process is applied by the scanning unit 360. The purpose of the scanning process is to rearrange the quantized conversion data so as to gather as many non-zero quantized conversion coefficients as possible, and thus, of course, to gather as many zero-value coefficients as possible. These features can enable the efficient application of so-called run-length encoding or similar techniques. Thus, the scanning process includes selecting coefficients from the quantized conversion data, particularly from a block of coefficients corresponding to a block of transformed and quantized image data, according to a "scanning order" such that (a) all coefficients are selected once as part of the scan, and (b) the scan provides the desired rearrangement easily. An example of a scanning order that gives useful results is the so-called orthogonal diagonal scanning order.
[0034] The scanning order may differ between conversion skip blocks and conversion blocks (blocks that have undergone at least one spatial frequency conversion).
[0035] Next, the scanned coefficients are passed to an entropy encoder (EE) 370. Again, various types of entropy coding may be used. As two examples, a variant of a so-called context adaptive binary arithmetic coding (CABAC) system and a variant of a so-called context adaptive variable-length coding (CAVLC) system can be mentioned. Generally, CABAC is considered to be more efficient, and some studies have shown that it reduces the amount of encoded output data for the same image quality by 10% to 20% compared to CAVLC. However, CAVLC is considered to have a much lower level of complexity (regarding its implementation) than CABAC. Note that the scanning process and the entropy coding process are shown as separate processes, but in reality, they may be combined or treated together. That is, the data can be read out to the entropy encoder in the scanning order. The same applies to each of the reverse processes described below.
[0036] The output of the entropy encoder 370 provides a compressed output video signal 380, together with additional data (described above and / or below) that defines, for example, how the predictor 320 generated the predicted image, whether the compressed data was transformed or the transformation was skipped.
[0037] However, since the operation of the predictor 320 itself depends on the data obtained by decompressing the compressed output data, a feedback path 390 is also provided.
[0038] The reason for this feature is as follows. At an appropriate stage of the decompression process (described later), data obtained by decompressing the residual data is generated. This decompressed residual data must be added to the predicted image in order to generate the output image (since the original residual data is the difference between the input image and the predicted image). For this process to be the same in the relationship between the compression side and the decompression side, the predicted image generated by the predictor 320 should be the same during the compression process and the decompression process. Of course, during decompression, the device cannot access the original input image and can only access the decompressed image. Therefore, during compression, the predictor 320 uses the decompressed image of the compressed image as the basis for its prediction (at least for inter-image coding).
[0039] The entropy encoding process performed by the entropy encoder 370 is considered to be "reversible" (at least in some examples), that is, it can be run in reverse to reach exactly the same data that was initially supplied to the entropy encoder 370. Therefore, in such examples, a return path may be implemented before the entropy encoding stage. In fact, since the scanning process performed by the scanning unit 360 is also considered to be reversible, in this embodiment, the return path 390 extends from the output of the quantizer 350 to the input of the complementary inverse quantizer 420. If a loss or potential loss is introduced by a stage, that stage (and its inverse) may be included in the feedback loop formed by the return path. For example, the entropy encoding stage can be made at least in principle irreversible by techniques such as bits being encoded within parity information. In such a case, the entropy encoding and decoding should form part of the feedback loop.
[0040] Generally, the entropy decoder 410, the inverse scanning unit 400, the inverse quantizer 420, and the inverse transformation unit / circuit 430 provide functions that are the inverse of those of the entropy encoder 370, the scanning unit 360, the quantizer 350, and the transformation unit 340, respectively. Here, the description continues through the compression process. The process for decompressing the input compressed video signal will be described separately below.
[0041] In the compression process, the scanned coefficients are passed from the quantizer 350 to the inverse quantizer 420 by the return path 390, and the inverse quantizer 420 performs the reverse operation of the operation of the scanning unit 360. The inverse quantization and inverse transform processes are performed by the inverse quantizer 420 and the inverse transform unit / circuit 430, generating the decompressed residual image signal 440.
[0042] The image signal 440 is added to the output of the predictor 320 in the adder 450 to generate the reconstructed output image 460 (however, this may be subject to so-called loop filtering and / or other filtering before being output. See below). Thereby, as will be described below, one input to the image predictor 320 is formed.
[0043] Next, consider the decoding process applied to decompress the received compressed video signal 470. The signal is supplied to the entropy decoder 410, from which it is supplied to the inverse scanning unit 400, the inverse quantizer 420, and the inverse transform unit 430 in this order, and then added to the output of the image predictor 320 by the adder 450. Therefore, on the decoder side, the decoder reconstructs the image in the residual image version and then applies this (block by block on the block basis) to the predicted version of the image to decode each block. Briefly speaking, the output 460 of the adder 450 forms the output decompressed video signal 480 (which is subject to the filtering process described below). In practice, further filtering may be optionally applied (for example, by the loop filter 565 shown in FIG. 8 but omitted from FIG. 7 for clarity) before the signal is output.
[0044] The apparatuses of FIGS. 7 and 8 can act as a compression (encoding) apparatus or a decompression (decoding) apparatus. The functions of these two types of apparatuses are substantially overlapping. The scanning unit 360 and the entropy encoder 370 are not used in the decompression mode, and the operations of the predictor 320 and other units (described in detail below) do not generate such information itself, but follow the mode and parameter information included in the received compressed bit stream.
[0045] FIG. 8 schematically shows the generation of a predicted image and, in particular, the operation of the image predictor 320.
[0046] There are two basic modes of prediction performed by the image predictor 320, namely, so-called intra-image prediction and so-called inter-image prediction or motion compensation (MC) prediction. On the encoder side, each includes detecting a prediction direction for the current block to be predicted and generating a predicted block of samples according to other samples (in the same image (intra-image) or another image (inter-image)). By the adder 310 or 450, the difference between the predicted block and the actual block is encoded or applied to encode or decode the blocks respectively.
[0047] (In the decoder or on the inverse decoding side of the encoder, the detection of the prediction direction may respond to data associated with the data encoded by the encoder indicating which direction was used in the encoder. Alternatively, the detection may respond to the same factors as those determined in the encoder.)
[0048] Intra-image prediction predicts the content of a block or region of an image based on data within the same image. This corresponds to so-called I-frame encoding in other video compression techniques. However, in contrast to I-frame encoding that encodes the entire image by intra encoding, in this embodiment, the selection between intra encoding and inter encoding can be made on a block-by-block basis, while in other embodiments, the selection is still made on an image-by-image basis.
[0049] Motion compensation prediction is an example of inter - picture prediction and utilizes motion information that attempts to define a source in other adjacent or nearby pictures of the picture details being encoded in the current picture. Thus, in an ideal example, the content of a block of image data within the predicted picture can be encoded very simply as a reference (motion vector) that points to a corresponding block at the same or slightly different location within an adjacent picture.
[0050] A technique known as "block copy" prediction is, in some respects, a hybrid of those two predictions. This is because a vector is used to indicate a block of samples located at a displaced position from the currently predicted block within the same picture, and the block is copied to form the currently predicted block.
[0051] Returning to FIG. 8, two image prediction configurations (corresponding to intra - picture prediction and inter - picture prediction) are shown, and the results of those predictions are selected by multiplexer 500 under the control of a mode signal 510 (e.g., from controller 343) to provide blocks of the predicted picture for supply to adders 310 and 450. The selection is made according to which selection gives the lowest "energy" (which can be regarded as the information content requiring encoding as described above), and is signaled to the decoder within the encoded output data stream. In this context, the image energy can be detected, for example, by tentatively subtracting regions of the two versions of the predicted picture from the input picture, squaring each pixel value of the difference picture, summing the squared values, and identifying which of the two versions gives a lower mean squared value of the difference picture associated with that image region. In other examples, tentative encoding can be performed for each selection or possible selection, and then the selection is made according to the cost of each possible selection with respect to one or both of the number of bits required for encoding and the distortion of the picture.
[0052] In an intra-coding system, the actual prediction is performed based on an image block received as part of signal 460 (the signal to be filtered by loop filtering, see below), i.e., the prediction is performed based on the coded and decoded image block so that the same prediction can be executed in the decoding device. However, data can be derived from the input video signal 300 by the intra-mode selector 520 to control the operation of the intra-predictor 530.
[0053] In the case of inter-prediction, the motion compensation (MC) predictor 540 uses motion information such as motion vectors derived by the motion estimator 550 from the input video signal 300. These motion vectors are applied to the image obtained by processing the reconstructed image 460 by the motion compensation predictor 540 to generate an inter-prediction block.
[0054] Therefore, each of the predictors 530 and 540 (operating with the estimator 550) acts as a detector for detecting the prediction direction for the current block to be predicted and as a generator for generating a predicted block of samples (forming part of the prediction passed to the adders 310 and 450) according to other samples defined by the prediction direction.
[0055] Next, the processing applied to signal 460 will be described.
[0056] First, the signal may be filtered by a so-called loop filter 565. Various types of loop filters can be used. One technique involves applying a "deblocking" filter to remove or at least reduce the impact of the block-based processing performed by the conversion unit 340 and subsequent operations. Further techniques may include applying a so-called sample adaptive offset (SAO) filter. Generally, in a sample adaptive offset filter, the filter parameter data (derived at the encoder and communicated to the decoder) defines one or more offset amounts that are selectively combined with a given intermediate video sample (a sample of signal 460) by the sample adaptive offset filter in accordance with (i) the value of a given intermediate video sample or (ii) the values of one or more intermediate video samples having a predetermined spatial relationship to the given intermediate video sample.
[0057] Also, an adaptive loop filter is optionally applied using coefficients derived by processing the reconstructed signal 460 and the input video signal 300. An adaptive loop filter is a type of filter that applies adaptive filter coefficients to the data being filtered using known techniques. That is, the filter coefficients can vary depending on various factors. The data defining which filter coefficients to use is included as part of the encoded output data stream.
[0058] The techniques described below relate to the handling of parameter data regarding the operation of the filter. The actual filtering operation (such as SAO filtering) may otherwise use known techniques.
[0059] The filtered output from the loop filter section 565 actually forms the output video signal 480 when the apparatus is operating as a decompression device. This output is also buffered in one or more image or frame memory sections 570, and the storage of consecutive images is a requirement for motion compensation prediction processing, particularly for the generation of motion vectors. To facilitate meeting the storage requirements, the stored images in the image memory section 570 may be held in a compressed state and then decompressed for use in generating motion vectors. For this particular purpose, any known compression / decompression system may be used. The stored images may be passed to an interpolation filter 580 that generates a higher-resolution image from the stored images. In this example, intermediate samples (sub-samples) are generated, and the resolution of the interpolated image output by the interpolation filter 580 is four times (in each dimension) that of the image stored in the image memory section 570 for the 4:2:0 luminance channel and eight times (in each dimension) that of the image stored in the image memory section 570 for the 4:2:0 chrominance channel. The interpolated image is passed as an input to the motion estimator 550 and also to the motion compensation predictor 540.
[0060] Next, a method for dividing an image for compression processing will be described. At a basic level, the image to be compressed is considered as an array of sample blocks or regions. Dividing the image into such blocks or regions can be performed by a decision tree as described in "Series H: Audiovisual and Multimedia Systems Infrastructure of audiovisual services - Coding of moving video High efficiency video coding Recommendation ITU-T H.265 12 / 2016" and "High Efficiency Video Coding (HEVC) Algorithms and Architectures, Chapter 3, Editors: Madhukar Budagavi, Gary J. Sullivan, Vivienne Sze; ISBN 978-3-319-06894-7; 2014" (each of which is hereby incorporated by reference in its entirety). Further background information is provided in Non-Patent Document 1. This Non-Patent Document 1 is also hereby incorporated by reference in its entirety.
[0061] In some examples, the resulting blocks or regions can generally have a size and, in some cases, a shape that follows the arrangement of image features within the image, as determined by a decision tree. This in itself can enable improved encoding efficiency. This is because such a configuration tends to group together samples that represent or follow similar image features. In some examples, square blocks or regions of different sizes (e.g., up to, for example, 64×64 or larger blocks such as 4×4 samples) are available for selection. In other exemplary configurations, blocks or regions of different shapes, such as rectangular blocks (e.g., oriented vertically or horizontally), can be used. Other non-square and non-rectangular blocks are also envisioned. As a result of dividing the image into such blocks or regions, each sample of the image (at least in this example) is assigned to only one such block or region.
[0062] (Coded Picture Buffer) Video decoding standards (and, indirectly, the corresponding encoders) may be defined to include a so-called Coded Picture Buffer (CPB). On the decoder side, the incoming video stream is timely stored in the CPB and read from the CPB for decoding. The standard assumes that the entire picture can be read from the CPB in a single (theoretically instantaneous) operation. On the encoder side, the encoded data may similarly be stored in the CPB (theoretically at least as a single instantaneous operation for writing the entire coded picture) and then output to the output encoded video stream. However, it is on the decoder side where the CPB is defined.
[0063] The CPB itself can be defined in various ways, including the data rate at which encoded data enters the CPB, the size of the CPB itself, and any latency potentially applied to the removal of data from the CPB (which then defines the time required to fill the CPB so that, as described above, an entire picture can be removed). These parameters are important to avoid overfilling (space shortage) of the CPB and underrunning (data shortage for providing to the next stage of processing) of the CPB. Examples of CPB usage are described below.
[0064] (Parameter set and encoding level) When encoding video data by the techniques described above for subsequent decoding, it is appropriate for the encoding side of the processing to communicate some parameters of the encoding process to the ultimate decoding side of the processing. Considering that these encoding parameters are needed whenever decoding the encoded video data, it is useful to associate the parameters with the encoded video data stream itself, for example, by embedding the parameters as a so-called parameter set (which can be "out-of-band" transmitted by a separate transmission channel, although not necessarily exclusive) into the encoded video data stream itself.
[0065] Parameter sets may be represented as a hierarchy of information, for example, as a video parameter set (VPS), a sequence parameter set (SPS), and a picture parameter set (PPS). The PPS occurs once per picture and contains information related to all the coded slices within that picture. The SPS occurs less frequently (once per sequence of pictures), and the VPS is expected to occur even less frequently. More frequently occurring parameter sets (such as PPS) can be implemented as references to previously coded instances of that parameter set to avoid the cost of re - coding. Each coded picture slice refers to a single active PPS, SPS, and VPS to provide the information used when decoding that slice. In particular, each slice header may include a PPS identifier to refer to the PPS, which in turn refers to the SPS, which in turn refers to the VPS.
[0066] Among these parameter sets, the SPS contains exemplary information related to some of the following descriptions, namely, data that defines the so - called profile, tier, and coding level used.
[0067] The profile defines a set of decoding tools or functions used. Examples of profiles include the "Main Profile" related to 8 - bit 4:2:0 video, and the "Main 10 Profile" that enables 10 - bit resolution and other extensions with respect to the Main Profile.
[0068] The coding level imposes restrictions on matters such as the maximum sampling rate and picture size. The tier defines the maximum data rate.
[0069] In the proposal for Versatile Video Coding (VVC) by the Joint Video Experts Team (JVET), as defined in the referenced JVET - Q2001 - vE standard (at the filing date), various levels are defined from 1 to 6.2.
[0070] (Example) Next, the examples will be described with reference to the drawings.
[0071] FIG. 9 schematically shows the use of the above-described video parameter set and sequence parameter set. In particular, these form part of the hierarchy of the above-described parameter sets such that a plurality of sequence parameter sets 900, 910, and 920 refer to the video parameter set 930, and then each of the sequences 902, 912, and 922 can refer to themselves. In an exemplary embodiment, level information applicable to each sequence is provided within the sequence parameter set.
[0072] However, it will be understood that in other embodiments, the level information may be provided in a different format or a different parameter set.
[0073] Similarly, the schematic diagram of FIG. 9 shows the sequence parameter set provided as part of the entire video data stream 940, but the sequence parameter set (or other data structure that transmits the level information) may alternatively be provided by a separate communication channel. In either case, the level information is associated with the video data stream 940.
[0074] (Operation Example - Decoder) FIG. 10 schematically shows an aspect of a decoding apparatus configured to receive an input (encoded) video data stream 1000 and generate and output a decoded video data stream 1010 using the decoder 1020 described above with reference to FIG. 7. For clarity of explanation, the control circuit / controller 343 of FIG. 7 is shown separately from the remainder of the decoder 1020.
[0075] Within the functionality of the controller / control circuit 343, there is a parameter set (PS) detector 1030 that detects various parameter sets including VPS, SPS, and PPS from appropriate fields of the input video data stream 1000. The parameter set detector 1030 derives information from the parameter sets including the levels as described above. This information is passed to the rest of the control circuit 343. Note that the parameter set detector 1030 can decode the levels or simply provide the encoded levels to the control circuit 343 for decoding.
[0076] The control circuit 343 also responds to at least one or more decoder parameters 1040 that define levels that, for example, the decoder 1020 can decode.
[0077] For a given or current input video data stream 1000, the control circuit 343 detects whether the decoder 1020 can decode the input video data stream and controls the decoder 1020 accordingly. The control circuit 343 can also provide various other operating parameters to the decoder 1020 in response to information obtained from the parameter sets detected by the parameter set detector 1030.
[0078] FIG. 10 also shows buffering the input video data stream 1000 using an encoded picture buffer (CPB) 1025 and then providing the input video data stream 1000 to the decoder 1020 for decoding. The control circuit 343 controls the parameters of the CPB 1025 according to the basic parameters of the decoder 1040 and the parameters derived from the parameter set decoder 1030.
[0079] Thus, FIG. 10 shows a video data decoder 1020, and an encoded picture buffer 1025 that buffers successive portions of the input video data stream and provides a portion to the video data decoder for decoding, The encoded picture buffer has an encoded picture buffer size, and the video data decoder responds to parameter data associated with the input video data stream, the parameter data indicating an encoding level selected from a plurality of encoding levels for a given input video data stream, each level defining at least a maximum luminance picture size, a minimum compression ratio, and a maximum value of the encoded picture buffer size An example of an encoding / decoding apparatus is provided.
[0080] (Operation Example - Encoder) Similarly, FIG. 11 schematically shows an aspect of an encoding apparatus including an encoder 1100 of the type described above with reference to FIG. 7, for example. The control circuit 343 of the encoder is shown separately for clarity of explanation. The encoder acts on an input video data stream 1110 to generate an output encoded video data stream 1120 under the control of the control circuit 343, and the control circuit 343 responds to encoding parameters 1130 including a definition of the encoding level to be applied.
[0081] The control circuit 343 also includes, or controls, a parameter set generator 1140 that generates a parameter set including, for example, VPS, SPS, and PPS to be included in the output encoded video data stream, and the SPS transmits the encoded level information as described above.
[0082] Similar to the above description, FIG. 11 also shows the use of an encoded picture buffer (CPB) 1125 for buffering an encoded video data stream 1120 as generated by the encoder 1100. The control circuit 343 controls the parameters of the CPB 1125 according to the encoding parameters 1130, and these parameters are communicated to the output encoded video data stream 1120 (and thus to the final decoding stage) by the parameter set generator 1140.
[0083] Thereby the video data encoder 1100, and An encoding picture buffer 1125 that buffers successive portions of the current output video data stream generated by the video data encoder, The encoding picture buffer has an encoding picture buffer size, and the video data encoder responds to parameter data associated with the output video data stream, the parameter data indicating an encoding level selected from a plurality of encoding levels for a given output video data stream, each level defining at least a maximum luminance picture size, a minimum compression ratio, and a maximum value of the encoding picture buffer size An example of an encoding / decoding apparatus is provided.
[0084] Embodiments of the present disclosure relate to potential changes in the parameters of the current VVC standard (applicable at the time of filing) referred to above, including a maximum encoding picture buffer (CPB) size and a MinCR scaling factor, to ensure that the buffer can always store a full picture when compressed at the minimum compression ratio (MinCR).
[0085] (Background of the Example) Table A.1 in the current standard specifies a maximum encoding picture buffer (CPB) size MaxCPB for each level. The encoded picture must also be entirely stored within the CPB and must be retrievable instantaneously, at least from Annex C1 of the standard.
[0086] Furthermore, Table A.2 of the current standard specifies MinCRBase, which, together with the MinCR scaling factor, enables the calculation of the minimum compression ratio (MinCR). MinCR is used to ensure that the picture is compressed by at least a specified amount.
[0087] However, MinCR is not necessarily a limiting factor. The exemplary embodiments described below propose either removing MinCR from the standard or adjusting the values in Annex A to make MinCR useful again.
[0088] For example, as shown by the analysis described below, an 8K image encoded at the minimum compression ratio of MinCR using the main 10 profile, level 6, main layer cannot be stored in the CPB, and in the case of the main 4:4:4 10 profile, this applies to some levels.
[0089] Figure 12 reproduces Table A.1 of the current standard regarding general layers and level limits. Here, the levels are indicated by the level numbers in the left column. The second column represents the maximum luma (or luminance) picture size. Following that are the definitions of the maximum CPB size (scaled by the scaling factor as shown in the table) for the main layer and high layer, the maximum number of picture slices, the maximum number of tile rows, and the maximum number of tile columns.
[0090] Figure 13 reproduces Table A.2 of the current standard, showing in detail for each level the maximum luma sampling rate, the maximum bitrates for the main layer and high layer, and the minimum compression ratios for the main layer and high layer.
[0091] In Figure 14, various scaling factors are reproduced from Table A.3.
[0092] The table in Figure 15 shows the maximum picture size, the maximum CPB size, the maximum encoded picture size at MinCR, and the minimum number of pictures in the CPB for the main 10 profile. Note that the second, fourth, and sixth columns are from Table A.1 or Table A.2 of the above standard, and the other fields are derived from CpbVcIFactor and MinCRScaleFactor (Figure 14).
[0093] Note that the column labeled "Exemplary maximum luma size" simply provides examples of luma picture configurations that conform to the number of samples defined by each maximum luma picture size and also conform to the aspect ratio constraints imposed by other parts of the current standard. This column is for illustrative purposes and does not form part of the standard.
[0094] By analyzing this information, the last column provides the minimum number of coded pictures that can be stored in the CPB, assuming the maximum CPB size defined by the standard and the minimum compression ratio defined by MinCRBase and MinCRScaleFactor.
[0095] Regarding Figure 15, it can be seen that the rows corresponding to level 6 are such that less than one coded picture can be stored in the CPB under these conditions.
[0096] Figure 16 shows the corresponding information for the Main 444 10 profile, and it can be seen that this problem occurs at levels 1, 3, 3.1, 4, 5, and 6.
[0097] Note that since the same problem does not occur at higher levels, only the main level is shown in Figures 15 and 16.
[0098] (Features of the exemplary embodiment) Since it is often not required, it is proposed to remove the MinCR constraint or adjust the values used to derive the CPB size so that MinCR is always appropriate again.
[0099] (Example 1 - Remove or ignore MinCR) In this example, the encoder and decoder control circuit 343 ignores the minimum compression ratio specification if the standard is such that the CPB of the maximum CPB size cannot hold the entire coded picture when considering all other current or general parameters.
[0100] Accordingly, an example is provided where the video data encoder is configured to generate the current output video data stream by applying a compression ratio greater than the compression ratio defined by the minimum compression ratio, when the minimum compression ratio defined by the encoding level applicable to the current output video stream is such that the encoding picture buffer size defines an encoding picture buffer that is insufficient to buffer the number of bits required to represent a picture at the maximum luminance picture size when encoded according to the minimum compression ratio.
[0101] (Example 2 - Modified Maximum CPB Size) The table in FIG. 17 shows the modified maximum CPB size for level 6. By allowing a larger maximum CPB size at level 6, the system can now store at least one encoded picture in the CPB for the most difficult combination of conditions (minimum compression ratio, maximum encoded picture size) described above.
[0102] Accordingly, an example is provided where, for each encoding level of the plurality of encoding levels, the maximum value of the encoding picture buffer size is greater than or equal to the number of bits required to represent a picture at the maximum luminance picture size when encoded according to the minimum compression ratio.
[0103] (Example 3 - Main 4:4:4 10 with Adjustment to Minimum Compression Ratio) Referring to FIG. 18, the measures taken for this profile relate to an adjustment of MinCRScaleFactor (equal to 0.75) and the same modification as for the maximum CPB at level 6 described above.
[0104] Accordingly, an example is provided where, for each of the plurality of encoding levels, the minimum compression ratio defines an encoded picture buffer that is sufficient to buffer the number of bits required to represent a picture at the maximum luminance picture size when the encoded picture buffer size is encoded according to the minimum compression ratio.
[0105] (Example 4 - CpbVcIFactor) Table A.3 shows that the CpbVcIFactor of Main 444 10 is 2.5 times that of Main 10. Therefore, in this example, it is proposed to change the CpbVcIFactor of Main 444 10 to 2 times that of Main 10 in order to match the CpbVcIFactor with the FormatCapabilityFactor. It is also proposed to keep the MinCRScaleFactor at 1.0 for both profiles. These aspects are shown in FIG. 19.
[0106] As background art, the FormatCapabilityFactor defines the bytes per pixel of the source, and thus, in 4:4:4, it is 2 times that of 4:2:0. It would seem to have the same ratio in the compressed data instead of the currently specified 2.5. This is based on the assumption that they compress equally well, which is possible at least in situations where there is no excessive chroma noise.
[0107] The CpbVcIFactor is 1000 in the HEVC main profile (8 - bit) and, in the HEVC main 10 profile, it should probably have increased but did not. VVC inherited these numbers from HEVC.
[0108] If this coefficient is reduced for 444, it may also be appropriate to align the MinCR for the two profiles in order to further adapt the full pictures within the CPB.
[0109] The table in FIG. 20 shows the effects of the variations proposed in this embodiment.
[0110] (Other embodiments) The following table provides further exemplary embodiments and may be considered as an alternative to the corresponding table provided in the drawings. Refer to Table 135 in Annex 4.1 of JVET-T2001-v1 of the 20th JVET meeting held in October 2020 (the content of which is incorporated herein by reference).
[0111] These are further exemplary configurations shown in FIGS. 15 and 16, showing situations similar to those described above, i.e., configurations in the main 10 profile, and the rows corresponding to level 6 are such that less than one coded picture can be stored in the CPB under these conditions.
[0112] [Table 1]
[0113] The following exemplary table shows similar situations for the main 444 10 profile, and it can be seen that this problem occurs at levels 1, 3, 3.1, 4, 5, and 6.
[0114] [Table 2]
[0115] The following table provides proposals to address these problems in a similar manner to FIGS. 17 and 18 (main 10 and main 444 10 respectively).
[0116] [Table 3]
[0117] [Table 4]
[0118] The following table provides further proposals to address these issues in a manner similar to FIGS. 17 and 18 (Main 10 and Main 44410, respectively).
[0119] [Table 5]
[0120] [Table 6]
[0121] Each of these alternative embodiments (and, in fact, each of the embodiments described above) can be considered to be within the scope of the appended claims.
[0122] (Encoded video data) Video data encoded by any of the techniques disclosed herein is also considered to represent an embodiment of the present disclosure.
[0123] (Overview of the method) FIG. 21 shows (in step 2100) buffering input video data in an encoded picture buffer to buffer consecutive portions of an input video data stream and providing a portion to the video data decoder for decoding, the encoded picture buffer having an encoded picture buffer size, (in step 2110) decoding a portion of the input video data stream in response to parameter data associated with the input video data stream An encoding / decoding method, wherein the parameter data indicates an encoding level selected from a plurality of encoding levels for a given input video data stream, each level defining at least a maximum luminance picture size, a minimum compression ratio, and a maximum value of the encoded picture buffer size, For each of the plurality of encoding levels, the maximum value of the encoding picture buffer size is greater than or equal to the number of bits required to represent a picture with the maximum luminance picture size when encoded according to the minimum compression ratio. It is a schematic flowchart showing an encoding / decoding method.
[0124] FIG. 22 shows (In step 2200) Buffer the input video data in the encoding picture buffer, buffer consecutive portions of the input video data stream, and provide a part to the video data decoder for decoding. The encoding picture buffer has an encoding picture buffer size. (In step 2210) Decode the input video data stream in response to parameter data associated with the input video data stream. An encoding / decoding method, The parameter data indicates an encoding level selected from a plurality of encoding levels for a given input video data stream. Each level defines at least a maximum luminance picture size, a minimum compression ratio, and a maximum value of the encoding picture buffer size. For each of the plurality of encoding levels, the minimum compression ratio is such that the encoding picture buffer size defines an encoding picture buffer sufficient to buffer the number of bits required to represent a picture with the maximum luminance picture size when encoded according to the minimum compression ratio. It is a schematic flowchart showing an encoding / decoding method.
[0125] FIG. 23 shows In step 2300, encode an input video data stream and generate the output encoded video data stream in response to parameter data associated with the output encoded video data stream, the parameter data indicating an encoding level selected from a plurality of encoding levels for a given output encoded video data stream, each level defining at least a maximum luminance picture size, a minimum compression ratio, and a maximum value of the encoded picture buffer size. In step 2310, buffer the output encoded video data stream in an encoded picture buffer configured to buffer successive portions of the output encoded video data stream generated by the encoding step. An encoding / decoding method, The encoded picture buffer has an encoded picture buffer size, When the minimum compression ratio defined by the encoding level applicable to the current output video stream is such that the encoded picture buffer size defines an encoded picture buffer insufficient to buffer the number of bits required to represent a picture at the maximum luminance picture size when encoded according to the minimum compression ratio, the encoding step includes generating the output encoded video data stream by applying a compression ratio greater than the compression ratio defined by the minimum compression ratio. It is a schematic flowchart showing an encoding / decoding method.
[0126] FIG. 24 is In step 2400, encode an input video data stream and generate the output encoded video data stream in response to parameter data associated with the output encoded video data stream, the parameter data indicating an encoding level selected from a plurality of encoding levels for a given output encoded video data stream, each level defining at least a maximum luminance picture size, a minimum compression ratio, and a maximum value of the encoded picture buffer size. Buffering, in an encoded picture buffer configured to buffer successive portions of the output encoded video data stream generated by the encoding step, the output encoded video data stream in (step 2410). An encoding / decoding method, wherein the encoded picture buffer has an encoded picture buffer size, for each encoding level of the plurality of encoding levels, the maximum value of the encoded picture buffer size is greater than or equal to the number of bits required to represent a picture at the maximum luminance picture size when encoded according to the minimum compression ratio. It is a schematic flowchart showing an encoding / decoding method.
[0127] FIG. 25 is (In step 2500) encoding an input video data stream to generate an output encoded video data stream in response to parameter data associated with the output encoded video data stream, the parameter data indicating an encoding level selected from a plurality of encoding levels for a given output encoded video data stream, each level defining at least a maximum luminance picture size, a minimum compression ratio, and a maximum value of the encoded picture buffer size, Buffering, in an encoded picture buffer configured to buffer successive portions of the output encoded video data stream generated by the encoding step, the output encoded video data stream in (step 2510). An encoding / decoding method, wherein the encoded picture buffer has an encoded picture buffer size, for each encoding level of the plurality of encoding levels, the minimum compression ratio is such that the encoded picture buffer size defines an encoded picture buffer sufficient to buffer the number of bits required to represent a picture at the maximum luminance picture size when encoded according to the minimum compression ratio. It is a schematic flowchart showing a symbolization / decoding method.
[0128] As long as the embodiments of the present disclosure are described as being implemented at least in part by a software-controlled data processing apparatus, computer-readable non-transitory media that hold such software, such as optical disks, magnetic disks, semiconductor memories, etc., will also be considered to represent an embodiment of the present disclosure. Similarly, data signals (regardless of whether they are embodied on a computer-readable non-transitory medium or not) that contain the encoded data generated according to the above-described method will also be considered to represent an embodiment of the present disclosure.
[0129] In light of the above teachings, it will be apparent that numerous modifications and variations of the present disclosure are possible. Therefore, it will be understood that the technology may be practiced in ways other than those specifically described herein within the scope of the appended claims.
[0130] In the above description for clarity of the invention, it will be understood that embodiments have been described with reference to different functional units, circuits, and / or processors. However, it is obvious that the functions may be arbitrarily and appropriately distributed among different functional units, circuits, and / or processors without departing from the embodiments of the present invention.
[0131] The embodiments described in this specification may be implemented in any suitable form, including hardware, software, firmware, or any combination thereof. The embodiments described in this specification may optionally be implemented at least in part as computer software executed on one or more data processors and / or digital signal processors. The elements and components in any embodiment are physically, functionally, and logically implemented in any suitable way. In fact, this functionality may be implemented in a single unit, in multiple units, or as part of other functional units. Thus, the embodiments of the present disclosure may be implemented in a single unit, or may be physically and functionally distributed among different units, circuits, and / or processors.
[0132] Although the present disclosure has been described in connection with some embodiments, it is not intended to be limited to the specific forms described herein. Further, although the features of the present disclosure may appear to be described in connection with specific embodiments, those skilled in the art will recognize that the various features of the described embodiments may be combined in any manner suitable for implementing the technology.
[0133] Each aspect and feature is defined by the following numbered clauses. 1. A video data decoder and, An encoded picture buffer that buffers successive portions of an input video data stream and provides a portion to the video data decoder for decoding, An encoding / decoding device comprising: The encoded picture buffer has an encoded picture buffer size, The video data decoder responds to parameter data associated with the input video data stream, the parameter data indicating an encoding level selected from a plurality of encoding levels for a given input video data stream, each level defining at least a maximum luminance picture size, a minimum compression ratio, and a maximum value of the encoded picture buffer size. For each of the encoding levels of the plurality of encoding levels, the maximum value of the encoded picture buffer size is greater than or equal to the number of bits required to represent a picture with the maximum luminance picture size when encoded according to the minimum compression ratio. Encoding / decoding device. 2. The encoding / decoding device according to 1. above, The above part is a part of the data representing the entire picture. Encoding / decoding device. 3. A video storage device, a capture device, a transmission device, or a reception device including the encoding / decoding device according to 1. above. 4. Buffer the input video data in an encoded picture buffer, buffer consecutive portions of the input video data stream, provide a part to the video data decoder for decoding, and the encoded picture buffer has an encoded picture buffer size. Decode a part of the input video data stream in response to parameter data associated with the input video data stream. An encoding / decoding method, The parameter data indicates an encoding level selected from a plurality of encoding levels for a given input video data stream, and each level defines at least a maximum luminance picture size, a minimum compression ratio, and a maximum value of the encoded picture buffer size. For each of the encoding levels of the plurality of encoding levels, the maximum value of the encoded picture buffer size is greater than or equal to the number of bits required to represent a picture with the maximum luminance picture size when encoded according to the minimum compression ratio. Encoding / decoding method. 5. Computer software that causes a computer to execute the encoding / decoding method according to 4. above when executed by the computer. 6. A computer-readable non-transitory storage medium storing the computer software according to 5. above. 7. A video data decoder, An encoded picture buffer for buffering successive portions of an input video data stream and providing a portion to the video data decoder for decoding An encoding / decoding apparatus comprising The encoded picture buffer has an encoded picture buffer size The video data decoder responds to parameter data associated with the input video data stream, the parameter data indicating an encoding level selected from a plurality of encoding levels for a given input video data stream, each level defining at least a maximum luminance picture size, a minimum compression ratio, and a maximum value of the encoded picture buffer size For each encoding level of the plurality of encoding levels, the minimum compression ratio is such that the encoded picture buffer size defines an encoded picture buffer sufficient to buffer the number of bits required to represent a picture at the maximum luminance picture size when the picture is encoded according to the minimum compression ratio Encoding / decoding apparatus 8. The encoding / decoding apparatus according to 7. above, wherein The portion is a portion of data representing the entire picture Encoding / decoding apparatus 9. A video storage device, a capture device, a transmission device, or a reception device comprising the encoding / decoding apparatus according to 7. above 10. A method for encoding / decoding, which buffers input video data in an encoded picture buffer, buffers successive portions of an input video data stream, provides a portion to the video data decoder for decoding, the encoded picture buffer having an encoded picture buffer size Decodes the input video data stream in response to parameter data associated with the input video data stream Encoding / decoding method, comprising The above parameter data indicates an encoding level selected from a plurality of encoding levels for a given input video data stream, and each level defines at least a maximum luminance picture size, a minimum compression ratio, and a maximum value of the above encoding picture buffer size. For each encoding level of the above plurality of encoding levels, the above minimum compression ratio is such that when the encoding picture buffer size is encoded according to the above minimum compression ratio, it defines an encoding picture buffer sufficient to buffer the number of bits required to represent a picture at the above maximum luminance picture size. Encoding / decoding method. 11. Computer software that causes the computer to execute the encoding / decoding method described in the above 10. when executed by a computer. 12. A computer-readable non-transitory storage medium storing the computer software described in the above 11. 13. A video data encoder, An encoding picture buffer that buffers consecutive portions of the current output video data stream generated by the above video data encoder, An encoding / decoding device comprising: The above encoding picture buffer has an encoding picture buffer size, The above video data encoder responds to parameter data associated with the above output video data stream, and the above parameter data indicates an encoding level selected from a plurality of encoding levels for a given output video data stream, and each level defines at least a maximum luminance picture size, a minimum compression ratio, and a maximum value of the above encoding picture buffer size. When the minimum compression ratio defined by the encoding level applicable to the current output video stream is such that the encoding picture buffer size is insufficient to buffer the number of bits required to represent a picture with the maximum luminance picture size when encoding according to the minimum compression ratio, the video data encoder is configured to generate the current output video data stream by applying a compression ratio greater than the compression ratio defined by the minimum compression ratio. Encoding / decoding device. 14. The encoding / decoding device according to 13. above, The above part is a part of data representing the entire picture. Encoding / decoding device. 15. A video storage device, a capture device, a transmission device, or a reception device including the encoding / decoding device according to 13. above. 16. Encoding an input video data stream and generating the output encoded video data stream in response to parameter data associated with the output encoded video data stream, the parameter data indicating an encoding level selected from a plurality of encoding levels for a given output encoded video data stream, each level defining at least a maximum luminance picture size, a minimum compression ratio, and a maximum value of the encoding picture buffer size, Buffering the output encoded video data stream in an encoding picture buffer configured to buffer consecutive portions of the output encoded video data stream generated by the encoding step An encoding / decoding method, The encoding picture buffer has an encoding picture buffer size, When the minimum compression ratio defined by the encoding level applicable to the current output video stream is such that the encoded picture buffer size is insufficient to buffer the number of bits required to represent a picture at the maximum luminance picture size when encoding according to the minimum compression ratio, the encoding step includes generating the output encoded video data stream by applying a compression ratio greater than the compression ratio defined by the minimum compression ratio. Encoding / decoding method. 17. Computer software which, when executed by a computer, causes the computer to execute the encoding / decoding method according to item 16 above. 18. A computer-readable non-transitory storage medium storing the computer software according to item 17 above. 19. A video data encoder and an encoded picture buffer for buffering successive portions of the current output video data stream generated by the video data encoder and an encoding / decoding apparatus comprising: the encoded picture buffer has an encoded picture buffer size, the video data encoder responds to parameter data associated with the output video data stream, the parameter data indicating an encoding level selected from a plurality of encoding levels for a given output video data stream, each level defining at least a maximum luminance picture size, a minimum compression ratio, and a maximum value of the encoded picture buffer size, for each encoding level of the plurality of encoding levels, the maximum value of the encoded picture buffer size is greater than or equal to the number of bits required to represent a picture at the maximum luminance picture size when encoding according to the minimum compression ratio Encoding / decoding apparatus. 20. The encoding / decoding apparatus according to item 19 above, wherein the portion is a portion of data representing the entire picture. Symbolization / decoding device. 21. A video storage device, a capture device, a transmission device, or a reception device including the symbolization / decoding device described in 19 above. 22. Encoding an input video data stream, generating the output encoded video data stream in response to parameter data associated with the output encoded video data stream, the parameter data indicating an encoding level selected from a plurality of encoding levels for a given output encoded video data stream, each level defining at least a maximum luminance picture size, a minimum compression ratio, and a maximum value of the encoding picture buffer size, Buffering the output encoded video data stream in an encoding picture buffer configured to buffer successive portions of the output encoded video data stream generated by the encoding step A symbolization / decoding method, The encoding picture buffer has an encoding picture buffer size, For each encoding level of the plurality of encoding levels, the maximum value of the encoding picture buffer size is greater than or equal to the number of bits required to represent a picture at the maximum luminance picture size when encoded according to the minimum compression ratio. Symbolization / decoding method. 23. Computer software that causes a computer to execute the symbolization / decoding method described in 22 above when executed by the computer. 24. A computer-readable non-transitory storage medium storing the computer software described in 23 above. 25. A video data encoder, An encoding picture buffer that buffers successive portions of the current output video data stream generated by the video data encoder A symbolization / decoding device comprising: The encoding picture buffer has an encoding picture buffer size, The video data encoder responds to parameter data associated with the output video data stream, and the parameter data indicates an encoding level selected from a plurality of encoding levels for a given output video data stream, and each level defines at least a maximum luminance picture size, a minimum compression ratio, and a maximum value of the encoded picture buffer size. For each encoding level of the plurality of encoding levels, the minimum compression ratio is such that the encoded picture buffer size defines an encoded picture buffer sufficient to buffer the number of bits required to represent a picture at the maximum luminance picture size when the encoded picture buffer size is encoded according to the minimum compression ratio. Encoding / decoding device. 26. The encoding / decoding device according to 25. above, The portion is a portion of data representing the entire picture. Encoding / decoding device. 27. A video storage device, a capture device, a transmission device, or a reception device including the encoding / decoding device according to 25. above. 28. Encoding an input video data stream to generate an output encoded video data stream in response to parameter data associated with the output encoded video data stream, the parameter data indicating an encoding level selected from a plurality of encoding levels for a given output encoded video data stream, and each level defining at least a maximum luminance picture size, a minimum compression ratio, and a maximum value of the encoded picture buffer size. Buffering the output encoded video data stream in an encoded picture buffer configured to buffer consecutive portions of the output encoded video data stream generated by the encoding step. Encoding / decoding method, The encoded picture buffer has an encoded picture buffer size. For each of the plurality of encoding levels, the minimum compression ratio defines an encoded picture buffer such that when the encoded picture buffer size is encoded according to the minimum compression ratio, it is sufficient to buffer the number of bits required to represent a picture with the maximum luminance picture size. Encoding / decoding method. 29. Computer software which, when executed by a computer, causes the computer to execute the encoding / decoding method described in 28. 30. A computer-readable non-transitory storage medium storing the computer software described in 29.
Claims
1. An apparatus for decoding video encoded according to the Versatile Video Coding (VVC) standard, comprising: a video data decoder; and an encoded picture buffer that buffers consecutive portions of a main layer input video data stream and provides a portion to the video data decoder for decoding; The encoded picture buffer has an encoded picture buffer size, The video data decoder responds to parameter data associated with the input video data stream, the parameter data indicating an encoding level selected from a plurality of encoding levels for a given input video data stream, each level defining at least a maximum luminance picture size, a minimum compression ratio, and a maximum value of the encoded picture buffer size, For each encoding level of the main layer of the plurality of encoding levels, the maximum value of the encoded picture buffer size is greater than or equal to the number of bits required to represent a picture at the maximum luminance picture size when encoded according to the minimum compression ratio, the minimum compression ratio being defined by a scaled base value, the base value for at least one or more of the encoding levels of the main layer being 8, In the case of an encoding level having a value of 6 as the main layer value, the scaled base value is 8 and the encoded picture buffer size is 80,000 bits. An apparatus.
2. The apparatus according to claim 1, wherein the portion is a portion of data representing the entire picture. An apparatus.
3. A video storage device, a capture device, a transmission device, or a reception device comprising the apparatus according to claim 1.
4. A method for decoding video encoded according to the Versatile Video Coding (VVC) standard, Buffering input video data of a main layer in an encoded picture buffer to buffer consecutive portions of an input video data stream and providing a portion to a video data decoder for decoding, wherein the encoded picture buffer has an encoded picture buffer size, Decoding a portion of the input video data stream in response to parameter data associated with the input video data stream A method, The parameter data indicates an encoding level selected from a plurality of encoding levels for a given input video data stream, each level defining at least a maximum luminance picture size, a minimum compression ratio, and a maximum value of the encoded picture buffer size, For each encoding level of the main layer of the plurality of encoding levels, the maximum value of the encoded picture buffer size is greater than or equal to the number of bits required to represent a picture at the maximum luminance picture size when encoded according to the minimum compression ratio, the minimum compression ratio being defined by a scaled base value, the base value for at least one or more of the encoding levels of the main layer being 8, For an encoding level having a value of 6 as the main layer, the scaled base value is 8 and the encoded picture buffer size is 80,000 bits A method. **Claim 5** Computer software which, when executed by a computer, causes the computer to perform the method according to claim 4. **Claim 6** A computer-readable non-transitory storage medium storing the computer software according to claim 5. **Claim 7** An apparatus for encoding video according to the Versatile Video Coding (VVC) standard, A video data encoder, and an encoded picture buffer that buffers successive portions of the output video data stream of the current main layer generated by the video data encoder The apparatus comprises: The encoded picture buffer has an encoded picture buffer size, The video data encoder responds to parameter data associated with the output video data stream, the parameter data indicating an encoding level selected from a plurality of encoding levels for a given output video data stream, each level defining at least a maximum luminance picture size, a minimum compression ratio, and a maximum value of the encoded picture buffer size, For each encoding level of the main layer of the plurality of encoding levels, the maximum value of the encoded picture buffer size is greater than or equal to the number of bits required to represent a picture at the maximum luminance picture size when encoded according to the minimum compression ratio, the minimum compression ratio being defined by a scaled base value, the base value for at least one or more of the encoding levels of the main layer being 8, In the case of an encoding level having a value of 6 as the main layer, the scaled base value is 8 and the encoded picture buffer size is 80,000 bits Apparatus.
8. The apparatus according to claim 7, wherein The portion is a portion of data representing the entire picture Apparatus.
9. A video storage device, a capture device, a transmission device, or a reception device comprising the apparatus according to claim 7.
10. A method for encoding video according to the Versatile Video Coding (VVC) standard, comprising: Encode the input video data stream of the main layer and generate the output encoded video data stream in response to parameter data associated with the output encoded video data stream, the parameter data indicating an encoding level selected from a plurality of encoding levels for a given output encoded video data stream, each level defining at least a maximum luminance picture size, a minimum compression ratio, and a maximum value of an encoded picture buffer size, Buffering the output encoded video data stream in an encoded picture buffer configured to buffer consecutive portions of the output encoded video data stream generated by the step of encoding the input data stream of the main layer A method comprising: The encoded picture buffer has the encoded picture buffer size, For each encoding level of the main layer of the plurality of encoding levels, the maximum value of the encoded picture buffer size is greater than or equal to the number of bits required to represent a picture at the maximum luminance picture size when encoded according to the minimum compression ratio, the minimum compression ratio being defined by a scaled base value, the base value for at least one or more of the encoding levels of the main layer being 8, In the case of an encoding level having a value of 6 as the main layer, the scaled base value is 8 and the encoded picture buffer size is 80,000 bits A method. Claim 11 Computer software which, when executed by a computer, causes the computer to perform the method according to claim 10. Claim 12 A computer-readable non-transitory storage medium storing the computer software according to claim 11.
Citation Information
Patent Citations
Decoding device and decoding method, and encoding device and encoding method
WO2015105003A1
Cited By
Video data encoding and decoding using a coded picture buffer whose size is defined by parameter data
US12563211B2