Video data encoding and decoding
By extending the encoding level to 15.15 and adjusting the decoding stream capability constraints, the problem that existing video encoding systems cannot support high frame rates and large picture sizes is solved, and effective encoding and decoding of high frame rates video data is realized, and compatibility with existing systems is maintained.
Patent Information
- Application Number
- CN202080099067.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Priority Date
- 2020-04-03
- Filing Date
- 2020-12-16
- Publication Date
- 2025-08-15
- Estimated Expiration
- 2040-12-16
AI Technical Summary
Existing video encoding and decoding systems have unreasonable restrictions when creating new levels, resulting in the inability to effectively support video data at high frame rates and large picture sizes, and existing decoders are not compatible with video streams at higher frame rates.
An alternative level encoding method is adopted to extend level encoding from 8.5 to 15.15, and new decoding stream capability constraints are proposed to allow video data at higher frame rates without adding higher-level decoders.
It realizes effective encoding and decoding of high frame rate and large picture size video data, is compatible with existing decoders, and does not affect the compatibility of existing systems.
Smart Images

Figure CN115336275B_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to video data encoding and decoding. Background Art
[0002] The "background" description provided herein is intended to generally present the context of the present disclosure. To the extent described in this background section, the work of the presently named inventors, as well as aspects of the specification that may not qualify as prior art at the time of filing, are not admitted, either explicitly or implicitly, as prior art to the present disclosure.
[0003] Several systems exist, such as video or image data encoding and decoding systems, that involve converting the video data into a frequency domain representation, quantizing the frequency domain coefficients, and then applying some form of entropy coding to the quantized coefficients. This allows for compression of the video data. Appropriate decoding or decompression techniques are then applied to recover a reconstructed version of the original video data. Summary of the Invention
[0004] The present disclosure solves or alleviates the problems caused by this process.
[0005] Various aspects and features of the present disclosure are defined in the following claims.
[0006] It is to be understood that both the foregoing general description and the following detailed description are illustrative and not restrictive of the present technology. BRIEF DESCRIPTION OF THE DRAWINGS
[0007] A more complete appreciation of the present disclosure, together with many of its attendant advantages, may be readily obtained by referring to the following detailed description when considered in conjunction with the accompanying drawings, wherein:
[0008] Figure 1 Schematically illustrates an audio / video (A / V) data transmission and reception system using video data compression and decompression;
[0009] Figure 2 schematically illustrates a video display system using video data decompression;
[0010] Figure 3 Schematically illustrates an audio / video storage system using video data compression and decompression;
[0011] Figure 4 schematically illustrates a video camera using video data compression;
[0012] Figure 5 and Figure 6 schematically illustrates a storage medium;
[0013] Figure 7 A schematic overview of video data compression and decompression equipment is provided;
[0014] Figure 8 A predictor is schematically shown;
[0015] Figure 9 and Figure 10 Schematically illustrating a set of coding levels;
[0016] Figure 11 Schematically illustrates the use of parameter sets;
[0017] Figure 12 Schematically illustrates various aspects of a decoding device;
[0018] Figure 13 is a schematic flow chart illustrating the method;
[0019] Figure 14 schematically illustrates various aspects of an encoding device; and
[0020] Figure 15 is a schematic flow chart illustrating the method. DETAILED DESCRIPTION
[0021] Referring now to the accompanying drawings, there is provided Figures 1 to 4 A schematic diagram of a device or system using the compression and / or decompression device described below in conjunction with an embodiment of the present technology is provided.
[0022] All data compression and / or decompression devices to be described below may be implemented in hardware, software running on a general-purpose data processing device such as a general-purpose computer, as programmable hardware such as an application-specific integrated circuit (ASIC) or a field-programmable gate array (FPGA), or as a combination thereof. Where an embodiment is implemented by software and / or firmware, it will be appreciated that such software and / or firmware as well as a non-transitory data storage medium are considered embodiments of the present technology, wherein such software or firmware is stored or otherwise provided by the non-transitory data storage medium.
[0023] Figure 1 Schematically illustrates an audio / video data transmission and reception system using video data compression and decompression.In this example, the data values to be encoded or decoded represent image data.
[0024] An input audio / video signal 10 is provided to a video data compression device 20, which compresses at least the video component of the audio / video data 10 for transmission along a transmission path 30 (e.g., a cable, optical fiber, wireless link, etc.). The compressed signal is processed by a decompression device 40 to provide an output audio / video signal 50. For the return path, a compression device 60 compresses the audio / video data for transmission along the transmission path 30 to a decompression device 70.
[0025] Thus, compression device 20 and decompression device 70 may form one node of the transmission link. Decompression device 40 and decompression device 60 may form another node of the transmission link. Of course, if the transmission link is unidirectional, only one node requires a compression device, while the other node only requires a decompression device.
[0026] Figure 2 Schematically, a video display system using video data decompression is shown. Specifically, a compressed audio / video signal 100 is processed by a decompression device 110 to provide a decompressed signal that can be displayed on a display 120. The decompression device 110 can be implemented as an integral part of the display 120, for example, being arranged in the same housing as the display device. Alternatively, the decompression device 110 can be arranged as, for example, a so-called set-top box (STB), noting that the expression "set-top box" does not mean that the set-top box is required to be placed in any particular orientation or position relative to the display 120; it is simply a term used in the art to indicate a device that can be connected to a display as a peripheral device.
[0027] Figure 3 An audio / video storage system using video data compression and decompression is schematically shown. An input audio / video signal 130 is provided to a compression device 140, which generates a compressed signal for storage by a storage device 150 (e.g., a magnetic disk, optical disk, tape, solid-state storage such as semiconductor memory, or other storage device). For playback, the compressed data is read from the storage device 150 and passed to a decompression device 160 for decompression to provide an output audio / video signal 170.
[0028] It should be understood that a compressed or encoded signal and a storage medium storing the signal (eg, a machine-readable non-transitory storage medium) are considered embodiments of the present technology.
[0029] Figure 4 A video camera using video data compression is schematically shown. Figure 4 , an image capture device 180 (e.g., a charge coupled device (CCD) image sensor and associated control and readout electronics) generates a video signal that is passed to a compression device 190. A microphone (or microphones) 200 generates an audio signal that is passed to the compression device 190. The compression device 190 generates a compressed audio / video signal 210 that is stored and / or transmitted (shown generally as schematic step 220).
[0030] The techniques that will be described below primarily relate to video data compression and decompression. It will be appreciated that many existing techniques can be used for audio data compression in conjunction with the video data compression techniques to be described to generate a compressed audio / video signal. Therefore, a separate discussion of audio data compression will not be provided. It will also be appreciated that the data rates associated with video data (particularly broadcast quality video data) are typically much higher than the data rates associated with audio data (whether compressed or uncompressed). It will therefore be appreciated that uncompressed audio data can accompany compressed video data to form a compressed audio / video signal. It will further be appreciated that although the present example (e.g., Figures 1 to 4 Although the embodiments (shown) relate to audio / video data, the techniques described below can be applied to systems that simply process (i.e., compress, decompress, store, display, and / or transmit) video data. That is, the embodiments can be applied to video data compression without necessarily having any associated audio data processing.
[0031] therefore, Figure 4 An example of a video capture device is provided that includes an image sensor and an encoding device of the type discussed below. Figure 2 Examples of decoding devices of the type discussed below are provided, as well as displays that output decoded images.
[0032] Figure 2 and Figure 4 The combination of can provide a video capture device including an image sensor 180 and an encoding device 190, a decoding device 110, and a display 120 that outputs a decoded image.
[0033] Figure 5 and Figure 6 A storage medium is schematically shown, which stores compressed data generated, for example, by the device 20 , 60 , the compressed data being input to the device 110 or the storage medium or steps 150 , 220 . Figure 5 schematically illustrates a disk storage medium, Figure 6 Schematically illustrates a solid-state storage medium, such as flash memory. Note that Figure 5 and Figure 6 Examples of non-transitory machine-readable storage media may also be provided that store computer software that, when executed by a computer, causes the computer to perform one or more of the methods discussed below.
[0034] The above arrangements thus provide examples of video storage, capture, transmission or reception devices embodying any of the present techniques.
[0035] Figure 7 An overview diagram of a video or image data compression (encoding) and decompression (decoding) apparatus is provided for encoding and / or decoding video or images representing one or more images.
[0036] The controller 343 controls the overall operation of the apparatus and, in particular, controls the test encoding process by acting as a selector to select various modes of operation, such as block size and shape, and whether the video data is to be losslessly encoded, when referring to the compression mode. The controller is considered to constitute part of the image encoder or image decoder, as the case may be. Successive images of the input video signal 300 are provided to the adder 310 and the image predictor 320. Figure 8 The image predictor 320 is described in more detail. The image encoder or decoder (as the case may be) is coupled Figure 8 The intra-image predictor can use the Figure 7 But this does not mean that the image encoder or decoder must Figure 7 Each feature.
[0037] The adder 310 actually performs a subtraction (negative addition) operation, since it receives the input video signal 300 on the "+" input and the output of the image predictor 320 on the "-" input, thereby subtracting the predicted image from the input image. The result is a so-called residual image signal 330 representing the difference between the actual image and the predicted image.
[0038] One reason for generating a residual image signal is as follows. The data encoding techniques to be described, i.e., the techniques to be applied to the residual image signal, tend to work more efficiently when there is less "energy" in the image to be encoded. Here, the term "efficiently" refers to generating a small amount of encoded data; for a particular level of image quality, it is desirable (and considered "efficient") to generate as little data as possible. The reference to "energy" in the residual image relates to the amount of information that the residual image contains. If the predicted image is identical to the actual image, then the difference between the two (i.e., the residual image) will contain zero information (zero energy) and will be very easy to encode into a small amount of encoded data. In general, if the prediction process can be made to work reasonably so that the predicted image content is similar to the image content to be encoded, then it is expected that the residual image data will contain less information (less energy) than the input image and will therefore be easier to encode into a small amount of encoded data.
[0039] Thus, encoding (using adder 310) involves predicting an image region of an image to be encoded; and generating a residual image region based on the difference between the predicted image region and a corresponding region of the image to be encoded. In conjunction with the techniques discussed below, the ordered array of data values includes data values representing the residual image region. Decoding involves predicting an image region of an image to be decoded; generating a residual image region indicating the difference between the predicted image region and the corresponding region of the image to be decoded; wherein the ordered array of data values includes data values representing the residual image region; and combining the predicted image region and the residual image region.
[0040] The remainder of the apparatus used as an encoder (encoding the residual or difference image) will now be described.
[0041] The residual image data 330 is provided to a conversion unit or circuit 340, which generates a discrete cosine transform (DCT) representation of the block or region of residual image data. DCT techniques themselves are well known and will not be described in detail here. It should also be noted that the use of DCT illustrates only one exemplary arrangement. Other conversions that may be used include, for example, discrete sine transforms (DSTs). Conversions may also include sequences or cascades of single conversions, such as an arrangement in which one conversion is followed (whether directly or not) by another conversion. The choice of conversion may be determined explicitly and / or may depend on side information used to configure the encoder and decoder. In other examples, a so-called "convert skip" mode may be selectively used in which no conversion is applied.
[0042] Thus, in an example, the encoding and / or decoding method comprises predicting an image region of an image to be encoded; and generating a residual image region based on a difference between the predicted image region and a corresponding region of the image to be encoded; wherein an ordered array of data values (to be discussed below) comprises data values representing the residual image region.
[0043] The output of the transform unit 340 (that is, in one example, a set of DCT coefficients for each transformed block of image data) is provided to a quantizer 350. Various quantization techniques are known in the art of video data compression, ranging from simple multiplication by a quantization scale factor to the application of complex lookup tables under the control of a quantization parameter. The overall goal is twofold. First, the quantization process reduces the number of possible values for the transformed data. Second, the quantization process can increase the probability that the transformed data value is zero. Both of these approaches allow the entropy encoding process, described below, to operate more efficiently when generating small amounts of compressed video data.
[0044] The scanning unit 360 applies a data scanning process. The purpose of the scanning process is to reorder the quantized transform data so as to group together as many non-zero quantized transform coefficients as possible, and therefore also to group together as many zero-valued coefficients as possible. These features may allow so-called run-length encoding or similar techniques to be applied efficiently. Thus, the scanning process involves selecting coefficients from the quantized transform data (specifically, selecting coefficients from coefficient blocks corresponding to blocks of converted and quantized image data) according to a "scan order" such that (a) all coefficients are selected once as part of the scan, and (b) the scan tends to provide the desired reordering. One example scanning order that may tend to give useful results is a so-called top-right diagonal scanning order.
[0045] The scanning order may be different between skipped transform blocks and transform blocks (blocks that have undergone at least one spatial frequency transform).
[0046] The scanned coefficients are then passed to an entropy encoder (EE) 370. Again, various types of entropy coding can be used. Two examples are variants of the so-called context-adaptive binary arithmetic coding (CABAC) system and the so-called context-adaptive variable length coding (CAVLC) system. In general, CABAC is considered to provide better efficiency, and in some studies it has been shown that the amount of coded output data for CABAC is reduced by 10% to 20% for comparable image quality compared to CAVLC. However, CAVLC is considered to represent a much lower complexity (in terms of its implementation) than CABAC. Note that the scanning process and the entropy coding process are shown as separate processes, but in fact can be combined or processed together. That is, the data can be read into the entropy encoder in the scanning order. Corresponding considerations apply to the various reverse processes to be described below.
[0047] The output of the entropy encoder 370, along with additional data (mentioned above and / or discussed below), such as defining how the predictor 320 generates the predicted image, whether the compressed data is converted or conversion is skipped, etc., provides a compressed output video signal 380.
[0048] However, a return path 390 is also provided because the operation of the predictor 320 itself depends on the decompressed version of the compressed output data.
[0049] The reason for this feature is as follows. At an appropriate stage in the scan-decompression process (described below), a decompressed version of the residual data is generated. This decompressed residual data must be added to the predicted image to generate the output image (because the original residual data is the difference between the input image and the predicted image). In order to make this process comparable between the compression and decompression sides, the predicted image generated by the predictor 320 should be the same during the compression process and during the scan-decompression process. Of course, during decompression, the device does not have access to the original input image, but only to the decompressed image. Therefore, during compression, the predictor 320 bases its predictions (at least for inter-image coding) on the decompressed version of the compressed image.
[0050] The entropy encoding process performed by the entropy encoder 370 is considered (in at least some examples) to be "lossless," that is, it can be reversed to obtain data that is identical to the data first provided to the entropy encoder 370. Therefore, in such examples, the return path can be implemented before the entropy encoding stage. In fact, the scanning process performed by the scanning unit 360 is also considered lossless, so in this embodiment, the return path 390 runs from the output of the quantizer 350 to the input of the complementary inverse quantizer 420. In the case where losses or potential losses are introduced by a stage, that stage (and its inverse stage) can be included in the feedback loop formed by the return path. For example, the entropy encoding stage can be lossy, at least in principle, such as by encoding bits within the parity information. In this case, the entropy encoding and decoding should form part of the feedback loop.
[0051] In general, the entropy decoder 410, the inverse scan unit 400, the inverse quantizer 420 and the inverse transform unit or circuit 430 provide corresponding inverse functions of the entropy encoder 370, the scan unit 360, the quantizer 350 and the transform unit 340; the process of decompressing the input compressed video signal will be discussed separately below.
[0052] In the compression process, the scanned coefficients are passed from the quantizer 350 via a return path 390 to an inverse quantizer 420 which performs the inverse operation of the scanning unit 360. The inverse quantization and inverse transform processes are performed by units 420, 430 to generate a decompressed compressed residual image signal 440.
[0053] The image signal 440 is added to the output of the predictor 320 at an adder 450 to generate a reconstructed output image 460 (although this may be subjected to so-called loop filtering and / or other filtering before output - see below). This forms one input to the image predictor 320, as described below.
[0054] Turning now to the decoding process applied to decompress the received compressed video signal 470, this signal is supplied to an entropy decoder 410 and from there to a chain of inverse scan unit 400, inverse quantizer 420 and inverse transform unit 430 before being added to the output of the picture predictor 320 by an adder 450. The decoder reconstructs a version of the residual picture which is then applied (by the adder 450) to the predicted version of the picture (block by block) in order to decode each block. Briefly, the output 460 of the adder 450 forms the output decompressed video signal 480 (subject to the filtering process discussed below). In practice, further filtering is optionally applied (e.g. by Figure 8 The loop filter 565 is shown, but for Figure 7 The clarity of the higher-level diagram, from Figure 7omitted).
[0055] Figure 7 and Figure 8 The device can be used as a compression (encoding) device or a decompression (decoding) device. The functions of the two types of devices are substantially overlapping. The scanning unit 360 and the entropy encoder 370 are not used in the decompression mode, and the operation of the predictor 320 (which will be described in detail below) and other units follows the mode and parameter information contained in the received compressed bit stream, rather than generating this information themselves.
[0056] Figure 8 The generation of a predicted image, in particular the operation of the image predictor 320 , is schematically illustrated.
[0057] There are two basic prediction modes performed by the image predictor 320: so-called intra-image prediction and so-called inter-image or motion-compensated (MC) prediction. On the encoder side, each prediction mode involves detecting the prediction direction for the current block to be predicted and generating a predicted block of samples based on other samples (in the same (intra-frame) or another (inter-frame) image). The difference between the predicted block and the actual block is encoded or applied by means of units 310 or 450 to encode or decode the block, respectively.
[0058] (At the decoder, or on the reverse decoding side of the encoder, detection of the prediction direction may be responsive to data associated with the encoded data of the encoder, indicating which direction was used at the encoder. Alternatively, the detection may be responsive to the same factors as those used to make the decision at the encoder).
[0059] Intra-image prediction predicts the content of a block or region of an image based on data from within the same image. This corresponds to so-called I-frame coding in other video compression techniques. However, unlike I-frame coding, which involves encoding the entire image via intra-frame coding, in this embodiment, the choice between intra-frame and inter-frame coding can be made on a block-by-block basis, although in other embodiments the choice is still made on a per-image basis.
[0060] Motion compensated prediction is an example of inter-image prediction and exploits motion information that attempts to define the source of image detail to be encoded in the current image in another adjacent or nearby image. Thus, in an ideal example, the contents of an image data block in the predicted image can be encoded very simply as a reference (motion vector) to a corresponding block at the same or slightly different location in an adjacent image.
[0061] A technique called "block copy" prediction is in some ways a hybrid of the two predictions, as it uses a vector to indicate the block of samples at a position in the same picture that is displaced from the current prediction block and that should be copied to form the current prediction block.
[0062] return Figure 8 , two image prediction arrangements (corresponding to intra-image and inter-image prediction) are shown, the results of which are selected by multiplexer 500 under the control of mode signal 510 (e.g., from controller 343) to set the blocks of the predicted image to be provided to adders 310 and 450, and this selection is signaled to the decoder within the encoded output data stream. In this case, the image energy can be detected, for example, by performing a test subtraction of the area of two versions of the predicted image from the input image, squaring each pixel value of the difference image, summing the squared values, and identifying which of the two versions results in a lower mean square value of the difference image associated with that image area. In other examples, trial encoding can be performed for each option or potential option, and then a selection can be made based on the cost of each potential option in terms of one or both of the number of bits required for encoding and picture distortion.
[0063] In an intra-coding system, the actual prediction is based on the image blocks received as part of signal 460 (filtered by loop filtering; see below), that is, the prediction is based on the coded decoded image blocks so that the exact same prediction can be made at the decompression device. However, data can be derived from the input video signal 300 by the intra-mode selector 520 to control the operation of the intra-image predictor 530.
[0064] For inter-image prediction, a motion compensated (MC) predictor 540 uses motion information, such as motion vectors derived from the input video signal 300 by a motion estimator 550. These motion vectors are applied by the motion compensated predictor 530 to a processed version of the reconstructed image 460 to generate inter-image predicted blocks.
[0065] Thus, units 530 and 540 (operating in conjunction with estimator 550) each act as a detector to detect a prediction direction with respect to a current block to be predicted, and as a generator to generate a prediction block of samples (forming part of the prediction passed to units 310 and 450) from other samples defined by the prediction direction.
[0066] The processing applied to signal 460 will now be described.
[0067] First, the signal may be filtered by a so-called loop filter 565. Various types of loop filters may be used. One technique involves applying a "deblocking" filter to remove, or at least tend to reduce, the effects of block-based processing and subsequent operations performed by the conversion unit 340. Another technique involving the application of a so-called sample adaptive offset (SAO) filter may also be used. Generally speaking, in a sample adaptive offset filter, filter parameter data (derived at the encoder and transmitted to the decoder) defines one or more offsets that are selectively combined by the sample adaptive offset filter with a given intermediate video sample (a sample of the signal 460) based on: (i) a given intermediate video sample; or (ii) one or more intermediate video samples that have a predetermined spatial relationship to the given intermediate video sample.
[0068] In addition, an adaptive loop filter is optionally applied using coefficients derived from processing the reconstructed signal 460 and the input video signal 300. An adaptive loop filter is a filter that applies adaptive filter coefficients to the data to be filtered using known techniques. In other words, the filter coefficients can vary depending on various factors. Data defining the filter coefficients to be used is included as part of the encoded output data stream.
[0069] The techniques discussed below relate to the processing of parameter data related to the operation of the filter. The actual filtering operation (eg, SAO filtering) may use other known techniques.
[0070] When the device operates as a decompression device, the filtered output from loop filter unit 565 actually forms the output video signal 480. It is also buffered in one or more image or frame memories 570; storage of consecutive images is a requirement for motion-compensated prediction processing, particularly the generation of motion vectors. To save on storage requirements, the images stored in image memory 570 can be kept in compressed form and then decompressed for use in generating motion vectors. Any known compression / decompression system can be used for this particular purpose. The stored images can be passed to interpolation filter 580, which generates a higher-resolution version of the stored images; in this example, intermediate samples (subsamples) are generated so that the resolution of the interpolated images output by interpolation filter 580 is four times the resolution (in each dimension) of the images stored in image memory 570 for the 4:2:0 luma channel and eight times the resolution (in each dimension) of the images stored in image memory 570 for the 4:2:0 chroma channels. The interpolated images are passed as input to motion estimator 550 and also to motion-compensated predictor 540.
[0071] The manner in which images are segmented for compression will now be described. At a basic level, the image to be compressed is viewed as an array of blocks or regions of samples. Segmenting an image into such blocks or regions can be done using a decision tree, as described in SERIES H: Audiovisual and Multimedia Systems Infrastructure of audiovisual services - Coding of moving video High efficiency video coding Recommendation ITU-T H.265 12 / 2016. Also: High Efficiency Video Coding (HEVC) Algorithms and Architectures, chapter 3, Editors: Madhukar Budagavi, Gary J. Sullivan, Vivienne Sze; ISBN 978-3-319-06894-7; 2014, each of which is incorporated herein by reference in its entirety. In addition, further background information is provided in [1] “Versatile Video Coding (Draft 8)”, JVET-Q2001-vE, B. Bross, J. Chen, S. Liu and Y-K. Wang, which is also incorporated herein by reference in its entirety.
[0072] In some examples, the generated blocks or regions have sizes, and in some cases shapes, that can generally follow the configuration of image features within the image using a decision tree. This in itself can improve coding efficiency because samples that represent or follow similar image features will tend to be grouped together by this arrangement. In some examples, square blocks or regions of different sizes can be selected (e.g., 4x4 samples, such as 64x64 or larger blocks). In other exemplary arrangements, blocks or regions of different shapes can be used, such as rectangular blocks (e.g., vertically or horizontally oriented). Other non-square and non-rectangular blocks are contemplated. The result of dividing the image into such blocks or regions is (at least in this example) that each sample of the image is assigned to one and only one such block or region.
[0073] Embodiments of the present disclosure discussed below relate to techniques for representing encoding levels at an encoder and a decoder.
[0074] Parameter Sets and Encoding Levels
[0075] When video data is encoded for subsequent decoding by the techniques described above, it is appropriate for the encoding side of the process to communicate some parameters of the encoding process to the eventual decoding side of the process. Given that these encoding parameters are required whenever the encoded video data is decoded, it is useful to associate these parameters with the encoded video data stream itself (although not necessarily uniquely, as these parameters can be sent "out of band" via a separate transmission channel), for example, by embedding these parameters as so-called parameter sets in the encoded video data stream itself.
[0076] Parameter sets can be represented as a hierarchy of information, such as a video parameter set (VPS), a sequence parameter set (SPS), and a picture parameter set (PPS). The PPS is expected to appear once per picture and contains information relevant to all coded slices in that picture, the SPS is less frequent (once per picture sequence), and the VPS is also less frequent. Parameter sets that appear more frequently (such as the PPS) can be implemented as references to previously coded instances of that parameter set to avoid the cost of re-encoding. Each coded image slice references a single active PPS, SPS, and VPS to provide information for decoding that slice. In particular, each slice header may contain a PPS identifier to reference the PPS, which in turn references the SPS, which in turn references the VPS.
[0077] Within these parameter sets, the SPS contains example information relevant to the following discussion, namely data defining the so-called profiles, layers, and encoding levels to be used.
[0078] A profile defines a set of decoding tools or features to use. Example profiles include the "Main Profile," which is specific to 8-bit 4:2:0 video, and the "Main 10 Profile," which allows for 10-bit resolution and other extensions related to the Main Profile.
[0079] The coding level provides constraints on issues such as maximum sampling rate and picture size. This layer specifies the maximum data rate.
[0080] In the Joint Video Experts Team (JVET) proposal for Versatile Video Coding (VVC), as defined in the above-referenced specification JVET-Q2001-vE (at the date of submission), various levels from 1 to 6.2 are defined.
[0081] The highest possible level is level 8.5. This limitation stems from the encoding used, since the level is encoded in 8 bits as level*30, which means that level 8.0 is encoded as the maximum 8-bit value 255.
[0082] The current specification also defines the condition that a decoder conforming to a given profile at a particular level of a particular layer shall be able to decode all bitstreams to which the following condition applies, that is, the bitstream is indicated as conforming to a level (of a given profile at a given layer) that is not level 8.5 and lower than or equal to the specified level.
[0083] Figure 9 A schematic diagram shows the currently defined levels with example picture sizes and frame rates, which represents the merger of the corresponding conditions imposed by the two example tables (A.1 and A.2) of the JVET file. The primary component of the level (indicating the maximum luminance picture size and monotonically increasing with increasing maximum luminance picture size) is represented by an integer in the left column; the secondary component of the level (indicating at least one of the maximum luminance sampling rate and the maximum frame rate, where for a given first component, the second component monotonically increases with increasing maximum luminance sampling rate and monotonically increases with increasing maximum frame rate) is represented by the digits after the decimal point.
[0084] Note that in principle, one of the maximum luma sampling rate and maximum frame rate can be derived from the other and the maximum luma picture size, so although both are specified in the table for clarity of explanation, in principle only one of them needs to be specified.
[0085] The data rate, picture size, etc. are defined for the luma (or bright) component; the corresponding data for chroma (color) can then be derived from the chroma subsampling format used (e.g. 4:2:0 or 4:4:4, which defines a corresponding number of chroma samples for a given number of luma samples).
[0086] Note that the column labeled "Example Maximum Luminance Size" only provides examples of luma image configurations that adhere to the number of samples defined by the corresponding maximum luma image size and to the aspect ratio constraints imposed by the rest of the current specification; this column is provided to aid this explanation and does not form part of this specification.
[0087] Potential issues to be resolved
[0088] There are several potential problems with current level specifications, which may be addressed, at least in part, by the example embodiments discussed below.
[0089] First, by definition, it's not worth creating a new level between different existing levels. For example, level 5.2 and level 6 have the same MaxLumaSr value. The only possible change would be to have some intermediate picture sizes, but that's not particularly useful.
[0090] Secondly, it is not possible to exceed 120fps (frames per second) for a given level's picture size. The only option is to use a level with a larger picture size. This is because a change in frame rate beyond 120fps would cause the maximum luma sampling rate to exceed the sampling rate applicable to the next level. For example, a new level (e.g., 5.3) cannot be defined for 4K at 240fps because this would require twice the sampling rate of the current level 6, and the current requirement is that a level 6 decoder must decode a level 5.X stream.
[0091] Third, the maximum level is 8.5, which is undefined in the current specification but implies (via the existing exponential relationship for levels above 4) that the maximum image size is at least 32K (e.g., 32768x17408). While this doesn't seem to be an immediate issue, it could cause problems for future omnidirectional applications using sub-images, for example.
[0092] In order to solve these problems, a method and apparatus for the encoding (and decoding) level are proposed. In addition, alternative constraints on the ability to decode the stream are proposed.
[0093] Level Encoding in Example Embodiments
[0094] An alternative level encoding is proposed as follows:
[0095] The level is encoded in 8 bits as major*16+minor.
[0096] Here, "major" refers to the integer value and "minor" refers to the value after the decimal point, so for (say) level 5.2, major=5 and minor=2.
[0097] So, for example, level 4.1 is encoded as 65, and the maximum level is 15.15.
[0098] This arrangement means that level codes that cannot be defined for the reasons mentioned above are not "wasted" from the level set, or at least not within the scope of the current system.
[0099] Furthermore, this in turn extends the highest possible value from 8.5 to 15.15.
[0100] On the other hand, the alternative encoding also makes the level code easily visible in the hex dump representation, since the first hex digit will represent the major component and the second the minor component.
[0101] Note that an alternative example is given below. Any one of these examples (or indeed other examples) may be considered within the scope of the appended claims.
[0102] Constraints on level decoding
[0103] To address the issue that higher frame rates are not currently possible without using higher levels (and thus forcing the use of longer line buffers, etc.), the exemplary embodiments change the constraints on the ability to decode the stream.
[0104] The alternative constraint definition is that a decoder conforming to a given profile at a particular level Dd (where D denotes the primary component and d denotes the secondary component) of a given layer shall be able to decode all bitstreams (of a given profile of a given layer) for which the following condition applies, i.e., the bitstream is indicated as conforming to level Ss(primary, secondary) that is not level 15.15, and S is less than or equal to D, and S*2+s is less than or equal to D*2+d.
[0105] Using these arrangements, new secondary levels for higher frame rates can be added without requiring the decoder for the next primary level to be able to decode the higher sampling rate.
[0106] This constraint is fully backwards compatible with all currently defined levels. This constraint also does not affect any decoders. However, it now allows the generation of high frame rate decoders at lower levels than is currently possible.
[0107] Example – Figure 10
[0108] Figure 10 An example is provided in Figure 9 A new level 5.3 is inserted into the list with the same maximum luminance picture size as the other main level = 5 levels, but the maximum luminance sampling rate and maximum frame rate are twice that of the previous level 5.2.
[0109] Applying the above representation, the new level 5.3 will occupy (16*5)+3=83 unused codes.
[0110] Applying the above decoding constraints, a level 5.3 video data stream will be decodable by a level 6.1 decoder because:
[0111] 5 is less than or equal to 6, and
[0112] 5*2+3 is less than or equal to 6*2+1.
[0113] Therefore, the example embodiments provide an alternative method for encoding the proposed levels that increases the maximum level from 8.5 to 15.15, along with an alternative constraint that allows levels with higher frame rates to be added without requiring all higher level decoders to be able to decode the increased bit rate.
[0114] Example Embodiments
[0115] Example embodiments will now be described with reference to the accompanying drawings.
[0116] Figure 11 The use of video parameter sets and sequence parameter sets is schematically illustrated, as described above. Specifically, these form part of the aforementioned hierarchy of parameter sets, such that multiple sequence parameter sets 1100, 1110, 1120 can reference a video parameter set 1130, which itself is in turn referenced by corresponding sequences 1102, 1112, 1122. In an exemplary embodiment, level information applicable to the corresponding sequence is provided in the sequence parameter set.
[0117] However, in other embodiments, it will be appreciated that the level information may be provided in a different form or a different set of parameters.
[0118] Similarly, although Figure 11 The diagram of 1140 shows a sequence parameter set provided as part of the entire video data stream 1140, but the sequence parameter set (or other data structure carrying level information) can be provided by a separate communication channel instead. In either case, the level information is associated with the video data stream 1140.
[0119] Example Operation - Decoder
[0120] Figure 12 Schematically illustrating aspects of a decoding device configured to receive an input (encoded) video data stream 1200 and using the above reference Figure 7 The decoder 1220 of the discussed form generates and outputs a decoded video data stream 1210. For clarity of this explanation, Figure 7 The control circuit or controller 343 is drawn separately from the rest of the decoder 1220.
[0121] Within the functionality of the controller or control circuit 343 is a parameter set (PS) detector 1230 that detects the various parameter sets, including the VPS, SPS, and PPS, from the appropriate fields of the input video data stream 1200. The parameter set detector 1230 derives information from the parameter sets, including the levels, as described above. This information is passed to the rest of the control circuit 343. Note that the parameter set detector 1230 can decode the levels, or can simply provide the encoded levels to the control circuit 344 for decoding.
[0122] The control circuit 343 is also responsive to one or more decoder parameters 1240, which define, for example, at least the levels (major, minor) that the decoder 1220 is capable of decoding using the numbering scheme described above.
[0123] The control circuit 343 detects whether the decoder 1220 is capable of decoding a given or current input video data stream 1200 and controls the decoder 1220 accordingly. The control circuit 343 may also provide various other operating parameters to the decoder 1220 in response to information obtained from the parameter set detected by the parameter set detector 1230.
[0124] Figure 13 is a schematic flow chart illustrating these operations on the decoder side.
[0125] At step 1300, the parameter set decoder 1230 detects the SPS and provides this information to the control circuit 343. The control circuit 343 also detects the decoder parameters 1240.
[0126] Based on the encoding level, the control circuit 343 detects (at step 1310) the level modulo N, where N is a first predetermined constant (16 in this example), and detects at step 1320 the remainder of the level divided by M, where N is a second predetermined constant (1 in this example, although it could be, for example, 2 or 4, or 3 in another example). The result of step 1310 provides the primary component, while the result of step 1320 provides the secondary component.
[0127] Then, at step 1330, the control circuit 343 checks whether the current input video data stream 1200 is decodable by applying the following test as described above to the decoder parameter value Dd and the input video stream parameter value Ss:
[0128] Is the encoding level = 255 (Ss stands for 15.15)? If yes, then it is not decodable because 15.15 is a special case indicating a non-standard level. If no, then pass the next test:
[0129] Is S less than or equal to D and is S*2+s less than or equal to D*2+d? If yes, then it is decodable; if no, then it is not decodable
[0130] The first part of this test can optionally be omitted.
[0131] If the answer is no, control passes to step 1340 where control circuit 343 instructs decoder 1220 to not decode the input video data stream.
[0132] However, if the answer is yes, then control passes to steps 1350 and 1360 where the maximum luma picture size (step 1350) and the maximum luma sampling rate and / or maximum frame rate (step 1360) are calculated using the values from the stored table (e.g., Figure 10 Based on these derived parameters, the control circuit 343 controls the decoder 1220 in step 1370 to decode the input video data stream 1200.
[0133] Thus, this provides an example of a method of operating a video data decoder, the method comprising:
[0134] detecting (1300) a parameter value associated with an input video data stream, the parameter value indicating a coding level selected from among a plurality of coding levels, each coding level defining at least a maximum luma picture size and a maximum luma sampling rate;
[0135] The encoding level defines a first numerical component and a second numerical component, the second numerical component being a numerical value greater than or equal to zero; wherein:
[0136] For coding levels where the second numerical component is zero, the first numerical component increases monotonically with increasing maximum luminance picture size; and
[0137] The second component varies with the maximum brightness sampling rate;
[0138] The parameter value is a numerical code of the coding level, which is: the sum of a first predetermined constant multiplied by a first numerical component and a second predetermined constant multiplied by a second numerical component;
[0139] performing (1330) predetermined testing of parameters associated with a given input video data stream relative to capability data of a video data decoder;
[0140] controlling (1370) the video data decoder to decode the given input video data stream when the parameter value associated with the given input video data stream passes a predetermined test relative to capability data of the video data decoder; and
[0141] When a parameter value associated with a given input video data stream fails a predetermined test with respect to capability data of the video data decoder, controlling (1340) the video data decoder to not decode the given input video stream.
[0142] according to Figure 13 The method of operation Figure 12 The arrangement provides an example of a device including:
[0143] a video data decoder 1220 configured to decode an input video data stream, the video data decoder being responsive to a parameter value associated with the input video data stream, the parameter value indicating a coding level selected from a plurality of coding levels, each coding level defining at least a maximum luma picture size and a maximum luma sampling rate;
[0144] The encoding level defines a first numerical component and a second numerical component, the second numerical component being a numerical value greater than or equal to zero; wherein:
[0145] For coding levels where the second numerical component is zero, the first numerical component increases monotonically with increasing maximum luminance picture size; and
[0146] The second component varies with the maximum brightness sampling rate;
[0147] The parameter value is a numerical code of the coding level, which is: a first predetermined constant multiplied by the first numerical component plus a second predetermined constant multiplied by the second numerical component;
[0148] a comparator 343 configured to perform a predetermined test of parameter values associated with a given input video data stream relative to capability data of a video data decoder; and
[0149] The control circuit 343 is configured to control the video data decoder to decode the given input video stream when the parameter value associated with the given input video data stream passes a predetermined test relative to the capability data of the video data decoder, and to control the video data decoder not to decode the given input video stream when the parameter value associated with the given input video data stream does not pass the predetermined test relative to the capability data of the video data decoder.
[0150] Example Operation - Encoder
[0151] In a similar way, for example, Figure 14 Schematically shows the above reference Figure 7Various aspects of an encoding device of the type discussed are described with reference to an encoder 1400. For clarity of explanation, the encoder's control circuitry 343 is depicted separately. The encoder operates on an input video data stream 1410 to generate an output encoded video data stream 1420 under the control of the control circuitry 343, which in turn responds to encoding parameters 1430 including a definition of the encoding level to be applied.
[0152] The control circuit 343 also comprises or controls a parameter set generator 1440 which generates a parameter set comprising, for example, a VP, an SPS and a PPS to be included in the output encoded video data stream, wherein the SPS carries the level information encoded as described above.
[0153] Now refer to Figure 15 The schematic flow charts describe the operational aspects of the device.
[0154] Step 1500 represents establishing the encoding level, for example by means of encoding parameters 1430, the encoding level being represented by (primary, secondary) or in other words by (first component, second component).
[0155] In step 1510 , the control circuit 343 controls the encoder 1400 according to the established encoding level.
[0156] Separately, to encode the encoding level, the parameter set generator 1440 multiplies the first component (primary component) by a first predetermined constant N (16 in this example) at step 1520, and multiplies the second component (secondary component) by a second predetermined constant M (1 in this example, but 3 in other examples discussed below) at step 1530, and then adds the two results to generate encoding level information at step 1540. Thus, the parameter set generator generates the required parameter set, including the encoding level information, at step 1550.
[0157] This therefore provides an example of a method that includes:
[0158] encoding (in response to 1510) an input video data stream to generate an output coded video data stream according to a coding level selected from among a plurality of coding levels, each coding level defining at least a maximum luma picture size and a maximum luma sampling rate, wherein the coding level defines a first numerical component and a second numerical component, the second numerical component being a numerical value greater than or equal to zero; wherein:
[0159] For coding levels where the second numerical component is zero, the first numerical component increases monotonically with increasing maximum luminance picture size; and
[0160] The second component varies with the maximum luminance sampling rate; and
[0161] Parameter values associated with the output coded video data stream are encoded (1520, 1530, 1540, 1550), the parameter values being numerical encodings of a coding level of a first predetermined constant multiplied by a first numerical component plus a second predetermined constant multiplied by a second numerical component.
[0162] according to Figure 15 The method operates Figure 14 The arrangement provides an example of a device including:
[0163] A video data encoder 1400 is configured to encode an input video data stream according to a coding level selected from a plurality of coding levels to generate an output coded video data stream, each coding level defining at least a maximum luminance picture size and a maximum luminance sampling rate, wherein the coding level defines a first numerical component and a second numerical component, the second numerical component being a numerical value greater than or equal to zero; wherein:
[0164] For coding levels where the second numerical component is zero, the first numerical component increases monotonically with increasing maximum luminance picture size; and
[0165] The second component varies with the maximum luminance sampling rate; and
[0166] The parameter value encoding circuit 1440 is configured to encode a parameter value associated with the output encoded video data stream, wherein the parameter value is a numerical encoding of the encoding level, which is: a first predetermined constant multiplied by a first numerical component plus a second predetermined constant multiplied by a second numerical component.
[0167] In the above encoding or decoding examples, the second component may increase monotonically with the maximum luminance sampling rate. In other examples, for a given first numerical component, the second component may vary with the maximum luminance sampling rate. In other examples, for a given first numerical component, the second component may increase monotonically with the maximum luminance sampling rate.
[0168] In other examples, for a threshold encoding level having at least a first component: the first component increases monotonically with increasing maximum luma picture size; and the second component indicates at least one of the maximum luma sampling rates, wherein, for a given first component, the second component varies with the maximum luma sampling rate.
[0169] A second numerical component of zero is indicated in the appendix by the typeset lacking the second digit (after the decimal point).
[0170] Regarding the text "For coding levels where the second numerical component is zero, the first numerical component increases monotonically with increasing maximum luminance picture size", this indicates that when m>n, the maximum luminance picture size of level m.0 (or "m") is at least the same as the maximum luminance picture size of level n.0 ("or "n").
[0171] Additional Examples
[0172] An alternative level encoding is proposed as follows:
[0173] The level is encoded in 8 bits as major*16 + minor*3.
[0174] Here, as mentioned above, "major" refers to the integer value and "minor" refers to the value after the decimal point, so for (say) level 5.2, major=5 and minor=2.
[0175] So, for example, level 4.1 is encoded as 67, and the maximum level that can be encoded by this technique is 15.5.
[0176] As mentioned above, this arrangement means that level sets are not "wasted" due to level codes that cannot be defined for the reasons mentioned above, or at least not to the extent of the current system.
[0177] Furthermore, this in turn extends the highest possible value from 8.5 to 15.5.
[0178] An example of encoding levels using this arrangement is as follows:
[0179]
[0180]
[0181] Encoded video data
[0182] Video data encoded by any of the techniques disclosed herein is also considered to represent embodiments of the present disclosure.
[0183] Appendix - JVET-Q2001-VES Draft Specification Changes to Reflect Examples
[0184] [A.3.1 Main 10 configuration files]
[0185] A bitstream conforming to the Main 10 profile shall adhere to the following constraints:
[0186] – The chroma_format_idc of the referenced SPS shall be equal to 0 or 1.
[0187] – The bit_depth_minus8 of the referenced SPS shall be in the range 0 to 2 (inclusive).
[0188] – The sps_palette_enabled_flag of the referenced SPS shall be equal to 0.
[0189] – general_level_idc and sublayer_level_idc[i] for all values of i in the VPS (if applicable) and the referenced SPS shall not be equal to 255 (indicating level 15.15 ).
[0190] – The layer and level constraints specified in the Main 10 profile in Section A.4 (if applicable) shall be met.
[0191] Conformance of the bitstream to the Main 10 profile is indicated by general_profile_idc equal to 1.
[0192] A decoder compliant with a specific Tier 10 profile at a specific level, a specific Layer Dd SHOULD be able to decode all bitstreams for which all of the following conditions apply:
[0193] – Indicates that the bitstream conforms to the Main 10 profile.
[0194] – Indicates that the bitstream conforms to a layer lower than or equal to the specified layer.
[0195] – Indicates that the bitstream conforms to a level other than 15.15 level Ss , and less than or equal to D and S*2+s is less than or equal to D*2 +d .
[0196] [A.3.2 Main 4:4:4 10 Profile]
[0197] Bitstreams conforming to the Main 4:4:4 10 profile shall adhere to the following constraints:
[0198] – The chroma_format_idc of the referenced SPS shall be in the range 0 to 3 (inclusive).
[0199] – The bit_depth_minus8 of the referenced SPS shall be in the range 0 to 2 (inclusive).
[0200] – general_level_idc and sublayer_level_idc[i] shall not be equal to 255 (meaning 15.15) for all values of i in the VPS (if applicable) and the reference SPS.
[0201] – The layer and level constraints specified in the main 4:4:4 10 profile in clause A.4 shall be satisfied (if applicable).
[0202] Conformance of the bitstream to the Main 4:4:4 10 profile is indicated by general_profile_idc equal to 2.
[0203] Comply with specific layers Dd A decoder for the Main 4:4:4 10 profile at a particular level shall be able to decode all bitstreams where all of the following conditions apply:
[0204] – Indicates that the bitstream conforms to the Main 4:4:4 10 or Main 10 profile.
[0205] – Indicates that the bitstream conforms to a layer lower than or equal to the specified layer.
[0206] – Indicates that the bitstream conforms to a level other than 15.15 Level Ss , and less than or equal to D and S*2+s is less than or equal to D*2 +d .
[0207] [A.4.1 General layer and level restrictions]
[0208] Table A.1 specifies the level of 15.15 The limit value for each level of each layer except .
[0209] The tier and level to which the bitstream conforms are indicated by the syntax elements general_tier_flag and general_level_idc, and the level to which the sublayer conforms is indicated by the syntax element sublayer_level_idc[i], as shown below:
[0210] – If the specified level is not a level 15.15 , then according to the level constraints specified in Table A.1, a general_tier_flag equal to 0 indicates compliance with the main layer, a general_tier_flag equal to 1 indicates compliance with the upper layer, and for levels lower than level 4 (corresponding to the entries marked with "-" in Table A.2), general_tier_flag shall be equal to 0. Otherwise (the specified level is 15.15 The bitstream conformance requirement general_tier_flag shall be equal to 1. The value of general_tier_flag 0 is reserved for future use by ITU-T | ISO / IEC. The decoder shall ignore the value of general_tier_flag.
[0211] – The values of general_level_idc and sublayer_level-idc[i] shall be set equal to those specified in Table A.1 16 times the major level number plus 1 times the minor level number .
[0212] All other references to level 8.5 were also changed to level 15.15.
[0213] As an alternative, the 15.15 designation can be changed to use variable names (such as "limitless_level_idc") to specify levels throughout the procedure, with that variable defined as 15.15.
[0214] With reference to the above alternative embodiments, the following refers to JVET-T2001-v1 Appendix 4.1 of the 20th JVET meeting in October 2020 (the contents of which are incorporated herein by reference):
[0215] The tier and level to which the bitstream conforms are indicated by the syntax elements general_tier_flag and general_level_idc, and the level to which the sublayer conforms is indicated by the syntax element sublayer_level_idc[i], as shown below:
[0216] – If the specified level is not level 15.5, general_tier_flag equal to 0 indicates conformance to the main tier and general_tier_flag equal to 1 indicates conformance to the upper tier, according to the level constraints specified in Table 135. For levels lower than level 4 (corresponding to entries marked with "-" in Table 135), general_tier_flag shall be equal to 0. Otherwise (the specified level is level 15.5), the bitstream conformance requirement is that general_tier_flag shall be equal to 1. The value of general_tier_flag 0 is reserved for future use by ITU-T | ISO / IEC and decoders shall ignore the value of general_tier_flag.
[0217] – For the level numbers specified in Table 135, general_level_idc and sublayer_level-idc[i] shall be set equal to the value of general_level_idc.
[0218] (Note that Table 135 is reproduced above in connection with the alternative embodiment).
[0219] To the extent that embodiments of the present disclosure have been described as being implemented at least in part by a data processing device controlled by software, it should be understood that a non-transitory machine-readable medium carrying such software, such as an optical disc, magnetic disk, semiconductor memory, etc., is also considered to represent an embodiment of the present disclosure. Similarly, a data signal (whether or not embodied on a non-transitory machine-readable medium) including encoded data generated according to the above-described method is also considered to represent an embodiment of the present disclosure.
[0220] Obviously, many modifications and variations of the present disclosure are possible in light of the above teachings. It should therefore be understood that, within the scope of the appended clauses, the present technology may be practiced otherwise than as specifically described herein.
[0221] It will be appreciated that for clarity, the above description has described embodiments with reference to different functional units, circuits and / or processors. However, it will be apparent that any suitable distribution of functionality between different functional units, circuits and / or processors may be used without departing from the embodiments.
[0222] The described embodiments may be implemented in any suitable form, including hardware, software, firmware, or any combination thereof. The described embodiments may optionally be implemented at least in part as computer software running on one or more data processors and / or digital signal processors. The elements and components of any embodiment may be implemented physically, functionally, and logically in any suitable manner. In practice, the functionality may be implemented in a single unit, multiple units, or as part of other functional units. Therefore, the disclosed embodiments may be implemented in a single unit, or may be physically and functionally distributed between different units, circuits, and / or processors.
[0223] Although the present disclosure has been described in conjunction with some embodiments, the present disclosure is not limited to the specific forms described herein. In addition, although features may be described in conjunction with specific embodiments, those skilled in the art will recognize that the various features of the described embodiments may be combined in any manner suitable for implementing the technology.
[0224] The various aspects and features are defined by the following numbered clauses:
[0225] 1. A device comprising:
[0226] a video data decoder configured to decode an input video data stream, the video data decoder being responsive to a parameter value associated with the input video data stream, the parameter value indicating a coding level selected from a plurality of coding levels, each coding level defining at least a maximum luminance picture size and a maximum luminance sampling rate;
[0227] The encoding level defines a first numerical component and a second numerical component, the second numerical component being a numerical value greater than or equal to zero; wherein:
[0228] For coding levels where the second numerical component is zero, the first numerical component increases monotonically with increasing maximum luminance picture size; and
[0229] The second component varies with the maximum brightness sampling rate;
[0230] The parameter value is a numerical code of the coding level, which is: the sum of a first predetermined constant multiplied by a first numerical component and a second predetermined constant multiplied by a second numerical component;
[0231] a comparator configured to perform a predetermined test of a parameter value associated with a given input video data stream relative to capability data of a video data decoder; and
[0232] The control circuit is configured to control the video data decoder to decode the given input video stream when the parameter value associated with the given input video data stream passes a predetermined test with respect to the capability data of the video data decoder, and to control the video data decoder not to decode the given input video stream when the parameter value associated with the given input video data stream fails to pass the predetermined test with respect to the capability data of the video data decoder.
[0233] 2. Apparatus according to clause 1, wherein the predetermined test comprises performing the following checks on a first numerical component (S) and a second numerical component (s) represented by a parameter value associated with a given input video data stream, and on a first numerical component (D) and a second numerical component (d) of capability data of a video data decoder:
[0234] Whether (i) S is less than or equal to D, and (ii) S*2+s is less than or equal to D*2+d.
[0235] 3. Apparatus according to clause 1 or clause 2, comprising a detector configured to detect parameter values from a parameter set associated with an input video data stream.
[0236] 4. Apparatus according to clause 3, wherein the parameter set is a sequence parameter set.
[0237] 5. The apparatus according to any of the preceding clauses, wherein the first predetermined constant is 16.
[0238] 6. The apparatus according to any of the preceding clauses, wherein the second predetermined constant is 1.
[0239] 7. The apparatus according to any one of clauses 1 to 5, wherein the second predetermined constant is 3.
[0240] 8. Apparatus according to any preceding clause, wherein the parameter value comprises an 8-bit value.
[0241] 9. Apparatus according to any of the preceding clauses, wherein the predetermined test comprises detecting whether a parameter value associated with the input video stream is not equal to 255.
[0242] 10. A video storage, capture, transmission or reception device comprising a device according to any of the preceding clauses.
[0243] 11. A device comprising:
[0244] A video data encoder configured to encode an input video data stream according to a coding level selected from a plurality of coding levels to generate an output coded video data stream, each coding level defining at least a maximum luminance picture size and a maximum luminance sampling rate, wherein the coding level defines a first numerical component and a second numerical component, the second numerical component being a numerical value greater than or equal to zero; wherein:
[0245] For coding levels where the second numerical component is zero, the first numerical component increases monotonically with increasing maximum luminance picture size; and
[0246] The second component varies with the maximum luminance sampling rate; and
[0247] The parameter value encoding circuit is configured to encode a parameter value associated with the output encoded video data stream, wherein the parameter value is a numerical encoding of an encoding level, which is: a first predetermined constant multiplied by a first numerical component plus a second predetermined constant multiplied by a second numerical component.
[0248] 12. Apparatus according to clause 11, wherein the parameter value encoding circuit is configured to encode the parameter value as at least part of a parameter set associated with the output encoded video data stream.
[0249] 13. Apparatus according to clause 12, wherein the parameter set is a sequence parameter set.
[0250] 14. The apparatus of any one of clauses 11 to 13, wherein the first predetermined constant is 16.
[0251] 15. The apparatus of any one of clauses 11 to 14, wherein the second predetermined constant is 1.
[0252] 16. The apparatus of any one of clauses 11 to 14, wherein the second predetermined constant is 1.
[0253] 17. Apparatus according to any of clauses 11 to 16, wherein the parameter value comprises an 8-bit value.
[0254] 18. A video storage, capture, transmission or reception device comprising a device according to any one of clauses 11 to 17.
[0255] 19. A method of operating a video data decoder, the method comprising:
[0256] detecting a parameter value associated with an input video data stream, the parameter value indicating a coding level selected from a plurality of coding levels, each coding level defining at least a maximum luminance picture size and a maximum luminance sampling rate;
[0257] The encoding level defines a first numerical component and a second numerical component, the second numerical component being a numerical value greater than or equal to zero; wherein:
[0258] For coding levels where the second numerical component is zero, the first numerical component increases monotonically with increasing maximum luminance picture size; and
[0259] The second component varies with the maximum brightness sampling rate;
[0260] The parameter value is a numerical code of the coding level, which is: the sum of a first predetermined constant multiplied by a first numerical component and a second predetermined constant multiplied by a second numerical component;
[0261] performing predetermined testing of parameter values associated with a given input video data stream relative to capability data of a video data decoder;
[0262] controlling the video data decoder to decode the given input video data stream when a parameter value associated with the given input video data stream passes a predetermined test relative to capability data of the video data decoder; and
[0263] When a parameter value associated with a given input video data stream fails a predetermined test with respect to capability data of the video data decoder, the video data decoder is controlled not to decode the given input video stream.
[0264] 20. A machine-readable non-transitory storage medium storing computer software which, when executed by a computer, causes the computer to perform the method of clause 19.
[0265] 21. A method comprising:
[0266] Encoding an input video data stream according to a coding level selected from a plurality of coding levels to generate an output coded video data stream, each coding level defining at least a maximum luminance picture size and a maximum luminance sampling rate, wherein the coding level defines a first numerical component and a second numerical component, the second numerical component being a numerical value greater than or equal to zero; wherein:
[0267] For coding levels where the second numerical component is zero, the first numerical component increases monotonically with increasing maximum luminance picture size; and
[0268] The second component varies with the maximum luminance sampling rate; and
[0269] A parameter value associated with the output coded video data stream is encoded, where the parameter value is a numerical encoding of the coding level, which is: a first predetermined constant multiplied by a first numerical component plus a second predetermined constant multiplied by a second numerical component.
[0270] 22. A machine-readable non-transitory storage medium storing computer software which, when executed by a computer, causes the computer to perform the method of clause 21.
Claims
1. A device for decoding video data, comprising: a video data decoder configured to decode an input video data stream, the video data decoder being responsive to a parameter value associated with the input video data stream, the parameter value indicating a coding level selected from a plurality of coding levels, each coding level defining at least a maximum luma picture size and a maximum luma sampling rate; The encoding level defines a first numerical component and a second numerical component, the second numerical component being a numerical value greater than or equal to zero; wherein: For coding levels where the second numerical component is zero, the first numerical component increases monotonically with increasing maximum luminance picture size; and The second component varies with the maximum brightness sampling rate; The parameter value is a numerical code of the coding level, which is: a first predetermined constant multiplied by the first numerical component plus a second predetermined constant multiplied by the second numerical component, wherein the first predetermined constant is 16; a comparator configured to perform a predetermined test of said parameter value associated with a given said input video data stream relative to capability data of said video data decoder; and The control circuit is configured to control the video data decoder to decode the given input video stream when the parameter value associated with the given input video data stream passes the predetermined test relative to the capability data of the video data decoder, and to control the video data decoder not to decode the given input video stream when the parameter value associated with the given input video data stream does not pass the predetermined test relative to the capability data of the video data decoder.
2. The device according to claim 1, wherein The predetermined test comprises performing the following checks on the first numerical component (S) and the second numerical component (s) represented by the parameter value associated with a given input video data stream, and on the first numerical component (D) and the second numerical component (d) of the capability data of the video data decoder: Whether (i) S is less than or equal to D, and (ii) S*2+s is less than or equal to D*2+d. 3 . The apparatus of claim 1 , comprising a detector configured to detect the parameter value from a parameter set associated with the input video data stream.
4. The device according to claim 3, wherein The parameter set is a sequence parameter set.
5. The apparatus according to claim 1, wherein The second predetermined constant is 1.
6. The apparatus according to claim 1, wherein The second predetermined constant is 3.
7. The apparatus according to claim 1, wherein The parameter value comprises an 8-bit value.
8. The apparatus according to claim 1, wherein The predetermined test comprises detecting whether the parameter value associated with the input video stream is not equal to 255.
9. A video storage, capture, transmission or reception device comprising the device according to claim 1.
10. A device for encoding video data, comprising: A video data encoder configured to encode an input video data stream according to a coding level selected from a plurality of coding levels to generate an output coded video data stream, each coding level defining at least a maximum luminance picture size and a maximum luminance sampling rate, wherein the coding level defines a first numerical component and a second numerical component, the second numerical component being a numerical value greater than or equal to zero; wherein: For coding levels where the second numerical component is zero, the first numerical component increases monotonically with increasing maximum luminance picture size; and The second component varies with the maximum luminance sampling rate; and A parameter value encoding circuit is configured to encode a parameter value associated with the output encoded video data stream, wherein the parameter value is a numerical encoding of an encoding level, which is: a first predetermined constant multiplied by the first numerical component plus a second predetermined constant multiplied by the second numerical component, wherein the first predetermined constant is 16.
11. The apparatus according to claim 10, wherein The parameter value encoding circuit is configured to encode the parameter value as at least part of a parameter set associated with the output encoded video data stream.
12. The apparatus according to claim 11, wherein The parameter set is a sequence parameter set.
13. The apparatus according to claim 10, wherein The second predetermined constant is 1.
14. The apparatus according to claim 10, wherein The second predetermined constant is 3.
15. The apparatus according to claim 10, wherein The parameter value comprises an 8-bit value.
16. A video storage, capture, transmission or reception device comprising the device according to claim 11.
17. A method of operating a video data decoder, the method comprising: detecting a parameter value associated with an input video data stream, the parameter value indicating a coding level selected from a plurality of coding levels, each coding level defining at least a maximum luma picture size and a maximum luma sampling rate; The encoding level defines a first numerical component and a second numerical component, the second numerical component being a numerical value greater than or equal to zero; wherein: For coding levels where the second numerical component is zero, the first numerical component increases monotonically with increasing maximum luminance picture size; and The second component varies with the maximum brightness sampling rate; The parameter value is a numerical code of the coding level, which is: a first predetermined constant multiplied by the first numerical component plus a second predetermined constant multiplied by the second numerical component, wherein the first predetermined constant is 16; performing predetermined testing of said parameter values associated with a given input video data stream relative to capability data of said video data decoder; controlling the video data decoder to decode the given input video data stream when the parameter value associated with the given input video data stream passes the predetermined test relative to the capability data of the video data decoder; and When the parameter value associated with a given input video data stream fails the predetermined test relative to the capability data of the video data decoder, the video data decoder is controlled not to decode the given input video stream.
18. A machine-readable non-transitory storage medium storing computer software which, when executed by a computer, causes the computer to perform the method according to claim 17.
19. A method for encoding video data, comprising: Encoding an input video data stream according to a coding level selected from a plurality of coding levels to generate an output coded video data stream, each coding level defining at least a maximum luminance picture size and a maximum luminance sampling rate, wherein the coding level defines a first numerical component and a second numerical component, the second numerical component being a numerical value greater than or equal to zero; wherein: For coding levels where the second numerical component is zero, the first numerical component increases monotonically with increasing maximum luminance picture size; and The second component varies with the maximum luminance sampling rate; and A parameter value associated with the output coded video data stream is encoded, wherein the parameter value is a numerical encoding of the coding level, which is: a first predetermined constant multiplied by the first numerical component plus a second predetermined constant multiplied by the second numerical component.
20. A machine-readable non-transitory storage medium storing computer software which, when executed by a computer, causes the computer to perform the method according to claim 19.
Citation Information
Patent Citations
Image decoding device
CN104620575A
Signaling decoder picture buffer information
US20140092961A1