Image data encoding and decoding

By using an entropy encoder and an attribute detector to selectively adjust coding constraints in video data encoding, the problems of insufficient video data encoding efficiency and quality in the prior art are solved, and efficient video data compression and decoding effects are achieved.

CN114788276BActive Publication Date: 2025-09-19SONY GROUP CORP
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202080066130.9
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Priority Date
2019-09-24
Filing Date
2020-09-23
Publication Date
2025-09-19
Estimated Expiration
2040-09-23

Smart Images

  • Figure CN114788276B_ABST
    Figure CN114788276B_ABST
Patent Text Reader

Abstract

An image data encoding device comprises: an entropy encoder configured to selectively encode data items representing image data so as to generate encoded binary symbols of consecutive output data units; the entropy encoder configured to generate a constrained output data stream, the constraint defining an upper limit on the number of binary symbols that can be represented by any individual output data unit relative to the byte size of the output data unit, wherein the entropy encoder is configured to provide padding data for each output data unit that does not satisfy the constraint so as to increase the byte size of the output data unit so as to satisfy the constraint; the device comprises: an attribute detector configured to detect an encoding attribute applicable to a given output data unit; and a selector configured to select a constraint for the given output data unit from two or more candidate constraints in response to the detected encoding attribute.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to image data encoding and decoding. Background Art

[0002] The "background" description provided herein is for the purpose of generally presenting the context of the present disclosure. To the extent described in this background section, the work of the presently named inventors and any description that may not qualify as prior art at the time of filing are neither explicitly nor implicitly admitted to be prior art with respect to the present disclosure.

[0003] Several video data encoding and decoding systems involve converting the video data into a frequency domain representation, quantizing the frequency domain coefficients, and then applying some form of entropy coding to the quantized coefficients. This allows for compression of the video data. A corresponding decoding or decompression technique is then applied to recover a reconstructed version of the original video data.

[0004] High Efficiency Video Coding (HEVC), also known as H.265 or MPEG-H Part 2, is the proposed successor to H.264 / MPEG-4 AVC. HEVC aims to improve video quality and double the data compression ratio compared to H.264, and is designed to be scalable from 128×96 to 7680×4320 pixel resolution, roughly equivalent to bit rates of 128 kbit / s to 800 Mbit / s. Summary of the Invention

[0005] The present disclosure solves or alleviates the problems caused by this process.

[0006] The present disclosure provides an image data encoding device, comprising:

[0007] an entropy encoder configured to selectively encode data items representing image data so as to generate encoded binary symbols of successive output data units;

[0008] An entropy encoder is configured to generate an output data stream subject to a constraint, the constraint defining an upper limit on the number of binary symbols that can be represented by any individual output data unit relative to the byte size of the output data unit, wherein the entropy encoder is configured to provide padding data for each output data unit that does not satisfy the constraint in order to increase the byte size of the output data unit so as to satisfy the constraint; the apparatus comprising:

[0009] a property detector configured to detect encoding properties applicable to a given output data unit; and

[0010] A selector is configured to select a constraint for a given output data unit from two or more candidate constraints in response to the detected encoding properties.

[0011] The present disclosure also provides an image data encoding method, comprising:

[0012] selectively encoding data items representing image data to generate encoded binary symbols of successive output data units;

[0013] generating a constrained output data stream, the constraint defining an upper limit on the number of binary symbols that can be represented by any individual output data unit relative to a byte size of the output data unit;

[0014] providing padding data for each output data unit that does not satisfy the constraint to increase the byte size of the output data unit so that the constraint is satisfied;

[0015] detecting encoding properties applicable to a given output data unit; and

[0016] In response to the detected encoding properties, a constraint for a given output data unit is selected from two or more candidate constraints.

[0017] The present disclosure also provides a suitable decoding device for decoding the data signal generated by the method or device defined above.

[0018] Further respective aspects and features of the disclosure are defined in the appended claims.

[0019] It is to be understood that both the foregoing general description and the following detailed description are exemplary and not restrictive of the present technology. BRIEF DESCRIPTION OF THE DRAWINGS

[0020] A more complete understanding of the present disclosure and many of its attendant advantages will be readily obtained as the present disclosure and many of its attendant advantages become better understood by reference to the following detailed description when considered in conjunction with the accompanying drawings, in which:

[0021] Figure 1 Schematically illustrates an audio / video (A / V) data transmission and reception system using video data compression and decompression;

[0022] Figure 2 schematically illustrates a video display system using video data decompression;

[0023] Figure 3 Schematically illustrates an audio / video storage system using video data compression and decompression;

[0024] Figure 4 schematically illustrates a video camera using video data compression;

[0025] Figure 5 and Figure 6 A storage medium is schematically shown;

[0026] Figure 7 Provides a schematic diagram of a video data compression and decompression device;

[0027] Figure 8 The predictor is shown schematically;

[0028] Figure 9 A partially encoded image is schematically shown;

[0029] Figure 10 A set of possible intra prediction directions is schematically shown;

[0030] Figure 11 A set of prediction modes is schematically shown;

[0031] Figure 12 Another set of prediction modes is schematically shown;

[0032] Figure 13 The intra prediction process is schematically shown;

[0033] Figure 14 A CABAC encoder is schematically shown;

[0034] Figure 15 and Figure 16 Schematically illustrates the CABAC coding technique;

[0035] Figure 17 and Figure 18 Schematically illustrates the CABAC decoding technique;

[0036] Figure 19 The partition image is schematically shown;

[0037] Figure 20 The device is schematically shown;

[0038] Figure 21 The controller is schematically shown;

[0039] Figure 22A and Figure 22B schematically represents an output data unit; and

[0040] Figure 23 is a schematic flow chart illustrating the method. DETAILED DESCRIPTION

[0041] Referring now to the accompanying drawings, there is provided Figures 1 to 4 A schematic diagram of a device or system utilizing the compression and / or decompression device described below in conjunction with embodiments of the present technology is provided.

[0042] All data compression and / or decompression devices to be described below may be implemented in hardware, in software running on a general-purpose data processing device (such as a general-purpose computer), as programmable hardware (such as an application-specific integrated circuit (ASIC) or a field-programmable gate array (FPGA)), or a combination of these. Where an embodiment is implemented by software and / or firmware, it should be understood that such software and / or firmware and the non-transitory data storage medium that stores or otherwise provides such software and / or firmware are considered embodiments of the present technology.

[0043] Figure 1 An audio / video data transmission and reception system using video data compression and decompression is schematically shown.

[0044] An input audio / video signal 10 is provided to a video data compression device 20, which compresses at least the video component of the audio / video signal 10 for transmission along a transmission path 30, such as a cable, optical fiber, wireless link, etc. The compressed signal is processed by a decompression device 40 to provide an output audio / video signal 50. For the return path, a compression device 60 compresses the audio / video signal for transmission along the transmission path 30 to a decompression device 70.

[0045] Compression device 20 and decompression device 70 can thus form one node of the transmission link. Decompression device 40 and compression device 60 can form another node of the transmission link. Of course, in the case where the transmission link is unidirectional, only one node requires a compression device, while the other node only requires a decompression device.

[0046] Figure 2 Schematically, a video display system using video data decompression is shown. Specifically, a compressed audio / video signal 100 is processed by a decompression device 110 to provide a decompressed signal that can be displayed on a display 120. The decompression device 110 can be implemented as an integral part of the display 120, for example, being arranged in the same housing as the display device. Alternatively, the decompression device 110 can be arranged as, for example, a so-called set-top box (STB), noting that the expression "set-top" does not mean that the box is required to be located in any particular orientation or position relative to the display 120, but is simply a term used in the art to indicate a device that can be connected to a display as a peripheral device.

[0047] Figure 3An audio / video storage system using video data compression and decompression is schematically shown. An input audio / video signal 130 is provided to a compression device 140, which generates a compressed signal for storage by a storage device 150 (such as a magnetic disk, optical disk, tape, solid-state storage (such as semiconductor memory), or other storage device). For playback, the compressed data is read from the storage device 150 and passed to a decompression device 160 for decompression to provide an output audio / video signal 170.

[0048] It should be understood that a compressed or encoded signal and a storage medium storing the signal (such as a machine-readable non-transitory storage medium) are considered embodiments of the present technology.

[0049] Figure 4 A video camera using video data compression is schematically shown. Figure 4 , an image capture device 180 (such as a charge coupled device (CCD) image sensor and associated control and readout electronics) generates a video signal that is transmitted to a compression device 190. A microphone (or microphones) 200 generates an audio signal that is transmitted to the compression device 190. The compression device 190 generates a compressed audio / video signal 210 (shown generally as schematic stage 220) to be stored and / or transmitted.

[0050] The techniques to be described below primarily relate to video data compression and decompression. It should be understood that many existing techniques can be used for audio data compression in conjunction with the video data compression techniques to be described to generate a compressed audio / video signal. Therefore, a separate discussion of audio data compression will not be provided. It should also be understood that the data rates associated with video data (especially broadcast quality video data) are typically much higher than the data rates associated with audio data (whether compressed or uncompressed). Therefore, it should be understood that uncompressed audio data can accompany compressed video data to form a compressed audio / video signal. It should also be understood that although this example ( Figures 1 to 4 Although the embodiments (shown) involve audio / video data, the techniques described below can find use in systems that simply process (that is, compress, decompress, store, display, and / or transmit) video data. In other words, the embodiments can be applied to video data compression without any associated audio data processing at all.

[0051] therefore, Figure 4 An example of a video capture device comprising an image sensor and an encoding device of the type discussed below is provided. Figure 2 Examples are provided of decoding devices of the types discussed below and displays to which decoded images are output.

[0052] Figure 2and Figure 4 The combination of may provide a video capture device including the image sensor 180 and the encoding device 190 , the decoding device 110 , and the display 120 to which the decoded image is output.

[0053] Figure 5 and Figure 6 A storage medium is schematically shown, storing, for example, compressed data generated by the device 20, 60, compressed data input to the device 110 or the storage medium or stage 150, 220. Figure 5 A disk storage medium such as a magnetic disk or optical disk is schematically shown, and Figure 6 A solid-state storage medium such as flash memory is schematically shown. Note that Figure 5 and Figure 6 Examples of non-transitory machine-readable storage media storing computer software that, when executed by a computer, causes the computer to perform one or more of the methods discussed below may also be provided.

[0054] The above arrangement thus provides an example of a video storage, capture, transmission or reception device embodying any of the present techniques.

[0055] Figure 7 Provides a schematic diagram of a video data compression and decompression device.

[0056] The controller 343 controls the overall operation of the device and, in particular, controls the attempted encoding process by acting as a selector to select various modes of operation, such as block size and shape, and whether the video data is to be losslessly encoded or otherwise encoded, as far as compression modes are concerned. The controller is considered to be part of the image encoder or image decoder, as the case may be. Successive images of the input video signal 300 are provided to the adder 310 and the image predictor 320. Reference will be made below to Figure 8 The image predictor 320 is described in more detail. The image encoder or decoder (as the case may be) is coupled Figure 8 The intra picture predictor can use the Figure 7 However, this does not mean that the image encoder or decoder must Figure 7 Each feature.

[0057] Adder 310 actually performs a subtraction (negative addition) operation because it receives the input video signal 300 at its "+" input and the output of the image predictor 320 at its "-" input, thereby subtracting the predicted image from the input image. The result is a so-called residual image signal 330, which represents the difference between the actual image and the projected image.

[0058] One reason for generating a residual image signal is as follows. The data encoding techniques to be described (that is, the techniques that will be applied to the residual image signal) tend to work more efficiently when there is less "energy" in the image to be encoded. Here, the term "efficiently" refers to generating a small amount of encoded data; for a particular level of image quality, it is desirable (and considered "efficient") to generate as little data as possible. The "energy" mentioned in the residual image refers to the amount of information contained in the residual image. If the predicted image is identical to the true image, then the difference between the two (that is, the residual image) will contain zero information (zero energy) and will be very easy to encode into a small amount of encoded data. In general, if the prediction process can be made to work reasonably well so that the predicted image content is similar to the image content to be encoded, then it is expected that the residual image data will contain less information (less energy) than the input image and will therefore be easier to encode into a small amount of encoded data.

[0059] The remainder of the apparatus that acts as an encoder (encoding the residual or difference image) will now be described. The residual image data 330 is provided to a transform unit or circuit 340 that generates a discrete cosine transform (DCT) representation of a block or region of the residual image data. DCT techniques are well known per se and will not be described in detail here. It is also noted that the use of the DCT is merely illustrative of one exemplary arrangement. Other transforms that may be used include, for example, the discrete sine transform (DST). Transforms may also comprise a sequence or concatenation of separate transforms, such as an arrangement in which one transform is followed (whether directly or not) by another transform. The choice of transform may be determined explicitly and / or may depend on side information used to configure the encoder and decoder.

[0060] The output of transform unit 340 (that is, a set of DCT coefficients for each transformed block of image data) is provided to quantizer (Q) 350. Various quantization techniques are known in the art of video data compression, ranging from simple multiplication by a quantization scale factor to the application of complex lookup tables under the control of a quantization parameter. The overall goal is twofold. First, the quantization process reduces the number of possible values ​​for the transformed data. Second, the quantization process can increase the probability that the transformed data value is zero. Both of these can enable the entropy encoding process, described below, to more efficiently generate a small amount of compressed video data.

[0061] The scanning unit 360 applies a data scanning process. The purpose of the scanning process is to reorder the quantized transform data so as to group together as many non-zero quantized transform coefficients as possible, and of course also to group together as many zero-valued coefficients as possible. These features may allow so-called run-length encoding or similar techniques to be applied efficiently. Thus, the scanning process comprises selecting coefficients from the quantized transform data, and in particular selecting coefficients from coefficient blocks corresponding to blocks of image data that have already been transformed and quantized, according to a "scan order" such that (a) all coefficients are selected once as part of the scan, and (b) the scan tends to provide the desired reordering. One example scan order that tends to give useful results is a so-called top-right diagonal scan order.

[0062] The scanned coefficients are then passed to an entropy encoder (EE) 370. Again, various types of entropy coding can be used. Two examples are variants of the so-called context-adaptive binary arithmetic coding (CABAC) system and variants of the so-called context-adaptive variable length coding (CAVLC) system. In general, CABAC is considered to provide better efficiency, and in some studies, CABAC has been shown to provide a 10% to 20% reduction in the amount of coded output data compared to CAVLC for comparable image quality. However, CAVLC is considered to represent lower complexity (in terms of its implementation) than CABAC. Note that the scanning process and the entropy coding process are shown as separate processes, but in practice can be combined or processed together. That is, the data can be read into the entropy encoder in the scanning order. Corresponding considerations apply to the corresponding inverse process to be described below.

[0063] The output of the entropy encoder 370 , along with additional data (mentioned above and / or discussed below), such as defining the manner in which the predictor 320 generates the predicted image, provides a compressed output video signal 380 .

[0064] However, a return path is also provided since the operation of the predictor 320 itself depends on a decompressed version of the compressed output data.

[0065] The reason for this feature is as follows. At an appropriate stage in the decompression process (described below), a decompressed version of the residual data is generated. This decompressed residual data must be added to the predicted image to generate the output image (because the original residual data is the difference between the input image and the predicted image). In order for the process to be comparable, the predicted image generated by the predictor 320 should be the same during the compression process and during the decompression process, both on the compression and decompression sides. Of course, during decompression, the device does not have access to the original input image, but only to the decompressed image. Therefore, during compression, the predictor 320 bases its predictions (at least for inter-frame image coding) on ​​the decompressed version of the compressed image.

[0066] The entropy encoding process performed by the entropy encoder 370 is considered (in at least some examples) to be "lossless," that is, it can be reversed to obtain data that is identical to the data that was first provided to the entropy encoder 370. Thus, in such examples, the return path can be implemented before the entropy encoding stage. In practice, the scanning process performed by the scanning unit 360 is also considered lossless, but in this embodiment, the return path 390 is from the output of the quantizer 350 to the complementary inverse quantizer (Q -1 ) 420. In cases where a stage introduces losses or potential losses, that stage can be included in the feedback loop formed by the return path. For example, the entropy encoding stage can, at least in principle, be lossy, for example, by encoding bits in parity information. In this case, entropy encoding and decoding should form part of the feedback loop.

[0067] In general, entropy decoder (ED) 410, inverse scan unit 400, inverse quantizer 420, and inverse transform unit or circuit 430 provide the respective inverse functions of entropy encoder 370, scan unit 360, quantizer 350, and transform unit 340. For the time being, the discussion will proceed through the compression process; the process of decompressing an input compressed video signal will be discussed separately below.

[0068] During compression, the scanned coefficients are passed from quantizer 350 via return path 390 to inverse quantizer 420, which performs the inverse operation of scanning unit 360. The inverse quantization and inverse transform processes are performed by units 420, 430 to generate a compressed-decompressed residual image signal 440.

[0069] The image signal 440 is added to the output of the predictor 320 at an adder 450 to generate a reconstructed output image 460. This forms one input to the image predictor 320 as described below.

[0070] Turning now to the process applied to decompress the received compressed video signal 470, which is supplied to the entropy decoder 410 and from there to a chain of inverse scan unit 400, inverse quantizer 420 and inverse transform unit 430 before being added to the output of the picture predictor 320 by adder 450. Thus, on the decoder side, the decoder reconstructs a version of the residual image and then applies it (via adder 450) to the predicted version of the image (on a block-by-block basis) in order to decode each block. In simple terms, the output 460 of adder 450 forms the output decompressed video signal 480. In practice, further filtering may optionally be applied (e.g. by Figure 8 filter 560 shown in FIG, but for Figure 7 For clarity of the high-level diagram, Figure 7This filter is omitted in the figure).

[0071] Figure 7 and Figure 8 The device can act as a compression (encoding) device or a decompression (decoding) device. The functions of the two types of devices are substantially overlapping. The scanning unit 360 and the entropy encoder 370 are not used in the decompression mode, and the operation of the predictor 320 (described in detail below) and other units follows the mode and parameter information contained in the received compressed bit stream rather than generating such information themselves.

[0072] Figure 8 The generation of a predicted image, and in particular the operation of the image predictor 320 , is schematically illustrated.

[0073] There are two basic modes of prediction performed by the image predictor 320: so-called intra-image prediction and so-called inter-image or motion-compensated (MC) prediction. On the encoder side, each involves detecting the prediction direction for the current block to be predicted and generating a predicted block of samples based on other samples (in the same (intra) or another (inter) image). The difference between the predicted block and the actual block is coded or applied, respectively, by means of units 310 or 450, to encode or decode the block, respectively.

[0074] (At the decoder, or on the reverse decoding side of the encoder, detection of the prediction direction may be responsive to data associated by the encoder with the encoded data, indicating which direction to use at the encoder. Or the detection may be responsive to the same factors as those used to make the decision at the encoder).

[0075] Intra-frame picture prediction predicts the content of a block or region of an image based on data from within the same picture. This corresponds to so-called I-frame coding in other video compression techniques. However, in contrast to I-frame coding, which involves encoding the entire image via intra-frame coding, in this embodiment, the choice between intra-frame coding and inter-frame coding can be made on a block-by-block basis, although in other embodiments, the choice is still made on a picture-by-picture basis.

[0076] Motion compensated prediction is an example of inter-image prediction and, using motion information, attempts to define the source of image details to be encoded in the current image in another adjacent or nearby image. Thus, in an ideal example, the contents of an image data block in the predicted image can be encoded very simply as a reference (motion vector) pointing to a corresponding block at the same or slightly different location in an adjacent image.

[0077] A technique known as "block copy" prediction is in some ways a hybrid of the two, in that a vector is used to indicate a block of samples at a position displaced from the currently predicted block within the same picture that should be copied to form the currently predicted block.

[0078] Back to Figure 8 , two image prediction settings (corresponding to intra-frame image prediction and inter-frame image prediction) are shown, the results of which are selected by multiplexer 500 under the control of mode signal 510 (e.g., from controller 343) to provide a block of predicted images to be provided to adders 310 and 450. The selection is made based on which option gives the lowest "energy" (as described above, this "energy" can be considered the information content to be encoded), and this selection is signaled to the decoder within the encoded output data stream. In this case, image energy can be detected, for example, by performing trial subtractions of a region of two versions of the predicted image from the input image, squaring each pixel value in the difference image, summing the squared values, and identifying which of the two versions produces the lower mean square value of the difference image associated with that image region. In other examples, trial encoding can be performed for each option or potential option, and then a selection can be made based on the cost of each potential option, one or both of the number of bits required for picture encoding and distortion.

[0079] In an intra-coding system, the actual prediction is performed based on the image blocks received as part of signal 460. That is, the prediction is based on the coded image blocks so that the same prediction can be performed in the decompression device. However, data can be derived from the input video signal 300 by the intra-mode selector 520 to control the operation of the intra-image predictor 530.

[0080] For inter picture prediction, a motion compensated (MC) predictor 540 uses motion information such as motion vectors derived from the input video signal 300 by a motion estimator 550. The motion compensated predictor 540 applies these motion vectors to a processed version of the reconstructed picture 460 to generate inter picture predicted blocks.

[0081] Thus, units 530 and 540 (operating in conjunction with estimator 550) each act as a detector to detect a prediction direction with respect to a current block to be predicted, and as a generator to generate a prediction block of samples (forming part of the prediction passed to units 310 and 450) from other samples defined by the prediction direction.

[0082] The processing applied to the signal 460 will now be described. First, the signal is optionally filtered by a filter unit 560, which will be described in more detail below. This includes applying a "deblocking" filter to eliminate or at least tend to reduce the effects of the block-based processing and subsequent operations performed by the transform unit 340. A sample adaptive offset (SAO) filter may also be used. In addition, an adaptive loop filter is optionally applied using coefficients obtained by processing the reconstructed signal 460 and the input video signal 300. An adaptive loop filter is a filter that applies adaptive filter coefficients to the data to be filtered using known techniques. That is, the filter coefficients can vary depending on various factors. Data defining which filter coefficients to use is included as part of the encoded output data stream.

[0083] When the device operates as a decompression device, the filtered output from filter unit 560 actually forms the output video signal 480. It is also buffered in one or more image or frame memories 570; storage of consecutive images is a requirement for motion-compensated prediction processing, and in particular, for the generation of motion vectors. To conserve storage requirements, the images stored in image memory 570 can be stored in compressed form and then decompressed for use in generating motion vectors. Any known compression / decompression system can be used for this purpose. The stored images are passed to interpolation filter 580, which generates a higher-resolution version of the stored images; in this example, intermediate samples (subsamples) are generated so that the resolution of the interpolated images output by interpolation filter 580 is four times (in each dimension) the resolution of the 4:2:0 images for the luma channels stored in image memory 570 and eight times (in each dimension) the resolution of the 4:2:0 images for the chroma channels stored in image memory 570. The interpolated images are passed as input to motion estimator 550 and also to motion-compensated predictor 540.

[0084] The manner in which an image is segmented for compression processing will now be described. Essentially, the image to be compressed is considered to be an array of blocks or regions of samples. The image can be segmented into such blocks or regions by a decision tree, such as that described in Bross et al., "High Efficiency Video Coding (HEVC) text specification draft 6", JCTVC-H1003_d0 (November 2011). 2011)), the contents of which are incorporated herein by reference. In some examples, the resulting blocks or regions have sizes, and in some cases shapes, that can generally follow a set of image features within the image by means of the decision tree. This in itself can allow for improved coding efficiency, since samples representing or following similar image features will tend to be grouped together by such a set. In some examples, square blocks or regions of different sizes (such as 4×4 samples up to 64×64 or larger blocks) are available for selection. In other example settings, blocks or regions of different shapes can be used, such as rectangular blocks (e.g., oriented vertically or horizontally). Other non-square and non-rectangular blocks are contemplated. The result of dividing the image into such blocks or regions is that (at least in this example) each sample of the image is assigned to one and only one such block or region.

[0085] The intra prediction process will now be discussed. Generally speaking, intra prediction involves generating a prediction of the current block of samples from previously encoded and decoded samples in the same picture.

[0086] Figure 9 A partially encoded image 800 is schematically shown. Here, the image is encoded on a block-by-block basis, from the upper left to the lower right. An example block that is encoded partway through processing the entire image is shown as block 810. The shaded area 820 above and to the left of block 810 has already been encoded. Intra-image prediction of the contents of block 810 can utilize any shaded area 820, but cannot utilize the unshaded area below it.

[0087] In some examples, images are encoded on a block-by-block basis so that larger blocks (called coding units or CUs) are encoded in a manner such as a reference Figure 9The order discussed is coded. Within each CU, it is possible that the CU is treated as a group of two or more smaller blocks or transform units (TUs) (depending on the block partitioning process that has occurred). This can give a hierarchical order of coding, so that the image is coded on a CU-by-CU basis, and each CU is potentially coded on a TU-by-TU basis. However, note that for a single TU within the current coding tree unit (the largest node in the tree structure of the block partitioning), the hierarchical order of coding discussed above (CU-by-CU, then TU-by-TU) means that there may be previously coded samples in the current CU and that can be used for the coding of that TU, for example, to the upper right or lower left of the TU.

[0088] Block 810 represents a CU; as described above, this can be subdivided into a set of smaller units for the purposes of intra picture prediction processing. An example of a current TU 830 is shown within the CU 810. More generally, a picture is partitioned into sample regions or groups of samples to allow efficient encoding of signaling information and transform data. The signaling of information may require a tree structure of sub-divisions that is different from the tree structure of the transform, and indeed the tree structure of the prediction information or prediction itself. To this end, a coding unit may have a different tree structure than a transform block or region, a prediction block or region, and the prediction information. In some examples such as HEVC, the structure may be a quadtree of so-called coding units, whose leaf nodes contain one or more prediction units and one or more transform units; a transform unit may contain multiple transform blocks corresponding to the luma and chroma representations of a picture, and the prediction may be considered to apply at the transform block level. In examples, the parameters applied to a particular sample group may be considered to be defined primarily at the block level, which may be different from the granularity of the transform structure.

[0089] Intra-image prediction considers samples that were coded before the current TU was considered, such as samples above and / or to the left of the current TU. The source samples for predicting the samples required can be located at different positions or directions relative to the current TU. To decide which direction is appropriate for the current prediction unit, the mode selector 520 of the example encoder can test all combinations of available TU structures for each candidate direction and select the prediction direction and TU structure with the best compression efficiency.

[0090] Pictures can also be coded on a "slice" basis. In one example, a slice is a group of horizontally adjacent CUs. More generally, however, the entire residual image can form a slice, or a slice can be a single CU, or a slice can be a row of CUs, and so on. Because slices are coded as independent units, they provide a certain degree of error resilience. The encoder and decoder states are completely reset at slice boundaries. For example, intra prediction is not performed across slice boundaries; for this purpose, slice boundaries are treated as picture boundaries.

[0091] Figure 10A set of possible (candidate) prediction directions is shown schematically. The prediction unit has access to the entire set of candidate directions. The direction is determined by the horizontal and vertical displacement relative to the current block position, but is encoded as a prediction "mode", a set of modes such as Figure 11 Note that the so-called DC mode represents the simple arithmetic mean of the surrounding upper left samples. Also note that Figure 10 The set of directions shown is just one example; in other examples, Figure 12 A set of (for example) 65 angular modes plus DC and planar (a total of 67 modes) schematically shown in composes the complete set. Other numbers of modes may be used.

[0092] In general, after detecting the prediction direction, the system is operable to generate a block of prediction samples based on other samples defined by the prediction direction. In an example, the image encoder is configured to encode data identifying the prediction direction selected for each sample or region of the image (and the image decoder is configured to detect such data).

[0093] Figure 13 The intra prediction process is schematically shown, where samples 900 of a block or region of samples 910 are derived from other reference samples 920 of the same image according to a direction 930 defined by the intra prediction mode associated with the sample. The reference samples 920 in this example are from blocks above and to the left of the block 910 in question, and the predicted value for the sample 900 is obtained by tracing the reference samples 920 along the direction 930. The direction 930 may point to a single individual reference sample, but in the more general case, interpolation between surrounding reference samples is used as the predicted value. Note that the block 910 may be such as Figure 13 A square is shown, but other shapes such as a rectangle are possible.

[0094] Figure 14 and Figure 15 The previously proposed reference sample projection process is schematically illustrated.

[0095] exist Figure 14 and Figure 15 In

[0045] , a block or region 1400 of samples to be predicted is surrounded by a linear array of reference samples from which intra prediction of the prediction samples is performed. Reference samples 1410 are Figure 14 and Figure 15 The samples to be predicted are shown as shaded blocks in the example, and the samples to be predicted are shown as unshaded blocks. Note that an 8×8 block or region of samples to be predicted is used in this example, but the technique is applicable to variable block sizes and actual block shapes.

[0096] As described above, the reference samples include at least two linear arrays at corresponding orientations relative to the current image area of ​​the sample to be predicted. For example, the linear arrays may be a sample array or row 1420 above the sample block to be predicted and a sample array or column 1430 to the left of the sample block to be predicted.

[0097] As referenced above Figure 13 As discussed, the reference sample array may extend beyond the range of the block to be predicted so that Figures 10 to 12 The prediction mode or direction is provided within the range shown. If necessary, if previously decoded samples cannot be used as reference samples for specific reference sample positions, other reference samples can be reused at these missing positions. The reference sample filtering process can be used for reference samples.

[0098] Figure 14 The operation of a CABAC entropy encoder is schematically illustrated.

[0099] The CABAC encoder operates on binary data (that is, data represented by only two symbols, 0 and 1). The encoder utilizes a so-called context modeling process, which selects a "context," or probability model, for subsequent data based on previously encoded data. The context selection is performed in a deterministic manner, so that the same determination can be made at the decoder based on previously decoded data, without the need to add further data (specifying the context) to the encoded data stream passed to the decoder.

[0100] refer to Figure 14 If the input data to be encoded is not already in binary form, it can be passed to a binary converter 1400; if the data is already in binary form, the converter 1400 is bypassed (via a schematic switch 1410). In this embodiment, the conversion to binary form is actually achieved by representing the quantized DCT coefficient data as a series of binary "maps", which will be further described below.

[0101] The binary data can then be processed by one of two processing paths: a "normal" path and a "bypass" path (these two paths are schematically shown as separate paths, but in the embodiments of the invention discussed below, they can actually be implemented by the same processing stages, just using slightly different parameters). The bypass path uses a so-called bypass encoder 1420, which does not necessarily use context modeling in the same form as the normal path. In some examples of CABAC encoding, this bypass path can be selected if particularly fast processing of a batch of data is required, but in this embodiment, two characteristics of the so-called "bypass" data are noted: first, the bypass data is processed by the CABAC encoder (1460) using only a fixed context model representing 50% probability; second, the bypass data involves certain categories of data, a specific example of which is coefficient sign data. Otherwise, the normal path is selected by schematic switches 1430, 1440 operating under the control of control circuit 1435. This includes data processed by the context modeler 1450 and then the encoding engine 1460.

[0102] If the block is formed entirely of zero-valued data, then Figure 14 The entropy encoder shown encodes this block of data (i.e., data corresponding to a block of coefficients associated with a residual image block, for example) as a single value. For each block that does not fall into this category, that is, a block that contains at least some non-zero data, a "significance map" is prepared. The significance map indicates, for each position in the block of data to be encoded, whether the corresponding coefficient in the block is non-zero. The importance map data in binary form is itself CABAC encoded. The use of the significance map aids compression because no data needs to be encoded for coefficients whose magnitude the significance map indicates is zero. In addition, the significance map can include a special code to indicate the final non-zero coefficient in the block, so that all final high frequency / trailing zero coefficients can be omitted from the encoding. In the encoded bitstream, the significance map is followed by data defining the values ​​of the non-zero coefficients specified by the significance map.

[0103] Further levels of mapping data are also prepared and encoded. One example is a mapping that defines as a binary value (1=yes, 0=no) whether the coefficient data at a mapping position that the significance map has indicated as "non-zero" actually has the value "1". Another mapping specifies whether the coefficient data at a mapping position that the significance map has indicated as "non-zero" actually has the value "2". Another mapping indicates, for those mapping positions where the significance map has indicated that the coefficient data is "non-zero", whether the data has a value "greater than 2". For data identified as "non-zero", another mapping indicates the sign of the data value (using a predetermined binary representation, such as 1 for +, 0 for -, and of course vice versa).

[0104] In an embodiment of the present invention, the significance map and other maps are assigned to a CABAC encoder or a bypass encoder in a predetermined manner and each represents a different corresponding attribute or value range for the same initial data item. In one example, at least the significance map is CABAC encoded and at least some of the remaining maps (such as sign data) are bypass encoded. Thus, each data item is divided into a corresponding data subset, and the corresponding subset is encoded by a first coding system (e.g., CABAC) and a second (e.g., bypass) coding system. The properties of the data and of CABAC and bypass coding are such that for a predetermined amount of CABAC-encoded data, a variable amount of zero or more bypass data is generated for the same initial data item. Thus, for example, if the quantized, reordered DCT data contains essentially all zero values, no bypass data or a very small amount of bypass data may be generated because the bypass data only relates to those mapping positions where the significance map has indicated a non-zero value. In another example, in quantized, reordered DCT data having many high-valued coefficients, a large amount of bypass data may be generated.

[0105] In an embodiment of the present invention, the significance map and other maps are generated, for example, by the scanning unit 360 from the quantized DCT coefficients and are subjected to a zigzag scanning process (or a scanning process selected from zigzag, horizontal raster and vertical raster scanning according to the intra prediction mode) before being subjected to CABAC encoding.

[0106] Generally speaking, CABAC coding involves predicting the context or probability model of the next bit to be encoded based on other previously encoded data. If the next bit is the same as the bit identified as "most likely" by the probability model, then the encoding of the information that "the next bit is consistent with the probability model" can be encoded very efficiently. The encoding efficiency of "the next bit is inconsistent with the probability model" is less, so the derivation of context data is important for the good operation of the encoder. The term "adaptive" refers to adjusting or changing the context or probability model during encoding in an attempt to provide a good match with the next data (not yet encoded).

[0107] To use a simple analogy, in written English, the letter "U" is relatively uncommon. However, in the position immediately following the letter "Q," it is very common. Therefore, the probability model might set the probability of "U" to a very low value, but if the current letter is "Q," the probability model might set the probability of "U" as the next letter to a very high value.

[0108] In this setting, CABAC encoding is used for at least the significance map and the mapping indicating whether a non-zero value is 1 or 2. In these embodiments, the bypass process is the same as CABAC encoding, but in fact, the probability model is fixed at an equal (0.5:0.5) probability distribution of 1s and 0s, and the bypass process is used for at least the symbol data and the mapping indicating whether the value is >2. For those data positions identified as >2, a separate so-called escape data encoding can be used to encode the actual value of the data. This may include Golomb-Rice encoding techniques.

[0109] The CABAC context modeling and coding process is described in more detail in WD4: Working Draft 4 of High-Efficiency Video Coding, JCTVC-F803_d5, Draft ISO / IEC 23008-HEVC; 201x(E)2011-10-28.

[0110] Now refer to Figure 15 and Figure 16 , an entropy encoder forming part of a video encoding device includes a first encoding system (e.g., an arithmetic coding encoding system, such as CABAC encoder 1500) and a second encoding system (such as bypass encoder 1510), which are arranged so that a particular data word or value is encoded into a final output data stream by either the CABAC encoder or the bypass encoder, rather than by both the CABAC encoder and the bypass encoder. In an embodiment of the present invention, the data values ​​passed to the CABAC encoder and the bypass encoder are respective subsets of ordered data values ​​separated or derived from the original input data (in this example, the reordered quantized DCT data), representing different mappings in a set of "mappings" generated from the input data.

[0111] Figure 15 The diagram in [1] treats the CABAC encoder and the bypass encoder as independent setups. This may be a good situation in practice, but in another possibility, such as Figure 16 As shown schematically, a single CABAC encoder 1620 is used as Figure 15 The encoder 1620 operates under the control of the coding mode selection signal 1630 to operate with an adaptive context model (as described above) when in the mode of the CABAC encoder 1500 and with a fixed 50% probability context model when in the mode of the bypass encoder 1510.

[0112] A third possibility combines the two, since two essentially identical CABAC encoders can be operated in parallel (similar to Figure 15 parallel setup), except that the CABAC encoder operating as a bypass encoder 1510 fixes its context model at a 50% probability context model.

[0113] The outputs of the CABAC encoding process and the bypass encoding process may be stored (at least temporarily) in respective buffers 1540, 1550. Figure 16 In the case of , the switch or demultiplexer 1660 operates under the control of the mode signal 1630 to route the CABAC encoded data to the buffer 1550 and bypass the encoded data to the buffer 1540.

[0114] Figure 17 and Figure 18 Schematically shows an example of an entropy decoder forming part of a video decoding device. Figure 17 , the corresponding buffers 1710, 1700 pass the data to the CABAC decoder 1730 and the bypass decoder 1720, which are arranged so that a particular encoded data word or value is decoded by the CABAC decoder or the bypass decoder but not by both the CABAC encoder and the bypass encoder. Logic 1740 reorders the decoded data into the appropriate order for the subsequent decoding stage.

[0115] Figure 17 The diagram in [1] treats the CABAC decoder and the bypass decoder as independent setups. This may be a good situation in practice, but in another possibility, such as Figure 18 As shown schematically, a single CABAC decoder 1850 is used as Figure 17 The decoder 1850 operates under the control of a decoding mode select signal 1860 to operate with an adaptive context model (as described above) when in the mode of the CABAC decoder 1730 and with a fixed 50% probability context model when in the mode of the bypass encoder 1720.

[0116] As mentioned before, a third possibility combines the two, since two essentially identical CABAC decoders can be operated in parallel (similar to Figure 17 ), except that the CABAC decoder operating as a bypass decoder 1720 fixes its context model at a 50% probability context model.

[0117] exist Figure 18In the case of mode signal 1860, switch or multiplexer 1870 acts under the control of mode signal 1860 to appropriately route CABAC encoded data from buffer 1700 or buffer 1710 to decoder 1850.

[0118] Figure 19 Picture 1900 is shown schematically and will be used to demonstrate various picture partitioning schemes relevant to the following discussion.

[0119] An example of a picture partition is a slice or "regular slice." Each regular slice is encapsulated in its own Network Abstraction Layer (NAL) unit. Prediction within a picture (e.g., intra-sample prediction, motion information prediction, coding mode prediction) and entropy coding dependencies across slice boundaries are not allowed. This means that a regular slice can be reconstructed independently of other regular slices in the same picture.

[0120] So-called tiles define horizontal and vertical boundaries to divide a picture into rows and columns of tiles. In a manner corresponding to regular slices, intra-picture prediction dependencies are not allowed across tile boundaries, nor are entropy decoding dependencies. However, tiles are not limited to being contained in a single NAL unit.

[0121] In general, there may be multiple tiles in a slice, or multiple slices in a tile, and there may be one or more slices or tiles in a picture. Figure 19 The illustrative example of shows four slices 1910, 1920, 1930, 1940, wherein the slice 1940 includes two tiles 1950, 1960. However, as mentioned above, this is only an arbitrary illustrative example.

[0122] In some example settings, there is a threshold on the number of bins (EP or CABAC) that can be coded in a slice or picture according to the following equation:

[0123] BinCountsinNalUnits<=(4 / 3)*NumBitsInVclNalUnits+(RawMinCuBits*PicSizelnMinCbsY) / 32 (Equation 1)

[0124] The right side of the equation depends on the sum of two parts: a constant value for a particular image region and related to the size of the slice or picture (RawMinCuBits*PicSizelnMinCbsY); and a dynamic value (NumBitsInVclNalUnits) that is the number of bits encoded in the output stream of the slice or picture. Note that the value 4 / 3 represents the number of binary digits per bit.

[0125] RawMinCuBits is the number of bits in the smallest size raw CU - usually 4*4; and PicSizelnMinCbsY is the number of smallest size CUs in a slice or picture.

[0126] If the threshold is exceeded, CABAC zero words (3 bytes with value 000003) will be appended to the stream until the threshold is reached. Each such zero word increases the dynamic value by 3.

[0127] This constraint can be expressed as:

[0128] N<=K1*B+(K2*CU)

[0129] in:

[0130] N = the number of binary symbols in the output data unit;

[0131] K1 is a constant;

[0132] B = number of coded bytes of the output data unit;

[0133] K2 is a variable that depends on the properties of the minimum-size coding unit employed by the image data encoding device; and

[0134] CU = the size of the picture, slice, or tile represented by the output data unit, expressed as the number of smallest-sized coding units.

[0135] In the examples presented previously, this threshold check was performed at the picture and slice level.

[0136] Alternatively, the following expression for the same threshold value may be used, except that bytes are referenced instead of bits:

[0137] BinCountsinNalUnits<=(32 / 3)*NumByteslnVclNalUnits+(RawMinCuBits*PicSizelnMinCbsY) / 32 (Equation 2)

[0138] The threshold in Equation 1 (or equivalently expressed in Equation 2) may be applied uniformly, that is, independently of other encoding or decoding parameters, properties, etc.

[0139] An example of a technical justification for using such a threshold value could be that the threshold value itself indirectly defines the maximum level of processing performance required of a real-time decoder. Real-time decoding relies on decoding each frame in a timely manner to achieve a specific output frame rate (e.g., a certain number of frames per second). The constraints on the processing performance of the real-time decoder are in turn based on the rate at which the binary numbers need to be decoded, rather than the amount of decoded data obtained. In order to allow relatively straightforward rate control on the encoder side, a technique for controlling this situation is to limit (e.g., an upper limit or threshold or constraint) the number of binary numbers per slice or picture. The maximum processing performance required of a CABAC decoder can be considered to be the product or function of this limit and the rate at which the slices or pictures must be decoded.

[0140] Therefore, it is important to set an appropriate threshold (or upper limit or constraint). A threshold that is too high may cause problems in implementing or successfully operating a real-time decoder. Similarly, using an incorrect or inappropriate threshold may lead to excessive use of filler data (e.g., dummy data or other data intended only to occupy part of the encoded data stream, e.g., to increase the size of the encoded output data unit), thereby reducing encoding efficiency.

[0141] As reference Figure 19 As noted, a picture or slice may be partitioned into multiple tiles. One example of why this may be done is to allow the use of multiple concurrent (parallel) decoders.

[0142] In the previously proposed setting, each tile does not necessarily meet the threshold calculation discussed above. For example, if the tiles are used or decoded independently like pictures, or if different tiles (e.g., with different quantization parameters or from different sources) are composited together, this may cause problems and there is no guarantee that the resulting composited slice or picture meets the above specifications.

[0143] To address this issue, in an example embodiment, the CABAC threshold is applied at the end of each tile, rather than at the end of each slice or picture. Thus, the threshold is applied at the end of encoding any of the tiles, slices, and pictures. However, if every tile in the image meets the threshold, it can be assumed that the entire picture must also meet the threshold, thus eliminating the need to reapply the threshold at the end of encoding the picture if the picture is divided into slices or tiles.

[0144] The terms "tile" and "slice" refer to independently decodable units and represent the names used on the priority date of this application. In the event of subsequent or other name changes, the same shall apply to other such independently decodable units.

[0145] Thus, in an example arrangement, the data units may be independently decodable data units.For example, the image portion (represented by the output data unit) may be one of a picture, a slice and a tile.

[0146] For the purpose of applying the equations discussed above, the dynamic value represents the number of bytes encoded in the output stream for the tile, while the fixed value depends on the number of minimum-sized coding units (CUs) in the tile.

[0147] Figure 20 The apparatus configured to perform the test is schematically shown. Figure 20 At input 2000, a CABAC / EP encoded stream is received from an encoder. A padding data detector 2010 detects whether the threshold calculation described above is met at a predetermined stage of completion of a reference slice or tile, such as at the end of encoding the slice or tile. In response to the detection by detector 2010, controller 2020 controls padding data generator 2030 to generate padding data 2040, such as the CABAC zero words described above, and appends it to the stream via combiner 2050 to form output stream 2060. The generation of zero words can also be signaled back to detector 2010, so that while zero words are being appended, padding data detector 2010 can continue to monitor whether the threshold is met and, once the threshold is met, cause controller 2020 to stop generating zero words.

[0148] In other examples, the controller 2020 may also function as a predictor configured to generate a prediction of whether an output data unit will satisfy a constraint during generation of the output data unit in conjunction with embodiments to be described herein.

[0149] For example, the predetermined phase may be every n encoded binary numbers (where n is an integer of 1 or greater), but in the example arrangement, the predetermined phase is the end of encoding the current output data unit.

[0150] Therefore, in Figure 20 In the present invention, examples of a padding data detector 2010 and a padding data generator 2030 are disclosed, wherein the padding data detector 2010 is configured to detect whether the current output data unit satisfies a constraint at a predetermined stage of encoding relative to the current output data unit, and the padding data generator 2030 is configured to generate sufficient padding data and insert it into the current output data unit so that the output data unit including the inserted padding data satisfies the constraint.

[0151] refer to Figure 21 , in some examples, the controller 2020 includes: an attribute detector 2070 configured to detect an encoding attribute applicable to a given output data unit; and

[0152] A selector 2080 is configured to select a constraint for a given output data unit from two or more candidate constraints 2082 in response to the detected encoding properties.

[0153] The controller 2020 may further include a comparator 2090 for comparing a threshold value derived from the currently selected constraint with the detection of the padding data detector 2010 in order to derive a control signal to control the operation of the padding data generator 2030 .

[0154] The property detected by the detector 2070 may be, for example, a coding property such as a coding mode or profile, e.g., enabling dependent quantization (discussed below), which is a mode of operation in which the selection of a quantization parameter for quantizing a current data value depends at least in part on properties of previously encoded data values, and is subsequently represented by flag data (schematically shown as 2072) included in or associated with the encoded data stream so that it can be later detected at a decoder. For example, the flag data may be included in header data such as output data unit header data (e.g., slice header data). The detector 2070 itself need not generate or insert the flag data; for the benefit of this specification, it is shown for illustrative purposes only. Figure 21 This aspect of the process is illustrated in The property may be considered to apply to a given output data unit even if (as in some examples) the property also applies to other (eg, preceding or subsequent) output data units.

[0155] Therefore, according to the Figure 19 and Figure 20 The discussed techniques (and those discussed below) operate Figure 7 and Figure 14 The device provides an example of an image data encoding device, the image data encoding device comprising:

[0156] a first data encoder 1450, 1460 and a second data encoder 1420, each encoder configured to generate output data bits representing binary symbols from successive symbols representing image data;

[0157] The first data encoder is configured to generate output data bits representing encoding symbols at a variable ratio of the number of data bits to the number of encoding symbols;

[0158] a second data encoder configured to generate a fixed number of output data bits to represent each encoded symbol;

[0159] The entropy encoder is configured (e.g., using controller 2020 and / or controller 1435) to generate a constrained output data stream, the constraint defining an upper limit on the number of binarized symbols that can be represented by any individual output data unit relative to the byte size of the output data unit, wherein the entropy encoder is configured to provide padding data for each output data unit that does not satisfy the constraint in order to increase the byte size of the output data unit so as to satisfy the constraint; the apparatus comprising:

[0160] an attribute detector 2070 configured to detect encoding attributes applicable to a given output data unit; and

[0161] The selector 2080 is configured to select a constraint for a given output data unit from two or more candidate constraints in response to the detected encoding properties.

[0162] use Figure 14 In the illustrated technique, the first data encoder / decoder may be a context-adaptive binary arithmetic coding (CABAC) encoder / decoder. The second data encoder / decoder may be a bypass encoder / decoder. The second data encoder / decoder may be a binary arithmetic encoder / decoder using a fixed 50% probability context model.

[0163] Examples of application of thresholds or constraints are as follows.

[0164] Generally speaking, the test is "Does the amount of output data generated satisfy a threshold test?" This may be performed, for example, at the end of encoding an output data unit.

[0165] Operated according to the techniques discussed herein Figure 7 The device provides an example of an image data encoding device, the image data encoding device comprising:

[0166] an entropy encoder configured to selectively encode data items representing image data so as to generate encoded binary symbols of successive output data units;

[0167] An entropy encoder is configured to generate an output data stream subject to a constraint, the constraint defining an upper limit on the number of binary symbols that can be represented by any individual output data unit relative to the byte size of the output data unit, wherein the entropy encoder is configured to provide padding data for each output data unit that does not satisfy the constraint in order to increase the byte size of the output data unit so as to satisfy the constraint; the apparatus comprising:

[0168] a property detector configured to detect encoding properties applicable to a given output data unit; and

[0169] A selector is configured to select a constraint for a given output data unit from two or more candidate constraints in response to the detected encoding properties.

[0170] The present disclosure also provides suitable decoding devices for decoding data signals generated by the methods or devices described herein.

[0171] Further examples of constraints will now be described.

[0172] In the examples discussed above, a single threshold derivation or expression is consistently used. In an alternative example to be described below, a selection is implemented between two or more candidate thresholds or constraints. The selection or selection can be responsive to one or more encoding attributes (e.g., parameters, properties, or modes, such as those signaled by a flag or the like from the encoder side to the decoder side, or parameters, properties, or modes that can be derived in a corresponding or matching manner at the encoder and decoder sides).

[0173] Another example of a candidate expression for a threshold or constraint is as follows:

[0174] BinCountsinNalUnits<=10*NumByteslnVclNalUnits+(RawMinCuBits*PicSizelnMinCbsY) / 16 (Equation 3)

[0175] Therefore, although Equation 3 may be used uniformly as previously described, in an exemplary embodiment, selection is made between two or more candidate expressions, for example, between Equation 1 / Equation 2 (which are equivalent expressions of the same thing) and Equation 3.

[0176] In some examples, the selection can be made based on whether a given coding property signaled in or with the coded data stream is in a first state or a second state, the corresponding state of the given property corresponding to the selection of the corresponding candidate expression (e.g., Equation 1 / Equation 2 or Equation 3) on the encoder side and the decoder side.

[0177] One example of such a property is the so-called “dep_quant_enabled_flag.” This indicates in a slice header whether a technique called dependent quantization is enabled for the slice to which the slice header applies.

[0178] The alternative to dep_quant_enabled_flag is actually to rely on the availability of the (dep_quant) tool, rather than whether the tool is enabled. So, for one example profile, the profile might define (for example) "dep_quant_enabled_flag must be off (not enabled)"; in another case, there might be no such constraint, allowing the dep_quant tool to be on (enabled) or off (not enabled). Thus, the choice can be made based on the profile constraint, rather than on whether the dep_quant tool is currently enabled or disabled.

[0179] Dependent quantization is defined in "Versatile Video Coding (Draft 5)," Bross et al., JVET-N1001-v10, July 2019, which is incorporated herein by reference; see, for example, Section 8.7.3. It relates to a technique by which a decoding process selects between a plurality of possible quantization parameters or groups of quantization parameters, for example, in response to properties of previously encoded and decoded sample values, such as parity properties. Thus, when dep_quant_enabled_flag = 1 (enabled), such persistent dependent quantization selection occurs. When dep_quant_enabled_flag = 0 (disabled), such persistent dependent quantization selection does not occur. As described above, for example, a flag dep_quant_enabled_flag is provided in the slice header (as an example of a coding property) so that enabling or disabling of dependent quantization applies to the entire slice.

[0180] When using correlated quantization, different constraints may be relevant, or at least possible, because it has been empirically observed that when correlated quantization is applied, it can change the expected relationship between the encoded and decoded bins. Different constraints (such as Equation 3) may be more suitable for use with correlated quantization.

[0181] Where multiple candidate constraints apply, it is desirable that the decoder be designed to provide sufficient processing power, speed, or capacity to process encoded data generated under the more challenging (or most challenging) of the different available constraints.

[0182] However, more generally, any such attribute (e.g., flag or parameter) may be used, whether or not explicitly signaled in or with the data stream. For example, different candidate expressions may be selected for different corresponding instances of a so-called "profile," where a profile in this context defines a set or basket of parameters such as bit depth, chroma sampling (e.g., 4:0:0 (monochrome), 4:2:0, 4:2:2, 4:4:4, etc.), restrictions on encoding types (such as only intra-picture encoding), etc.

[0183] In an example embodiment, it may be convenient to use an attribute that defines an aspect of an output data unit, as the relevant threshold derivation can then be applied to that output data unit. In this case, examples of output data units may include output data units that represent corresponding image portions, such as slices, tiles, or pictures. Other examples of suitable attributes may include: for slices: whether the slice type is intra-only or unrestricted (which may include inter-). As an example of tiles, consider the case where a picture is synthesized from multiple sources (one tile for each source). Each tile may have its own threshold derivation, depending on how it was originally encoded. Alternatively, the attribute value of the previous tile / output data unit at the same position in the picture may be used to determine the attribute.

[0184] In some examples, as described above, the constraint or threshold is defined by the following expression:

[0185] N<=K1*B+(K2*CU)

[0186] in:

[0187] N = the number of binary symbols in the output data unit;

[0188] K1 is a constant;

[0189] B = number of coded bytes of the output data unit;

[0190] K2 is a variable that depends on the properties of the minimum-size coding unit employed by the image data encoding device; and

[0191] CU = the size of the picture, slice, or tile represented by the output data unit, expressed as the number of smallest-sized coding units.

[0192] It should be noted that this is actually a generalization of Equations 1, 2, and 3 discussed above. An example list of candidate constraints involving K1 and K2 is as follows:

[0193]

[0194] Regarding the example of Equation 5, the variable vcIByteScaleFactor can be expressed as:

[0195] vclByteScaleFactor=(32+4*general_tier_flag) / 3

[0196] Here, general_tier_flag is an indicator of a coding layer and (in at least some examples) varies with flag values ​​0 or 1, where 0 indicates a so-called main layer and 1 indicates a so-called high layer. For a particular coding level (indicating the maximum size of an image to be encoded), the high layer typically corresponds to a higher bitrate representation than the main layer. Thus, in this example, the image data encoding device is configured to operate on a coding layer selected from at least two candidate coding layers and generate layer parameters defining the currently selected coding layer (e.g., encoded in or at least associated with the coded image data or bitstream), wherein at least the constant K1 depends on the layer parameters. For example, the high layer parameters may indicate a higher quality encoded output for a given image size, and the parameter K1 may increase with the layer parameters.

[0197] This allows the processing, circuitry, code, or logic used to generate the thresholds to be conveniently the same or substantially the same in each case, with parameters K1 and K2 simply changing with respect to each of the example Equations 1 / 2 through 7. However, it should be understood that one or more different equations or expressions (or even potentially different fixed thresholds) may be used, such as between the different candidate thresholds or constraints described above.

[0198] Thus, the candidate constraints 2082 can be stored or represented as (K1, K2) pairs selected by the selector 2080. The (K1, K2) pairs can be predetermined and stored, and the encoder and decoder or an indication of these pairs (or a subset of a larger set of predetermined pairs) can be sent from the encoder to the decoder as part of, for example, a profile or parameter set data. The expression N <= K1*B + (K2*CU) can then be tested at the comparator 2090 using the primary selected (K1, K2).

[0199] Figure 22A and Figure 22B Two illustrative examples are provided of output data units that extend across a page and include binary data 2200 provided and inserted by the techniques described above, and padding data 2210. The amount of padding data may depend on the specific image data being encoded (in terms of how efficiently it can be encoded) and the constraints in use.

[0200] Figure 23 is a schematic flow chart showing a method for encoding image data, comprising:

[0201] selectively encoding (at step 2300) data items representing image data to generate encoded binary symbols of successive output data units;

[0202] generating (at step 2310) a constrained output data stream, the constraint defining an upper limit on the number of binary symbols representable by any individual output data unit relative to a byte size of the output data unit;

[0203] providing (at step 2320) padding data for each output data unit that does not satisfy the constraint to increase the byte size of the output data unit so that the constraint is satisfied;

[0204] detecting (at step 2330) encoding properties applicable to a given output data unit; and

[0205] In response to the detected encoding properties, a constraint for a given output data unit is selected (at step 2340) from two or more candidate constraints.

[0206] Example embodiments also provide an image decoder including circuitry configured to interpret an encoded signal generated by controlling the image data encoding apparatus of any one or more of the embodiments described herein and output a decoded video image.

[0207] Example embodiments also provide an image decoder including circuitry configured to interpret encoded signals generated by controlling any one or more of the first and second data encoders and the controller of the embodiments described herein, and output a decoded video image.

[0208] The described embodiments may be implemented in any suitable form (including hardware, software, firmware or any combination thereof). The described embodiments may optionally be implemented at least in part as computer software running on one or more data processors and / or digital signal processors. The elements and components of any embodiment may be implemented physically, functionally and logically in any suitable manner. In fact, the function may be implemented as a single unit, as a plurality of units or as a part of other functional units. In this way, the disclosed embodiments may be implemented as a single unit, or may be physically and functionally distributed between different units, circuits and / or processors. Similarly, a data signal comprising coded data generated according to the above method (whether or not contained on a non-transitory machine-readable medium) is also considered to represent an embodiment of the present disclosure.

[0209] Obviously, many modifications and variations of the present disclosure are possible in light of the above teachings.It is therefore to be understood that, within the scope of the appended claims, the present technology may be practiced in ways other than as specifically described herein.

[0210] The corresponding aspects and features are defined by the following numbered clauses:

[0211] 1. An image data encoding device, comprising:

[0212] an entropy encoder configured to selectively encode data items representing image data so as to generate encoded binary symbols of successive output data units;

[0213] An entropy encoder is configured to generate an output data stream subject to a constraint, the constraint defining an upper limit on the number of binary symbols that can be represented by any individual output data unit relative to the byte size of the output data unit, wherein the entropy encoder is configured to provide padding data for each output data unit that does not satisfy the constraint in order to increase the byte size of the output data unit so as to satisfy the constraint; the apparatus comprising:

[0214] a property detector configured to detect encoding properties applicable to a given output data unit; and

[0215] A selector is configured to select a constraint for a given output data unit from two or more candidate constraints in response to the detected encoding properties.

[0216] 2. A device according to clause 1, wherein the entropy encoder is configured to selectively encode data items representing image data to be encoded by a first context-adaptive binary arithmetic coding (CABAC) coding system or a second bypass coding system to generate encoded binary symbols.

[0217] 3. The apparatus of clause 1, wherein the image data represents one or more pictures;

[0218] Each image includes data representing:

[0219] (i) one or more slices within a corresponding Network Abstraction Layer (NAL) unit, each slice of a picture being decodable independently of any other slice of the same picture; and

[0220] (ii) zero or more tiles, defining respective horizontal and vertical boundaries of the picture area and not limited to being encapsulated within respective NAL units, that are decodable independently of other tiles of the same picture; and

[0221] The output data unit includes one or more of a picture, a slice, and a tile.

[0222] 4. Apparatus according to clause 1 or clause 2, wherein the second coding system is a binary arithmetic coding system using a fixed 50% probability context model.

[0223] 5. The apparatus according to any one of the preceding clauses, wherein the constraint is defined by the following equation:

[0224] N<=K1*B+(K2*CU) (Constraint Equation 1)

[0225] in:

[0226] N = the number of binary symbols in the output data unit;

[0227] K1 is a constant;

[0228] B = number of coded bytes of the output data unit;

[0229] K2 is a variable that depends on the properties of the minimum-size coding unit employed by the image data encoding device; and

[0230] CU = the size of the picture, slice, or tile represented by the output data unit, expressed as the number of smallest-sized coding units.

[0231] 6. Apparatus according to clause 5, wherein:

[0232] At least two candidate constraints are defined by constraint equation 1, a corresponding set (K1, K2) is associated with each of the at least two candidate constraints; and

[0233] The selector is configured to select one of the sets (K1, K2) for a given output data unit.

[0234] 7. Apparatus according to any of the preceding clauses, wherein:

[0235] The controller is configured to encode, in association with the output data stream representing the given output data unit, a representation of encoding properties applicable to the given output data unit.

[0236] 8. Apparatus according to clause 7, wherein:

[0237] The image data encoding apparatus comprises: a quantizer configured to selectively operate in a related quantization mode; and

[0238] The encoding attribute indicates whether the associated quantization mode is enabled or disabled for a given output data unit.

[0239] 9. The apparatus of any preceding clause, wherein the entropy encoder comprises:

[0240] a detector configured to detect, at a predetermined stage relative to encoding of the current output data unit, whether the current output data unit will satisfy the constraint; and

[0241] The padding data generator is configured to generate sufficient padding data and insert the padding data into the current output data unit so that the output data unit including the inserted padding data satisfies the constraint.

[0242] 10. Apparatus according to clause 6, wherein the predetermined phase is the end of encoding a current output data unit.

[0243] 11. A video storage, capture, transmission or reception device comprising a device according to any of the preceding clauses.

[0244] 12. A method for encoding image data, comprising:

[0245] selectively encoding data items representing image data to generate encoded binary symbols of successive output data units;

[0246] generating a constrained output data stream, the constraint defining an upper limit on the number of binary symbols that can be represented by any individual output data unit relative to a byte size of the output data unit;

[0247] providing padding data for each output data unit that does not satisfy the constraint to increase the byte size of the output data unit so that the constraint is satisfied;

[0248] detecting encoding properties applicable to a given output data unit; and

[0249] In response to the detected encoding properties, a constraint for a given output data unit is selected from two or more candidate constraints.

[0250] 13. Computer software which, when executed by a computer, causes the computer to perform the method according to clause 12.

[0251] 14. A machine-readable non-transitory storage medium having stored thereon computer software according to clause 13.

[0252] 15. A data signal comprising encoded data generated according to the method of clause 12.

[0253] 16. An image data decoder configured to decode a data signal according to clause 15.

Claims

1. An image data encoding device, comprising: an entropy encoder configured to selectively encode data items representing image data so as to generate encoded binary symbols of successive output data units; the entropy encoder being configured to generate an output data stream subject to a constraint, the constraint defining an upper limit on the number of binary symbols that can be represented by any individual output data unit relative to a byte size of the output data unit, wherein the entropy encoder is configured to provide padding data for each output data unit that does not satisfy the constraint, so as to increase the byte size of the output data unit so as to satisfy the constraint; a property detector configured to detect a coding property applicable to a given output data unit, the coding property being a parameter or mode signaled by the encoding side; and A selector is configured to select a constraint for the given output data unit from two or more candidate constraints in response to the detected encoding properties.

2. The image data encoding device according to claim 1, wherein The entropy encoder is configured to selectively encode data items representing image data to be encoded by the first context adaptive binary arithmetic coding encoding system or by the second bypass encoding system so as to generate encoded binary symbols.

3. The image data encoding device according to claim 1, wherein The image data represents one or more pictures; Each image includes data representing: (i) one or more slices within the corresponding network abstraction layer unit, each slice of a picture being decodable independently of any other slice of the same picture; and (ii) zero or more tiles, defining respective horizontal and vertical boundaries of a picture region and not limited to being encapsulated within respective network abstraction layer units, said zero or more tiles being decodable independently of other tiles of said same picture; and The output data unit includes one or more of a picture, a slice, and a tile.

4. The image data encoding device according to claim 1, wherein The second coding system is a binary arithmetic coding system using a fixed 50% probability context model.

5. The image data encoding device according to claim 1, wherein The constraint is defined by the following constraint equation 1: N<=K1*B+(K2*CU), in: N = the number of binary symbols in the output data unit; K1 is a constant; B = the number of coded bytes of the output data unit; K2 is a variable that depends on the properties of the minimum-size coding unit employed by the image data encoding device; and CU = size of the picture, slice, or tile represented by this output data unit, expressed as the number of smallest-sized coding units.

6. The image data encoding device according to claim 5, wherein: At least two candidate constraints are defined by the constraint equation 1, a corresponding set (K1, K2) being associated with each of the at least two candidate constraints; and The selector is configured to select one of the sets (K1, K2) for the given output data unit.

7. The image data encoding device according to claim 1, wherein: The controller is configured to encode, in association with an output data stream representing the given output data unit, a representation of the encoding properties applicable to the given output data unit.

8. The image data encoding device according to claim 7, wherein: The image data encoding apparatus includes: a quantizer configured to selectively operate in a related quantization mode; and The encoding attribute indicates whether the associated quantization mode is enabled or disabled for the given output data unit.

9. The image data encoding device according to claim 1, wherein The entropy encoder comprises: a detector configured to detect, at a predetermined stage relative to encoding of a current output data unit, whether the current output data unit will satisfy the constraint; and A padding data generator is configured to generate sufficient padding data and insert the padding data into the current output data unit, so that the output data unit including the inserted padding data satisfies the constraint.

10. The image data encoding device according to claim 9, wherein The predetermined phase is the end of encoding the current output data unit.

11. A video storage, capture, transmission or reception device, comprising the image data encoding device according to claim 1.

12. A method for encoding image data, comprising: selectively encoding data items representing image data to generate encoded binary symbols of successive output data units; generating a constrained output data stream, said constraints defining an upper limit on the number of said base symbols that can be represented by any individual output data unit relative to the byte size of the output data unit; providing padding data for each output data unit that does not satisfy the constraint to increase the byte size of the output data unit so as to satisfy the constraint; detecting coding properties applicable to a given output data unit, said coding properties being parameters or modes signaled by the encoding side; and In response to the detected encoding properties, a constraint for the given output data unit is selected from two or more candidate constraints.

13. A machine-readable non-transitory storage medium storing computer software which, when executed by a computer, causes the computer to perform the method according to claim 12.

14. An image data decoding device, comprising: a buffer circuit configured to receive entropy-encoded binary symbols of data units subject to different constraints, said different constraints defining an upper limit on the number of said binary symbols that can be represented by any individual data unit relative to the byte size of the data unit, and in case said data unit does not satisfy the applied constraints, said data unit being provided with padding data to satisfy said applied constraints, said applied constraints being selected from said different constraints based on encoding properties representing an encoding mode or profile; A circuit comprising a context-adaptive binary arithmetic coding decoder and a bypass decoder is configured to receive and decode the entropy-encoded binary symbols of the data unit from the buffer circuit and output image data from the binary symbols.