Video data decoding method, apparatus, and video data encoding method
Discrete cosine transforms with 14-bit precision and integer arithmetic improve video data encoding and decoding efficiency, addressing inefficiencies in existing systems and reducing encoded data by up to 10-20% for high bit-depth video data.
Patent Information
- Application Number
- JP2025078241
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2020-04-03
- Filing Date
- 2025-05-08
- Publication Date
- 2025-08-26
AI Technical Summary
Existing video/image data encoding and decoding systems face challenges in efficiently compressing and decompressing video data due to inefficiencies in entropy coding and transform processes, particularly when dealing with high bit-depth video data.
The use of discrete cosine transforms (DCT) with 14-bit precision and integer arithmetic for video data encoding and decoding, along with reversible entropy encoding, to improve compression efficiency and accuracy.
Enhances compression efficiency by up to 10-20% for high bit-depth video data, reducing the amount of encoded data while maintaining image quality.
Smart Images

Figure 2025124666000001_ABST
Abstract
Description
[Technical Field]
[0001] TECHNICAL FIELD This disclosure relates to encoding and decoding video data. [Background technology]
[0002] The "Background" discussion provided herein is intended to generally present the context for the present disclosure. The inventors' work, to the extent that it is described in this Background section, is not expressly or implicitly admitted as prior art to the present disclosure, as are aspects of the description that may not otherwise qualify as prior art at the time of filing.
[0003] There are several video / image data encoding and decoding systems that involve transforming video data into a frequency domain representation, quantizing the frequency domain coefficients, and then applying some form of entropy coding to the quantized coefficients. This allows the video data to be compressed. To recover a reconstruction of the original video data, a corresponding decoding or decompression technique is applied. [Prior art documents] [Non-patent literature]
[0004] [Non-Patent Document 1] "Versatile Video Coding (Draft 8)", JVET-G2001-vE, B. Brass, J. Chen, S. Liu and YK. Wang Summary of the Invention [Problem to be solved by the invention]
[0005] The present disclosure addresses or alleviates the challenges that arise from this process. [Means for solving the problem]
[0006] Each aspect and feature of the present disclosure is defined in the accompanying claims.
[0007] It is to be understood that both the foregoing general description and the following detailed description are exemplary, but are not restrictive, of the present technology.
[0008] A more complete understanding of the present disclosure and many of the attendant advantages thereof will be readily obtained as the same becomes better understood by reference to the following detailed description when considered in connection with the accompanying drawings. [Brief explanation of the drawings]
[0009] [Figure 1] FIG. 1 shows a schematic diagram of an audio / video (A / V) data transmission and reception system using video data compression and decompression. [Figure 2] FIG. 2 shows a schematic diagram of a video display system that uses video data decompression. [Figure 3] FIG. 3 shows a schematic diagram of an audio / video storage system that uses video data compression and decompression. [Figure 4] FIG. 4 shows a schematic diagram of a video camera that uses video data compression. [Figure 5] FIG. 5 shows a schematic representation of a storage medium. [Figure 6] FIG. 6 shows a schematic representation of a storage medium. [Figure 7] FIG. 7 is a schematic diagram showing an overview of a video data compression and decompression device. [Figure 8] FIG. 8 shows a schematic diagram of the predictor. [Figure 9] FIG. 9 shows a schematic representation of a transform or inverse transform unit. [Figure 10a] Figure 10a shows the coefficients of the DTC2 transform schematically. [Figure 10b] Figure 10b shows the coefficients of the DTC2 transform schematically. [Figure 10c] Figure 10c shows the coefficients of the DTC2 transform schematically. [Figure 10d]Figure 10d shows the coefficients of the DTC2 transform schematically. [Figure 10e] Figure 10e shows the coefficients of the DTC2 transform schematically. [Figure 11] FIG. 11 shows the coefficients of the DTC2 transformation diagrammatically. [Figure 12] FIG. 12 shows the coefficients of the DCT8 transform schematically. [Figure 13] FIG. 13 shows the coefficients of the DST7 transform schematically. [Figure 14a] FIG. 14a shows a schematic representation of the flag structure. [Figure 14b] FIG. 14b is a schematic flow chart illustrating the encoding method. [Figure 14c] FIG. 14c is a schematic flow chart illustrating the decoding method. [Figure 15] FIG. 15 is a schematic flow chart illustrating a video data encoding method. DETAILED DESCRIPTION OF THE INVENTION
[0010] Referring now to the drawings, FIGS. 1 through 4 are provided to provide a schematic diagram of an apparatus or system that utilizes the compression and / or decompression apparatus described below in connection with embodiments of the present technology.
[0011] All of the data compression and / or decompression devices described below may be implemented in hardware, in software running on a general-purpose data processing device such as a general-purpose computer, as programmable hardware such as an application specific integrated circuit (ASIC) or field programmable gate array (FPGA), or as a combination thereof. It will be understood that when an embodiment is implemented by software and / or firmware, such software and / or firmware, as well as non-transitory data storage media on which such software and / or firmware is stored or otherwise provided, are considered embodiments of the present technology.
[0012] Figure 1 shows a schematic diagram of an audio / video data transmission and reception system using video data compression and decompression. In this example, the data values being encoded or decoded represent image data.
[0013] An input audio / video signal 10 is provided to a video data compressor 20, which compresses at least the video component of the audio / video signal 10 for transmission along a transmission path 30, such as a cable, optical fiber, or wireless link. The compressed signal is processed by a decompressor 40 to provide an output audio / video signal 50. On the return path, a compressor 60 compresses the audio / video signal for transmission along the transmission path 30 to a decompressor 70.
[0014] Thus, compressor 20 and decompressor 70 may form one node of a transmission link, and decompressor 40 and compressor 60 may form the other node of the transmission link. Of course, if the transmission link is unidirectional, only one of these nodes needs a compressor and the other a decompressor.
[0015] 2 illustrates a schematic diagram of a video display system employing video data decompression. In particular, a compressed audio / video signal 100 is processed by a decompressor 110 to provide a decompressed signal that can be displayed on a display 120. The decompressor 110 can be implemented as an integral part of the display 120, for example, provided in the same housing as the display device. Alternatively, the decompressor 110 can be provided (for example) as a so-called set-top box (STB). Note that the term "set-top" does not imply a requirement that the box be positioned in any particular orientation or position relative to the display 120, but is simply a term used in the art to denote a device that can be connected to a display as a peripheral device.
[0016] 3 shows a schematic diagram of an audio / video storage system using video data compression and decompression. An input audio / video signal 130 is provided to a compressor 140, which generates a compressed signal for storage by a storage device 150, such as a magnetic disk drive, an optical disk drive, a magnetic tape drive, a solid-state storage device such as a semiconductor memory, or other storage device. For playback, the compressed data is read from the storage device 150 and passed to a decompressor 160 for decompression, providing an output audio / video signal 170.
[0017] It will be understood that compressed or encoded signals and storage media, such as computer-readable non-transitory storage media, that store the signals are considered embodiments of the present technology.
[0018] Figure 4 shows a schematic diagram of a video camera that uses video data compression. In Figure 4, an imaging device 180, such as a charge-coupled device (CCD) image sensor and associated control and readout electronics, generates a video signal that is passed to a compression device 190. A microphone (or microphones) 200 generates an audio signal that is passed to the compression device 190. The compression device 190 generates a compressed audio / video signal 210 (shown generally as schematic stage 220) that is stored and / or transmitted.
[0019] The techniques described below relate primarily to video data compression and decompression. It will be understood that many existing techniques for audio data compression may be used in conjunction with the described video data compression techniques to generate compressed audio / video signals. Accordingly, a separate discussion of audio data compression will not be provided. It will also be understood that data rates associated with video data, particularly broadcast-quality video data, are generally much higher than data rates associated with audio data (whether compressed or uncompressed). It will therefore be understood that uncompressed audio data may accompany compressed video data to form a compressed audio / video signal. While the present embodiment (shown in FIGS. 1 through 4) relates to audio / video data, it will further be understood that the techniques described below have application in systems that simply process (i.e., compress, decompress, store, display, and / or transmit) video data. That is, the embodiments may be applied to video data compression without any processing of the associated audio data.
[0020] Thus, Figure 4 provides an example of a video capture device that includes an image sensor and encoding device of the type described below, and Figure 2 provides an example of a decoding device of the type described below and a display on which the decoded image is output.
[0021] The combination of Figures 2 and 4 provides a video capture device comprising an image sensor 180, an encoding device 190, a decoding device 110, and a display 120 on which the decoded image is output.
[0022] Figures 5 and 6 schematically illustrate storage media for storing compressed data generated by (for example) apparatus 20, 60, or input to apparatus 110, storage media, or stages 150, 220. Figure 5 schematically illustrates a disk storage medium, such as a magnetic or optical disk, while Figure 6 schematically illustrates a solid-state storage medium, such as flash memory. Note that Figures 5 and 6 may also provide examples of computer-readable non-transitory storage media for storing computer software that, when executed by a computer, causes the computer to perform one or more of the methods described below.
[0023] Thus, the above configurations provide examples of video storage, capture, transmission, or reception devices embodying any of the present techniques.
[0024] FIG. 7 is a schematic diagram illustrating an overview of a video / image data compression (encoding) and decompression (decoding) apparatus for encoding and / or decoding video / image data representing one or more images.
[0025] The controller 343 controls the operation of the entire apparatus, and in particular the trial encoding process by acting as a selector for selecting various modes of operation, such as block size and shape when referring to compression modes, and whether the video data should be losslessly encoded or otherwise encoded. The controller is considered to form part of the image encoder or image decoder (as the case may be). Successive images of the input video signal 300 are provided to the adder 310 and the image predictor 320. The image predictor 320 is described in more detail below with reference to FIG. 8. The addition of the intra-image predictor of FIG. 8 to the image encoder / decoder (as the case may be) can take advantage of the features of the apparatus of FIG. 7. However, this does not mean that the image encoder / decoder necessarily requires all the features of FIG. 7.
[0026] The adder 310 actually receives the input video signal 300 at its "+" input and the output of the image predictor 320 at its "-" input, and performs a subtraction (negative addition) operation to subtract the predicted image from the input image, resulting in a so-called residual image signal 330 representing the difference between the actual image and the predicted image.
[0027] One reason a residual image signal is generated is as follows: The described data encoding techniques, i.e., techniques to be applied to the residual image signal, tend to work more efficiently when there is less "energy" in the image to be encoded. Here, the term "efficiently" refers to the generation of a small amount of encoded data; at a particular image quality level, it is desirable (and considered "efficient") to generate as little data as possible. References to "energy" in the residual image relate to the amount of information contained in the residual image. If the predicted image is identical to the actual image, the difference between the two images (i.e., the residual image) contains zero information (zero energy) and is very easy to encode into a small amount of encoded data. In general, if the prediction process can be performed in a way that works reasonably well so that the content of the predicted image is similar to the content of the image to be encoded, it is expected that the residual image data will contain less information (less energy) than the input image and therefore be easier to encode into a small amount of encoded data.
[0028] Thus, encoding (using adder 310) involves predicting an image region of the image to be coded and generating a residual image region that depends on the difference between the predicted image region and the corresponding region of the image to be coded. In conjunction with the techniques described below, an ordered array of data values comprises data values of a representation of the residual image region. Decoding involves predicting an image region of the image to be decoded, generating a residual image region that indicates the difference between the predicted image region and the corresponding region of the image to be decoded, the ordered array of data values comprising data values of a representation of the residual image region, and combining the predicted and residual image regions.
[0029] We now describe the remaining part of the device, which acts as an encoder (for encoding the residual or difference image).
[0030] The residual image data 330 is provided to a transform unit / circuit 340, which generates a discrete cosine transform (DCT) representation of the block or region of residual image data. DCT technology itself is well known and will not be described in detail here. It should also be noted that the use of a DCT is merely illustrative of one exemplary configuration. Examples of other transforms that may be used include, for example, a discrete sine transform (DST). Transforms may also include sequences or cascades of individual transforms, such as one transform followed (directly or otherwise) by another. The choice of transform may be explicitly determined and / or may depend on side information used to configure the encoder and decoder. In another example, a so-called "transform skip" mode may be selectively used, in which no transform is applied.
[0031] Thus, in various embodiments, the encoding and / or decoding method comprises predicting an image region of the image to be coded and generating a residual image region that depends on the difference between the predicted image region and the corresponding region of the image to be coded, and the ordered array of data values (described below) comprises data values of a representation of the residual image region.
[0032] The output of transform unit 340, i.e., in one example, a set of DCT coefficients for each transformed block of image data, is provided to quantizer 350. A variety of quantization techniques are known in the field of video data compression, ranging from simple multiplication by a quantization scaling factor to the application of complex look-up tables under the control of a quantization parameter. Their general purpose is twofold: first, the quantization process reduces the number of possible values of the transformed data; and second, the quantization process can increase the likelihood that the value of the transformed data is zero. Both of these allow the entropy coding process described below to work more efficiently in generating small amounts of compressed video data.
[0033] A data scanning process is applied by the scanner 360. The objective of the scanning process is to reorder the quantized transformed data to collect as many non-zero quantized transform coefficients as possible, and therefore, of course, to collect as many zero-valued coefficients as possible. These characteristics can allow so-called run-length coding or similar techniques to be applied efficiently. The scanning process therefore involves selecting coefficients from the quantized transformed data, particularly from blocks of coefficients corresponding to blocks of transformed and quantized image data, in a "scan order" such that (a) all coefficients are selected once as part of the scan, and (b) the scan is likely to provide the desired reordering. One example of a scan order that is likely to give useful results is the so-called orthogonal diagonal scan order.
[0034] The scanning order may differ between transform-skip blocks and transform blocks (blocks that have undergone at least one spatial frequency transform).
[0035] The scanned coefficients are then passed to an entropy encoder (EE) 370. Again, various types of entropy coding may be used. Two examples include variations of the so-called Context Adaptive Binary Arithmetic Coding (CABAC) system and variations of the so-called Context Adaptive Variable-Length Coding (CAVLC) system. CABAC is generally considered more efficient, with some studies showing a 10% to 20% reduction in the amount of coded output data for comparable image quality compared to CAVLC. However, CAVLC is considered to have a much lower level of complexity (in terms of its implementation) than CABAC. Note that although the scanning and entropy coding processes are shown as separate processes, in practice they may be combined or treated together. That is, data may be read into the entropy encoder in scan order. The same applies to the respective inverse processes described below.
[0036] The output of the entropy encoder 370 provides a compressed output video signal 380, along with additional data (described above and / or below) that defines, for example, how the predictor 320 generated the predicted image, whether the compressed data was transformed or transform-skipped, etc.
[0037] However, since the operation of the predictor 320 itself depends on the decompression of the compressed output data, a return path 390 is also provided.
[0038] The reason for this feature is as follows: At an appropriate stage in the decompression process (described below), decompressed residual data is generated. This decompressed residual data must be added to the predicted image to generate the output image (since the original residual data is the difference between the input image and the predicted image). For this process to be similar between the compression and decompression sides, the predicted image generated by predictor 320 should be the same during the compression and decompression processes. Of course, during decompression, the device does not have access to the original input image, only the decompressed image. Therefore, during compression, predictor 320 bases its predictions (at least for inter-image coding) on the decompressed version of the compressed image.
[0039] The entropy encoding process performed by entropy encoder 370 is considered (at least in some instances) to be "reversible," i.e., it can be performed in reverse to arrive at the exact same data originally fed to entropy encoder 370. Therefore, in such instances, a return path may be implemented before the entropy encoding stage. Indeed, since the scanning process performed by scanner 360 is also considered to be reversible, in this embodiment, return path 390 is from the output of quantizer 350 to the input of complementary inverse quantizer 420. If a loss or potential loss is introduced by a stage, that stage (and vice versa) may be included in the feedback loop formed by the return path. For example, the entropy encoding stage may be made at least in principle irreversible, e.g., by techniques in which bits are encoded within parity information. In such cases, entropy encoding and decoding should form part of the feedback loop.
[0040] Generally, the entropy decoder 410, inverse scan unit 400, inverse quantizer 420, and inverse transform unit / circuitry 430 provide the inverse functions of the entropy encoder 370, scan unit 360, quantizer 350, and transform unit 340, respectively. We now proceed through the compression process. The process for decompressing the input compressed video signal is described separately below.
[0041] In the compression process, the scanned coefficients are passed by return path 390 from quantizer 350 to inverse quantizer 420, which performs the inverse of the operations of scanner 360. The inverse quantization and inverse transform processes are performed by inverse quantizer 420 and inverse transform unit / circuit 430 to produce a decompressed residual image signal 440.
[0042] Image signal 440 is added to the output of predictor 320 in adder 450 to produce a reconstructed output image 460 (although this may undergo so-called loop filtering and / or other filtering before being output, see below), which forms one input to image predictor 320, as described below.
[0043] Consider now the decoding process applied to decompress a received compressed video signal 470. The signal is fed to an entropy decoder 410, from which it is fed sequentially to an inverse scan unit 400, an inverse quantizer 420, and an inverse transform unit 430, before being added to the output of the image predictor 320 by an adder 450. Thus, on the decoder side, the decoder reconstructs a residual image version of the image and then applies this (by adder 450) to a predicted version of the image (block by block) to decode each block. In brief, the output 460 of the adder 450 forms the output decompressed video signal 480 (subject to a filtering process described below). In practice, further filtering may optionally be applied before the signal is output (e.g., by a loop filter 565 shown in FIG. 8 but omitted from FIG. 7 for clarity).
[0044] The apparatus of Figures 7 and 8 can act as a compression (encoding) apparatus or a decompression (decoding) apparatus. The functions of these two types of apparatus substantially overlap. Scanner 360 and entropy encoder 370 are not used in decompression mode, and the operation of predictor 320 and other units (described in more detail below) depends on mode and parameter information contained in the received compressed bitstream, rather than generating such information themselves.
[0045] FIG. 8 illustrates the generation of a predicted image, and in particular the operation of the image predictor 320.
[0046] There are two basic modes of prediction performed by the image predictor 320: so-called intra-picture prediction and so-called inter-picture or motion compensated (MC) prediction. On the encoder side, each involves finding a prediction direction for the current block to be predicted and generating a predictive block of samples according to other samples (in the same picture (intra-picture) or other pictures (inter-picture)). By adder 310 or 450, the difference between the predictive block and the actual block is coded or applied to code or decode the block, respectively.
[0047] (At the decoder, or at the inverse decoding side of the encoder, the detection of the prediction direction may be responsive to data associated with the data encoded by the encoder that indicates which direction was used at the encoder, or the detection may be responsive to the same factors as the decision was made at the encoder.)
[0048] Intra-picture prediction predicts the content of blocks or regions of an image based on data from within the same image. This corresponds to so-called I-frame coding in other video compression techniques. However, in contrast to I-frame coding, which codes the entire image with intra-coding, in this embodiment the choice between intra-coding and inter-coding can be made on a block-by-block basis, while in other embodiments the choice is still made on a picture-by-picture basis.
[0049] Motion compensated prediction is an example of inter-picture prediction, which makes use of motion information that attempts to define the source in other adjacent or nearby pictures of the picture details to be coded in the current picture. Thus, in an ideal case, the content of a block of image data in a predicted picture can be coded very simply as a reference (motion vector) that points to a corresponding block at the same or slightly different position in an adjacent picture.
[0050] A technique known as "block copy" prediction is in some ways a hybrid of the two predictions, since it uses a vector to indicate a block of samples that is displaced from the currently predicted block within the same image and is then copied to form the currently predicted block.
[0051] Returning to FIG. 8, two image prediction configurations (corresponding to intra-picture prediction and inter-picture prediction) are shown, the results of which are selected by multiplexer 500 under control of mode signal 510 (e.g., from controller 343) to provide a block of predicted images for feeding to adders 310 and 450. The selection is made according to which selection yields the lowest "energy" (which, as discussed above, can be thought of as the information content requiring encoding) and is signaled to the decoder in the encoded output data stream. In this context, image energy can be found, for example, by trial subtracting regions of two versions of the predicted image from the input image, squaring each pixel value of the difference image, summing the squared values, and identifying which of the two versions yields a lower mean-squared value of the difference image associated with that image region. In another example, trial encoding can be performed for each selection or potential selection, and then a selection is made according to the cost of each potential selection in terms of the number of bits required for encoding and / or image distortion.
[0052] In an intra-coding system, the actual prediction is made based on the image blocks received as part of signal 460 (the signal as filtered by loop filtering, see below), i.e., the prediction is made based on the coded and decoded image blocks so that the exact same prediction can be performed in the decompressor. However, the intra mode selector 520 can derive data from the input video signal 300 to control the operation of the intra-picture predictor 530.
[0053] For inter-picture prediction, a motion compensated (MC) predictor 540 uses motion information such as motion vectors derived by a motion estimator 550 from the input video signal 300. These motion vectors are applied by the motion compensated predictor 540 to a processed version of the reconstructed image 460 to generate blocks of inter-picture prediction.
[0054] Thus, each of the predictors 530 and 540 (operating with the estimator 550) acts as a detector for detecting the prediction direction for the current block being predicted, and as a generator for generating a predictive block of samples (forming part of the prediction passed to the adders 310 and 450) according to other samples defined by the prediction direction.
[0055] The processing applied to signal 460 will now be described.
[0056] First, the signal may be filtered by a so-called loop filter 565. Various types of loop filters can be used. One technique involves applying a “deblocking” filter to remove, or at least tend to reduce, the effects of the block-based processing and subsequent operations performed by the transform unit 340. Further techniques may be used, including applying a so-called sample adaptive offset (SAO) filter. Generally, in a sample adaptive offset filter, filter parameter data (derived in the encoder and communicated to the decoder) defines one or more offset amounts to be selectively combined by the sample adaptive offset filter with a given intermediate picture sample (a sample of the signal 460) depending on (i) the value of the given intermediate picture sample, or (ii) the values of one or more intermediate picture samples having a predetermined spatial relationship to the given intermediate picture sample.
[0057] An adaptive loop filter is also optionally applied, using coefficients derived by processing the reconstructed signal 460 and the input video signal 300. An adaptive loop filter is a type of filter that applies adaptive filter coefficients to the data to be filtered, using known techniques. That is, the filter coefficients may change depending on various factors. Data defining which filter coefficients to use is included as part of the encoded output data stream.
[0058] The techniques described below relate to the handling of parameter data related to the operation of the filter. The actual filtering operation (such as SAO filtering) may otherwise use known techniques.
[0059] The filtered output from loop filter 565 actually forms output video signal 480 when the device is operating as a decompressor. This output is also buffered in one or more image or frame stores 570, as storing successive images is a requirement for motion compensation prediction, particularly for generating motion vectors. To facilitate storage requirements, the stored images in image store 570 may be kept compressed and then decompressed for use in generating motion vectors. Any known compression / decompression system may be used for this specific purpose. The stored images may be passed to interpolation filter 580, which generates higher resolution images from the stored images. In this example, intermediate samples (subsamples) are generated such that the resolution of the interpolated image output by interpolation filter 580 is four times (in each dimension) the resolution of the image stored in image store 570 for the 4:2:0 luminance channel and eight times (in each dimension) the resolution of the image stored in image store 570 for the 4:2:0 chrominance channel. The interpolated image is passed as an input to a motion estimator 550 and also to a motion compensated predictor 540 .
[0060] Next, we will discuss how to partition an image for compression processing. At a basic level, the image to be compressed is considered as an array of blocks or regions of samples. Partitioning an image into such blocks or regions can be performed by a decision tree, such as those described in "SERIES H: AUDIOVISUAL AND MULTIMEDIA SYSTEMS Infrastructure of audiovisual services - Coding of moving video High efficiency video coding Recommendation ITU-T H.265 12 / 2016" and "High Efficiency Video Coding (HEVC) Algorithms and Architectures, Chapter 3, edited by Madhukar Budagavi, Gary J. Sullivan, Vivienne Sze; ISBN 978-3-319-06894-7; 2014," each of which is incorporated herein by reference in its entirety. Further background information is provided in Non-Patent Document 1. This non-patent document 1 is also incorporated herein by reference in its entirety.
[0061] In some examples, the resulting blocks or regions have a size and, in some cases, a shape that can be determined by the decision tree to generally follow the arrangement of image features within the image. This in itself can improve coding efficiency, because such configurations tend to group together samples that represent or follow similar image features. In some examples, square blocks or regions of different sizes (e.g., 4x4 samples, up to, e.g., 64x64 or larger blocks) are available for selection. In other exemplary configurations, blocks or regions of different shapes, such as rectangular blocks (e.g., vertically or horizontally oriented), can be used. Other non-square and non-rectangular blocks are also contemplated. As a result of dividing the image into such blocks or regions, each sample of the image (at least in this example) is assigned to only one such block or region.
[0062] (Transformation matrix) The following description relates to aspects of the transform unit 340 and the inverse transform unit 430. Note that, as noted above, the transform unit resides within the encoder, and the inverse transform unit resides in the return decoding path of the encoder and the decoding path of the decoder.
[0063] The transform and inverse transform units are intended to provide complementary transforms to and from the spatial frequency domain. That is, the transform unit 340 operates on a set of video data (or data derived from the video data, such as the difference or residual data described above) to generate a corresponding set of spatial frequency coefficients. The inverse transform unit 430 operates on a set of spatial frequency coefficients to generate a corresponding set of video data.
[0064] In practice, the transformation is implemented as a matrix calculation. A forward transform is defined by a transformation matrix, and the transformation is implemented by matrix-multiplying the transformation matrix by an array of sample values to generate a corresponding array of spatial frequency coefficients. In some embodiments, an array of samples M is left-multiplied by a transformation matrix T, and then the transformation matrix T is left-multiplied by the left-multiplied array of sample values M.T The output coefficient array is then right-multiplied by the transpose of TMT T It is defined as follows.
[0065] A property of matrices that represent this type of spatial frequency transform is that the transpose of a matrix is the same as its inverse, so in principle the forward and inverse transform matrices are related by a simple transpose.
[0066] In the proposed VVC standard, as it exists at the filing date of this application (see above), the transform is defined only as the inverse transform matrix used in the decoder function. The encoder transform matrix is not defined as such. The decoder (inverse) transform is defined to 6-bit precision.
[0067] As mentioned above, in principle, a suitable set of forward (encoder) matrices can be obtained as the transpose of the inverse matrix. However, the relationship between forward and inverse matrices only applies when values are represented with infinite precision. Representing values to a limited precision, such as 6 bits, means that the transpose is not necessarily the proper relationship between a forward matrix and its inverse.
[0068] It is recognized that creating or implementing integer arithmetic transform units and integer arithmetic inverse transform units may be easier, cheaper, faster, and / or less processor-intensive than implementing similar units using floating-point arithmetic. Accordingly, the matrix coefficients described below are represented as integers scaled up from the actual values required to perform the transform by a sufficient number of bits (powers of two) to allow the required coefficient precision. If necessary, the scale-up can be removed by bit shifting (division by a power of two) at another stage in the process. In other words, the actual transform coefficients are proportional to the values described below.
[0069] (Provide transformation matrix) 9 illustrates schematically an example of the (forward) transform unit 340 or the inverse transform unit 430, which includes a matrix data store 900. The matrix data store 900 stores appropriate data (see below), for example, in response to one or more of the current video data bit depth, the current block size, and the type of transform in use, from which a transform matrix (or, in other embodiments, an inverse transform matrix) is derived by a matrix generator 910.
[0070] The generated matrix is provided to matrix processor 920, which performs the matrix multiplication associated with the required transform (or inverse transform) that operates on processed block 930 to generate processing block 940. In the case of a forward transform, processed block 930 is, for example, the residual data from subtraction stage 310, and processing block 940 is provided to quantization stage 350. In the case of an inverse transform, processed block 930 is, for example, the dequantized block output by inverse quantizer 420, and processing block 940 will be provided to adder 450.
[0071] (Conversion tool) In some exemplary video data processing systems, a variety of transform tools are available for selection, either individually or as a series of transforms. Examples of these transform tools include the Discrete Cosine Transform (DCT) Type II (DCT2), the DCT Type VIII (DCT8), and the Discrete Sine Transform (DST) Type VII (DST7). Other exemplary tools are available but will not be described in detail herein.
[0072] The selection between tools can be made by any one or more of a variety of techniques, such as (i) a direct mapping between any one of a set of parameters including, for example, block size, prediction direction, prediction type (intra, inter), etc.; (ii) the results of one or more trial encoding processes; (iii) configuration data provided by a set of parameters; (iv) configuration carried over from an adjacent or temporarily preceding block, slice, or other region; (v) predetermined configuration settings, etc.
[0073] Whatever technique is used to select a transformation tool, this is not important to the teachings of this disclosure regarding the generation of matrix coefficients to implement a given transformation.
[0074] (High precision forward transformation matrix) This section, read in conjunction with the above description, describes a video data encoding method and apparatus for encoding an array of video data values using a forward transform matrix at, for example, a 14-bit level of precision.
[0075] The basis for this explanation is that improved results can be obtained by using a forward matrix matched to a standard inverse matrix with coefficients represented with a resolution higher than 6 bits. This is especially true when the system encodes high bit-depth video data (i.e., video data with a large "bit depth," or data precision expressed as a number of bits, e.g., 16-bit video data).
[0076] The matrix values described below are applicable, for example, to frequency transforming video data values according to a frequency transform, to generating an array of frequency transformed values by a matrix multiplication process using a transform matrix having a data precision greater than 6 bits and having coefficients defined by at least a subset of the values described below, and to a frequency transform unit (such as transform unit 340) that performs such operations.
[0077] (DCT2 matrix) Figures 10a to 10e show the 64 × 64 fundamental matrix M for the DCT2. 64collectively define . In particular, Figure 10a shows the relative configuration of the four matrix portions shown in Figures 10b through 10e. Figure 11 shows the coefficient values and an alternative naming protocol, where the left column of names (a through E) is used for matrix sizes up to 32x32, and the second column of names is used if a 64x64 matrix is required. The "14-bit" values in the third column are values provided as part of this disclosure. The corresponding "6-bit" values are previously proposed values provided for comparison purposes only.
[0078] Generally, for an NxN transform, where N is 2, 4, 8, or 16, a 64x64 transform matrix M 64 is subsampled to select a subset of N × N values, and the subset of values M N [x][y] is For x, y = 0 to (N-1), M N [x][y]=M 64 [x][(2 (6―log2(N)) )y] is defined as:
[0079] Thus, the operation of the configuration of FIG. 7, operating according to FIG. 9 and using the data of FIGS. 10a to 10e and 11, is: 1. A video data encoding method for encoding an array of video data values, comprising: a generating step of frequency transforming the video data values according to a frequency transform to produce an array of frequency transformed values by a matrix multiplication process using a transform matrix having 14 bits of data precision, the frequency transform being a discrete cosine transform; 64x64 transformation matrix M for 64x64 DCT transformation 64 A definition step of defining the matrix M 64 The definition steps are defined by the accompanying drawings 10a to 10e and 11. For an N × N transform, where N is 2, 4, 8, or 16, a 64 × 64 transform matrix M is used to select a subset of N × N values. 64 a sub-sampling step of sub-sampling said subset M N[x][y] is For x, y = 0 to (N-1), M N [x][y]=M 64 [x][(2 (6―log2(N)) )y] and the subsampling step defined by Contains 1 provides an example of a video data encoding method.
[0080] Similarly, the configurations of Figures 7 and 9, which operate as described above, 1. A data encoding apparatus for encoding an array of video data values, comprising: frequency transforming the video data values according to a frequency transform to generate an array of frequency transformed values by a matrix multiplication process using a transform matrix having 14-bit data precision, the frequency transform being a discrete cosine transform, and using a 64×64 transform matrix M for a 64×64 DCT transform; 64 Define the above matrix M 64 is defined by the accompanying drawings 10a to 10e and 11, and for an N×N transformation where N is 2, 4, 8, or 16, the N×N transformation matrix is a 64×64 transformation matrix M 64 and the subset M of the above values N [x][y] is For x, y = 0 to (N-1), M N [x][y]=M 64 [x][(2 (6―log2(N)) )y] A frequency conversion circuit configured as defined by Equipped with 1 provides an example of a data encoding device.
[0081] Applying the relationship defined in the above formula, the following example is provided:
[0082] (2x2DCT2) Matrix M2 is the coupling matrix M 64 is defined as the first two coefficients of every 32nd row in
[0083] [Table 1]
[0084] (4x4DCT2) Matrix M4 is the coupling matrix M 64 is defined as the first four coefficients of every 16th row in
[0085] [Table 2]
[0086] (8x8DCT2) Matrix M8 is the coupling matrix M 64 is defined as the first eight coefficients of every eighth row in
[0087] [Table 3]
[0088] 16x16DCT2 matrix M 16 is the coupling matrix M 64 is defined as the first 16 coefficients of every fourth row in
[0089] [Table 4]
[0090] (32x32DCT2) matrix M 32 is the coupling matrix M 64 is defined as the first 32 coefficients of every second row in
[0091] [Table 5]
[0092] (DCT2 64x64) 64x64DCT2 is the whole matrix M 64Use.
[0093] (DCT8 matrix) Figure 12 shows a schematic set of values for use in generating a transformation matrix for the DCT8 transform tool. As mentioned above, the "14-bit" column relates to newly proposed values, and the "6-bit" values are provided solely for comparison with previously proposed configurations.
[0094] The value is selected as follows: The operation of the arrangement of FIG. 7, operating according to FIG. 9 and using the data of FIG. 12, is: 1. A video data encoding method for encoding an array of video data values, comprising: a generating step of frequency transforming the video data values according to a frequency transform to produce an array of frequency transformed values by a matrix multiplication process using a transform matrix having 14 bits of data precision, the frequency transform being a discrete cosine transform; A definition step of defining a set of values as shown in Figure 12 of the accompanying drawings; For an N×N transformation where N is 4, 8, 16, or 32, the N×N transformation matrix M N a selection step of selecting a value of from the set of values defined by Tables 6 to 9 below; Contains 1 provides an example of a video data encoding method. (i) Regarding the 4x4 DCT8 transform,
[0095] [Table 6] (ii) Regarding the 8x8DCT8 transform,
[0096] [Table 7] (iii) Regarding the 16x16DCT8 transform,
[0097] [Table 8] (iv) Regarding the 32x32DCT8 transform,
[0098] [Table 9]
[0099] Similarly, the configurations of Figures 7 and 9, which operate as described above, 1. A data encoding apparatus for encoding an array of video data values, comprising: a frequency transformation circuit configured to frequency transform video data values according to a frequency transformation to generate an array of frequency transformed values by a matrix multiplication process using a transformation matrix having 14 bits of data precision, said frequency transformation being a discrete cosine transform and defining a set of values as shown in accompanying Figure 12, wherein for an NxN transformation where N is 4, 8, 16, or 32, the NxN transformation matrix contains selected values from a set of values defined by a subset of the enumerated values; Equipped with 1 provides an example of a data encoding device.
[0100] (DST7 conversion) Figure 13 shows exemplary data for use in generating a DST7 transformation matrix. As mentioned above, the "6-bit" column is provided solely for comparison purposes.
[0101] The operation of the arrangement of FIG. 7, operating according to FIG. 9 and using the data of FIG. 13, is: 1. A video data encoding method for encoding an array of video data values, comprising: a generating step of frequency transforming the video data values according to a frequency transform to generate an array of frequency transformed values by a matrix multiplication process using a transform matrix having 14 bits of data precision, the frequency transform being a discrete sine transform; A definition step for defining a set of values as shown in Figure 13 attached herewith; For an N×N transformation where N is 4, 8, 16, or 32, the N×N transformation matrix M Na selection step of selecting the value of from the set of values defined by Tables 10 to 13 below; Contains 1 provides an example of a video data encoding method. (i) Regarding 4x4DST7 conversion,
[0102] [Table 10] (ii) For 8x8DST7 transformation,
[0103] [Table 11] (iii) Regarding 16x16DST7 conversion,
[0104] [Table 12] (iv) Regarding 32x32DST7 conversion,
[0105] [Table 13]
[0106] Similarly, the configurations of Figures 7 and 9, which operate as described above, 1. A data encoding apparatus for encoding an array of video data values, comprising: a frequency conversion circuit configured to frequency convert the video data values according to a frequency conversion to generate an array of frequency converted values by a matrix multiplication process using a conversion matrix having 14 bits of data precision, said frequency conversion being a discrete sine conversion and defining a set of values as shown in Figure 13 attached hereto, wherein for an NxN conversion where N is 4, 8, 16 or 32, the NxN conversion matrix includes selected values from said enumerated set of values; Equipped with 1 provides an example of a data encoding device.
[0107] (flag encoding) Previous exemplary configurations aimed at operating at high bit depths have proposed a so-called "extended precision flag" as a parameter that can be set in association with a video data stream. In some previous examples, the effect of setting this flag was to enable a potential increase in the dynamic range used during the computation of the forward and inverse spatial frequency transforms, so as to increase the precision of the transform values passed (forward) from the transform unit to the entropy encoder. The flag may be coded as part of a so-called sequence parameter set (SPS).
[0108] In an exemplary embodiment of the present disclosure, an extended precision flag may be provided among one or more sets of other flags, for example, in an SPS. Figure 14 provides an example of a hierarchy of such flags, including a high bit-depth control flag. If that flag is not set, the extended precision flag is not available, but if that flag is set, the extended precision flag may be set or unset. The overall high bit-depth control flag allows other features related to high bit-depth (e.g., involving more than 10 bits of video data) operation (e.g., at least in principle, such as high precision quantizers or alternative coefficient encoders for transform or non-transform blocks) to be switched on or off, thus providing a mechanism for encoding and allowing for future modifications of encoding tools.
[0109] Figure 14b shows 1. A data encoding method for encoding video data values, comprising: (at step 1400) selectively encoding a high bit-depth control flag, and when said high bit-depth control flag is set to indicate high bit-depth operation, selectively encoding an extended precision flag to indicate extended precision operation of at least the spatial frequency transform stage; (at step 1410) encoding the video data values in accordance with an operating mode defined by the encoded high bit-depth control flag and, if the extended precision flag is encoded, the encoded extended precision flag; 1 is a schematic flow chart illustrating a data encoding method.
[0110] Similarly, on the decoding side, Fig. 14c shows 1. A data decoding method for decoding video data values, comprising: (at step 1420) selectively decoding a high bit-depth control flag, and when said high bit-depth control flag is set to indicate high bit-depth operation, selectively decoding an extended precision flag indicating extended precision operation of at least the spatial frequency transform stage; (at step 1430) encode the video data values according to a mode of operation defined by the decoded high bit depth control flag and, if the extended precision flag is decoded, the decoded extended precision flag. 1 is a schematic flow chart illustrating a data decoding method.
[0111] The high bit depth control flag and (optionally) the extended precision flag may be encoded or decoded from a sequence parameter set of the video data stream.
[0112] Thus, the apparatus of FIG. 7, operating according to these methods, 1. A data encoding apparatus for encoding video data values, comprising: a parameter encoder (343) configured to selectively encode a high bit-depth control flag and, when said high bit-depth control flag is set to indicate high bit-depth operation, to selectively encode an extended precision flag indicating extended precision operation of at least the spatial frequency transform stage; an encoder (FIG. 7) configured to encode the video data values in accordance with an operating mode defined by the encoded high bit-depth control flag and, when the extended precision flag is encoded, the encoded extended precision flag; Equipped with 1 provides an example of a data encoding device.
[0113] Similarly, the apparatus of FIG. 7, which operates according to these methods, 1. A data decoding apparatus for decoding video data values, comprising: a parameter decoder (343) configured to selectively decode a high bit-depth control flag and, when said high bit-depth control flag is set to indicate high bit-depth operation, to selectively decode an extended precision flag indicating extended precision operation of at least the spatial frequency transform stage; a decoder (FIG. 7) configured to decode the video data values according to an operating mode defined by the decoded high bit-depth control flag and, when the extended precision flag is decoded, the decoded extended precision flag; Equipped with 1 provides an example of a data decoding device.
[0114] (Method Overview - Matrix Generation) Figure 15 shows (at step 1500) frequency transforming the video data values according to a frequency transform to produce an array of frequency transformed values by a matrix multiplication process using one or more of the transforms defined above; (at step 1510) defining a set of values as shown in the figure that are associated with the transformation; (at step 1520) For an N×N transformation, the N×N transformation matrix M N Select a value for from a set of values provided using the above techniques. 1 is a schematic flow chart illustrating a video data encoding method.
[0115] (Image data) Image / video data encoded or decoded using these techniques and carriers bearing such image data are considered to represent embodiments of the present disclosure.
[0116] Insofar as embodiments of the present disclosure are described as being implemented, at least in part, by a software-controlled data processing apparatus, it will be understood that a computer-readable, non-transitory medium bearing such software, such as an optical disk, magnetic disk, semiconductor memory, etc., is also considered to represent an embodiment of the present disclosure. Similarly, a data signal (whether embodied on a computer-readable, non-transitory medium or not) containing encoded data generated according to the above-described method is also considered to represent an embodiment of the present disclosure.
[0117] Obviously, numerous modifications and variations of the present disclosure are possible in light of the above teachings. It is therefore understood that, within the scope of the appended clauses, the present technology may be practiced other than as specifically described herein.
[0118] It will be appreciated that in the above description, for purposes of clarity, embodiments have been described with reference to different functional units, circuits and / or processors. However, it will be apparent that functionality may be distributed in any suitable manner between different functional units, circuits and / or processors without detracting from embodiments of the invention.
[0119] The embodiments described herein may be implemented in any suitable form including hardware, software, firmware, or any combination of these. The embodiments described herein may optionally be implemented at least in part as computer software running on one or more data processors and / or digital signal processors. The elements and components in any embodiment may be physically, functionally, and logically implemented in any suitable way. Indeed, functionality may be implemented in a single unit, in multiple units, or as part of other functional units. Thus, embodiments of the present disclosure may be implemented in a single unit, or may be physically and functionally distributed between different units, circuits, and / or processors.
[0120] Although the present disclosure has been described in connection with some embodiments, it is not intended to be limited to the specific form set forth herein. Moreover, while features of the present disclosure may be described in connection with particular embodiments, those skilled in the art will recognize that various features of the described embodiments may be combined in any manner suitable for practicing the present technology.
[0121] Each aspect and feature is defined by the following numbered clauses: 1. A video data encoding method for encoding an array of video data values, comprising: a generating step of frequency transforming the video data values according to a frequency transform to produce an array of frequency transformed values by a matrix multiplication process using a transform matrix having 14 bits of data precision, the frequency transform being a discrete cosine transform; 64x64 transformation matrix M for 64x64 DCT transformation 64 A definition step of defining the matrix M 64 the definition steps defined by the accompanying drawings 10a to 10e and 11; For an N × N transform, where N is 2, 4, 8, or 16, a 64 × 64 transform matrix M is used to select a subset of N × N values. 64 a sub-sampling step of sub-sampling said subset M N [x][y] is For x, y = 0 to (N-1), M N [x][y]=M 64 [x][(2 (6―log2(N)) )y] and the subsampling step defined by Contains Video data encoding method. 2. Image data encoded by the video data encoding method described in 1 above. 3. Computer software which, when executed by a computer, causes said computer to carry out the video data encoding method described in 1. above. 4. A computer-readable non-transitory storage medium storing the computer software described in 3 above. 5. A data encoding apparatus for encoding an array of video data values, comprising: frequency transforming the video data values according to a frequency transform to generate an array of frequency transformed values by a matrix multiplication process using a transform matrix having 14-bit data precision, the frequency transform being a discrete cosine transform, and using a 64×64 transform matrix M for a 64×64 DCT transform; 64 Define the above matrix M 64 is defined by the accompanying drawings 10a to 10e and 11, and for an N×N transformation where N is 2, 4, 8, or 16, the N×N transformation matrix is a 64×64 transformation matrix M 64 and the subset M of the above values N [x][y] is For x, y = 0 to (N-1), M N [x][y]=M 64 [x][(2 (6―log2(N)) )y] A frequency conversion circuit configured as defined by Equipped with Data encoding device. 6. A video data capture, transmission, display and / or storage device comprising the data encoding device described in 5 above. 7. A video data encoding method for encoding an array of video data values, comprising: a generating step of frequency transforming the video data values according to a frequency transform to produce an array of frequency transformed values by a matrix multiplication process using a transform matrix having 14 bits of data precision, the frequency transform being a discrete cosine transform; A definition step of defining a set of values as shown in Figure 12 of the accompanying drawings; For an N×N transformation where N is 4, 8, 16, or 32, the N×N transformation matrix M N a selection step of selecting the value of from the set of values defined by Tables 14 to 17 below; Contains Video data encoding method. (i) Regarding the 4x4 DCT8 transform,
[0122] [Table 14] (ii) Regarding the 8x8DCT8 transform,
[0123] [Table 15] (iii) Regarding the 16x16DCT8 transform,
[0124] [Table 16] (iv) Regarding the 32x32DCT8 transform,
[0125] [Table 17] 8. Image data encoded by the video data encoding method described in 7 above. 9. Computer software which, when executed by a computer, causes said computer to carry out the video data encoding method set forth in 7 above. 10. A computer-readable non-transitory storage medium storing the computer software described in 9 above. 11. A data encoding apparatus for encoding a sequence of video data values, comprising: a frequency conversion circuit configured to frequency convert video data values according to a frequency conversion to generate an array of frequency converted values by a matrix multiplication process using a conversion matrix having 14 bits of data precision, said frequency conversion being a discrete cosine transform and defining a set of values as shown in Figure 12 attached hereto, and for an NxN conversion where N is 4, 8, 16, or 32, the NxN conversion matrix is configured to include selected values from the set of values defined by the following Tables 18 to 21: Equipped with Data encoding device. (i) Regarding the 4x4 DCT8 transform,
[0126] [Table 18] (ii) Regarding the 8x8DCT8 transform,
[0127] [Table 19] (iii) Regarding the 16x16DCT8 transform,
[0128] [Table 20] (iv) Regarding the 32x32DCT8 transform,
[0129] [Table 21] 12. A video data capture, transmission, display and / or storage device comprising the data encoding device described in 11 above. 13. A video data encoding method for encoding an array of video data values, comprising: a generating step of frequency transforming the video data values according to a frequency transform to generate an array of frequency transformed values by a matrix multiplication process using a transform matrix having 14 bits of data precision, the frequency transform being a discrete sine transform; A definition step for defining a set of values as shown in Figure 13 attached herewith; For an N×N transformation where N is 4, 8, 16, or 32, the N×N transformation matrix M N a selection step of selecting the value of from the set of values defined by Tables 22 to 25 below; Equipped with Video data encoding method. (i) Regarding 4x4DST7 conversion,
[0130] [Table 22] (ii) For 8x8DST7 transformation,
[0131] [Table 23] (iii) Regarding 16x16DST7 conversion,
[0132] [Table 24] (iv) For 32x32DST7 conversion,
[0133] [Table 25] 14. Image data encoded by the video data encoding method described in 13 above. 15. Computer software which, when executed by a computer, causes said computer to carry out the video data encoding method set forth in 13 above. 16. A computer-readable non-transitory storage medium storing the computer software described in 15 above. 17. A data encoding apparatus for encoding an array of video data values, comprising: a frequency conversion circuit configured to frequency convert video data values according to a frequency conversion to generate an array of frequency converted values by a matrix multiplication process using a conversion matrix having 14 bits of data precision, said frequency conversion being a discrete sine conversion and defining a set of values as shown in Figure 13 attached hereto, and for an NxN conversion where N is 4, 8, 16, or 32, the NxN conversion matrix is configured to include selected values from the set of values defined by the following Tables 26 to 29: Equipped with Data encoding device. (i) Regarding 4x4DST7 conversion,
[0134] [Table 26] (ii) For 8x8DST7 transformation,
[0135] [Table 27] (iii) Regarding 16x16DST7 conversion,
[0136] [Table 28] (iv) For 32x32DST7 conversion,
[0137] [Table 29] 18. A video data capture, transmission, display and / or storage device comprising the data encoding device described in 17 above. 19. A data encoding method for encoding video data values, comprising: selectively encoding a high bit-depth control flag, and when said high bit-depth control flag is set to indicate high bit-depth operation, selectively encoding an extended precision flag indicating extended precision operation of at least the spatial frequency transform stage; encoding the video data values in accordance with an operating mode defined by the encoded high bit-depth control flag and, if the extended precision flag is encoded, the encoded extended precision flag; Data encoding method. 20. The data encoding method according to claim 19, Selectively encoding the high bit depth control flag and the extended precision flag into a sequence parameter set of a video data stream. Data encoding method. 21. Computer software which, when executed by a computer, causes said computer to carry out the data encoding method set forth in 19 above. 22. A computer-readable non-transitory storage medium storing the computer software described in 21 above. 23. A data decoding method for decoding video data values, comprising: selectively decoding a high bit-depth control flag, and when said high bit-depth control flag is set to indicate high bit-depth operation, selectively decoding an extended precision flag indicating extended precision operation of at least the spatial frequency transform stage; encoding the video data values according to an operating mode defined by the decoded high bit depth control flag and, if the extended precision flag is decoded, the decoded extended precision flag; Data decryption method. 24. The data decoding method according to claim 23, Selectively decoding the high bit depth control flag and the extended precision flag from a sequence parameter set of a video data stream. Data decryption method. 25. Computer software which, when executed by a computer, causes said computer to carry out the data decryption method set forth in 23 above. 26. A computer-readable non-transitory storage medium storing the computer software described in 25 above. 27. A data encoding device for encoding video data values, comprising: a parameter encoder configured to selectively encode a high bit-depth control flag and, when said high bit-depth control flag is set to indicate high bit-depth operation, to selectively encode an extended precision flag indicating extended precision operation of at least the spatial frequency transform stage; an encoder configured to encode the video data values in accordance with an operating mode defined by the encoded high bit-depth control flag and, when the extended precision flag is encoded, the encoded extended precision flag; Equipped with Data encoding device. 28. A video data capture, transmission, display and / or storage device comprising a data encoding device as described in 27 above. 29. A data decoding device for decoding video data values, comprising: a parameter decoder configured to selectively decode a high bit-depth control flag and, when said high bit-depth control flag is set to indicate high bit-depth operation, to selectively decode an extended precision flag indicating extended precision operation of at least the spatial frequency transform stage; a decoder configured to decode the video data values according to an operational mode defined by the decoded high bit depth control flag and, when the extended precision flag is decoded, the decoded extended precision flag; Equipped with Data decoding device. 30. A video data capture, transmission, display and / or storage device comprising the data decoding device described in 29 above.
Claims
1. 1. A video data decoding method for decoding video data values, comprising: selectively decoding a control flag in a control flag hierarchy including a first control flag and a second control flag, the second control flag being an extended precision flag; and decoding, by a circuit, the extended precision flag and the video data value according to functions defined by the extended precision flag for a first extended precision function for a conversion process and a second extended precision function different from the first extended precision function, depending on whether the first control flag is set. Video data decoding method.
2. 2. A video data decoding method according to claim 1, comprising: The first control flag is a flag related to an extended operation of the decoding device. Video data decoding method.
3. 3. A video data decoding method according to claim 2, comprising: The first control flag is set to indicate an operation greater than 10-bit operation. Video data decoding method.
4. 2. A video data decoding method according to claim 1, comprising: The second extended precision function relates to a value that replaces the decoding function. Video data decoding method.
5. 2. A video data decoding method according to claim 1, comprising: The second control flag can be set to a set state or an unset state. Video data decoding method.
6. 2. A video data decoding method according to claim 1, comprising: The second control flag is decoded only if the first control flag is decoded as set in the video data value. Video data decoding method.
7. 2. A video data decoding method according to claim 1, comprising: Parameters related to functions defined by the extended precision flag are decoded if the first control flag is decoded as set in the video data value. Video data decoding method.
8. 2. A video data decoding method according to claim 1, comprising: The decoded video data value includes a first control flag that is not set if the decoding operation is in an operation mode equal to or less than 10-bit operation. Video data decoding method.
9. 2. A video data decoding method according to claim 1, comprising: and selectively decoding the control flags in the hierarchy of control flags from a sequence parameter set of the video data stream. Video data decoding method.
10. A computer-readable non-transitory storage medium storing computer software that, when executed by a computer, causes the computer to perform the video data decoding method of claim 1.
11. 1. A video data decoding apparatus for decoding video data values, comprising: selectively decoding a control flag in a control flag hierarchy including a first control flag and a second control flag, the second control flag being an extended precision flag; and decoding, by a circuit, the extended precision flag and the video data value according to functions defined by the extended precision flag for a first extended precision function for a conversion process and a second extended precision function different from the first extended precision function, depending on whether the first control flag is set. The circuit is configured as follows: Video data decoding device.
12. 12. A video data decoding device according to claim 11, The first control flag is set to indicate an operation extending beyond 10-bit operation. Video data decoding device.
13. 12. A video data decoding device according to claim 11, The second extended precision function relates to a value that replaces the decoding function. Video data decoding device.
14. 12. A video data decoding device according to claim 11, configured to detect whether the second control flag is set or not set. Video data decoding device.
15. 12. A video data decoding device according to claim 11, configured to decode the second control flag only if the first control flag is decoded as set in the video data value. Video data decoding device.
16. 12. A video data decoding device according to claim 11, configured to decode parameters related to functions defined by the extended precision flag when the first control flag is decoded as set in the video data value. Video data decoding device.
17. 12. A video data decoding device according to claim 11, The decoded video data value includes a first control flag that is not set if the decoding operation is in an operation mode equal to or less than 10-bit operation. Video data decoding device.
18. 12. An apparatus for capturing, transmitting, displaying and / or storing video data, comprising a video data decoding apparatus according to claim 11.
19. 1. A video data encoding method for encoding video data values, comprising the steps of: selectively encoding a control flag in a control flag hierarchy including a first control flag and a second control flag, the second control flag being an extended precision flag; encoding the extended precision flag and the video data value according to functions defined by the extended precision flag for a first extended precision function related to a conversion process and a second extended precision function different from the first extended precision function, depending on whether the first control flag is set; Video data encoding method.