Video data encoding and decoding
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-03-12
- Publication Date
- 2026-08-11
AI Technical Summary
[0004]本公开解决或减轻了由该处理引起的问题
[0006] It should be understood that the foregoing general description and the following detailed description are exemplary and not intended to limit the technology.
Smart Images

Figure CN115380533B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to video data encoding and decoding. Background Technology
[0002] The “Background” description provided herein is intended to generally present the background of this disclosure. To the extent described in this Background section, the work of the currently named inventors, and aspects of the description that may not constitute prior art at the time of filing, are neither explicitly nor implicitly considered to be prior art to this disclosure.
[0003] Several systems exist, such as video or image data encoding and decoding systems, which involve transforming video data into a frequency domain representation, quantizing the frequency domain coefficients, and then applying some form of entropy encoding to the quantized coefficients. This enables the compression of video data. Appropriate decoding or decompression techniques are then applied to recover a reconstructed version of the original video data. Summary of the Invention
[0004] This disclosure resolves or mitigates the problems arising from this treatment.
[0005] The various aspects and features of this disclosure are defined in the appended claims.
[0006] It should be understood that the foregoing general description and the following detailed description are exemplary and not intended to limit the technology. Attached Figure Description
[0007] A more complete and better understanding of this disclosure and its many accompanying advantages will be readily obtained when considered in conjunction with the accompanying drawings and by referring to the following detailed description, in which:
[0008] Figure 1 An audio / video (A / V) data transmission and reception system using video data compression and decompression is illustrated schematically;
[0009] Figure 2 A video display system that uses video data decompression is illustrated schematically;
[0010] Figure 3 An audio / video storage system using video data compression and decompression is illustrated schematically;
[0011] Figure 4 A camera using video data compression is illustrated schematically;
[0012] Figure 5 and Figure 6 The storage medium is shown schematically;
[0013] Figure 7 A schematic diagram of a video data compression and decompression device is provided;
[0014] Figure 8 The predictor is illustrated schematically;
[0015] Figure 9 The transformation unit or inverse transformation unit is schematically shown;
[0016] Figures 10a to 10e and Figure 11 schematically show the coefficients of the DTC2 transform;
[0017] Figure 12 schematically illustrates the coefficients of the DCT8 transform;
[0018] Figure 13 schematically shows the coefficients of the DST7 transform;
[0019] Figure 14a The symbol structure is illustrated schematically; and
[0020] Figure 14b , Figure 14c and Figure 15 This is a schematic flowchart illustrating the corresponding method. Detailed Implementation
[0021] Now, referring to the attached diagram, we provide... Figures 1 to 4 A schematic diagram of an apparatus or system utilizing compression and / or decompression devices is provided, which will be described below in conjunction with embodiments of the present technology.
[0022] All data compression and / or decompression devices described below can be implemented in hardware, software running on general-purpose data processing equipment (such as a general-purpose computer), programmable hardware (such as application-specific integrated circuits (ASICs) or field-programmable gate arrays (FPGAs)), or combinations thereof. Where the implementation is implemented by software and / or firmware, it is understood that such software and / or firmware, as well as non-transitory data storage media storing or otherwise providing such software and / or firmware, are considered implementations of this technology.
[0023] Figure 1 This diagram schematically illustrates a system for transmitting and receiving audio / video data using video data compression and decompression. In this example, the data values to be encoded or decoded represent image data.
[0024] Input audio / video signal 10 is provided to video data compression device 20, which compresses at least the video component of the audio / video signal 10 for transmission along transmission path 30 (e.g., cable, fiber optic, wireless link, etc.). The compressed signal is processed by decompression device 40 to provide output audio / video signal 50. For the return path, compression device 60 compresses the audio / video signal for transmission along transmission path 30 to decompression device 70.
[0025] Therefore, compression device 20 and decompression device 70 can form one node of the transmission link. Decompression device 40 and compression device 60 can form another node of the transmission link. Of course, in the case of a unidirectional transmission link, only one node needs a compression device, while the other node only needs a decompression device.
[0026] Figure 2 A video display system using video data decompression is schematically illustrated. Specifically, compressed audio / video signal 100 is processed by decompression device 110 to provide a decompressed signal that can be displayed on display 120. Decompression device 110 can be implemented as part of display 120, for example, provided within the same housing as the display device. Alternatively, decompression device 110 can be provided as, for example, a so-called set-top box (STB). Note that the expression "set-top box" does not imply that the set-top box is located in any particular orientation or position relative to display 120; it is simply a term used in the art to refer to a device that can be connected to the display as a peripheral device.
[0027] Figure 3 An audio / video storage system using video data compression and decompression is schematically illustrated. Input audio / video signal 130 is provided to compression device 140, which generates a compressed signal for storage by storage device 150, such as a disk drive, optical disk drive, magnetic tape drive, solid-state storage device (e.g., semiconductor memory), or other storage device. For playback, compressed data is read from storage device 150 and passed to decompression device 160 for decompression to provide output audio / video signal 170.
[0028] It is understood that compressed or encoded signals and storage media storing signals (e.g., machine-readable non-transitory storage media) are considered implementations of this technology.
[0029] Figure 4 A camera using video data compression is illustrated schematically. Figure 4 In this configuration, an image capture device 180 (such as a charge-coupled device (CCD) image sensor) and associated control and readout electronics generate a video signal that is transmitted to a compression device 190. A microphone (or multiple microphones) 200 generates an audio signal to be transmitted to the compression device 190. The compression device 190 generates a compressed audio / video signal 210 (generally shown as schematic stage 220) to be stored and / or transmitted.
[0030] The techniques described below primarily relate to video data compression and decompression. It is understood that many existing techniques can be used in conjunction with the video data compression techniques described herein for audio data compression to generate compressed audio / video signals. Therefore, a separate discussion of audio data compression will not be provided. It is also understood that the data rate associated with video data (especially broadcast-quality video data) is typically much higher than the data rate associated with audio data (whether compressed or uncompressed). Therefore, it is understood that uncompressed audio data can be accompanied by compressed video data to form a compressed audio / video signal. It should also be understood that, although this example (… Figures 1 to 4 The diagrams shown involve audio / video data, but the techniques described for use in the system may be found to handle only video data (i.e., compress, decompress, store, display, and / or transmit). In other words, these implementations can be applied to video data compression without necessarily having any associated audio data processing.
[0031] therefore, Figure 4 Examples of video capture devices, including image sensors and encoding devices of the types discussed below, are provided. Therefore, Figure 2 Examples of the types of decoding devices and the displays to which the decoded images are output are provided below.
[0032] Figure 2 and Figure 4 The combination can provide a video capture device, including an image sensor 180 and an encoding device 190, a decoding device 110, and a display 120 on which the decoded image is output.
[0033] Figure 5 and Figure 6 The storage medium is schematically shown, which stores, for example, compressed data generated by device 20, device 60, compressed data input to device 110, or storage medium or stage 150, stage 220. Figure 5 The illustration schematically shows a disk storage medium such as a magnetic disk or optical disk, and Figure 6 This schematically illustrates a solid-state storage medium such as flash memory. Note that... Figure 5 and Figure 6 Examples may also be provided of non-transitory machine-readable storage media for storing computer software that, when executed by a computer, causes the computer to perform one or more of the methods discussed below.
[0034] Therefore, the above arrangement provides examples of video storage, capture, transmission, or reception devices embodying any of the present technologies.
[0035] Figure 7A schematic diagram of a video or image data compression (encoding) and decompression (decoding) apparatus for encoding and / or decoding video or image data representing one or more images is provided.
[0036] The controller 343 controls the overall operation of the device, particularly when compression modes are involved. The controller 343 acts as a selector to control the trial encoding process, selecting various operating modes such as block size and shape, and whether the video data should be losslessly encoded. The controller is considered as part of an image encoder or image decoder (depending on the situation). A continuous image of the input video signal 300 is provided to the adder 310 and the image predictor 320. References will follow below. Figure 8 A more detailed description of the image predictor 320. Image encoder or decoder (depending on the situation) plus... Figure 8 Intra-image predictors can use data from... Figure 7 The characteristics of the device. However, this does not mean that an image encoder or decoder must require... Figure 7 Each feature.
[0037] Adder 310 actually performs a subtraction (negative addition) operation because it receives the input video signal 300 at the "+" input and the output of image predictor 320 at the "-" input, so that the predicted image is subtracted from the input image. The result is the generation of a so-called residual image signal 330, representing the difference between the actual image and the predicted image.
[0038] One reason for generating residual image signals is as follows. The data encoding technique to be described, that is, the technique to be applied to residual image signals, tends to work more efficiently when there is less "energy" in the image to be encoded. Here, the term "efficiently" refers to generating a small amount of encoded data; for a given image quality level, it is desirable (and considered "efficient") to generate as little data as possible. The "energy" mentioned in residual images relates to the amount of information contained in the residual image. If the predicted image is the same as the true image, the difference between the two (that is, the residual image) will contain zero information (zero energy) and will be very easy to encode into a small amount of encoded data. Generally, if the prediction process can be made to work fairly well such that the content of the predicted image is similar to the content of the image to be encoded, it is expected that the residual image data will contain less information (less energy) than the input image, and will therefore be easier to encode into a small amount of encoded data.
[0039] Therefore, encoding (using adder 310) involves predicting image regions of the image to be encoded; and generating residual image regions based on the difference between the predicted image regions and the corresponding regions of the image to be encoded. In conjunction with the techniques discussed below, an ordered array of data values includes data values represented by the residual image regions. Decoding involves predicting image regions of the image to be decoded; generating residual image regions that represent the difference between the predicted image regions and the corresponding regions of the image to be decoded; wherein an ordered array of data values includes data values represented by the residual image regions; and combining the predicted image regions and the residual image regions.
[0040] The remainder of the apparatus will now be described as an encoder (to encode residual or poor images).
[0041] Residual image data 330 is provided to a transform unit or circuit 340, which generates a Discrete Cosine Transform (DCT) representation of the blocks or regions of the residual image data. DCT techniques are well-known and will not be described in detail here. It should also be noted that the use of DCT is merely an illustration of an example arrangement. Other transforms that can be used include, for example, the Discrete Sine Transform (DST). Transforms can also include sequences or cascades of individual transforms, such as an arrangement where one transform follows (whether directly or not) another transform. The selection of transforms can be explicitly determined and / or depends on lateral information used to configure the encoder and decoder. In other examples, a so-called "transform skip" mode can be selectively used, where no transform is applied.
[0042] Therefore, in the example, an encoding and / or decoding method includes: predicting an image region of the image to be encoded; and generating a residual image region based on the difference between the predicted image region and the corresponding region of the image to be encoded; wherein an ordered array of data values (discussed below) includes data values representing the residual image region.
[0043] The output of transform unit 340, that is (in this example) the set of DCT coefficients for each transform block of the image data, is provided to quantizer 350. Various quantization techniques are known in the field of video data compression, ranging from simply multiplying by a quantization scaling factor to applying complex lookup tables under the control of quantization parameters. The overall goal is twofold. First, the quantization process reduces the number of possible values for the transform data. Second, quantization can increase the probability that the transform data value is zero. Both enable the entropy coding process, which will be described below, to work more efficiently when generating small amounts of compressed video data.
[0044] The scanning unit 360 applies data scanning processing. The purpose of the scanning processing is to reorder the quantized transform data so as to gather as many non-zero quantized transform coefficients as possible together, and of course, as many zero-value coefficients as possible together. These characteristics allow for the efficient application of so-called run-length encoding or similar techniques. Therefore, the scanning processing involves selecting coefficients from the quantized transform data, and in particular from coefficient blocks corresponding to the transformed and quantized image data blocks, according to the “scan order” such that: (a) all coefficients are selected at once as part of the scan, and (b) the scan tends to provide the desired reordering. An example scan order that may tend to give useful results is the so-called top-right diagonal scan order.
[0045] The scanning order can be different, such as between transform skip blocks and transform blocks (blocks that have undergone at least one spatial frequency transformation).
[0046] The scanned coefficients are then passed to the entropy encoder (EE) 370. Similarly, various types of entropy coding can be used. Two examples are variants of the so-called CABAC (Context Adaptive Binary Arithmetic Coding) system and the so-called CAVLC (Context Adaptive Variable Length Coding) system. Generally, CABAC is considered to offer better efficiency, and some studies have shown that for comparable image quality compared to CAVLC, CABAC provides a 10-20% reduction in the amount of encoded output data. However, CAVLC is considered to represent a much lower level of complexity than CABAC (in terms of its implementation). Note that the scan processing and entropy coding processes are shown as separate processes, but in practice they can be combined or processed together. That is, data can be read into the entropy encoder in the order of the scan. The corresponding considerations apply to the various inverse processes described below.
[0047] The output of the entropy encoder 370, together with additional data (mentioned above and / or discussed below), provides a compressed output video signal 380. The additional data includes, for example, defining how the predictor 320 generates the predicted image, whether the compressed data is transformed or skipped, etc.
[0048] However, a return path 390 is also provided because the operation of the predictor 320 itself depends on the decompressed version of the compressed output data.
[0049] The reason for this function is as follows. At an appropriate stage in the decompression process (described below), a decompressed version of the residual data is generated. This decompressed residual data must be added to the prediction image to generate the output image (because the original residual data is the difference between the input image and the prediction image). To make this process comparable between the compression and decompression sides, the prediction image generated by predictor 320 should be the same during the compression and decompression processes. Of course, during decompression, the device cannot access the original input image, but only the decompressed image. Therefore, during compression, predictor 320 makes its prediction (at least for inter-image coding) based on the decompressed version of the compressed image.
[0050] The entropy encoding process performed by entropy encoder 370 is considered (in at least some examples) to be "lossless," meaning it can be reversed to obtain data exactly the same as the data previously provided to entropy encoder 370. Therefore, in such examples, the return path can be implemented before the entropy encoding stage. In fact, the scan processing performed by scan unit 360 is also considered lossless; therefore, in this embodiment, the return path 390 is from the output of quantizer 350 to the input of complementary inverse quantizer 420. If loss or potential loss is introduced at some stage, that stage (and its inverse stage) can be included in the feedback loop formed by the return path. For example, the entropy encoding stage can be at least in principle lossy, for example, by techniques such as bit-by-bit encoding in parity information. In this case, entropy encoding and decoding should form part of the feedback loop.
[0051] In summary, the entropy decoder 410, inverse scan unit 400, inverse quantizer 420, and inverse transform unit or circuit 430 provide the corresponding inverse functions of the entropy encoder 370, scan unit 360, quantizer 350, and transform unit 340. The discussion will now continue with the compression process; the decompression process of the input compressed video signal will be discussed separately below.
[0052] During compression, the scanned coefficients are passed from the quantizer 350 to the inverse quantizer 420, which performs the inverse operation of the scan unit 360, via return path 390. The inverse quantization and inverse transform processes are performed by units 420 and 430 to generate the compressed-decompressed residual image signal 440.
[0053] Image signal 440 is added to the output of predictor 320 at adder 450 to generate reconstructed output image 460 (although this may be subjected to so-called loop filtering and / or other filtering before the output—see below). This forms an input to image predictor 320, which will be described below.
[0054] Now we turn to the decoding process applied to decompress the received compressed video signal 470, which is provided to the entropy decoder 410 and, before being added to the output of the image predictor 320 by the adder 450, is provided from the entropy decoder 410 to a chain of inverse scan unit 400, inverse quantizer 420, and inverse transform unit 430. Thus, on the decoder side, the decoder reconstructs a version of the residual image and then (via adder 450) applies it to the predicted version of the image (on a block-by-block basis) to decode each block. Simply put, the output 460 of adder 450 forms the output decompressed video signal 480 (subject to the filtering process discussed below). In practice, further filtering (e.g., by...) can be applied before the output signal. Figure 8 The loop filter 565 shown is, but in order to Figure 7 For clarity, from the high-level diagram Figure 7 The loop filter is omitted in the text.
[0055] Figure 7 and Figure 8 The device can be used as a compression (encoding) device or a decompression (decoding) device. The functions of these two types of devices are largely overlapping. In decompression mode, the scanning unit 360 and the entropy encoder 370 are not used, and the operation of the predictor 320 (described in detail below) and other units follows the mode and parameter information contained in the received compressed bitstream, rather than generating that information themselves.
[0056] Figure 8 The generation of the predicted image is illustrated schematically, and the operation of the image predictor 320 is shown in detail.
[0057] There are two basic prediction modes performed by the image predictor 320: so-called intra-image prediction and so-called inter-image or motion-compensated (MC) prediction. On the encoder side, each involves detecting the prediction direction with respect to the current block to be predicted and generating a predicted block of samples based on other samples (in the same (intra-frame) or another (inter-frame) image). The difference between the predicted block and the actual block is encoded or applied by means of unit 310 or unit 450 to encode or decode the block respectively.
[0058] (At the decoder or on the inverse decoding side of the encoder, the detection of the predicted direction may be in response to data associated with the data encoded by the encoder, which indicates which direction to use at the encoder. Alternatively, the detection may be in response to the same factors that make the decision at the encoder.)
[0059] Intra-image prediction predicts the content of a block or region of an image based on data within the same image. This corresponds to the so-called I-frame coding in other video compression techniques. However, compared to I-frame coding, which involves encoding the entire image through intra-frame coding, this embodiment allows for a choice between intra-frame and inter-frame coding on a block-by-block basis, although in other embodiments, the choice is still made on a per-image basis.
[0060] Motion-compensated prediction is an example of inter-image prediction, and it utilizes motion information that attempts to define a source of image details to be encoded in the current image from another neighboring or nearby image. Thus, in an ideal example, the contents of image data blocks in the predicted image can be encoded very simply as references (motion vectors) pointing to corresponding blocks at the same or slightly different locations in adjacent images.
[0061] The technique known as “block copying” prediction is in some respects a hybrid of the two, as it uses vectors to indicate the location of a sample block within the same image that is shifted from the current predicted block and should be copied to form the current predicted block.
[0062] Return to Figure 8 The diagram illustrates two image prediction arrangements (corresponding to intra-image and inter-image predictions), the results of which are selected by multiplexer 500 under the control of mode signal 510 (e.g., from controller 343) to provide blocks of predicted images for adder 310 and adder 450. The selection is based on which choice yields the lowest "energy" (which, as described above, can be considered the information content to be encoded), and this selection is signaled to the decoder within the encoded output data stream. In this paper, image energy can be detected, for example, by performing trial subtraction on regions of the two versions of the predicted images of the input image, squaring each pixel value of the difference image, summing the squared values, and identifying which of the two versions generates the lower mean square value of the difference image associated with that image region. In other examples, trial encoding can be performed on each selection or potential selection, and then the selection can be made based on the cost of the number of bits required to encode each potential selection and one or both of image distortion.
[0063] In an intra-frame coding system, the actual prediction is based on the image blocks received as part of the signal 460 (e.g., filtered by a loop filter; see below), that is, the prediction is based on the encoded-decoded image blocks so that the exact same prediction can be performed at the decompression device. However, data can be derived from the input video signal 300 by the intra-frame mode selector 520 to control the operation of the intra-frame predictor 530.
[0064] For inter-image prediction, the motion compensation (MC) predictor 540 uses motion information, such as motion vectors derived from the input video signal 300 by the motion estimator 550. The motion compensation predictor 540 applies these motion vectors to the processed version of the reconstructed image 460 to generate inter-image predicted blocks.
[0065] Therefore, units 530 and 540 (operating together with estimator 550) each act as detectors to detect the prediction direction of the current block to be predicted, and act as generators to generate a prediction block of samples (forming part of the prediction passed to units 310 and 450) based on other samples defined by the prediction direction.
[0066] The processing applied to signal 460 will now be described.
[0067] First, the signal can be filtered by a so-called loop filter 565. Various types of loop filters can be used. One technique involves applying a "deblocking" filter to remove or at least tend to reduce the effects of block-based processing and subsequent operations performed by the transform unit 340. Another technique involves applying a so-called sample adaptive offset (SAO) filter. Generally, in a sample adaptive offset filter, filter parameter data (derived at the encoder and transmitted to the decoder) defines one or more offsets that will be selectively combined by the sample adaptive offset filter with a given intermediate video sample (a sample of signal 460) according to: (i) the given intermediate video sample; or (ii) one or more intermediate video samples that have a predetermined spatial relationship with the given intermediate video sample.
[0068] Alternatively, an adaptive loop filter is applied using coefficients derived from processing the reconstructed signal 460 and the input video signal 300. An adaptive loop filter is a filter that uses known techniques to apply adaptive filter coefficients to the data to be filtered. That is, the filter coefficients can vary depending on various factors. The data defining which filter coefficients to use is included as part of the encoded output data stream.
[0069] The techniques discussed below involve the processing of parameter data related to filter operation. Actual filtering operations (such as SAO filtering) can utilize other known techniques.
[0070] When the device operates as a decompression unit, the filtered output from the loop filter unit 565 effectively forms the output video signal 480. It is also buffered in one or more image or frame memories 570; the storage of consecutive images is required for motion compensation prediction processing, particularly for the generation of motion vectors. To save storage space, the images stored in the image memory 570 can be saved in compressed form and then decompressed for use in generating motion vectors. For this specific purpose, any known compression / decompression system can be used. The stored images can be passed to an interpolation filter 580, which generates a higher-resolution version of the stored images; in this example, for the 4:2:0 luma channel, intermediate samples (subsamples) are generated such that the resolution of the interpolated image output by the interpolation filter 580 is four times the resolution of the image stored in the image memory 570 (in each dimension), and for the 4:2:0 chroma channel, intermediate samples (subsamples) are generated such that the resolution of the interpolated image output by the interpolation filter 580 is eight times the resolution of the image stored in the image memory 570 (in each dimension). The interpolated image is passed as input to the motion estimator 550 and also to the motion compensation predictor 540.
[0071] The method of segmenting an image for compression will now be described. At a basic level, the image to be compressed is considered as an array of blocks or regions of samples. Segmenting an image into such blocks or regions can be achieved using decision trees, as described, for example, in SERIES H: AUDIOVISUAL AND MULTIMEDIA SYSTEMS Infrastructure of audiovisual services - Coding of moving video High efficiency video coding Recommendation ITU-T H.26512 / 2016 and: High Efficiency Video Coding (HEVC) Algorithms and Architectures, chapter 3, editors: Madhukar Budagavi, Gary J. Sullivan, Vivienne Sze; ISBN 978-3-319-06894-7; 2014, the entire contents of which are incorporated herein by reference. Further background information can be found in [1] “Versatile Video Coding (Draft 8)”, JVET-G2001-vE, B. Brass, J. Chen, S. Liu and YK. Wang, the full text of which is incorporated herein by reference.
[0072] In some examples, the resulting blocks or regions have a size, and in some cases, a shape, which, by means of a decision tree, can generally follow the arrangement of image features within the image. This in itself can allow for improved encoding efficiency, as samples representing or following similar image features will tend to be grouped together by such an arrangement. In some examples, square blocks or regions of different sizes (e.g., 4x4 samples up to (say) 64x64 or larger blocks) are available for selection. In other exemplary arrangements, blocks or regions of different shapes can be used, such as rectangular blocks (e.g., vertically or horizontally oriented). Other non-square and non-rectangular blocks are conceivable. The result of dividing the image into such blocks or regions is (at least in this example) that each sample of the image is assigned to one and only one such block or region.
[0073] Transformation matrix
[0074] The following discussion pertains to the transform unit 340 and the inverse transform unit 430. Note that, as mentioned above, the transform unit exists within the encoder. The inverse transform unit exists in the encoder's return decoding path and the decoder's decoding path.
[0075] The transform unit and inverse transform unit are designed to provide complementary transforms to and from the spatial frequency domain. That is, transform unit 340 operates on the video data set (or data derived from the video data, such as the difference or residual data discussed above) and generates a corresponding set of spatial frequency coefficients. Inverse transform unit 430 operates on the set of spatial frequency coefficients and generates a corresponding set of video data.
[0076] In practice, the transformation is implemented as a matrix computation. The forward transformation is defined by the transformation matrix and is achieved by multiplying the transformation matrix by the array of sample values to generate the corresponding array of spatial frequency coefficients. In some examples, the sample array M is left-multiplied by the transformation matrix T, and the result is right-multiplied by the transpose of the transformation matrix T. T Therefore, the output coefficient array is defined as:
[0077] TMT T
[0078] The property of matrices representing this type of spatial frequency transformation is that the transpose of the matrix is the same as its inverse. Therefore, in principle, the forward and inverse transformation matrices are related by a simple transpose.
[0079] In existing VVC specifications as of the filing date of this application (see the references above), the transform is defined only as the inverse transform matrix used in the decoder function. The encoder transform matrix is not defined in this way. The decoder (inverse) transform is defined with 6-bit precision.
[0080] As mentioned above, in principle, a suitable set of forward (encoder) matrices can be obtained as the transpose of the inverse matrix. However, the relationship between the forward and inverse matrices only applies when the values are represented with infinite precision. Representing these values with finite precision (e.g., 6 bits) means that the transpose is not necessarily an appropriate relationship between the forward and inverse matrices.
[0081] It should be recognized that manufacturing or implementing integer arithmetic transformation and inverse transformation units is easier, cheaper, faster, and / or less processor-intensive than implementing similar units using floating-point calculations. Therefore, the matrix coefficients discussed below are represented as integers, proportionally increased by a sufficient number of bits (powers of 2) from the actual values required to implement the transformation to allow the desired coefficient precision. If necessary, this amplification can be removed at another stage of the process by shifting (dividing by powers of 2). In other words, the actual transformation coefficients are proportional to the values discussed below.
[0082] Provide transformation matrix
[0083] Figure 9 An example of a (forward) transform unit 340 or an inverse transform unit 430 is illustrated, and includes a matrix data memory 900 storing appropriate data (see below), wherein these transform matrices (or, in other embodiments, inverse transform matrices) are derived by a matrix generator 910, for example in response to the current video data bit depth, the current block size, and the transform type in use.
[0084] The generated matrix is provided to matrix processor 920, which performs matrix multiplication associated with the desired transformation (or inverse transformation) on block 930 to generate processed block 940. In the case of a forward transformation, block 930 to be processed will be, for example, from subtraction stage 310, and processed block 940 is provided to quantization stage 350. In the case of an inverse transformation, for example, block 930 to be processed will be an inverse-quantized block output by inverse quantizer 420, and processed block 940 will be provided to adder 450.
[0085] Transformation tools
[0086] In some exemplary video data processing systems, various transform tools are available for selection, either individually or as a series of consecutive transforms. Examples of these transform tools include Discrete Cosine Transform (DCT) Type II (DCT2); DCT Type VIII (DCT8); and Discrete Sine Transform (DST) Type VII (DST7). Other example tools are also available, but will not be described in detail here.
[0087] The choice between tools can be made through various techniques, such as any one or more of the following: (i) direct mapping between any one of a set of parameters, including, for example, block size, prediction direction, prediction type (inside, between), etc.; (ii) the result of one or more trial coding processes; (iii) configuration data provided by parameter sets, etc.; (iv) continuation of configuration from adjacent or temporarily preceding blocks, segments or other regions; (v) predetermined configuration settings, etc.
[0088] Regardless of the technique used to select the transformation tool, it is not important to the teachings of this disclosure concerning the generation of matrix coefficients to achieve a given transformation.
[0089] Improved accuracy of the forward transformation matrix
[0090] Based on the above discussion, this section discusses a video data encoding method and apparatus for encoding an array of video data values using a forward transform matrix at, for example, a 14-bit precision level.
[0091] The basis of this discussion is that improved results can be obtained by using a forward matrix that matches the standard inverse matrix, with coefficients represented at a higher resolution than 6 bits. This is especially true when the system is encoding high bit-depth video data (i.e., video data with a large “bit depth” or data precision, such as 16-bit video data, expressed in bits).
[0092] The matrix values discussed below can be applied, for example, to frequency transform video data values according to a frequency transform, to generate an array of frequency transform values by using matrix multiplication of a transform matrix with a data precision greater than 6 bits, the transform matrix having coefficients defined by at least a subset of the values discussed below, and frequency transform units (such as unit 340).
[0093] DCT2 matrix
[0094] Figures 10a to 10e together define the 64×64 basis matrix M of DCT2. 64 Specifically, Figure 10a shows the relative configuration of the four matrix portions shown in Figures 10b through 10e, respectively. Figure 11 illustrates the values of the coefficients and the alternative naming protocol; here, the name column on the left (a...E) is used for matrices of a maximum size of 32x32, and the second column name is used if a 64x64 matrix is required. The "14-bit" values in the third column are those provided as part of this disclosure. The corresponding "6-bit" values are those previously proposed for comparison purposes only.
[0095] Typically, for an N×N transformation, where N is 2, 4, 8, or 16, the transformation matrix M is 64×64. 64Subsampling is performed to select a subset of N×N values, the subset of values M N [x][y] are defined by the following formula:
[0096] M N [x][y]=M 64 [x][(2 (6-log2(N)) [y] where x, y = 0..(N-1).
[0097] Therefore, operation Figure 7 The arrangement, according to Figure 9 Using the data from Figures 10a to 10e and Figure 11, an example of a video data encoding method for encoding an array of video data values is provided, the method comprising the steps of:
[0098] The video data values are frequency transformed according to the frequency transform, which generates an array of frequency transform values by matrix multiplication of the transform matrix with 14-bit data precision. The frequency transform is the discrete cosine transform.
[0099] Define the 64×64 transformation matrix M for the 64×64 DCT transformation. 64 The matrix M 64 Defined by Figures 10a to 10e and Figure 11;
[0100] For an N×N transformation, where N is 2, 4, 8, or 16, the 64×64 transformation matrix M... 64 Subsampling is performed to select a subset M of N×N values. N [x][y] are defined by the following formula:
[0101] M N [x][y]=M 64 [x][(2 (6-log2(N)) [y] where x, y = 0..(N-1).
[0102] Similarly, as described in the operation Figure 7 and Figure 9 The arrangement provides an example of a data encoding apparatus for encoding an array of video data values, the apparatus comprising:
[0103] A frequency transformation circuit is configured to perform a frequency transformation on video data values according to a frequency transformation, generating an array of frequency transformation values by matrix multiplication of a transformation matrix with 14-bit data precision. The frequency transformation is a discrete cosine transform. The frequency transformation circuit defines a 64×64 transformation matrix M for 64×64 DCT transformation. 64 Matrix M 64Defined by Figures 10a to 10e and Figure 11, where for an N×N transformation, N is 2, 4, 8, or 16, the N×N transformation matrix includes a 64×64 transformation matrix M. 64 A subset of, a subset of values M N [x][y] are defined by the following formula:
[0104] M N [x][y]=M 64 [x][(2 (6-log2(N)) [y] where x, y = 0..(N-1).
[0105] The following example demonstrates how the relationship defined by the formulas given above can be applied.
[0106] 2×2DCT2
[0107] Matrix M2 is defined as the combination matrix M 64 The first two coefficients in each of the 32nd rows.
[0108]
[0109] 4×4DCT2
[0110] Matrix M4 is defined as a combination matrix M 64 The first four coefficients in each of the 16th rows.
[0111]
[0112] 8×8DCT2
[0113] Matrix M8 is defined as a combination matrix M 64 The first 8 coefficients in each of the 8th rows.
[0114]
[0115] 16×16DCT2
[0116] Matrix M 16 Defined as a combination matrix M 64 The first 16 coefficients in each of the 4th rows.
[0117]
[0118] 32×32DCT2
[0119] Matrix M 32 Defined as a combination matrix M 64 The first 32 coefficients in each of the second rows.
[0120]
[0121]
[0122] DCT2 64×64
[0123] 64x64 DCT2 uses the full matrix M 64 .
[0124] DCT 8 matrix
[0125] Figure 12 schematically illustrates the set of values used by the DCT8 transform tool to generate the transform matrix. As mentioned earlier, the "14-bit" column relates to the newly proposed values, where the "6-bit" values are provided only for comparison with the previously proposed arrangement.
[0126] The value selection is performed as follows.
[0127] operate Figure 7 The arrangement, according to Figure 9 The manipulation and use of the data in Figure 12 provide an example of a video data encoding method for encoding an array of video data values, which includes the following steps:
[0128] The video data values are frequency transformed according to the frequency transform, which generates an array of frequency transform values by matrix multiplication of the transform matrix with 14-bit data precision. The frequency transform is the discrete cosine transform.
[0129] Define the set of values as shown in Figure 12;
[0130] For an N×N transformation, where N is 4, 8, 16, or 32, the N×N transformation matrix M is selected from the set of values defined by the following formula. N Value:
[0131] (i) For 4×4DCT8 transform:
[0132]
[0133] (ii) For 8×8 DCT8 transform:
[0134]
[0135] (iii) For 16×16 DCT8 transform:
[0136]
[0137] (iv) For 32×32 DCT8 transform:
[0138]
[0139]
[0140] Similarly, as described in the operation Figure 7 and 9 The arrangement provides an example of a data encoding apparatus for encoding an array of video data values, the apparatus comprising:
[0141] A frequency transformation circuit is configured to perform frequency transformation on video data values according to a frequency transformation to generate an array of frequency transformation values by using matrix multiplication of a transformation matrix with 14-bit data precision. The frequency transformation is a discrete cosine transformation. The frequency transformation circuit defines a set of values as shown in Figure 12, where for an N×N transformation, N is 4, 8, 16, or 32. The N×N transformation matrix includes selected values from the set of values defined by a subset of the above values.
[0142] DST7 transformation
[0143] Figure 13 provides sample data for generating the DST7 transformation matrix. As mentioned earlier, the "6-bit" column is provided only for comparison.
[0144] operate Figure 7 The arrangement, according to Figure 9 The manipulation and use of the data in Figure 13 provides an example of a video data encoding method for encoding an array of video data values, which includes the following steps:
[0145] The video data values are frequency transformed according to the frequency transformation, in order to generate an array of frequency transformation values by using matrix multiplication of the transformation matrix with 14-bit data precision. The frequency transformation is a discrete sine transformation.
[0146] Define the set of values as shown in Figure 13;
[0147] For an N×N transformation, where N is 4, 8, 16, or 32, the N×N transformation matrix M is selected from the set of values defined by the following formula. N Value:
[0148] (i) For 4×4 DST7 transformation
[0149]
[0150] (ii) For 8×8 DST7 transformation
[0151]
[0152] (iii) For 16×16 DST7 transformation
[0153]
[0154] (iv) For 32×32 DST7 transformation
[0155]
[0156] Similarly, as described in the operation Figure 7 and 9 The arrangement provides an example of a data encoding apparatus for encoding an array of video data values, the apparatus comprising:
[0157] A frequency conversion circuit is configured to perform frequency conversion on video data values according to a frequency conversion to generate an array of frequency conversion values by using matrix multiplication of a conversion matrix with 14-bit data precision. The frequency conversion is a discrete sine conversion. The frequency conversion circuit defines a set of values as shown in Figure 13, where for an N×N conversion, N is 4, 8, 16, or 32, and the N×N conversion matrix includes selected values from the set of values mentioned above.
[0158] Identification code
[0159] For previous example arrangements operating at high bit depth, a so-called "extended precision flag" has been proposed as a parameter that can be set in conjunction with the video data stream. In some previous examples, setting this flag allows for a potential increase in the dynamic range used during the computation of forward and reverse spatial frequency transforms, thereby increasing the precision of the transform values passed from the transform unit to the entropy encoder (in the forward direction). This flag can be encoded as part of a so-called sequence parameter set (SPS).
[0160] In an exemplary embodiment of the invention, an extended precision flag may be provided among a set of one or more other flags (e.g., in SPS). Figure 14 provides an example of a hierarchy of such flags, including a high bit depth control flag. If this flag is not set, the extended precision flag is unavailable; however, if it is set, the extended precision flag may or may not be set. The overall high bit depth control flag allows other functions related to high bit depth (e.g., for video data greater than 10 bits) operations (e.g., at least in principle, higher precision quantizers or alternative coefficient encoders for transform or non-transform blocks), and thus provides a mechanism for encoding and enabling future modifications to the encoding tool.
[0161] Figure 14b This is a schematic flowchart illustrating a method for encoding video data values, the method comprising:
[0162] Selectively encode (in step 1400) the high-depth control flag, and when the high-depth control flag is set to indicate high-depth operation, selectively encode the extended precision flag to indicate at least extended precision operation in the spatial frequency conversion phase; and
[0163] The video data value is encoded (in step 1410) according to the operating mode defined by the high-bit depth control flag and when the extended precision flag is encoded.
[0164] Similarly, on the decoding side, Figure 14c This is a schematic flowchart illustrating a method for decoding video data values, the method including:
[0165] The high-depth control flag is selectively decoded (in step 1420), and when the high-depth control flag is set to indicate high-depth operation, the extended precision flag is selectively decoded to indicate at least extended precision operation in the spatial frequency conversion phase.
[0166] The video data value is decoded (in step 1430) according to the operating mode defined by the high-bit depth control flag and when the extended precision flag is decoded.
[0167] High-bit depth control flags and (optionally) extended precision flags can be encoded into or decoded from the sequence parameter set of the video data stream.
[0168] Therefore, operations based on these methods Figure 7 The apparatus provides an example of a data encoding apparatus for encoding video data values, the apparatus comprising:
[0169] The parameter encoder (343) is configured to selectively encode a high-depth control flag, and when the high-depth control flag is set to indicate high-depth operation, selectively encode an extended precision flag to indicate at least extended precision operation during the spatial frequency conversion phase; and
[0170] encoder ( Figure 7 It is configured to encode video data values according to the operating mode defined by the encoding high bit depth control flag and when the extended precision flag is encoded.
[0171] Similarly, operations based on these methods Figure 7 The device provides an example of a data decoding apparatus for decoding video data values, the apparatus comprising:
[0172] The parameter decoder (343) is configured to selectively decode the high-bit depth control flag, and when the high-bit depth control flag is set to indicate high-bit depth operation, selectively decode the extended precision flag to indicate at least extended precision operation in the spatial frequency conversion phase; and
[0173] decoder ( Figure 7 The device is configured to decode the video data values according to an operating mode defined by the high-bit depth control flag and when the extended precision flag is decoded.
[0174] Summary Methods - Matrix Generation
[0175] Figure 15 A schematic flowchart of a method is shown, which includes: performing a frequency transformation on video data values according to a frequency transformation (in step 1500) to generate an array of frequency transformation values by using matrix multiplication of one or more transformations defined above;
[0176] Define (in step 1510) the set of values associated with this transformation as shown in the figure;
[0177] For an N×N transformation, using the above technique, the N×N transformation matrix M is selected from the set of provided values (in step 1520). N The value of .
[0178] Image data
[0179] Image or video data encoded or decoded using these techniques, as well as data carriers carrying such image data, are considered to represent embodiments of this disclosure.
[0180] Therefore, since the embodiments of this disclosure have been described as being implemented at least in part by a software-controlled data processing apparatus, it should be understood that non-transitory machine-readable media carrying such software, such as optical discs, magnetic disks, and semiconductor memories, are also considered to represent embodiments of this disclosure. Similarly, data signals including encoded data generated according to the methods described above (whether or not they are contained on a non-transitory machine-readable medium) are also considered to represent embodiments of this disclosure.
[0181] Clearly, many modifications and variations of this disclosure are possible based on the foregoing teachings. Therefore, it should be understood that this technology may be practiced in ways other than those specifically described herein, within the scope of the appended provisions.
[0182] It should be understood that, for clarity, the above description has referenced various functional units, circuits, and / or processors in describing the implementation. However, it will be apparent that any suitable functional distribution among the various functional units, circuits, and / or processors can be used without departing from the implementation.
[0183] The described embodiments can be implemented in any suitable form, including hardware, software, firmware, or any combination thereof. The described embodiments can alternatively be implemented, at least in part, as computer software running on one or more data processors and / or digital signal processors. Elements and components of any embodiment can be implemented physically, functionally, and logically in any suitable manner. In practice, the function can be implemented in a single unit, in multiple units, or as part of other functional units. Therefore, the disclosed embodiments can be implemented in a single unit or can be physically and functionally distributed among different units, circuits, and / or processors.
[0184] Although this disclosure has been described in conjunction with some embodiments, it is not intended to be limited to the specific forms set forth herein. Furthermore, while features may appear to be described in conjunction with specific embodiments, those skilled in the art will recognize that the various features of the described embodiments can be combined in any manner suitable for implementing the technology.
[0185] The various aspects and characteristics are defined by the following numbered clauses:
[0186] 1. A method for encoding video data into an array of video data values, the method comprising the following steps:
[0187] The video data values are frequency transformed according to a frequency transform to generate an array of frequency transform values by matrix multiplication of a transform matrix with 14-bit data precision, wherein the frequency transform is a discrete cosine transform.
[0188] Define a 64×64 transformation matrix M64 for the 64×64 DCT transform, wherein the matrix M64 is defined by Figures 10a to 10e and Figure 11;
[0189] For an N×N transformation, where N is 2, 4, 8, or 16, the 64×64 transformation matrix M64 is subsampled to select a subset of N×N values, the subset MN[x][y] of which is defined by the following formula:
[0190] M N [x][y]=M 64 [x][(2 (6-log2(N)) [y] where x, y = 0..(N-1).
[0191] 2. Image data, encoded by means of the method described in Clause 1.
[0192] 3. Computer software, when executed by a computer, causes the computer to perform the method described in accordance with Clause 1.
[0193] 4. A non-transitory machine-readable storage medium storing computer software as described in Clause 3.
[0194] 5. An apparatus for encoding an array of video data values, the apparatus comprising:
[0195] A frequency transformation circuit is configured to perform a frequency transformation on the video data values according to a frequency transformation, thereby generating an array of frequency transformation values by using matrix multiplication of a transformation matrix with 14-bit data precision. The frequency transformation is a discrete cosine transform. The frequency transformation circuit defines a 64×64 transformation matrix M64 for a 64×64 DCT transform, defined by Figures 10a to 10e and Figure 11. For an N×N transform, where N is 2, 4, 8, or 16, the N×N transform matrix includes a subset of the 64×64 transform matrix M64, and the subset of values MN[x][y] is defined by the following formula:
[0196] M N [x][y]=M 64 [x][(2 (6-log2(N)) [y] where x, y = 0..(N-1).
[0197] 6. A device for capturing, transmitting, displaying and / or storing video data, including the device described in Clause 5.
[0198] 7. A method for encoding video data into an array of video data values, the method comprising the following steps:
[0199] The video data values are frequency transformed according to a frequency transform to generate an array of frequency transform values by matrix multiplication of a transform matrix with 14-bit data precision, wherein the frequency transform is a discrete cosine transform.
[0200] Define the set of values as shown in Figure 12;
[0201] For an N×N transformation, where N is 4, 8, 16, or 32, the value of the N×N transformation matrix MN is selected from the set of said values, defined by the following formula:
[0202] (i) For 4×4DCT8 transform:
[0203]
[0204] (ii) For 8×8 DCT8 transform:
[0205]
[0206] (iii) For 16×16 DCT8 transform:
[0207]
[0208] (iv) For 32×32 DCT8 transform:
[0209]
[0210]
[0211] 8. Image data, encoded by means of the method described in accordance with Clause 7.
[0212] 9. Computer software, when executed by a computer, causes the computer to perform the method described in accordance with Clause 7.
[0213] 10. A non-transitory machine-readable storage medium storing computer software as described in Clause 9.
[0214] 11. An apparatus for encoding an array of video data values, the apparatus comprising:
[0215] A frequency transformation circuit is configured to perform a frequency transformation on the video data values according to a frequency transformation to generate an array of frequency transformation values by using matrix multiplication of a transformation matrix with 14-bit data precision. The frequency transformation is a discrete cosine transform. The frequency transformation circuit defines a set of values as shown in Figure 12, where, for an N×N transformation, N is 4, 8, 16, or 32, the N×N transformation matrix includes selected values from the set of values, defined by the following formula:
[0216] (i) For 4×4DCT8 transform:
[0217]
[0218] (ii) For 8×8 DCT8 transform:
[0219]
[0220] (iii) For 16×16 DCT8 transform:
[0221]
[0222] (iv) For 32×32 DCT8 transform:
[0223]
[0224]
[0225] 12. A device for capturing, transmitting, displaying and / or storing video data, including the device described in Clause 11.
[0226] 13. A method for encoding video data into an array of video data values, the method comprising the following steps:
[0227] The video data values are frequency transformed according to a frequency transform to generate an array of frequency transform values by matrix multiplication of a transform matrix with 14-bit data precision, wherein the frequency transform is a discrete sine transform.
[0228] Define the set of values as shown in Figure 13;
[0229] For an N×N transformation, where N is 4, 8, 16, or 32, the value of the N×N transformation matrix MN is selected from the set of said values, defined by the following formula:
[0230] (i) For 4×4 DST7 transformation
[0231]
[0232] (ii) For 8×8 DST7 transformation
[0233]
[0234] (iii) For 16×16 DST7 transformation
[0235]
[0236] (iv) For 32×32 DST7 transformation
[0237]
[0238]
[0239] 14. Image data, encoded by means of the method described in accordance with Clause 13.
[0240] 15. Computer software, when executed by a computer, causes the computer to perform the method described in accordance with clause 13.
[0241] 16. A non-transitory machine-readable storage medium storing computer software as described in Clause 15.
[0242] 17. An apparatus for encoding an array of video data values, the apparatus comprising:
[0243] A frequency transformation circuit is configured to perform a frequency transformation on the video data values according to a frequency transformation to generate an array of frequency transformation values by using matrix multiplication of a transformation matrix with 14-bit data precision. The frequency transformation is a discrete sine transformation. The frequency transformation circuit defines a set of values as shown in Figure 13, where, for an N×N transformation, N is 4, 8, 16, or 32, the N×N transformation matrix includes selected values from the set of values, defined by the following formula:
[0244] (i) For 4×4 DST7 transformation
[0245]
[0246] (ii) For 8×8 DST7 transformation
[0247]
[0248] (iii) For 16×16 DST7 transformation
[0249]
[0250]
[0251] (iv) For 32×32 DST7 transformation
[0252]
[0253]
[0254] 18. A means for capturing, transmitting, displaying and / or storing video data, including the means described in Clause 17.
[0255] 19. A method for encoding video data values, the method comprising:
[0256] Selectively encode the high-depth control flag, and when the high-depth control flag is set to indicate high-depth operation, selectively encode the extended precision flag to indicate at least extended precision operation during the spatial frequency conversion phase; and
[0257] The video data values are encoded according to the high-bit depth control flag and the operating mode defined by the extended precision flag when encoding.
[0258] 20. The method according to Clause 19, comprising selectively encoding the high-bit depth control flag and selectively encoding the extended precision flag into a sequence parameter set of the video data stream.
[0259] 21. Computer software, when executed by a computer, causes the computer to perform the method described in accordance with clause 19.
[0260] 22. A non-transitory machine-readable storage medium storing computer software as described in Clause 21.
[0261] 23. A method for decoding video data values, the method comprising:
[0262] The high-depth control flag is selectively decoded, and when the high-depth control flag is set to indicate high-depth operation, the extended precision flag is selectively decoded to indicate at least extended precision operation in the spatial frequency conversion phase.
[0263] The video data values are decoded according to the encoded high-bit depth control flag and the operating mode defined by the extended precision flag when decoding.
[0264] 24. The method according to Clause 23 includes selectively decoding the high-bit depth control flag and selectively encoding the extended precision flag from a set of sequence parameters of a video data stream.
[0265] 25. Computer software, when executed by a computer, causes the computer to perform the method described in accordance with clause 23.
[0266] 26. A non-transitory machine-readable storage medium storing computer software as described in Clause 25.
[0267] 27. An apparatus for encoding video data values, the apparatus comprising:
[0268] A parameter encoder is configured to selectively encode a high-depth control flag, and, when the high-depth control flag is set to indicate high-depth operation, selectively encode an extended precision flag to indicate at least extended precision operation during the spatial frequency conversion phase; and
[0269] The encoder is configured to encode the video data values according to an operating mode defined by the high bit depth control flag and the extended precision flag when encoding.
[0270] 28. A means for capturing, transmitting, displaying and / or storing video data, including the means described in Clause 27.
[0271] 29. An apparatus for decoding video data values, the apparatus comprising:
[0272] The parameter decoder is configured to selectively decode high-bit depth control flags, and when the high-bit depth control flags are set to indicate high-bit depth operation, selectively decode extended precision flags to indicate at least extended precision operation during the spatial frequency conversion phase; and
[0273] The decoder is configured to decode the video data values according to an operating mode defined by the encoded high-bit depth control flag and the extended precision flag when decoding.
[0274] 30. A device for capturing, transmitting, displaying and / or storing video data, including the device described in Clause 29.
Claims
1. A method for encoding video data values, the method comprising: The high-depth control flag is selectively encoded. If the high-depth control flag is not set, the extended precision flag is unavailable. If the high-depth control flag is set to indicate high-depth operation, the extended precision flag is selectively encoded to indicate at least extended precision operation in the spatial frequency conversion phase. The video data values are encoded according to the high-bit depth control flag and the operating mode defined by the extended precision flag when encoding; as well as The high-depth control flag and the extended precision flag are provided in a flag hierarchy, the high-depth control flag allowing control of other functions for the high-depth operation according to the flag hierarchy.
2. The method of claim 1, further comprising selectively encoding the high-bit depth control flag and selectively encoding the extended precision flag into a sequence parameter set of the video data stream.
3. A computer program product comprising instructions that, when executed by a computer, cause the computer to perform the method according to claim 1.
4. A non-transitory machine-readable storage medium for storing a computer program product according to claim 3.
5. A method for decoding video data values, the method comprising: The high-depth control flag is selectively decoded. If the high-depth control flag is not set, the extended precision flag is unavailable. If the high-depth control flag is set to indicate high-depth operation, the extended precision flag is selectively decoded to indicate at least extended precision operation in the spatial frequency conversion phase. The video data values are decoded according to the encoded high-bit depth control flag and the operating mode defined by the extended precision flag when decoding; as well as The high-depth control flag and the extended precision flag are provided in a flag hierarchy, the high-depth control flag allowing control of other functions for the high-depth operation according to the flag hierarchy.
6. The method of claim 5, further comprising selectively decoding the high-bit depth control flag and selectively encoding the extended precision flag from a set of sequence parameters of the video data stream.
7. A computer program product comprising instructions that, when executed by a computer, cause the computer to perform the method according to claim 5.
8. A non-transitory machine-readable storage medium storing a computer program product according to claim 7.
9. An apparatus for encoding video data values, the apparatus comprising: The parameter encoder is configured to selectively encode the high-bit depth control flag. If the high-bit depth control flag is not set, the extended precision flag is unavailable. If the high-bit depth control flag is set to indicate high-bit depth operation, the extended precision flag is selectively encoded to indicate at least extended precision operation during the spatial frequency conversion phase. as well as The encoder is configured to encode the video data values according to an operating mode defined by a high-bit depth control flag and an extended precision flag during encoding. The high-depth control flag and the extended precision flag are provided in a flag hierarchy, the high-depth control flag allowing control of other functions for the high-depth operation according to the flag hierarchy.
10. A video data capture, transmission, display and / or storage device, including the device according to claim 9.
11. An apparatus for decoding video data values, the apparatus comprising: The parameter decoder is configured to selectively decode the high-bit depth control flag. If the high-bit depth control flag is not set, the extended precision flag is unavailable. If the high-bit depth control flag is set to indicate high-bit depth operation, the extended precision flag is selectively decoded to indicate at least extended precision operation in the spatial frequency conversion phase. as well as The decoder is configured to decode the video data values according to an operating mode defined by the encoded high-bit depth control flag and the extended precision flag when decoding. The high-depth control flag and the extended precision flag are provided in a flag hierarchy, the high-depth control flag allowing control of other functions for the high-depth operation according to the flag hierarchy.
12. A video data capture, transmission, display and / or storage device, including the device according to claim 11.
Citation Information
Patent Citations
Data encoding and decoding
CN105453566A