Image decoding device, image encoding device, image decoding method, and image encoding method
The image decoding device simplifies the decoding process by using transform skip residual coding flags to reduce complexity and costs, addressing the issue of complicated quantization combinations in existing devices.
Patent Information
- Application Number
- JP2024059268
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Filing Date
- 2024-04-02
- Publication Date
- 2025-05-21
- Estimated Expiration
- 2040-02-21
AI Technical Summary
Existing image decoding and encoding devices are complicated due to the presence of combinations of quantization and dependent quantization, leading to increased implementation costs without significant benefits.
An image decoding device that decodes transform skip residual coding flags and switches between different derivation methods for transform coefficients based on specific flags, eliminating the need for combinations of transform skip, sign hiding, and dependent quantization.
This configuration reduces the implementation costs of hardware and software by simplifying the decoding process, making it more efficient and cost-effective.
Smart Images

Figure 0007681150000001 
Figure 0007681150000002 
Figure 0007681150000003
Abstract
Description
[Technical field]
[0001] An embodiment of the present invention relates to an image decoding device, an image encoding device, an image decoding method, and an image encoding method. [Background technology]
[0002] In order to efficiently transmit or record images, an image encoding device is used that generates encoded data by encoding an image, and an image decoding device is used that generates a decoded image by decoding the encoded data.
[0003] Specific image coding methods include, for example, H.264 / AVC and HEVC (High-Efficiency Video Coding method, etc.
[0004] In such an image coding method, images (pictures) constituting a video are divided into slices obtained by dividing the image, coding tree units (CTUs) obtained by dividing the slices, coding tree units (CTUs) obtained by dividing the coding tree units, and so on. A coding unit (sometimes called a coding unit (CU)) to be encoded, and The coding unit is divided into transform units (TUs) and managed in a hierarchical structure, and encoding / decoding is performed for each CU.
[0005] In such an image coding method, a predicted image is usually generated based on a locally decoded image obtained by coding / decoding an input image, and a prediction error (sometimes called a "difference image" or "residual image") obtained by subtracting the predicted image from the input image (original image) is coded. Methods of generating a predicted image include inter-prediction and intra-prediction.
[0006] Non-Patent Document 1 is an example of a recent image coding and decoding technique. Non-Patent Document 1 discloses "sign data hiding (SDH)", a technique for estimating the positive or negative sign of some transform coefficients calculated by inverse quantizing quantized transform coefficients related to prediction errors without coding them in order to improve coding efficiency. The paper discloses a "dependent quantization (DQ)" technique that performs quantization and inverse quantization by switching between two quantizers having different levels.
[0007] On the other hand, in applications where video is used in medicine, art, etc., there are cases where a decoded image with no or almost no degradation due to encoding is required. To achieve such lossless or near-lossless, techniques have been disclosed that do not use transform (inverse transform) or quantization (inverse quantization) in encoding (decoding). One of these techniques is called "Transform Skip (TS)". Transform skip is also considered to be one type of transform (Identical Transform), and whether it is a transform skip or not is coded as the type of transform. Also, in the transform skip, "Transform Skip Residual Coding (TSRC)" is known which performs coding different from regular residual coding (RRC) of the prediction error. [Prior art documents] [Non-patent literature]
[0008] [Non-Patent Document 1] "Versatile Video Coding (Draft 8)", JVET-Q2001, Joint Video Exploration Team (JVET) of ITU-T SG 16 WP 3 and ISO / IEC JTC 1 / SC 29 / WG 11, 7-17 January 2020 Summary of the Invention [Problem to be solved by the invention]
[0009] In the technology of Non-Patent Document 1, transform skip, sign hiding, dependent quantization, transform skip prediction error quantization (TSRC), and normal prediction error coding (RRC) are combined to perform transform skipping. In skip prediction error quantization (TSRC), sign hiding and dependent quantization are not performed. This makes transform skip, sign hiding, and dependent quantization exclusive. However, in normal prediction error coding (RRC), transform skip and sign hiding are not performed. However, there is still a problem that image decoding devices and image encoding devices are complicated due to the presence of combinations of quantization and dependent quantization. In particular, there is a problem that the implementation of hardware and software that realizes combinations with small effects is not worth the increased cost. [Means for solving the problem]
[0010] In order to solve the above problem, an image decoding device according to one aspect of the present invention includes a header decoding unit that decodes a transform skip residual coding disable flag indicating that normal prediction error coding is used to decode prediction errors of a transform skip block of a current slice, or that transform skip prediction error coding is used to decode prediction errors of a transform skip block of a current slice, based on a value of a transform skip enable flag indicating whether a transform skip flag is present in a transform unit; the transform skip flag indicating whether a transform is applied to a corresponding block based on a block size and a corresponding block size; and an inverse quantization and inverse transform unit deriving a transform coefficient value by using a scale value, wherein if the value of the transform skip flag is 0 or the value of the transform skip residual coding disable flag is 1, the normal prediction error coding is used, and otherwise the transform skip prediction error coding is used, and the scale value is derived by switching between at least two derivation methods based on the value of the transform skip flag. In addition, the image decoding device according to an aspect of the present invention includes a flag ph_dep_quant_enabled_flag indicating whether or not dependent quantization is available from the encoded data, and a flag indicating whether or not transform skip prediction error quantization is disabled. and a TU decoding unit that decodes transform coefficients in an RRC mode in which a LAST position that is a decoding start position of transform coefficients in a TU block is coded or in a TSRC mode in which the LAST position is not coded, In the case where the transform coefficients of the transform skip are decoded in the TSRC mode (slice_ts_residual_coding_disabled_flag=0), the residual quantization is performed in the RRC mode, and the transform coefficients of the transform skip are decoded in the TSRC mode (slice_ts_residual_coding_disabled_flag=0). When the number is decoded in RRC mode (slice_ts_residual_coding_disabled_flag=1), dependent quantization is not performed in RRC mode.
[0011] The header decoding unit decodes a flag ph_dep_quant_enabled_flag indicating whether or not dependent quantization is available from the encoded data, and a flag slice_ts_residual_coding_disabled_flag indicating that transform skip prediction error quantization is disabled, and a TU decoding unit decodes transform coefficients; The TU decoding unit performs dependent quantization and decodes transform coefficients when ph_dep_quant_enabled_flag is 1 and slice_ts_residual_coding_disabled_flag is 0, Otherwise, if ph_dep_quant_enabled_flag is 0 or slice_ts_residual_coding_disabled_flag is 1, the transform coefficients are decoded without performing dependent quantization. do.
[0012] The header decoding unit decodes a flag ph_dep_quant_enabled_flag indicating whether or not dependent quantization is available from the encoded data, a flag pic_sign_data_hiding_enabled_flag indicating whether or not sign hiding is available, and a flag slice_ts_residual_coding_disabled_flag prohibiting transform skip prediction error quantization, and a TU decoding unit decodes transform coefficients; The TU decoding unit is configured to decode the ph_dep_quant_enabled_flag when ph_dep_quant_enabled_flag is set to 1 or when pic_sign_data_hiding_ena When bled_flag is 0 or slice_ts_residual_coding_disabled_flag is 1, It is characterized by not performing in-hiding. By adopting such a configuration, it is possible to eliminate the combination of transform skip, sign hiding, and dependent quantization, which has the effect of reducing the implementation costs of hardware and software. Effect of the Invention
[0013] According to the above configuration, any of the above problems can be solved. [Brief description of the drawings]
[0014] [Figure 1] 1 is a schematic diagram showing a configuration of an image transmission system according to an embodiment of the present invention. [Diagram 2] 1 is a diagram showing the configuration of a transmitting device equipped with a video encoding device according to the present embodiment, and a receiving device equipped with a video decoding device, in which PROD_A indicates the transmitting device equipped with the video encoding device, and PROD_B indicates the receiving device equipped with the video decoding device. [Diagram 3] 1 is a diagram showing the configuration of a recording device equipped with a video encoding device according to the present embodiment, and a playback device equipped with a video decoding device, in which PROD_C indicates a recording device equipped with a video encoding device, and PROD_D indicates a playback device equipped with a video decoding device. [Figure 4] FIG. 2 is a diagram showing a hierarchical structure of data in an encoded stream. [Diagram 5] FIG. 13 is a diagram illustrating an example of division of a CTU. [Figure 6] FIG. 1 is a schematic diagram showing a configuration of a video decoding device. [Figure 7] 11 is a flowchart illustrating a schematic operation of the video decoding device. [Figure 8] FIG. 13 is a schematic diagram showing the configuration of a TU decoding unit. [Figure 9] FIG. 1 illustrates syntax elements for quantized transform coefficients. [Figure 10] FIG. 1 illustrates syntax elements for quantized transform coefficients. [Figure 11a] 1 is a syntax table for RRC. [Figure 11b] 1 is a syntax table for RRC. [Figure 11c] 1 is a syntax table for RRC. [Figure 12a] 1 is a syntax table for TSRC. [Figure 12b] 1 is a syntax table for TSRC. [Figure 13] 13 is a syntax table relating to a slice header according to the present embodiment. [Figure 14a] 1 is a syntax table relating to RRC according to the present embodiment. [Figure 14b]1 is a syntax table relating to RRC according to the present embodiment. [Figure 14c] 1 is a syntax table relating to RRC according to the present embodiment. [Figure 15] 13 is a syntax table relating to a TU according to the present embodiment. [Figure 16a] 1 is a syntax table relating to RRC according to the present embodiment. [Figure 16b] 1 is a syntax table relating to RRC according to the present embodiment. [Figure 16c] 1 is a syntax table relating to RRC according to the present embodiment. [Figure 17] FIG. 1 is a block diagram showing a configuration of a video encoding device. DETAILED DESCRIPTION OF THE PREFERRED EMBODIMENTS
[0015] [Embodiment 1] Hereinafter, an embodiment of the present invention will be described with reference to the drawings.
[0016] FIG. 1 is a schematic diagram showing the configuration of an image transmission system 1 according to this embodiment.
[0017] The image transmission system 1 is a system that transmits an encoded stream obtained by encoding an image to be encoded, decodes the transmitted encoded stream, and displays an image. The image transmission system 1 includes a video encoding device (image encoding device) 11, a network 21, a video decoding device (image decoding device), and a video decoding device. The image display device (image display device) 41 is included in the image display device (image display device) 31 .
[0018] An image T is input to the video encoding device 11 .
[0019] The network 21 transmits the encoded stream Te generated by the video encoding device 11 to the video decoding device 31. The network 21 may be the Internet, a wide area network (WAN), a local area network (LAN), or any of these. The network 21 is not necessarily limited to a two-way communication network, but may be a one-way communication network that transmits broadcast waves such as terrestrial digital broadcasting and satellite broadcasting. The network 21 may also be replaced by a storage medium on which the encoded stream Te is recorded, such as a DVD (Digital Versatile Disc: registered trademark) or a BD (Blue-ray Disc: registered trademark).
[0020] The video decoding device 31 decodes each of the coded streams Te transmitted by the network 21, and generates one or more decoded images Td.
[0021] The image display device 41 displays all or part of one or more decoded images Td generated by the video decoding device 31. The image display device 41 includes a display device such as a liquid crystal display or an organic EL (Electro-luminescence) display. The display may be in the form of a stationary display, a mobile display, an HMD, or the like. In addition, the video decoding device 31 has a high processing capability. When the device has a higher processing power, it displays a high quality image, and when the device has a lower processing power, it displays an image that does not require high processing power or display power.
[0022] <operator> The operators used in this specification are listed below.
[0023] >> is a right bit shift, << is a left bit shift, & is a bitwise AND, | is a bitwise OR , |= is the OR assignment operator, and || indicates logical OR.
[0024] x?y:z is a ternary operator that takes y when x is true (non-zero) and z when x is false (0).
[0025] Clip3(a, b, c) is a function that clips c to a value between a and b, returning a if c < a, b if c > b, and c otherwise (where a <= b).
[0026] abs(a) is a function that returns the absolute value of a.
[0027] Int(a) is a function that returns the integer value of a.
[0028] floor(a) is a function that returns the smallest integer less than or equal to a.
[0029] ceil(a) is a function that returns the largest integer greater than or equal to a.
[0030] a / d represents the division of a by d (truncating the fractional part).
[0031] <Structure of the Encoded Stream Te> Prior to the detailed description of the moving image encoding device 11 and the moving image decoding device 31 according to this embodiment, the data structure of the encoded stream Te generated by the moving image encoding device 11 and decoded by the moving image decoding device 31 will be described.
[0032] Figure 4 is a diagram showing the hierarchical structure of the data in the encoded stream Te. The encoded stream Te exemplarily includes a sequence and a plurality of pictures constituting the sequence. In Figure 4, there are shown, respectively, an encoded video sequence that defines the sequence SEQ, an encoded picture that defines the picture PICT, an encoded slice that defines the slice S, encoded slice data that defines the slice data, an encoded tree unit included in the encoded slice data, and an encoded unit included in the encoded tree unit. Figure 4 shows diagrams indicating an encoded video sequence that defines the sequence SEQ, an encoded picture that defines the picture PICT, an encoded slice that defines the slice S, encoded slice data that defines the slice data, an encoded tree unit included in the encoded slice data, and an encoded unit included in the encoded tree unit.
[0033] (Coded Video Sequence) In the case of a coded video sequence, a video decoder is used to decode the sequence SEQ to be processed. A set of data to be referred to by the device 31 is defined. The sequence SEQ includes a Video Parameter Set, a Sequence A sequence parameter set SPS, a picture parameter set PPS, a picture header, a picture PICT, and an attached It includes supplemental enhancement information SEI (Supplemental Enhancement Information).
[0034] The video parameter set VPS is a set of parameters that can be used to A set of coding parameters common to the image and sets of coding parameters related to the layers and individual layers contained in the image are defined.
[0035] The sequence parameter set SPS is used to decode the target sequence. A set of coding parameters to be referenced by the PPS is specified. For example, the width and height of a picture are specified. Note that there may be multiple SPSs. In that case, one of the multiple SPSs can be selected from the PPS. Select .
[0036] The picture parameter set PPS specifies the number of pictures to be decoded for each picture in the target sequence. A set of coding parameters to be referred to by the video decoding device 31 is defined. For example, the reference value of the quantization width (pic_init_qp_minus26) used for decoding a picture and the application of weighted prediction are defined. A flag (weighted_pred_flag) indicating the weighting factor and a scaling list (quantization matrix) are included. Note that there may be multiple PPSs. In that case, each picture in the target sequence is Select one of multiple PPS from the list.
[0037] The picture header defines coding parameters common to all slices included in one coded picture, such as POC (Picture Order Count) and coding parameters related to partitioning.
[0038] (Encoded Picture) The coded picture defines a set of data to be referenced by the video decoding device 31 in order to decode the picture PICT to be processed. The picture PICT is, as shown in the coded picture of FIG. As shown in the figure, slices 0 to NS-1 are included (NS is the total number of slices included in the picture PICT). .
[0039] In the following, when there is no need to distinguish between slices 0 to NS-1, the symbols are The subscripts may be omitted in the description, and the same applies to other data to which subscripts are added that are included in the coded stream Te described below.
[0040] (Coded Slice) In the case of the coded slice, the video decoding device 31 refers to the slice S to be processed in order to decode the slice S. A slice is a set of data that is encoded as shown in the coding slice of Figure 4. It includes a slice header and slice data.
[0041] The slice header includes a set of coding parameters that the video decoding device 31 refers to in order to determine a decoding method for the current slice. The information (slice_type) is an example of a coding parameter included in the slice header.
[0042] Slice types that can be specified by the slice type specification information include (1) an I slice that uses only intra prediction when encoding, (2) a P slice that uses unidirectional prediction or intra prediction when encoding, and (3) a B slice that uses unidirectional prediction, bidirectional prediction, or intra prediction when encoding. Note that inter prediction is not limited to uni-prediction or bi-prediction, and a predicted image may be generated using more reference pictures. Hereinafter, when referring to P or B slice, it refers to a slice including a block that can use inter prediction.
[0043] In addition, the slice header may include a reference to the picture parameter set PPS (pic_parameter_set_id).
[0044] (Encoded slice data) The coded slice data defines a set of data to be referenced by the video decoding device 31 in order to decode the slice data to be processed. As shown in the slice header, the slice includes a CTU, which is a block of a fixed size (for example, 64x64) that constitutes a slice, and is also called a Largest Coding Unit (LCU).
[0045] (coding tree unit) The coding tree unit in FIG. 4 defines a set of data to be referenced by the video decoding device 31 in order to decode the CTU to be processed. The CTU is divided into recursive quad trees (QTs) partitioning), binary tree partitioning (BT (Binary Tree) partitioning) or ternary tree partitioning (TT (Ternary Tree) The BT and TT partitions are collectively called multi-tree partitions (MT (Multi Tree) partitions). The nodes of the tree structure obtained by partitioning are called coding nodes. The intermediate nodes of the quadtree, binary tree, and ternary tree are coding nodes, and the CTU itself is also a top coding node. It is specified as a code.
[0046] The CT includes a division flag indicating whether or not to perform QT division as CT information. An example of division is shown in FIG. show.
[0047] A CU is a terminal node of a coding node and is not divided any further. A CU is the basic unit of coding processing.
[0048] (Encoding Unit) As shown in the coding unit of FIG. 4, a video is generated to decode the coding unit to be processed. The CU header defines a set of data to be referenced by the image decoding device 31. Specifically, a CU is composed of a CU header CUH, prediction parameters, transformation parameters, quantization transformation coefficients, etc. A prediction mode, etc. are defined.
[0049] The prediction process may be performed in units of CUs, or in units of sub-CUs that are obtained by further dividing a CU. When the sizes of a CU and a sub-CU are equal, there is one sub-CU in a CU. If the size of the CU is larger than the size of the sub-CU, the CU is divided into sub-CUs. For example, if the CU is 8x8 and the sub-CU is 4x4, the CU is divided into 2 horizontal divisions and 2 vertical divisions, into 4 sub-CUs.
[0050] The prediction type (prediction mode CuPredMode) includes at least two types: intra prediction (MODE_INTRA) and inter prediction (MODE_INTER). In addition, intra block copy prediction (MODE_IBC) Intra prediction and intra block copy prediction are predictions within the same picture. Here, inter prediction refers to a prediction process performed between different pictures (for example, between display times, between layer images).
[0051] Transformation and quantization are performed in units of CUs, but quantization coefficients are stored in units of subblocks such as 4x4. It may be entropy coded.
[0052] The predicted image is derived from prediction parameters associated with the block, which include intra-prediction and inter-prediction parameters.
[0053] (Configuration of a video decoding device) The configuration of a video decoding device 31 (FIG. 6) according to this embodiment will be described.
[0054] The video decoding device 31 includes an entropy decoding unit 301, a parameter decoding unit (prediction image decoding device The image processing unit 302 includes a loop filter 305, a reference picture memory 306, a prediction parameter memory 307, a prediction image generating unit (prediction image generating device) 308, an inverse quantization and inverse transformation unit 311, and an adder 312. In addition, in accordance with a video encoding device 11 described later, the video decoding device 31 may have a configuration in which the loop filter 305 is not included.
[0055] The parameter decoding unit 302 further includes a header decoding unit 3020, a CT information decoding unit 3021, and a CU decoding unit 3022. The CU decoding unit 3022 includes a CU decoding unit 3022 (prediction mode decoding unit), and the CU decoding unit 3022 includes a TU decoding unit 3024. These may be collectively referred to as a decoding module. The header decoding unit 3020 decodes parameter set information such as VPS, SPS, and PPS, and slice headers (slice information) from the encoded data. The CT information decoding unit 3021 decodes the CT from the encoded data. The CU decoding unit 3022 decodes the CU from the encoded data. The TU decoding unit 3024 decodes the QP update information (quantization correction value) and the quantized transform coefficient (residual_coding) from the encoded data.
[0056] When a TU includes a prediction error, the TU decoding unit 3024 decodes the QP update information and the quantized transform coefficients from the encoded data. The quantized transform coefficients are derived in a number of modes (e.g., RRC mode In particular, different processes may be performed for derivation of a normal prediction error using a transform (Regular Residual Coding: RRC) and derivation of a prediction error in a transform skip mode without using a transform (Transform Skip Residual Coding: TSRC). The QP update information is a difference value from a quantization parameter predicted value qPpred, which is a predicted value of the quantization parameter QP.
[0057] In the following, we will describe an example in which CTU and CU are used as processing units, but this is not limiting. Alternatively, the processing may be performed in units of sub-CUs. and processing may be performed in units of blocks or sub-blocks.
[0058] The entropy decoding unit 301 performs entropy decoding on the encoded stream Te input from the outside. Entropy coding is a method of variable-length coding syntax elements using a context (probability model) adaptively selected according to the type of syntax element and surrounding circumstances, and a method of variable-length coding syntax elements using a predefined table or formula. An example of the former is CABAC (Context Adaptive Binary Arithmetic Coding). The parsed codes include prediction information for generating a predicted image and prediction errors for generating a difference image.
[0059] The entropy decoding unit 301 outputs the separated code to the parameter decoding unit 302. The separated code is, for example, a prediction mode CuPredMode. Control of which code to decode is performed based on an instruction from the parameter decoding unit 302.
[0060] (Basic flow) FIG. 7 is a flowchart illustrating a schematic operation of the video decoding device 31.
[0061] (S1100: Decoding Parameter Set Information) The header decoding unit 3020 decodes parameter set information such as the VPS, SPS, PPS, SEI, and PH from the encoded data.
[0062] (S1200: Decode slice information) The header decoding unit 3020 decodes slice header information from the encoded data. Decode (slice information).
[0063] Hereinafter, the video decoding device 31 performs steps S1300 to S5000 for each CTU included in the target picture. By repeating the above process, a decoded image of each CTU is derived.
[0064] (S1300: Decode CTU Information) The CT information decoding unit 3021 decodes the CTU from the encoded data.
[0065] (S1400: Decode CT information) The CT information decoding unit 3021 decodes the CT from the encoded data.
[0066] (S1500: CU Decoding) The CU decoding unit 3022 performs S1510 and S1520 to decode the CU from the encoded data. Furthermore, the CU decoding unit 3022 decodes the quantization parameter difference CuQpDeltaVal from the encoded data on a CU basis to derive the quantization parameter.
[0067] (S1510: Decode CU information) The CU decoding unit 3022 decodes CU information, prediction information, and TU division from the encoded data. The flag split_transform_flag, the CU residual flags cbf_cb, cbf_cr, cbf_luma, etc. are decoded.
[0068] (S1520: Decode TU information) When a TU includes a prediction error, the TU decoding unit 3024 decodes the encoded TU information. 2. Decode QP update information and quantized transform coefficients from the data.
[0069] (S2000: Generation of predicted image) The predicted image generation unit 308 generates a predicted image for each block included in the current CU based on prediction information.
[0070] (S3000: Inverse Quantization and Inverse Transformation) The inverse quantization and inverse transform unit 311 executes inverse quantization and inverse transform processing for each TU included in the target CU.
[0071] (S4000: Decoded image generation) The adder 312 receives a predicted image from the predicted image generator 308 and The prediction error supplied from the inverse quantization and inverse transform unit 311 is added to the A decoded image is generated.
[0072] (S5000: Loop Filter) The loop filter 305 applies a loop filter such as a deblocking filter, SAO (Sample Adaptive Filter), or ALF (Adaptive Loop Filter) to the decoded image to generate a decoded image. (Derivation of quantized transform coefficients, residual coding) In lossless coding or when the correlation between pixels in the original image is small, coding efficiency may be higher if no transform is performed. The technology of not performing transform is called transform skip. Transform skip is also called identical transform, and it is a method of changing the transform coefficient according to the quantization parameter. Only scaling of the Y, Cb, and Cr components is performed. Whether or not to skip the transformation is indicated by using the syntax element transform_skip_flag. transform_skip_flag may be indicated for each color component (cIdx) of Y, Cb, and Cr.
[0073] Regular residual coding (RRC) and skip-transformation In contrast to the TSRC (Transform Skip Residual Coding) method, the prediction error is derived The encoding and decoding methods are different.
[0074] FIG. 8 is a block diagram of the TU decoding unit 3024, which includes an RRC unit 30241 and a TSRC unit 30242. The RRC unit 30241 is a processing unit that derives a normal prediction error using a transform, and the TSRC unit 30242 is a processing unit that derives a transform skip error. A processing unit that derives a prediction error in a multi-mode.
[0075] In FIG. 9(a), sps_transform_skip_enabled_flag is a flag indicating whether or not transform_skip_flag is notified in each TU. sps_transform_skip_enabled_flag=1 indicates that transform_skip_flag is notified in each TU. sps_transform_skip_enabled_flag=0 indicates that transform_skip_flag is not notified in each TU. If not, it is assumed to be 0.
[0076] In the figure, min_qp_prime_ts_minus4 is a parameter for deriving the minimum quantization parameter QpPrimeTsMin in the transform skip mode. QpPrimeTsMin is derived as 4+min_qp_prime_ts_minus4.
[0077] In Figure 10(b), transform_skip_flag[x0][y0][cIdx] indicates whether a transform is applied to the block at position (x0, y0) relative to the top-left coordinate of the picture for color component cIdx. If transform_skip_flag=1 (transform skip mode), no transform is applied to this block. If transform_skip_flag=0, whether a transform is applied to this block depends on other parameters. Remains.
[0078] Figures 11 and 12 show the syntax for coding the transform coefficients (prediction errors) in RRC and TSRC. It is.
[0079] (RRC section, RRC mode) In a normal prediction error coding method without transform skip (FIG. 11), the RRC unit 30241 decodes a syntax element (not shown) indicating the LAST position, and derives the LAST position (LastSignificantCoeffX, LastSignificantCoeffY). The LAST position is the position of the last non-zero coefficient when the transform coefficients of a TU are scanned in the direction from low frequency components to high frequency components. When transform coefficients are coded and decoded in order from high frequency components, the LAST position indicates the position of the quantized transform coefficient to be decoded first. Next, the RRC unit 30241 decodes the coded_sub_block_flag by referring to the LAST position. The coded_sub_block_flag is a flag indicating whether or not a non-zero coefficient is included in the sub-block. The sub-block is an area obtained by dividing the TU into 4x4 units. If coded_sub_block_flag=1 (the sub-block includes a non-zero coefficient), the RRC unit 30241 decodes the sig_coeff_flag. sig_coeff_flag is a flag indicating whether the coefficient value is non-zero. If sig_coeff_flag=1 (coefficient value is non-zero), the RRC unit 30241 decodes abs_level_gtx_flag, par_level_flag, abs_remainder, and dec_abs_level. These are syntax elements indicating the absolute value of the coefficient. The RRC unit 30241 derives the absolute value of the coefficient from these syntax elements. Also, the RRC unit 30241 checks whether coeff_sign_flag has been notified by referring to signHidden (described later) and the absolute value and position of the coefficient. If it is notified, coeff_sign_flag is decoded. coeff_sign_flag[n] is skipped. The RRC unit 30241 derives the coefficient value from the absolute value of the coefficient and coeff_sign_flag.
[0080] As described above, the RRC unit 30241 is characterized by decoding the last position in a sub-block of a TU.
[0081] In the prediction error coding method (FIG. 12) when no transform is performed (transform skip mode), the TSRC unit 30242 decodes the coded_sub_block_flag of each subblock. If coded_sub_block_flag=1 (subblock contains a non-zero coefficient), the TSRC unit 30242 Decode sig_coeff_flag[xC][yC] of the transform coefficient at position (xC, xC). sig_coeff_flag=1 (coefficient If the value is non-zero, the TSRC unit 30242 decodes coeff_sign_flag, abs_level_gtx_flag, par_level_flag, and abs_remainder. These are syntax elements that indicate the absolute values of the coefficients. The TSRC unit 30242 derives the absolute values of the coefficients from these syntax elements.
[0082] As described above, the TSRC unit 30242 does not decode the LAST position in the sub-block of the TU. It is characterized by the following.
[0083] (Deriving quantized transform coefficients, Sign Data Hiding) The sign of a non-zero quantized transform coefficient may be estimated by referring to other parameters, other than being derived from a flag notified for each coefficient (sign hiding). sps_sign_data_hiding_enabled_flag in FIG. 9(a) is a flag indicating whether or not the sign bit of a coefficient can be estimated in a picture that refers to a certain SPS. sps_sign_data_hiding_enabled_flag=0 indicates that the sign bit cannot be estimated. sps_sign_data_hiding_enabled_flag=1 indicates that the sign bit may be estimated. If sps_sign_data_hiding_enabled_flag is not notified, it is estimated to be 0.
[0084] The pic_sign_data_hiding_enabled_flag in FIG. 9(b) indicates whether the quantization transform coefficient A flag indicating whether or not sign bit estimation of numbers is possible. pic_sign_data_hiding_enabled_flag=0 indicates that sign bit estimation is not possible. pic_sign_data_hiding_enabled_flag=1 indicates that sign bit estimation may be possible. If pic_sign_data_hiding_enabled_flag is not signaled, it is inferred to be 0.
[0085] In Figure 11, coeff_sign_flag[n] is a flag indicating the sign of the quantized transform coefficient value at scan position n. If coeff_sign_flag[n]=0, the coefficient value is positive. Otherwise (coeff_sign_flag[n]=1), the coefficient value is negative. If coeff_sign_flag[n] is not signaled, it is assumed to be 0. can be.
[0086] The RRC unit 30241 derives a flag signHidden indicating whether or not to derive the sign of a specific transform coefficient by estimation. For example, as shown in the following formula, the RRC unit 30241 derives signHidden by referring to pic_sign_data_hiding_enabled_flag, ph_dep_quant_enabled_flag (described later), and the distribution of non-zero coefficients. The distribution of non-zero coefficients is, for example, the difference value between firstSigScanPosSb and lastSigScanPosSb. If this difference value is greater than a specific value, the RRC unit 30241 sets SignHidden to 1. otherwise, SignHidden may be set to 0. The specified value may be, for example, 3.
[0087] if(ph_dep_quant_enabled_flag || !pic_sign_data_hiding_enabled_flag) signHidden = 0 else signHidden = (lastSigScanPosSb - firstSigScanPosSb > 3 ? 1 : 0) Here, lastSigScanPosSb is the position of the last high-frequency transform coefficient in the subblock (the position of the transform coefficient to be decoded first), and firstSigScanPosSb is the This is the position of the transform coefficient on the low frequency side located first.
[0088] As shown in FIG. 11, when a coefficient value is non-zero (AbsLevel[xC][yC]>0) and any one of the following conditions is satisfied, coeff_sign_flag is notified, and the RRC unit 30241 decodes coeff_sign_flag to derive the sign of the coefficient. -signHidden is 0. The coefficient is not the last decoded position in the sub-block (n != firstSigScanPosSb). Otherwise (if the coefficient value is zero or signHidden is non-zero and the coefficient position is a subblock coeff_sign_flag is not signaled if the bit is the last decoded position in the block; and The RRC unit 30241 determines the sign of the coefficient in question depending on whether the sumAbsLevel is an odd number or an even number (SYN1102). The sumAbsLevel is the sum of the absolute values of the transform coefficients already decoded in the subblock. Specifically, if (sumAbsLevel%2)==1 is true, the sign of the coefficient is assumed to be negative.
[0089] (Derivation of quantized transform coefficients, dependent quantization) There are two quantization methods: scalar quantization and dependent quantization. The inverse quantization and inverse transform unit 311 inverse quantizes the transform coefficients. In the case of dependent quantization, the RRC unit 30241 Even if the quantization is performed, part of the inverse quantization process may be performed.
[0090] The inverse quantization and inverse transform unit 311 derives the linear scale values ls[x][y] according to the values of the quantization parameter qP, rectNonTsFlag, and the quantization matrix m[][] as follows: The inverse transform unit 311 performs the following operations: when the dependent quantization is enabled and the transform skip is disabled; and when the other operations are disabled. Switches the method of deriving ls[][].
[0091] if (ph_dep_quant_enabled_flag && !transform_skip_flag) ls[x][y] = (m[x][y]*levelScale[rectNonTsFlag][(qP+1)%6]) << ((qP+1) / 6) else ls[x][y] = (m[x][y]*levelScale[rectNonTsFlag][qP%6]) << (qP / 6) rectNonTsFlag = transform_skip_flag==1 || (((Log2(nTbW)+Log2(nTbH)) & 1)==1) ? 1:0 bdShift = transform_skip_flag==1 ? 10 : BitDepth+rectNonTsFlag+((Log2(nTbW)+Log2(nTbH)) / 2)-5+ph_dep_quant_enabled_flag First, when scalar quantization is enabled, a transform coefficient is derived uniquely from the quantized transform coefficient and the quantization parameter. For example, in FIG. 11 or FIG. 12, the transform coefficient value d is derived by the following formula: .
[0092] TransCoeffLevel[x0][y0][cIdx][xC][yC] = AbsLevel[xC][yC] * (1-2*coeff_sign_flag[n]) d[x0][y0][cIdx][xC][yC] = (TransCoeffLevel[x0][y0][cIdx][xC][yC]*ls[xC][yC]+((1<<bdShift)> >1)) >> bdShift Here, AbsLevel is the quantized transform coefficient value, and ls, bdShift are parameters derived from the quantization parameter qP.
[0093] On the other hand, dependent quantization has two quantizers with different levels. It switches between four states QState using the parity of the intermediate values (AbsLevelPass1, AbsLevel) of the quantized transform coefficients. Then, quantization and dequantization are performed according to the QState. Note that scaling, which is quantization and dequantization using the quantization parameter, is performed separately.
[0094] QState = QStateTransTable[QState][AbsLevelPass1[xC][yC] & 1] TransCoeffLevel[x0][y0][cIdx][xC][yC] = (2*AbsLevel[xC][yC]-(QState>1 ? 1 : 0)) * (1-2*coeff_sign_flag[n]) d[x0][y0][cIdx][xC][yC] = (TransCoeffLevel[x0][y0][cIdx][xC][yC]*ls[xC][yC]+((1<<bdShift)> >1)) >> bdShift Here, QState is a state, and QStateTransTable[][] is a table used for state transitions, for example, QStateTransTable[][]={{0,2},{2,0},{1,3},{3,1}}.
[0095] QState may be derived using the following formula without using QStateTransTable[][]. QState =(32040 >>((QState << 2)+((AbsLevelPass1[xC][yC] & 1)<< 1)))& 3 Depending on the QState value, a different TransCoeffLevel (or d) is derived even if the AbsLevel is the same. Since the QState is derived by referring to the previously decoded quantized transform coefficient value, it is possible to perform (inverse) quantization with better coding efficiency by utilizing the correlation between coefficients compared to general scalar (inverse) quantization.
[0096] The sps_dep_quant_enabled_flag in FIG. 9(b) indicates whether a picture that references a certain SPS is dependent on the A flag indicating whether dependent quantization is enabled. sps_dep_quant_enabled_flag=0 indicates that dependent quantization is disabled. sps_dep_quant_enabled_flag=1 indicates that dependent quantization may be enabled.
[0097] In the figure, ph_dep_quant_enabled_flag is a flag that indicates whether dependent quantization is enabled for the current picture. lag. ph_dep_quant_enabled_flag=0 indicates that dependency transformation is not enabled. ph_dep_quant_enabled_flag=1 indicates that dependency quantization is enabled. If ph_dep_quant_enabled_flag is not signaled, it is inferred to be 0.
[0098] In the conventional prediction error coding method (RRC), non-zero transform coefficients are present in some areas of low frequency components. To suit cases where non-zero transform coefficients are concentrated in a certain area, the RRC unit 30241 encodes the decoding start position of non-zero transform coefficients, called the LAST position, and scans the low frequency component side from the LAST position to decode the quantized transform coefficients. Also, since sign hiding and dependent quantization are used, the RRC unit 30241 derives parameters such as the sum of absolute values of the transform coefficients required for these processes and the state of dependent quantization (QState). On the other hand, in predictive error coding using transform skip mode (TSRC), non-zero transform coefficients are not concentrated in a certain area, so the LAST position is not encoded, and the entire block is scanned to decode the quantized transform coefficients. Also, sign hiding and dependent quantization are not used, and the TSRC unit 30242 derives parameters such as the sum of absolute values of the transform coefficients required for these processes. The above parameters required for processing are not derived.
[0099] However, even in the case of the transform skip mode, there are cases where the coding efficiency is better when the normal prediction error coding method is used. For such a case, the normal prediction error may be used in the case of the transform skip mode. For example, slice_ts_residual_coding_disabled_flag shown in FIG. 10(a) may be notified.
[0100] slice_ts_residual_coding_disabled_flag indicates whether residual_ts_coding(), which is a TSRC mode, is used to decode the prediction error of a block to which transform skip is applied (transform skip block) in the current slice. slice_ts_residual_coding_disabled_flag=1 indicates that residual_coding(), which is an RRC mode, is used to parse the prediction error of the transform skip block. slice_ts_residual_coding_disabled_flag=0 indicates that residual_ts_coding() is used to parse the prediction error of the transform skip block. If slice_ts_residual_coding_disabled_flag is not notified, it is estimated to be 0.
[0101] As shown in FIG. 10B, in the case of !transform_skip_flag[x0][y0][0] || slice_ts_residual_coding_disabled_flag, the RRC unit 30241 may process the TSRC unit 30242. That is, in addition to the case where transform_skip_flag is 0, the transform coefficients (prediction errors) may be notified by RRC even when slice_ts_residual_coding_disabled_flag is not 0.
[0102] By referring to slice_ts_residual_coding_disabled_flag, the normal prediction error coding method (RRC mode) can be used even in the transform skip mode. (Embodiment 2) As shown in FIG. 13, slice_ts_residual_coding_disabled_flag may refer to sps_transform_skip_enabled_flag, and when sps_transform_skip_enabled_flag=1, slice_ts_residual_coding_disabled_flag may be notified. slice_ts_residual_coding_disabled_flag is a flag indicating whether or not the TSRC mode is used to decode the prediction error of a block to which transform skip is applied in the current slice. Therefore, in the case of transform skip in the SPS, When it is notified that the mode is not used, the TSRC mode is not used and there is no need to notify slice_ts_residual_coding_disabled_flag, which makes it possible to reduce the amount of coding, thereby improving the coding efficiency. (Embodiment 3) By operating transform skip and dependent quantization, and also operating sign hiding and dependent quantization exclusively, it is possible to reduce the implementation costs of hardware and software.
[0103] As shown in FIG. 14, the RRC unit 30241 may switch the four states QState of dependent quantization by referring to slice_ts_residual_coding_disabled_flag (SYN1401, SYN1402, SYN1404, SYN1405). Specifically, when ph_dep_quant_enabled_flag is 1 and slice_ts_residual_coding_disabled_flag is 0, the QState is updated (SYN1401, SYN1402). Conversely, when ph_dep_quant_enabled_flag is 0 or slice_ts_residual_coding_disabled_flag is 1, the QState is not updated. As a result, the quantized transform coefficients of the transform skip are also updated in the RRC mode. Therefore, even if the transform skip is enabled, the RRC unit 30241 does not perform dependent quantization if slice_ts_residual_coding_disabled_flag=1. Under the above conditions, it is possible to operate the transform skip and the dependent quantization exclusively.
[0104] Alternatively, when updating QState without using QStateTransTable[][], the variable stateVal may be used. stateVal = (ph_dep_quant_enabled_flag && !slice_ts_residual_coding_disabled_flag) ? 32040 : 0 QState =(stateVal >>((QState << 2)+((AbsLevelPass1[xC][yC] & 1)<< 1)))& 3 When stateVal is 0, QState is always 0, which means that dependent quantization is not performed, so it is possible to operate transform skip and dependent quantization exclusively.
[0105] As shown in FIG. 14, slice_ts_residual_coding_disabled_flag may be used to derive the value of signHidden, which determines whether to apply sign hiding (SYN1403). If ph_dep_quant_enabled_flag is 1, or pic_sign_data_hiding_enabled_flag is 0, or slice_ts_residual_coding_disabled_flag is 1, set signHidden to 0.
[0106] if(ph_dep_quant_enabled_flag || !pic_sign_data_hiding_enabled_flag || slice_ts_residual_coding_disabled_flag) signHidden = 0 else signHidden = (lastSigScanPosSb-firstSigScanPosSb>3 ? 1 : 0) By doing this, when ph_dep_quant_enabled_flag is 0, pic_sign_data_hiding_enabled_flag is 1, and slice_ts_residual_coding_disabled_flag is 0, The distribution of non-zero coefficients is, for example, the difference between firstSigScanPosSb and lastSigScanPosSb. It is possible to have unhiding operate exclusively with
[0107] Also, the RRC unit 30241 uses the transform_skip_flag instead of the slice_ts_residual_coding_disabled_flag to exclusively perform the transform skip, the sign hiding, and the dependent quantization. may be operated in the same manner.
[0108] As shown in FIG. 15, first, the TU decoding unit 3024 notifies the RRC unit 30241 of transform_skip_flag. Next, as shown in FIG. 16, the RRC unit 30241 If ant_enabled_flag is 1 and transform_skip_flag is 0, update the QState. On the other hand, if ph_dep_quant_enabled_flag is 0 or slice_ts_residual_coding_disabled_flag is 1, the RRC unit 30241 does not update the QState. When the SYN1601, SYN1602, SYN1604, SYN1605 are enabled, no dependent quantization is performed.
[0109] Furthermore, transform_skip_flag may be used to derive the value of signHidden. If ph_dep_quant_enabled_flag is 1, or pic_sign_data_hiding_enabled_flag is 0, or transform_skip_flag is 1, the RRC unit 30241 sets signHidden to 0 (SYN1603).
[0110] if(ph_dep_quant_enabled_flag || !pic_sign_data_hiding_enabled_flag || transform_skip_flag) signHidden = 0 else signHidden = (lastSigScanPosSb-firstSigScanPosSb>3? 1 : 0) By adopting such a configuration, it is possible to eliminate the combination of transform skip, sign hiding, and dependent quantization, which has the effect of reducing the implementation costs of hardware and software.
[0111] The inverse quantization and inverse transform unit 311 scales (inverse quantizes) the quantized transform coefficients input from the entropy decoding unit 301 to obtain transform coefficients d[][]. In the process, the prediction error is transformed by DCT (Discrete Cosine Transform), DST (Discrete Sine Transform), etc., and then quantized. The inverse quantization and inverse transform unit 311 performs inverse frequency transform such as inverse DCT and inverse DST on the transform coefficients d[][] to calculate the prediction error res[][]. When transform_skip_flag is 1, the inverse quantization and inverse transform unit 311 sets res[x][y]=d[x][y]. The inverse quantization and inverse transform unit 311 outputs the prediction error to the adder 312.
[0112] In addition, since the inverse transformation and the transformation are paired processes, the transformation and the inverse transformation may be interpreted as being interchangeable. Alternatively, when the inverse transformation is called the transformation, the transformation may be called the forward transformation. For example, when the inverse non-separable transformation is called the non-separable transformation, the non-separable transformation may be called the forward non-separable transformation. Moreover, the separable transformation is simply called the transformation.
[0113] The adder 312 adds, for each pixel, the predicted image of the block input from the predicted image generation unit 308 and the prediction error input from the inverse quantization and inverse transform unit 311 to generate a decoded image of the block. The adder 312 stores the decoded image of the block in the reference picture memory 306 , and also outputs it to the loop filter 305 .
[0114] (Configuration of a video encoding device) Next, the configuration of the video encoding device 11 according to this embodiment will be described. Fig. 17 is a block diagram showing the configuration of the video encoding device 11 according to this embodiment. The video encoding device 11 includes a prediction image generating unit 101, a subtraction unit 102, a transformation and quantization unit 103, an inverse quantization and inverse transformation unit 105, an addition unit 106, a loop filter 107, a prediction parameter memory (prediction parameter storage unit, frame memory) 108, a reference picture memory (reference image storage unit, frame memory) 109, an encoding parameter determining unit 110, a parameter encoding unit 111, and an entropy encoding unit 104.
[0115] The predicted image generating unit 101 generates a predicted image for each CU, which is an area obtained by dividing each picture of the image T. The predicted image generating unit 101 operates in the same manner as the predicted image generating unit 308 already described, and therefore a description thereof will be omitted.
[0116] The subtraction unit 102 subtracts the pixel values of the prediction image of the block input from the prediction image generation unit 101 from the pixel values of the image T to generate a prediction error. Output to.
[0117] The transform / quantization unit 103 calculates transform coefficients by frequency transforming the prediction error input from the subtraction unit 102, and derives quantized transform coefficients by quantizing the prediction error. The quantized transform coefficients are output to the entropy coding unit 104 and the inverse quantization and inverse transform unit 105 .
[0118] The inverse quantization and inverse transform unit 105 is the same as the inverse quantization and inverse transform unit 311 (FIG. 10) in the video decoding device 31, and a description thereof will be omitted. The calculated prediction error is output to the addition unit .
[0119] The entropy coding unit 104 receives the quantized transform coefficients from the transform / quantization unit 103 and the coding parameters from the parameter coding unit 111. The coding parameters are, for example, , and predMode indicating a prediction mode. predMode may be either MODE_INTRA indicating intra prediction or MODE_INTER indicating inter prediction, or may be MODE_INTRA, MODE_INTER, or MODE_IBC indicating intra block copy prediction in which a block in a screen is copied and used as a predicted image.
[0120] The entropy coding unit 104 entropies the division information, prediction parameters, quantized transform coefficients, etc. The encoded stream Te is then generated and output by performing bit-wise encoding.
[0121] The parameter encoding unit 111 includes a header encoding unit 1110, a CT information encoding unit 1111, and a CU A coding unit 1112 (prediction mode coding unit), and an inter-prediction parameter coding unit 112 and an inter The CU encoding unit 1112 further includes a TU encoding unit 1114. It is equipped with:
[0122] The operation of each module will be described below. The parameter coding unit 111 encodes header information, The coding process is performed on parameters such as division information, prediction information, and quantized transform coefficients.
[0123] The CT information encoding unit 1111 encodes the QT, MT (BT, TT) division information and the like from the encoded data.
[0124] The CU encoding unit 1112 encodes the CU information, prediction information, TU division flag, CU residual flag, and the like.
[0125] When a prediction error is included in a TU, the TU encoding unit 1114 encodes the QP update information and the quantized transform coefficients.
[0126] The CT information encoding unit 1111 and the CU encoding unit 1112 supply syntax elements such as inter prediction parameters, intra prediction parameters, and quantized transform coefficients to the entropy encoding unit 104.
[0127] The adder 106 adds, for each pixel, the pixel value of the predicted image of the block input from the predicted image generation unit 101 and the prediction error input from the inverse quantization and inverse transform unit 105 to generate a decoded image. The unit 106 stores the generated decoded image in a reference picture memory 109 .
[0128] The loop filter 107 performs deblocking filtering, SAO, and ALF on the decoded image generated by the adder 106. Note that the loop filter 107 does not necessarily include the above three types of filters. For example, the filter may be configured with only a deblocking filter.
[0129] SAO is a filter that adds an offset according to the classification result on a sample-by-sample basis, and ALF is a filter that uses the sum of products of the transmitted filter coefficients and a reference image (or the difference between the reference image and the target pixel).
[0130] The prediction parameter memory 108 stores the prediction parameters generated by the coding parameter determination unit 110 in a predetermined location for each current picture and CU.
[0131] The reference picture memory 109 stores the decoded image generated by the loop filter 107 at a predetermined position for each current picture and CU.
[0132] The coding parameter determination unit 110 selects one set from among a plurality of sets of coding parameters. The coding parameters are the above-mentioned QT, BT or TT division information, prediction parameters, or parameters to be coded that are generated in relation to these. The predicted image generation unit 101 generates a predicted image using these coding parameters.
[0133] The coding parameter determination unit 110 determines the size of the amount of information and the coding parameter for each of the plurality of sets. The coding parameter determination unit 110 calculates an RD cost value indicating an error. The entropy coding unit 104 then selects the set of coding parameters that results in the smallest coding rate. The entropy coding unit 104 then outputs the selected set of coding parameters as the coded stream Te. The parameter determination unit 110 stores the determined coding parameters in the prediction parameter memory 108 .
[0134] In addition, in the above-described embodiment, a part of the video encoding device 11 and the video decoding device 31, for example, the entropy decoding unit 301, the parameter decoding unit 302, the loop filter 305, the predicted image generating unit 306, 308, an inverse quantization and inverse transform unit 311, an addition unit 312, a predicted image generation unit 101, a subtraction unit 102, a transform and The quantization unit 103, the entropy coding unit 104, the inverse quantization and inverse transform unit 105, the loop filter 107, the coding parameter determination unit 110, and the parameter coding unit 111 may be realized by a computer. In this case, a program for realizing these control functions may be recorded on a computer-readable recording medium, and the program recorded on the recording medium may be read into a computer system and executed. Note that the "computer system" referred to here is a computer system built into either the video coding device 11 or the video decoding device 31, and includes hardware such as an OS and peripheral devices. Also, the "computer-readable recording medium" refers to a flexible disk, a magneto-optical disk, a ROM, etc. The term "computer-readable recording medium" refers to a portable medium such as a CD-ROM, or a storage device such as a hard disk built into a computer system. Furthermore, "computer-readable recording medium" may also include a device that dynamically stores a program for a short period of time, such as a communication line when transmitting a program via a network such as the Internet or a communication line such as a telephone line, or a device that stores a program for a certain period of time, such as a volatile memory inside a computer system that serves as a server or client in such a case. Furthermore, the above program may be one that realizes part of the above-mentioned functions, or may be one that can realize the above-mentioned functions in combination with a program already recorded in the computer system.
[0135] In addition, a part or the whole of the video encoding device 11 and the video decoding device 31 in the above-mentioned embodiment may be realized as an integrated circuit such as an LSI (Large Scale Integration). Each functional block of the video encoding device 11 and the video decoding device 31 may be individually processed, or may be integrated into a processor in part or in whole. The integrated circuit method is not limited to LSI, and may be realized by a dedicated circuit or a general-purpose processor. In addition, when an integrated circuit technology that replaces LSI appears due to the advancement of semiconductor technology, an integrated circuit based on that technology may be used.
[0136] Although one embodiment of the present invention has been described in detail above with reference to the drawings, the specific configuration is not limited to the above, and various design changes, etc. are possible within the scope that does not deviate from the gist of the present invention.
[0137] [Application example] The above-mentioned video encoding device 11 and video decoding device 31 can be mounted on various devices that transmit, receive, record, and play videos. The video may be a natural video captured by a camera or the like, or an artificial video (including CG and GUI) generated by a computer or the like.
[0138] First, it will be described with reference to FIG. 2 that the above-mentioned video encoding device 11 and video decoding device 31 can be used for transmitting and receiving videos.
[0139] FIG. 2 is a block diagram showing the configuration of a transmission device PROD_A equipped with a video encoding device 11. As shown in the figure, the transmitting device PROD_A includes an encoding unit PROD_A1 that obtains encoded data by encoding a moving image, a modulation unit PROD_A2 that obtains a modulated signal by modulating a carrier wave with the encoded data obtained by the encoding unit PROD_A1, and a transmitting unit PROD_A3 that transmits the modulated signal obtained by the modulation unit PROD_A2. The above-mentioned moving image encoding device 11 is used as this encoding unit PROD_A1.
[0140] The transmission device PROD_A captures moving images as a supply source of moving images to be input to the encoding unit PROD_A1. The apparatus may further include a camera PROD_A4 for recording moving images, a recording medium PROD_A5 for recording moving images, an input terminal PROD_A6 for inputting moving images from the outside, and an image processing unit A7 for generating or processing images. In the figure, a configuration in which the transmitting device PROD_A includes all of these components is illustrated, but some of them may be omitted.
[0141] The recording medium PROD_A5 may be one in which unencoded moving images are recorded. Alternatively, the recording medium PROD_A5 may be a recording medium that has recorded thereon a moving image encoded by a recording encoding method different from the transmission encoding method. In the latter case, a decoding unit ( It is advisable to use a conductor (not shown) between the two.
[0142] FIG. 2 is a block diagram showing the configuration of a receiving device PROD_B equipped with a video decoding device 31. FIG. 1 shows a receiving device PROD_B. As shown in the figure, the receiving device PROD_B includes a receiving unit PROD_B1 that receives a modulated signal, a demodulation unit PROD_B2 that obtains coded data by demodulating the modulated signal received by the receiving unit PROD_B1, and a decoding unit PROD_B3 that obtains a moving image by decoding the coded data obtained by the demodulation unit PROD_B2. The above-mentioned moving image decoding device 31 is used as this decoding unit PROD_B3.
[0143] The receiving device PROD_B is a supply destination of the video output by the decoding unit PROD_B3, and displays the video. The display PROD_B4 may further include a recording medium PROD_B5 for recording moving images, and an output terminal PROD_B6 for outputting the moving images to the outside. Although the receiving device PROD_B is illustrated as having the above configuration, some of these may be omitted.
[0144] The recording medium PROD_B5 is for recording unencoded moving images. Alternatively, the signal may be encoded by a coding method for recording that is different from the coding method for transmission. In the latter case, a signal from the decoding unit PROD_B3 to the recording medium PROD_B5 may be provided between the decoding unit PROD_B3 and the recording medium PROD_B5. It is preferable to interpose an encoding unit (not shown) that encodes the acquired video image according to an encoding method for recording.
[0145] The transmission medium for transmitting the modulated signal may be wireless or wired. The transmission mode for transmitting the modulated signal may be broadcast (here, this refers to a transmission mode in which the destination is not specified in advance) or communication (here, this refers to a transmission mode in which the destination is specified in advance). That is, the transmission of the modulated signal may be realized by any of wireless broadcasting, wired broadcasting, wireless communication, and wired communication.
[0146] For example, a broadcasting station (broadcasting equipment, etc.) / receiving station (television receiver, etc.) of terrestrial digital broadcasting is an example of a transmitting device PROD_A / receiving device PROD_B that transmits and receives modulated signals by wireless broadcasting. Also, a broadcasting station (broadcasting equipment, etc.) / receiving station (television receiver, etc.) of cable television broadcasting is an example of a transmitting device PROD_A / receiving device PROD_B that transmits and receives modulated signals by cable broadcasting.
[0147] Also, a server (such as a workstation) / client (such as a television receiver, a personal computer, a smartphone, etc.) of a VOD (Video On Demand) service or a video sharing service using the Internet is an example of a transmitting device PROD_A / receiving device PROD_B that transmits and receives modulated signals by communication (usually, in a LAN, either wireless or wired is used as a transmission medium, and in a WAN, wired is used as a transmission medium). Here, personal computers include desktop PCs, laptop PCs, and tablet PCs. Also, smartphones include multi-function mobile phone terminals.
[0148] The client of the video hosting service has a function to decode the encoded data downloaded from the server and display it on a display, as well as a function to encode the video captured by the camera and upload it to the server. That is, the client of the video hosting service functions as both the transmitting device PROD_A and the receiving device PROD_B.
[0149] Next, it will be described with reference to FIG. 3 that the above-mentioned video encoding device 11 and video decoding device 31 can be used for recording and reproducing videos.
[0150] FIG. 3 is a block diagram showing the configuration of a recording device PROD_C equipped with the above-mentioned video encoding device 11. As shown in the figure, the recording device PROD_C includes an encoding unit PROD_C1 that encodes a moving image to obtain encoded data, and a writing unit PROD_C2 that writes the encoded data obtained by the encoding unit PROD_C1 onto a recording medium PROD_M. is used as this encoding unit PROD_C1.
[0151] The recording medium PROD_M may be (1) a type built into the recording device PROD_C, such as a hard disk drive (HDD) or a solid state drive (SSD), or (2) a type that is attached to the recording device PROD_C, such as an SD memory card or a universal serial bus (USB) flash memory. (3) a drive built into the recording device PROD_C, such as a DVD (Digital Versatile Disc: registered trademark) or a BD (Blu-ray Disc: registered trademark) It may be loaded into a device (not shown).
[0152] The recording device PROD_C supplies a video image to the encoding unit PROD_C1. The recording device PROD_C may further include a camera PROD_C3 for capturing an image, an input terminal PROD_C4 for inputting a moving image from the outside, a receiving unit PROD_C5 for receiving the moving image, and an image processing unit PROD_C6 for generating or processing an image. In the figure, the recording device PROD_C is illustrated as having all of these components, but some of them may be omitted.
[0153] The receiving unit PROD_C5 may receive uncoded video. Alternatively, the receiving unit PROD_C5 may receive encoded data that has been encoded by a transmission encoding method different from the recording encoding method. In the latter case, a transmission decoding unit (not shown) that decodes the encoded data that has been encoded by the transmission encoding method may be provided between the receiving unit PROD_C5 and the encoding unit PROD_C1.
[0154] Examples of such a recording device PROD_C include a DVD recorder, a BD recorder, and a HDD (Hard Drive (In this case, the input terminal PROD_C4 or the receiving unit PROD_C5 is The main source of video is the camcorder (in this case, camera PROD_C3). The main source of supply), a personal computer (in this case, the receiving unit PROD_C5 or the image processing unit (in this case, the camera PROD_C3 or the receiving unit PROD_C6 is the main source of video images), A receiving unit PROD_C5 is a main source of video images. do.
[0155] FIG. 3 shows the configuration of a playback device PROD_D equipped with the above-mentioned video decoding device 31. As shown in the figure, the playback device PROD_D includes a reading unit PROD_D1 that reads out the coded data written on the recording medium PROD_M, and a decoding unit PROD_D2 that obtains a moving image by decoding the coded data read by the reading unit PROD_D1. The image decoding device 31 is used as this decoding unit PROD_D2.
[0156] The recording medium PROD_M may be (1) a type built into the playback device PROD_D, such as an HDD or SSD, or (2) a type that can be stored in a portable device, such as an SD memory card or a USB flash memory. (3) a type that can be connected to a playback device PROD_D, such as a DVD or BD, Alternatively, the disc may be loaded into a drive device (not shown) built into the playback device PROD_D.
[0157] In addition, the playback device PROD_D receives the video output from the decoding unit PROD_D2 and outputs the video to the playback device PROD_D. The apparatus may further include a display PROD_D3 for displaying the moving image, an output terminal PROD_D4 for outputting the moving image to the outside, and a transmission unit PROD_D5 for transmitting the moving image. Although the configuration of the playback device PROD_D is illustrated, some of the configuration may be omitted.
[0158] The transmission unit PROD_D5 may transmit uncoded video. Alternatively, the decoder unit PROD_D2 may transmit encoded data encoded by a transmission encoding method different from the recording encoding method. In the latter case, it is preferable to interpose an encoding unit (not shown) between the decoder unit PROD_D2 and the transmitter unit PROD_D5, which encodes the video image by the transmission encoding method.
[0159] Examples of such a playback device PROD_D include a DVD player, a BD player, and a HDD player. The main source of image data is a television receiver (in this case, the display PROD_D3 is the main destination of the moving image), digital signage (also called electronic signboard or electronic bulletin board, etc., the display PROD_D3 or the transmission unit PROD_D5 is the main destination of the moving image), desktop PC (in this case, the output terminal PROD_D4 or the transmission unit PROD_D5 is the main destination of the moving image), laptop or tablet PC (in this case, the display PROD_D3 or the transmission unit PROD_D5 is the main destination of the moving image), smartphone (in this case, the display PROD_D3 or The transmitting unit PROD_D5 is the main supply destination of the moving image) is also an example of such a playback device PROD_D. be.
[0160] (Hardware and Software Realization) Furthermore, each block of the above-mentioned video decoding device 31 and video encoding device 11 may be realized in hardware by a logic circuit formed on an integrated circuit (IC chip), or may be realized in software by using a CPU (Central Processing Unit).
[0161] In the latter case, each of the above devices includes a CPU that executes the instructions of a program that realizes each function, ROM (Read Only Memory) stores the program, and RAM (Random Access Memory) expands the program. The object of the embodiment of the present invention is to provide a control program for each of the above devices, which is software for implementing the above-mentioned functions, and a storage device (recording medium) such as a control program for each of the above devices. This can also be achieved by supplying a recording medium on which the program code (such as a program code, intermediate code program, or source program) is recorded in a computer-readable manner to each of the above-mentioned devices, and having the computer (or CPU or MPU) read and execute the program code recorded on the recording medium.
[0162] Examples of the recording medium include tapes such as magnetic tapes and cassette tapes, magnetic disks such as floppy disks (registered trademark) and hard disks, and CD-ROMs (Compact Disc Read-Only Memory), MO disks (Magneto-Optical discs), MDs (Mini Discs), DVDs (Digital Versatile Discs: registered trademark), CD-Rs (CD Recordable), and Blu-ray discs (Blu-ray Discs, including optical discs such as Disc: registered trademark, IC cards (including memory cards) / Cards such as optical cards, mask ROM / EPROM (Erasable Programmable Read-Only Memory) / EEPROM (Electrically Erasable and Programmable Read-Only Memory: registered trademark) Semiconductor memories such as flash ROM, or logic circuits such as PLD (Programmable logic device) or FPGA (Field Programmable Gate Array) can be used.
[0163] Moreover, each of the above devices may be configured to be connectable to a communication network, and the above program code may be supplied via the communication network. This communication network is not particularly limited as long as it is capable of transmitting the program code. For example, the Internet, an intranet, an extranet, a LAN (Local Area Network), an ISDN (Integrated Services Digital Network), a VAN (Value-Added Network), a CATV (Community Antenna television / Cable Television) communication network, a virtual private network, a telephone line network, a mobile communication network, a satellite communication network, etc. may be used. Moreover, the transmission medium constituting this communication network is not limited to a specific configuration or type as long as it is a medium capable of transmitting the program code. For example, it can be used in wired communication such as IEEE (Institute of Electrical and Electronic Engineers) 1394, USB, power line carrier, cable TV line, telephone line, ADSL (Asymmetric Digital Subscriber Line) line, etc., or in wireless communication such as infrared such as IrDA (Infrared Data Association) or remote control, BlueTooth (registered trademark), IEEE802.11 wireless, HDR (High Data Rate), NFC (Near Field Communication), DLNA (Digital Living Network Alliance: registered trademark), mobile phone network, satellite line, terrestrial digital broadcasting network, etc. The embodiment of the present invention can also be realized in the form of a computer data signal embedded in a carrier wave in which the program code is embodied by electronic transmission.
[0164] The present invention is not limited to the above-described embodiment, and various modifications are possible within the scope of the claims. In other words, the technical scope of the present invention also includes embodiments obtained by combining technical means that are appropriately modified within the scope of the claims. [Industrial Applicability]
[0165] The embodiments of the present invention can be suitably applied to a video decoding device that decodes coded data in which image data is coded, and a video coding device that generates coded data in which image data is coded, and can also be suitably applied to the data structure of coded data that is generated by a video coding device and referenced by the video decoding device. [Explanation of symbols]
[0166] 31 Video Decoding Device 301 Entropy Decoding Unit 302 Parameter Decoding Unit 3020 Header Decoding Unit 308 Prediction Image Generation Unit 311 Inverse quantization and inverse transformation unit 312 Addition section 11 Video Encoding Device 101 Prediction image generation unit 102 Subtraction section 103 Transformation and Quantization Section 104 Entropy coding unit 105 Inverse quantization and inverse transformation unit 107 Loop Filter 110 Encoding parameter determination unit 111 Parameter Encoding Unit 1110 Header encoding part 1111 CT information encoder 1112 CU encoding unit (prediction mode encoding unit) 1114 TU encoding section 311 Inverse quantization and inverse transformation unit 3111 Inverse quantization section 3112 Inverse conversion unit
Claims
1. An image decoding device that decodes encoded data, comprising: a header decoding unit for decoding a transform skip residual coding disable flag indicating that normal prediction error coding is used to decode the prediction error of the transform skip block of the current slice, or that transform skip prediction error coding is used to decode the prediction error of the transform skip block of the current slice, based on a value of a transform skip enable flag indicating whether a transform skip flag exists in the transform unit; Based on the value of the transform skip valid flag, the value of BdpcmFlag, and the transform block size, a TU decoding unit for decoding the transform skip flag indicating whether a transform is applied to a corresponding block; an inverse quantization and inverse transform unit that derives transform coefficient values using the scale value; If the transform skip flag has a value of 0 or the transform skip residual coding disable flag has a value of 1, the normal prediction error coding is used; Otherwise, the transform skip prediction error coding described above is used, An image decoding device, comprising: a scale value deriving section for deriving the scale value by switching between at least two derivation methods based on a value of the transform skip flag.
2. An image encoding device for encoding an image, comprising: a header encoding unit for encoding a transform skip residual coding disable flag indicating that normal prediction error coding is used to code the prediction error of the transform skip block of the current slice, or that transform skip prediction error coding is used to code the prediction error of the transform skip block of the current slice, based on a value of a transform skip enable flag indicating whether a transform skip flag exists in the transform unit; Based on the value of the transform skip valid flag, the value of BdpcmFlag, and the transform block size, a TU encoding unit for encoding the transform skip flag indicating whether a transform is applied to a corresponding block; an inverse quantization and inverse transform unit that derives transform coefficient values using the scale value; If the transform skip flag has a value of 0 or the transform skip residual coding disable flag has a value of 1, the normal prediction error coding is used; Otherwise, the transform skip prediction error coding described above is used, deriving the scale value by switching between at least two derivation methods based on a value of the transform skip flag.
3. 1. An image decoding method for decoding encoded data, comprising the steps of: decoding a transform skip residual coding disable flag indicating that normal prediction error coding is used to decode the prediction error of the transform skip block of the current slice, or that transform skip prediction error coding is used to decode the prediction error of the transform skip block of the current slice, based on a value of the transform skip enable flag indicating whether a transform skip flag exists in the transform unit; Based on the value of the transform skip valid flag, the value of BdpcmFlag, and the transform block size, decoding said transform skip flag indicating whether a transform is applied to a corresponding block; and deriving transform coefficient values using the scale values, If the transform skip flag has a value of 0 or the transform skip residual coding disable flag has a value of 1, the normal prediction error coding is used; Otherwise, the transform skip prediction error coding described above is used, An image decoding method, comprising the steps of: deriving the scale value by switching between at least two derivation methods based on a value of the transform skip flag.
4. 1. An image encoding method for encoding an image, comprising: encoding a transform skip residual coding disable flag indicating that normal prediction error coding is used to code the prediction error of the transform skip block of the current slice, or that transform skip prediction error coding is used to code the prediction error of the transform skip block of the current slice, based on a value of the transform skip enable flag indicating whether a transform skip flag exists in the transform unit; Based on the value of the transform skip valid flag, the value of BdpcmFlag, and the transform block size, encoding said transform skip flag indicating whether a transform is applied to a corresponding block or not; and deriving transform coefficient values using the scale values, If the transform skip flag has a value of 0 or the transform skip residual coding disable flag has a value of 1, the normal prediction error coding is used; Otherwise, the transform skip prediction error coding described above is used, An image coding method, comprising the steps of: deriving the scale value by switching between at least two derivation methods based on a value of the transform skip flag.
Citation Information
Patent Citations
Signaling a coding scheme for residual values in transform skips for video coding
JP2022552173A
Image decoding method and apparatus for residual coding
JP2023504940A
JPP7469904B