Image Processing Apparatus and Method
By setting the minimum encoding cost transformation type in the non-joint chromatic aberration coding mode in the image processing device to the joint chromatic aberration coding mode, the problem of degradation of encoding efficiency caused by the redundancy of the transformation skip mark in the joint chromatic aberration coding mode is solved, and the effect of suppressing the increase in encoding load is achieved.
Patent Information
- Application Number
- CN202080083856.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Priority Date
- 2019-12-12
- Filing Date
- 2020-12-11
- Publication Date
- 2025-06-17
- Estimated Expiration
- 2040-12-11
AI Technical Summary
In the joint color difference coding mode, in the prior art, the coding amount is unnecessary to increase and the encoding efficiency is reduced due to the need to notify the transformation skip flag.
The encoding mode of the image is set by setting the transform type with the minimum encoding cost in the non-joint chromatic aberration coding mode to the transform type in the joint chromatic aberration coding mode and obtaining the encoding cost in the joint chromatic aberration coding mode.
The increase in coding load is suppressed, and the increase in coding complexity and load is avoided, thereby improving coding efficiency.
Smart Images

Figure CN114762327B_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to an image processing apparatus and method, and more particularly, to an image processing apparatus and method capable of suppressing an increase in encoding load. Background Art
[0002] In the past, encoding methods for obtaining a prediction residual of a moving image, performing coefficient transformation, quantization, and encoding have been proposed (for example, see Non-Patent Document 1 and Non-Patent Document 2). In the Versatile Video Coding (VVC) working draft described in Non-Patent Document 1, chroma transform skip can be applied regardless of the joint chroma coding mode (joint CbCr mode). Meanwhile, in the implementation of the VVC VTM software described in Non-Patent Document 2, the application of chroma transform skip is restricted in the joint chroma coding mode.
[0003] In the case where the application of chroma transform skip is restricted in the joint chroma coding mode as described in Non-Patent Document 2, it is unnecessary to signal a transform skip flag in the joint chroma coding mode. That is, since the transform skip flag is signaled in the joint chroma coding mode, the amount of encoding may increase unnecessarily, and the encoding efficiency may decrease. That is, there is a possibility that the encoding efficiency decreases. In contrast, in the case of the method described in Non-Patent Document 1, the application of chroma transform skip is not restricted in the joint chroma coding mode, and thus a decrease in encoding efficiency due to redundancy of the transform skip flag is suppressed.
[0004] Citation List
[0005] Non-Patent Document
[0006] Non-Patent Document 1: Benjamin Bross, Jianle Chen, Shan Liu, Ye-Kui Wang, "Versatile Video Coding (Draft 7)", JVET-P2001-vE, 16th meeting of the Joint Video Exploration Team (JVET) of ITU-T SG 16 WP 3 and ISO / IEC JTC 1 / SC 29 / WG 11: Geneva, Switzerland, October 1 - 11, 2019.
[0007] Non-Patent Document 2: Jianle Chen, Yan Ye, Seung Hwan Kim, "Algorithm description for Versatile Video Coding and Test Model 7 (VTM 7)", JVET-P2002-v1, 16th Meeting of the Joint Video Exploration Team (JVET) of ITU-T SG 16 WP 3 and ISO / IEC JTC 1 / SC 29 / WG 11: Geneva, Switzerland, October 1 - 11, 2019. Summary of the Invention
[0008] Problems to be Solved by the Invention
[0009] However, in the case of the method described in Non-Patent Document 1, it is necessary to evaluate two cases: the case where transform skip is applied to the joint chrominance coding mode and the case where transform skip is not applied. Therefore, the coding complexity may increase and the coding load may increase.
[0010] The present disclosure is made in view of the foregoing and aims to suppress an increase in the coding load.
[0011] Solutions to the Problems
[0012] An image processing apparatus according to one aspect of the present technology is an image processing apparatus including: an encoding mode setting unit configured to set an encoding mode for encoding an image by setting a transform type having the minimum encoding cost in a non-joint chrominance coding mode as the transform type in the joint chrominance coding mode and obtaining the encoding cost in the joint chrominance coding mode.
[0013] An image processing method according to one aspect of the present technology is an image processing method including: setting an encoding mode for encoding an image by setting a transform type having the minimum encoding cost in a non-joint chrominance coding mode as the transform type in the joint chrominance coding mode and obtaining the encoding cost in the joint chrominance coding mode.
[0014] In the image processing apparatus and the image processing method according to one aspect of the present technology, an encoding mode for encoding an image is set by setting a transform type having the minimum encoding cost in a non-joint chrominance coding mode as the transform type in the joint chrominance coding mode and obtaining the encoding cost in the joint chrominance coding mode. Brief Description of the Drawings
[0015] Figure 1 It is a diagram for describing the setting of a transform skip flag.
[0016] Figure 2 It is a diagram showing an example of setting a transform skip flag to obtain an encoding cost.
[0017] Figure 3 It is a block diagram showing an example of the main configuration of an image encoding apparatus.
[0018] Figure 4 It is a flowchart showing an example of the flow of an image encoding process.
[0019] Figure 5 It is a flowchart showing an example of the flow of an encoding mode setting process.
[0020] Figure 6 It is Figure 5 The following flowchart shows an example of the flow of an encoding mode setting process.
[0021] Figure 7 It is a block diagram showing an example of the main configuration of an image decoding apparatus.
[0022] Figure 8 It is a flowchart showing an example of the flow of an image decoding process.
[0023] Figure 9 It is a block diagram showing an example of the main configuration of a computer. Detailed Implementation Modes
[0024] Hereinafter, modes for implementing the present disclosure (hereinafter referred to as "implementation modes") will be described. Note that the description will be given in the following order.
[0025] 1. Setting of Encoding Mode
[0026] 2. First Implementation Mode (Image Encoding Apparatus)
[0027] 3. Second Implementation Mode (Image Decoding Apparatus)
[0028] 4. Supplementary
[0029] <1. Setting of Encoding Mode>
[0030] <Literature Supporting Technical Contents and Technical Terms, etc.>
[0031] The scope disclosed in the present technology includes not only the content described in the implementation modes but also the content described in the following non-patent literatures, etc. known at the time of filing the application and the content of other literatures referred to in the following non-patent literatures.
[0032] Non-Patent Literature 1: (as described above)
[0033] Non-Patent Literature 2: (as described above)
[0034] Non - Patent Document 3: ITU - T Recommendation H.264 (04 / 2017) "Advanced video coding for generic audiovisual services", April 2017
[0035] Non - Patent Document 4: ITU - T Recommendation H.265 (02 / 18) "High efficiency video coding", February 2018
[0036] That is, the content described in the above non - patent documents can be used as a basis for determining support requirements. For example, even if the quadtree block structure and the quadtree plus binary tree (QTBT) block structure described in the above non - patent documents are not directly described in the examples, these contents fall within the scope of the disclosure of the present technology and meet the support requirements of the claims. In addition, for example, even if technical terms such as parsing, syntax, and semantics are not directly described in the examples, these technical terms similarly fall within the scope of the disclosure of the present technology and meet the support requirements of the claims.
[0037] In addition, in this specification, unless otherwise specified, the "block" (not the block indicating the processing unit) used to describe a partial area or a processing unit of an image (picture) indicates any partial area in the picture, and the size, shape, characteristics, etc. of the block are not restricted. For example, the "block" includes any partial area (processing unit) such as the transform block (TB), transform unit (TU), prediction block (PB), prediction unit (PU), minimum coding unit (SCU), coding unit (CU), largest coding unit (LCU), coding tree block (CTB), coding tree unit (CTU), sub - block, macro - block, tile, or slice described in the above non - patent documents.
[0038] In addition, when specifying the size of such a block, not only the block size can be directly specified, but also the block size can be indirectly specified. For example, the block size can be specified using identification information for identifying the size. In addition, for example, the block size can be specified by the ratio or difference from the size of a reference block (e.g., LCU, SCU, etc.). For example, when the information for specifying the block size is sent as a syntax element or the like, the information for indirectly specifying the size as described above can be used as this information. With this configuration, the amount of information can be reduced, and in some cases, the coding efficiency can be improved. In addition, the specification of the block size also includes the specification of the range of the block size (e.g., the specification of the allowable range of the block size, etc.).
[0039] In addition, in this specification, encoding includes not only the entire process of transforming an image into a bitstream, but also a part of this process. For example, encoding includes not only a process including prediction processing, orthogonal transformation, quantization, arithmetic coding, etc., but also a process collectively referred to as quantization and arithmetic coding, a process including prediction processing, quantization, and arithmetic coding, etc. Similarly, decoding includes not only the entire process of transforming a bitstream into an image, but also a part of this process. For example, decoding includes not only a process including inverse arithmetic decoding, inverse quantization, inverse orthogonal transformation, prediction processing, etc., but also a process including inverse arithmetic decoding and inverse quantization, a process including inverse arithmetic decoding, inverse quantization, and prediction processing, etc.
[0040] <Combined Chrominance Coding Mode and Transform Skip>
[0041] In the Versatile Video Coding (VVC) described in Non-Patent Document 1 or Non-Patent Document 2, a transform skip flag is defined. The transform skip flag is flag information indicating whether to apply transform skip, which is a mode for skipping (omitting) orthogonal transformation. Figure 1 FIG. A shows an example of the syntax of the transform skip flag related to the chrominance component Cb. Figure 1 FIG. B shows an example of the syntax of the transform skip flag related to the chrominance component Cr.
[0042] In addition, in the VVC described in Non-Patent Document 1 or Non-Patent Document 2, a combined chrominance coding mode (combined CbCr mode) is prepared. The combined chrominance coding mode is a mode for encoding the residual samples of both Cb and Cr into a single transform block. In other words, the combined chrominance coding mode is a mode for encoding orthogonal transform coefficients from which the residuals of both Cb and Cr can be obtained. For example, in the combined chrominance coding mode, the coefficients of Cb are encoded. Then, at the time of decoding, the coefficients of Cr are obtained using the decoded coefficients of Cb. By doing so, an improvement in coding efficiency can be expected.
[0043] <Increase in Load in Setting of Coding Mode>
[0044] Incidentally, in the implementation manner of the VVC VTM software described in Non-Patent Document 2, the type of transform applied to chrominance is restricted in the combined chrominance coding mode. An example thereof is shown in Figure 1It is shown in C of . tu_joint_cbcr_residual_flag is flag information indicating whether to apply the joint chrominance coding mode. The case where tu_joint_cbcr_residual_flag = 1 indicates that the joint chrominance coding mode is applied. The case where tu_joint_cbcr_residual_flag = 0 indicates that the joint chrominance coding mode is not applied (also referred to as the non-joint chrominance coding mode).
[0045] As Figure 1 shown in C of , in the case of the non-joint chrominance coding mode, the applicable transform types are discrete cosine transform 2 (DCT2) and transform skip (TS). In contrast, in the case of the joint chrominance coding mode, the applicable transform type is only DCT2. When the application of transform skip is restricted in this way, it is unnecessary to signal the transform skip flag in the joint chrominance coding mode. That is, since signaling the transform skip flag in the joint chrominance coding mode may unnecessarily increase the coding amount and may reduce the coding efficiency.
[0046] In contrast, in the common video coding (VVC) working draft described in Non-Patent Document 1, chrominance transform skip can be applied regardless of the joint chrominance coding mode (joint CbCr mode). An example thereof is shown in Figure 1 D of . As Figure 1 shown in D of , in this case, DCT2 and transform skip can be applied regardless of the joint chrominance coding mode. Therefore, compared with the method described in Non-Patent Document 2, a reduction in coding efficiency due to redundancy of the transform skip flag can be suppressed.
[0047] However, in the VVC described in Non-Patent Document 1 or Non-Patent Document 2, multiple coding modes are prepared, and the coding mode with the lowest coding cost is selected and applied from the coding modes. That is, in the case of the method described in Non-Patent Document 1, it is necessary to evaluate the coding costs of both the case where transform skip is applied and the case where transform skip is not applied for each of the joint chrominance coding mode and the non-joint chrominance coding mode at the time of coding. Therefore, the coding complexity may increase and the coding load may increase.
[0048] <Transfer of transform type setting>
[0049] Therefore, when setting the coding mode, the transform type with the minimum coding cost in the non-Joint CbCr coding mode is set as the transform type in the Joint CbCr coding mode, and the coding cost in the Joint CbCr coding mode is obtained. Here, as described above, the transform type can be DCT2 or transform skip. In this case, it is only necessary to set the value of the chrominance transform skip flag with the minimum coding cost in the non-Joint CbCr coding mode as the chrominance transform skip flag in the Joint CbCr coding mode. Figure 2 shows an example of the syntax. In Figure 2 ,"bestTsFlag[codedCIdx] in non-JointCbCr mode" indicates the chrominance transform skip flag with the minimum coding cost in the non-Joint CbCr coding mode. In addition, "transform_skip_flag[codedCIdx] in JointCbCr mode" indicates the chrominance transform skip flag in the Joint CbCr coding mode.
[0050] For example, in an image processing method, the coding mode of image coding is set by setting the transform type with the minimum coding cost in the non-Joint CbCr coding mode as the transform type in the Joint CbCr coding mode and obtaining the coding cost in the Joint CbCr coding mode.
[0051] For example, an image processing apparatus includes a coding mode setting unit that sets the coding mode of image coding by setting the transform type in the chrominance coding mode and obtaining the coding cost in the Joint CbCr coding mode.
[0052] By doing so, in the Joint CbCr mode, the transform type of the Joint CbCr coding mode can be set without searching both the DCT2 and transform skip modes. Therefore, compared with the case of obtaining the coding costs for both the case where transform skip is applied and the case where transform skip is not applied for each of the Joint CbCr coding mode and the non-Joint CbCr coding mode, an increase in coding complexity can be suppressed and an increase in coding load can be suppressed. Therefore, for example, the transform type can be set at high speed. In addition, an increase in the cost of the encoder can be suppressed.
[0053] In addition, compared with the case where the application of transform skip is restricted in the Joint CbCr coding mode as described in Non-Patent Document 2, a decrease in coding efficiency can be suppressed.
[0054] <2. First Embodiment>
[0055] <Image Coding Apparatus>
[0056] Figure 3 is a block diagram showing an example of the configuration of an image coding apparatus as a mode of an image processing apparatus to which this technology is applied.Figure 3 The image encoding device 300 shown in [Figure] is a device that encodes the image data of a moving image. For example, the image encoding device 300 can encode the image data of a moving image by an encoding method described in any of the non-patent documents.
[0057] Note that Figure 3 the main processing units (blocks), data flows, etc. are shown, and Figure 3 these shown in [Figure] are not necessarily all. That is, in the image encoding device 300, there may be processing units that are not shown as blocks in [Figure], or there may be processing or data flows that are not shown as arrows, etc. in [Figure]. Figure 3 in [Figure], Figure 3 in [Figure].
[0058] As Figure 3 shown, the image encoding device 300 includes a control unit 301, a rearrangement buffer 311, a calculation unit 312, an orthogonal transformation unit 313, a quantization unit 314, an encoding unit 315, an accumulation buffer 316, an inverse quantization unit 317, an inverse orthogonal transformation unit 318, a calculation unit 319, an in-loop filter unit 320, a frame memory 321, a prediction unit 322, and a rate control unit 323.
[0059] <Control Unit>
[0060] The control unit 301 divides the moving image data stored in the rearrangement buffer 311 into blocks (CU, PU, transform blocks, etc.) by processing unit based on an external processing unit or a pre-specified block size. In addition, the control unit 301 determines encoding parameters (header information Hinfo, prediction mode information Pinfo, transform information Tinfo, filter information Finfo, etc.) to be provided to each block based on, for example, rate-distortion optimization (RDO).
[0061] The details of these encoding parameters will be described below. After determining the above encoding parameters, the control unit 301 provides the encoding parameters to each block. For example, the header information Hinfo is provided to each block. The prediction mode information Pinfo is provided to the encoding unit 315 and the prediction unit 322. The transform information Tinfo is provided to the encoding unit 315, the orthogonal transformation unit 313, the quantization unit 314, the inverse quantization unit 317, and the inverse orthogonal transformation unit 318. The filter information Finfo is provided to the in-loop filter unit 320.
[0062] <Rearrangement Buffer>
[0063] Each field of the moving image data (input image) is input to the image encoding device 300 in the reproduction order (display order). The rearrangement buffer 311 acquires and stores (saves) each input image in its reproduction order (display order). Based on the control of the control unit 301, the rearrangement buffer 311 rearranges the input images in the encoding order (decoding order), or divides the input images into blocks by processing unit. The rearrangement buffer 311 supplies the processed input images to the calculation unit 312. In addition, the rearrangement buffer 311 also supplies the input images (original images) to the prediction unit 322 and the in-loop filter unit 320.
[0064] <Calculation unit>
[0065] The calculation unit 312 receives, as inputs, the image I corresponding to the blocks by processing unit and the prediction image P supplied from the prediction unit 322. As shown in the following expression, the calculation unit 312 subtracts the prediction image P from the image I to obtain the prediction residual D, and supplies the prediction residual D to the orthogonal transformation unit 313.
[0066] D = I - P
[0067] <Orthogonal transformation unit>
[0068] The orthogonal transformation unit 313 performs processing for coefficient transformation. For example, the orthogonal transformation unit 313 acquires the prediction residual D supplied from the calculation unit 312. In addition, the orthogonal transformation unit 313 acquires the transformation information Tinfo supplied from the control unit 301.
[0069] The orthogonal transformation unit 313 performs an orthogonal transformation on the prediction residual D based on the transformation information Tinfo to obtain the transformation coefficients Coeff. For example, the orthogonal transformation unit 313 performs a primary transformation on the prediction residual D to generate primary transformation coefficients. Then, the orthogonal transformation unit 313 performs a secondary transformation on the primary transformation coefficients to generate secondary transformation coefficients. The orthogonal transformation unit 313 supplies the obtained secondary transformation coefficients as the transformation coefficients Coeff to the quantization unit 314.
[0070] Note that the orthogonal transformation is an example of coefficient transformation and is not limited to this example. That is, the orthogonal transformation unit 313 may perform any coefficient transformation on the prediction residual D. In addition, the orthogonal transformation unit 313 may perform any coefficient transformation as the primary transformation and the secondary transformation.
[0071] <Quantization unit>
[0072] Quantization unit 314 performs processing regarding quantization. For example, quantization unit 314 obtains the transform coefficient Coeff provided from orthogonal transform unit 313. In addition, quantization unit 314 obtains the transform information Tinfo provided from control unit 301. Further, quantization unit 314 scales (quantizes) the transform coefficient Coeff based on the transform information Tinfo. Note that this quantization method is arbitrary. In addition, the rate of this quantization is controlled by rate control unit 323. Quantization unit 314 provides the quantized transform coefficient (i.e., quantized transform coefficient level) obtained through quantization to encoding unit 315 and inverse quantization unit 317.
[0073] <Encoding unit>
[0074] Encoding unit 315 performs processing regarding encoding. For example, encoding unit 315 obtains the quantized transform coefficient level provided from quantization unit 314. In addition, encoding unit 315 obtains various encoding parameters (header information Hinfo, prediction mode information Pinfo, transform information Tinfo, filter information Finfo, etc.) provided from control unit 301. Further, encoding unit 315 obtains information regarding the filter (e.g., filter coefficients) provided from in-loop filter unit 320. In addition, encoding unit 315 obtains information regarding the best prediction mode provided from prediction unit 322.
[0075] Encoding unit 315 performs variable length encoding (e.g., arithmetic coding) on the quantized transform coefficient level to generate a bit string (encoded data). In addition, encoding unit 315 derives residual information Rinfo from the quantized transform coefficient level. Then, encoding unit 315 encodes the derived residual information Rinfo to generate a bit string.
[0076] Encoding unit 315 includes the information regarding the filter provided from in-loop filter unit 320 in the filter information Finfo. In addition, encoding unit 315 includes the information regarding the best prediction mode provided from prediction unit 322 in the prediction mode information Pinfo. Then, encoding unit 315 encodes the above various encoding parameters (header information Hinfo, prediction mode information Pinfo, transform information Tinfo, filter information Finfo, etc.) to generate a bit string.
[0077] Encoding unit 315 multiplexes the bit strings of the various types of information generated as described above to generate encoded data. Encoding unit 315 provides the encoded data to accumulation buffer 316.
[0078] <Accumulation buffer>
[0079] The accumulation buffer 316 temporarily stores the encoded data obtained by the encoding unit 315. The accumulation buffer 316 outputs the stored encoded data as a bitstream or the like to the outside of the image encoding device 300 at a predetermined timing. For example, the encoded data is sent to the decoding side via an arbitrary recording medium, an arbitrary transmission medium, an arbitrary information processing device, or the like. That is, the accumulation buffer 316 is also a transmission unit that transmits the encoded data (bitstream).
[0080] <Inverse quantization unit>
[0081] The inverse quantization unit 317 performs a process of inverse quantization. For example, the inverse quantization unit 317 acquires the quantization transform coefficient level provided from the quantization unit 314. In addition, the inverse quantization unit 317 acquires the transform information Tinfo provided from the control unit 301.
[0082] The inverse quantization unit 317 scales (inverse quantizes) the value of the quantization transform coefficient level based on the transform information Tinfo. Note that the inverse quantization is the inverse process of the quantization performed in the quantization unit 314. The inverse quantization unit 317 provides the transform coefficient Coeff_IQ obtained by the inverse quantization to the inverse orthogonal transform unit 318.
[0083] <Inverse orthogonal transform unit>
[0084] The inverse orthogonal transform unit 318 performs a process of inverse coefficient transformation. For example, the inverse orthogonal transform unit 318 acquires the transform coefficient Coeff_IQ provided from the inverse quantization unit 317. In addition, the inverse orthogonal transform unit 318 acquires the transform information Tinfo provided from the control unit 301.
[0085] The inverse orthogonal transform unit 318 performs an inverse orthogonal transform on the transform coefficient Coeff_IQ based on the transform information Tinfo to obtain a prediction residual D'. Note that the inverse orthogonal transform is the inverse process of the orthogonal transform performed in the orthogonal transform unit 313. For example, the inverse orthogonal transform unit 318 performs an inverse secondary transform on the transform coefficient Coeff_IQ (secondary transform coefficient) to generate a primary transform coefficient. In addition, the inverse orthogonal transform unit 318 performs an inverse primary transform on the primary transform coefficient to generate a prediction residual D'. Note that the inverse secondary transform is the inverse process of the secondary transform performed in the orthogonal transform unit 313. In addition, the inverse primary transform is the inverse process of the primary transform performed in the orthogonal transform unit 313.
[0086] The inverse orthogonal transform unit 318 provides the prediction residual D' obtained by the inverse orthogonal transform to the calculation unit 319. Note that since the inverse orthogonal transform unit 318 is similar to the inverse orthogonal transform unit on the decoding side (to be described below), the description given for the decoding side (to be described below) can be applied to the inverse orthogonal transform unit 318.
[0087] <Computing unit>
[0088] The computing unit 319 uses the prediction residual D' provided by the inverse orthogonal transform unit 318 and the prediction image P provided by the prediction unit 322 as inputs. The computing unit 319 adds the prediction residual D' and the prediction image P corresponding to the prediction residual D' to obtain the local decoded image Rlocal. The computing unit 319 provides the obtained local decoded image Rlocal to the in-loop filter unit 320 and the frame memory 321.
[0089] <In-loop filter unit>
[0090] The in-loop filter unit 320 performs processing regarding in-loop filtering. For example, the in-loop filter unit 320 uses the local decoded image Rlocal provided by the computing unit 319, the filter information Finfo provided by the control unit 301, and the input image (original image) provided by the rearrangement buffer 311 as inputs. Note that the information input to the in-loop filter unit 320 is arbitrary, and information other than the above information can be input. For example, as needed, information such as the prediction mode, motion information, coding amount target value, quantization parameter QP, picture type, and blocks (CU, CTU, etc.) can be input to the in-loop filter unit 320.
[0091] The in-loop filter unit 320 appropriately performs a filtering process on the local decoded image Rlocal based on the filter information Finfo. The in-loop filter unit 320 also uses the input image (original image) and other input information for the filtering process as needed.
[0092] For example, the in-loop filter unit 320 can apply a bilateral filter as the filtering process. In addition, the in-loop filter unit 320 can apply a deblocking filter (DBF) as the filtering process. In addition, the in-loop filter unit 320 can apply an adaptive offset filter (sample adaptive offset (SAO)) as the filtering process. In addition, the in-loop filter unit 320 can apply an adaptive loop filter (ALF) as the filtering process. In addition, the in-loop filter unit 320 can apply multiple of the above filters in combination as the filtering process. Note that which filter to apply and in what order to apply the filter are arbitrary and can be appropriately selected. For example, the in-loop filter unit 320 sequentially applies four in-loop filters, namely a bilateral filter, a deblocking filter, an adaptive offset filter, and an adaptive loop filter, as the filtering process.
[0093] Of course, the filtering process performed by the in-loop filter unit 320 is arbitrary and is not limited to the above example. For example, the in-loop filter unit 320 may apply a Wiener filter or the like.
[0094] The in-loop filter unit 320 supplies the filtered local decoded image Rlocal to the frame memory 321. Note that when information about the filter (e.g., filter coefficients) is sent to the decoding side, the in-loop filter unit 320 supplies information about the filter to the encoding unit 315.
[0095] <Frame Memory>
[0096] The frame memory 321 performs a process of storing data related to the image. For example, the frame memory 321 uses the local decoded image Rlocal provided from the calculation unit 319 and the filtered local decoded image Rlocal provided from the in-loop filter unit 320 as inputs and saves (stores) the inputs. Further, the frame memory 321 reconstructs and saves the decoded image R for each picture element using the local decoded image Rlocal (stores the decoded image R in a buffer in the frame memory 321). The frame memory 321 supplies the decoded image R (or a part thereof) to the prediction unit 322 in response to a request from the prediction unit 322.
[0097] <Prediction Unit>
[0098] The prediction unit 322 performs a process of generating a prediction image. For example, the prediction unit 322 acquires prediction mode information Pinfo provided from the control unit 301. Further, the prediction unit 322 acquires the input image (original image) provided from the rearrangement buffer 311. Further, the prediction unit 322 acquires the decoded image R (or a part thereof) read from the frame memory 321.
[0099] The prediction unit 322 performs prediction processing such as inter prediction or intra prediction using the prediction mode information Pinfo and the input image (original image). That is, the prediction unit 322 generates a prediction image P by performing prediction and motion compensation with reference to the decoded image R as a reference image.
[0100] The prediction unit 322 supplies the generated prediction image P to the calculation unit 312 and the calculation unit 319. Further, the prediction unit 322 supplies the prediction mode (i.e., information about the best prediction mode) selected by the above processing to the encoding unit 315 as needed.
[0101] <Rate Control Unit>
[0102] The rate control unit 323 performs processing related to rate control. For example, the rate control unit 323 controls the rate of the quantization operation of the quantization unit 314 based on the amount of code of the encoded data accumulated in the accumulation buffer 316 so that overflow or underflow does not occur.
[0103] <Control of Encoding Mode>
[0104] The present technology described in <1. Setting of Encoding Mode> is applied to the image encoding apparatus 300 having the above configuration. That is, as described above, in <Transfer of Transform Type Setting>, it is assumed that when setting the encoding mode, chrominance transform skip can be applied regardless of the joint chrominance encoding mode. Then, the transform type having the minimum encoding cost in the non-joint chrominance encoding mode is set as the transform type in the joint chrominance encoding mode, and the encoding cost in the joint chrominance encoding mode is obtained.
[0105] For example, the control unit 301 serves as an encoding mode setting unit that sets the encoding mode of image encoding. Then, in the setting of the encoding mode, the control unit 301 can apply chrominance transform skip regardless of the joint chrominance encoding mode. In addition, the control unit 301 sets the encoding mode of image encoding by setting the transform type having the minimum encoding cost in the non-joint chrominance encoding mode as the transform type in the joint chrominance encoding mode and obtaining the encoding cost in the joint chrominance encoding mode.
[0106] By doing so, in the joint chrominance mode, the control unit 301 can set the transform type of the joint chrominance encoding mode without searching for both the DCT2 and transform skip modes. Therefore, compared with the case where the application of chrominance transform skip is not restricted in the joint chrominance encoding mode as described in Non-Patent Document 1, the image encoding apparatus 300 can suppress an increase in encoding complexity and can suppress an increase in encoding load. Therefore, for example, the image encoding apparatus 300 can set the transform type at high speed. In addition, an increase in the cost of the image encoding apparatus 300 can be suppressed.
[0107] In addition, compared with the case where the application of transform skip is restricted in the joint chrominance encoding mode as described in Non-Patent Document 2, the image encoding apparatus 300 can suppress a decrease in encoding efficiency.
[0108] Note that, in the example shown in D of Figure 1 , the setting of the joint chrominance encoding mode and the setting of the transform type are performed as the setting of the encoding mode. Similar to this example, the control unit 301 can perform the setting of the joint chrominance encoding mode and the setting of the transform type as the setting of the encoding mode.
[0109] For example, the control unit 301 can set whether to apply the joint chrominance difference coding mode (i.e., whether to apply the joint chrominance difference coding mode or whether to apply the non-joint chrominance difference coding mode). In addition, in the case of applying the joint chrominance difference coding mode, the control unit 301 can also set which one of multiple candidate modes (contents of the joint chrominance difference coding) to apply. For example, multiple modes can be provided, such as a mode for applying the same coefficient as that of Cb to Cr, a mode for applying a coefficient with the inverted sign of the coefficient of Cb to Cr, and a mode for applying a value obtained by multiplying the coefficient of Cb by 1 / 2 to Cr, as candidates for the joint chrominance difference coding mode. In addition, the control unit 301 can set what the transform type will be.
[0110] In addition, in Figure 1 the example shown in D of
[0111] the transform skip can also be applied to the joint chrominance difference coding mode. Similar to this example, the control unit 301 can set whether to apply the transform skip as the transform type in the joint chrominance difference coding mode. That is, in the case of applying the joint chrominance difference coding mode, the control unit 301 can set the value of the transform skip flag, which is flag information indicating whether to apply the transform skip. Figure 2 In this case, as
[0112] shown, the control unit 301 can set the value of the transform skip flag (bestTsFlag) having the minimum coding cost in the non-joint chrominance difference coding mode to the transform skip flag (transform_skip_flag) in the joint chrominance difference coding mode. Figure 1 In addition, in the case of the example shown in D of
[0113] when the joint chrominance difference coding mode is applied and the transform skip is not applied (in the case of non-transform skip), DCT2 is applied as the transform type. Similar to this example, the control unit 301 can apply DCT2 as the transform type when the transform skip is not applied in the joint chrominance difference coding mode.
[0114] The orthogonal transformation unit 313 performs an orthogonal transformation on the prediction residual D obtained by the calculation unit 312 based on this information (i.e., according to the set coding mode). The quantization unit 314 quantizes the transform coefficients Coeff obtained by the orthogonal transformation unit 313. In addition, the coding unit 315 encodes the quantized transform coefficient level level obtained by the quantization unit 314 based on this information (i.e., according to the set coding mode) to generate coded data. In addition, the coding unit 315 encodes information (such as a transform skip flag, etc.), and includes the coded information in the coded data of the quantized transform coefficient level level.
[0115] <Configuration example>
[0116] Note that these processing units (such as Figure 3 the processing units of the control unit 301 shown) have an arbitrary configuration. For example, each processing unit can be configured by a logic circuit that implements the above processing. In addition, each processing unit can include, for example, a CPU, a ROM, a RAM, etc., and implements the above processing by executing a program using the above resources. Of course, each processing unit can have both of these configurations, and implements a part of the above processing by a logic circuit, and implements another part of the processing by executing a program. The configurations of the processing units can be independent of each other. For example, some of the processing units can implement a part of the above processing by a logic circuit, some of the processing units can implement the above processing by executing a program, and some of the processing units can implement the above processing by both a logic circuit and the execution of a program.
[0117] <Flow of image coding processing>
[0118] Next, an example of the flow of image coding processing executed by the image coding apparatus 300 having the above configuration will be described with reference to Figure 4 the flowchart of.
[0119] When the image coding processing starts, in step S301, the rearrangement buffer 311 is controlled by the control unit 301, and the frames of the input moving image data are rearranged from the display order to the coding order.
[0120] In step S302, the control unit 301 sets a processing unit (performing block partitioning) for the input image stored in the rearrangement buffer 311.
[0121] In step S303, the control unit 301 determines (sets) the coding parameters of the input image stored in the rearrangement buffer 311.
[0122] In step S304, the prediction unit 322 performs prediction processing and generates a prediction image, etc. in the best prediction mode. For example, in this prediction processing, the prediction unit 322 performs intra prediction to generate a prediction image, etc. in the best intra prediction mode. In addition, the prediction unit 322 performs inter prediction to generate a prediction image, etc. in the best inter prediction mode. In addition, the prediction unit 322 selects the best prediction mode from the above modes based on the cost function value, etc.
[0123] In step S305, the calculation unit 312 calculates the difference between the input image and the prediction image in the best mode selected by the prediction processing in step S304. That is, the calculation unit 312 generates a prediction residual D between the input image and the prediction image. The prediction residual D obtained in this way has a reduced data volume compared to the original image data. Therefore, the data volume can be compressed compared to the case of encoding the image as it is.
[0124] In step S306, the orthogonal transformation unit 313 performs an orthogonal transformation process on the prediction residual D generated by the process in step S305 to obtain transformation coefficients Coeff. For example, the orthogonal transformation unit 313 performs a primary transformation on the prediction residual D to generate primary transformation coefficients. In addition, the orthogonal transformation unit 313 performs a secondary transformation on the primary transformation coefficients to generate secondary transformation coefficients (transformation coefficients Coeff).
[0125] In step S307, the quantization unit 314 quantizes the transformation coefficients Coeff obtained by the process in step S306 by using the quantization parameter calculated by the control unit 301, etc. to obtain a quantized transformation coefficient level level.
[0126] In step S308, the inverse quantization unit 317 performs inverse quantization on the quantized transformation coefficient level level generated by the process in step S307 by using a characteristic corresponding to the quantization in step S307 to obtain transformation coefficients Coeff_IQ.
[0127] In step S309, the inverse orthogonal transformation unit 318 performs an inverse orthogonal transformation on the transformation coefficients Coeff_IQ obtained by the process in step S308 by a method corresponding to the orthogonal transformation process in step S306 to obtain a prediction residual D'. For example, the inverse orthogonal transformation unit 318 performs an inverse secondary transformation on the transformation coefficients Coeff_IQ (secondary transformation coefficients) to generate primary transformation coefficients. In addition, the inverse orthogonal transformation unit 318 performs an inverse primary transformation on the primary transformation coefficients to generate a prediction residual D'.
[0128] Note that the inverse orthogonal transform process is similar to the inverse orthogonal transform process performed on the decoding side. Therefore, the description of the decoding side to be described below can be applied to the inverse orthogonal transform process in step S309.
[0129] In step S310, the calculation unit 319 adds the predicted image obtained through the prediction process in step S304 to the prediction residual D' obtained through the process in step S309 to generate a local decoded image.
[0130] In step S311, the in-loop filter unit 320 performs an in-loop filtering process on the local decoded image obtained through the process in step S310.
[0131] In step S312, the frame memory 321 stores the local decoded image obtained through the process in step S310 and the locally decoded image filtered in step S311.
[0132] In step S313, the encoding unit 315 encodes the quantized transform coefficient level obtained through the process in step S307. For example, the encoding unit 315 encodes the quantized transform coefficient level, which is information about the image, through arithmetic coding or the like to generate encoded data. In addition, at this time, the encoding unit 315 encodes various encoding parameters (header information Hinfo, prediction mode information Pinfo, and transform information Tinfo). In addition, the encoding unit 315 derives residual information RInfo from the quantized transform coefficient level and encodes the residual information RInfo.
[0133] In step S314, the accumulation buffer 316 accumulates the encoded data thus obtained and outputs the encoded data, for example, as a bitstream to the outside of the image encoding device 300. For example, the bitstream is sent to the decoding side via a transmission path or a recording medium. In addition, the rate control unit 323 performs rate control as needed.
[0134] When the process in step S314 ends, the image encoding process ends.
[0135] <Control of Encoding Mode>
[0136] The present technology described in <1. Setting of Encoding Mode> is applied to the image encoding process of the above process. That is, as described above, in <Transfer of Transform Type Setting>, it is assumed that when setting the encoding mode, chrominance transform skip can be applied regardless of the joint chrominance coding mode. Then, the transform type with the minimum encoding cost in the non-joint chrominance coding mode is set as the transform type in the joint chrominance coding mode, and the encoding cost in the joint chrominance coding mode is obtained.
[0137] For example, in step S303, the control unit 301 performs encoding mode setting processing and sets an encoding mode for encoding an image. In setting the encoding mode, chrominance transform skip can be applied regardless of the joint chrominance encoding mode. Further, the control unit 301 sets the encoding mode by setting the transform type having the minimum encoding cost in the non-joint chrominance encoding mode as the transform type in the joint chrominance encoding mode and obtaining the encoding cost in the joint chrominance encoding mode.
[0138] In step S306, the orthogonal transform unit 313 performs orthogonal transform on the prediction residual D according to the set encoding mode. Further, in step S313, the encoding unit 315 encodes the quantization transform coefficient level according to the set encoding mode to generate encoded data. Further, the encoding unit 315 encodes information related to the encoding mode (e.g., a transform skip flag, etc.) and includes the encoded information in the encoded data of the quantization transform coefficient level.
[0139] By doing so, in the joint chrominance mode, the control unit 301 can set the transform type of the joint chrominance encoding mode without searching for both the DCT2 and transform skip modes. Therefore, compared with the case where the application of chrominance transform skip is not restricted in the joint chrominance encoding mode as described in Non-Patent Document 1, the image encoding apparatus 300 can suppress an increase in encoding complexity and can suppress an increase in encoding load. Therefore, for example, the image encoding apparatus 300 can set the transform type at high speed. Further, an increase in the cost of the image encoding apparatus 300 can be suppressed.
[0140] Further, compared with the case where the application of transform skip is restricted in the joint chrominance encoding mode as described in Non-Patent Document 2, the image encoding apparatus 300 can suppress a decrease in encoding efficiency.
[0141] <Flow of encoding mode setting processing>
[0142] Reference will be made to Figure 5 and Figure 6 the flowchart of Figure 4 to describe an example of the flow of the encoding mode setting processing performed in step S303 of
[0143] When the encoding mode setting processing starts, the control unit 301 sets the non-joint chrominance encoding mode in step S351. For example, the control unit 301 sets tu_joint_cbcr_residual_flag to false (e.g., "0") and sets TuCResMode[xTbY][yTbY] to "0".
[0144] In step S352, the control unit 301 obtains the encoding cost for each transform type in the non-united chrominance coding mode. For example, in the non-united chrominance coding mode, the control unit 301 obtains the encoding cost for the case where the transform type is DCT2 and the case where the transform type is transform skip (TS). The control unit 301 performs this process for each of the chrominance components Cb and Cr.
[0145] In step S353, the control unit 301 sets the transform type having the minimum encoding cost among the encoding costs obtained in the process of step S352. For example, the control unit 301 sets the value of the transform skip flag corresponding to the transform type having the minimum encoding cost obtained in the process of step S352 to bestFlag[cIdx]. The control unit 301 performs this process for each of the chrominance components Cb and Cr.
[0146] In step S354, the control unit 301 sets the united chrominance coding mode based on the chrominance cbf (coding block flag) in the non-united chrominance coding mode. The chrominance cbf is flag information indicating whether to encode the transform coefficients of a block. In other words, the chrominance cbf is flag information indicating whether a block includes non-zero transform coefficients.
[0147] For example, the control unit 301 sets tu_joint_cbcr_residual_flag to true (e.g., "1"). Then, the control unit 301 sets TuCResMode[xTbY][yTbY] based on tu_cbf_cb (which is the cbf of the TU to be processed for the chrominance component Cb) and tu_cbf_cr (which is the cbf of the TU to be processed for the chrominance component Cr).
[0148] For example, when tu_cbf_cb == 1 and tu_cbf_cr == 0, the control unit 301 sets TuCResMode[xTbY][yTbY] to "1". In addition, when tu_cbf_cb == 1 and tu_cbf_cr == 1, the control unit 301 sets TuCResMode[xTbY][yTbY] to "2". In addition, when tu_cbf_cb == 0 and tu_cbf_cr == 1, the control unit 301 sets TuCResMode[xTbY][yTbY] to "3".
[0149] When the process of step S354 ends, the process proceeds to Figure 6 . In Figure 6In step S361, the control unit 301 sets the coded component identifier codedCIdx based on the combined chrominance difference coding mode set in step S354. For example, the control unit 301 sets codedCIdx to "1" (i.e., Cb) when TuCResMode[xTbY][yTbY] is "1" or "2", and sets codedCIdx to "2" (i.e., Cr) in other cases.
[0150] In step S362, the control unit 301 sets the bestTsFlag[cIdx] set in step S353 to the transform skip flag tsFlag[codedCIdx] in the combined chrominance difference coding mode (tsFlag[codedCIdx] = bestTsFlag[cIdx]).
[0151] In step S363, the control unit 301 obtains the coding cost of the combined chrominance difference coding mode. As described above, in step S362, the value of the transform skip flag corresponding to the transform type with the minimum coding cost in the non-combined chrominance difference coding mode is set to the transform skip flag in the combined chrominance difference coding mode. Therefore, for the combined chrominance difference coding mode, the control unit 301 only needs to obtain the coding cost of the mode corresponding to the value of the transform skip flag. That is, in this case, the control unit 301 does not need to obtain the coding costs for both the case where transform skip is applied and the case where transform skip is not applied. Therefore, the control unit 301 can more easily obtain the coding cost of the combined chrominance mode.
[0152] In step S364, the control unit 301 compares the minimum coding cost of the non-combined chrominance difference coding mode with the coding cost of the combined chrominance difference coding mode, and selects the mode with the minimum coding cost.
[0153] When the processing of step S364 ends, the coding mode setting process ends, and the process returns to Figure 4 .
[0154] By doing so, compared with the case where the application of chrominance difference transform skip is not restricted in the combined chrominance difference coding mode as described in Non-Patent Document 1, the image coding device 300 can suppress an increase in coding complexity and can suppress an increase in coding load. Therefore, for example, the image coding device 300 can set the transform type at high speed. In addition, an increase in the cost of the image coding device 300 can be suppressed.
[0155] In addition, compared with the case where the application of transform skip is restricted in the combined chrominance difference coding mode as described in Non-Patent Document 2, the image coding device 300 can suppress a decrease in coding efficiency.
[0156] <3. Second Embodiment>
[0157] <Image Decoding Device>
[0158] Figure 7 It is a block diagram showing an example of the configuration of an image decoding device as a mode of an image processing device to which this technology is applied. Figure 7 The image decoding device 400 shown in the figure is a device that decodes encoded data of a moving image. For example, the image decoding device 400 can decode the encoded data by a decoding method described in any of the above non-patent documents. For example, the image decoding device 400 decodes the encoded data (bitstream) generated by the above image encoding device 300.
[0159] Note that Figure 7 it shows the main processing units (blocks), data flows, etc., and Figure 7 these shown in the figure are not necessarily all. That is, in the image decoding device 400, there may be processing units that are not shown as blocks in the figure, or there may be processing or data flows that are not shown as arrows, etc. in the figure. Figure 7 in the figure Figure 7 in the figure
[0160] In Figure 7 the figure, the image decoding device 400 includes an accumulation buffer 411, a decoding unit 412, an inverse quantization unit 413, an inverse orthogonal transformation unit 414, a calculation unit 415, an in-loop filter unit 416, a rearrangement buffer 417, a frame memory 418, and a prediction unit 419. Note that the prediction unit 419 includes an intra-frame prediction unit and an inter-frame prediction unit (not shown). The image decoding device 400 is a device that generates moving image data by decoding encoded data (bitstream).
[0161] <Accumulation Buffer>
[0162] The accumulation buffer 411 acquires the bitstream input to the image decoding device 400 and stores (saves) the bitstream. For example, the accumulation buffer 411 provides the accumulated bitstream to the decoding unit 412 at a predetermined timing or when a predetermined condition is satisfied.
[0163] <Decoding Unit>
[0164] The decoding unit 412 performs processing for image decoding. For example, the decoding unit 412 acquires the bitstream provided from the accumulation buffer 411. For example, the decoding unit 412 performs variable length decoding on the syntax value of each syntax element from the bit string according to the definition of the syntax table to obtain parameters.
[0165] Parameters derived from syntax elements and the syntax values of syntax elements include, for example, information such as header information Hinfo, prediction mode information Pinfo, transform information Tinfo, residual information Rinfo, and filter information Finfo. That is, the decoding unit 412 parses (analyzes and obtains) such information from the bitstream. This information will be described below.
[0166] <Header information Hinfo>
[0167] The header information Hinfo includes, for example, header information such as a video parameter set (VPS), a sequence parameter set (SPS), a picture parameter set (PPS), and a slice header (SH). For example, the header information Hinfo includes information defining the following: image size (width PicWidth and height PicHeight), bit depth (luminance bitDepthY and chrominance bitDepthC), chroma array type ChromaArrayType, maximum value MaxCUSize and minimum value MinCUSize of the CU size, maximum depth MaxQTDepth and minimum depth MinQTDepth of the quadtree partition, maximum depth MaxBTDepth and minimum depth MinBTDepth of the binary tree partition, maximum value MaxTSSize of the transform skip block (also referred to as the maximum transform skip block size), on / off flag (also referred to as the enable flag) of each coding tool, etc.
[0168] For example, examples of the on / off flags of the coding tools included in the header information Hinfo include on / off flags related to the following transform processing and quantization processing. Note that the on / off flag of the coding tool can also be interpreted as a flag indicating whether there is syntax related to the coding tool in the coded data. The case where the value of the on / off flag is 1 (true) indicates that the coding tool can be used. The case where the value of the on / off flag is 0 (false) indicates that the coding tool cannot be used. Note that the interpretation of the flag value can be reversed.
[0169] For example, the header information Hinfo may include an inter-component prediction enable flag (ccp_enabled_flag). The inter-component prediction enable flag is flag information indicating whether inter-component prediction (cross-component prediction (CCP), also referred to as CC prediction) is available. For example, in the case where the flag information is "1" (true), the flag information indicates that inter-component prediction is available. In the case where the flag information is "0" (false), the flag information indicates that inter-component prediction is not available.
[0170] Note that this CCP is also referred to as inter-component linear prediction (CCLM or CCLMP).
[0171] <Prediction mode information Pinfo>
[0172] The prediction mode information Pinfo includes, for example, information such as the size information PBSize (prediction block size) of a prediction block (PB) to be processed, intra prediction mode information IPinfo, and motion prediction information MVinfo.
[0173] The intra prediction mode information IPinfo includes, for example, prev_intra_luma_pred_flag, mpm_idx, and rem_intra_pred_mode in JCTVC-W1005, 7.3.8.5 coding unit syntax, the luma intra prediction mode IntraPredModeY derived from the syntax, and the like.
[0174] In addition, the intra prediction mode information IPinfo may include, for example, an inter-component prediction flag (ccp_flag (cclmp_flag)). The inter-component prediction flag (ccp_flag (cclm_flag)) is flag information indicating whether inter-component linear prediction is applied. For example, ccp_flag == 1 indicates that inter-component prediction is applied, and ccp_flag == 0 indicates that inter-component prediction is not applied.
[0175] In addition, the intra prediction mode information IPinfo may include a multi-class linear prediction mode flag (mclm_flag). The multi-class linear prediction mode flag (mclm_flag) is information about the linear prediction mode (linear prediction mode information). More specifically, the multi-class linear prediction mode flag (mclm_flag) is flag information indicating whether a multi-class linear prediction mode is set. For example, "0" indicates a single-class mode (e.g., CCLMP), and "1" indicates a multi-class mode (e.g., MCLMP).
[0176] In addition, the intra prediction mode information IPinfo may include a chroma sample position type identifier (chroma_sample_loc_type_idx). The chroma sample position type identifier (chroma_sample_loc_type_idx) is an identifier for identifying the type of pixel position of a chroma component (also referred to as the chroma sample position type). For example, when the chroma array type (ChromaArrayType), which is information about the color format, indicates a 420 format, the chroma sample position type identifier is assigned as in the following expressions.
[0177] chroma_sample_loc_type_idx == 0: type 2
[0178] chroma_sample_loc_type_idx == 1: type 3
[0179] chroma_sample_loc_type_idx == 2: Type 0
[0180] chroma_sample_loc_type_idx == 3: Type 1
[0181] Note that the chroma sample location type identifier (chroma_sample_loc_type_idx) is sent as (stored in) the information (chroma_sample_loc_info()) about the pixel position of the chroma component, i.e., stored in the information about the pixel position of the chroma component.
[0182] In addition, the intra prediction mode information IPinfo may include a chroma MPM identifier (chroma_mpm_idx). The chroma MPM identifier (chroma_mpm_idx) is an identifier indicating which prediction mode candidate in the chroma intra prediction mode candidate list (intraPredModeCandListC) will be designated as the chroma intra prediction mode.
[0183] In addition, the intra prediction mode information IPinfo may include the luma intra prediction mode (IntraPredModeC) derived from these syntaxes.
[0184] The motion prediction information MVinfo includes information such as, for example, merge_idx, merge_flag, inter_pred_idc, ref_idx_LX, mvp_lX_flag, X = {0, 1}, mvd, etc. (see, for example, JCTVC-W1005, 7.3.8.6 Prediction Unit Syntax).
[0185] Of course, the information included in the prediction mode information Pinfo is arbitrary and may also include information other than the above information.
[0186] <Transformation information Tinfo>
[0187] The transformation information Tinfo may include, for example, the width size TBWSize and the height TBHSize of the transform block to be processed. Note that the logarithm to the base 2 log2TBWSize may be applied instead of the width size TBWSize of the transform block to be processed. In addition, the logarithm to the base 2 log2TBHSize may be applied instead of the height TBHSSize of the transform block to be processed.
[0188] In addition, the transform information Tinfo may include a transform skip flag (transform_skip_flag (or ts_flag)). The transform skip flag is a flag indicating whether to skip the coefficient transform (or inverse coefficient transform). Note that the transform skip flag (transform_skip_flag[0], transform_skip_flag[1], and transform_skip_flag[2]) may be signaled for each of the Y, Cb, and Cr components.
[0189] In addition, the transform information Tinfo may include parameters such as a scan identifier (scanIdx), a quantization parameter (qp), and a quantization matrix (scaling_matrix (e.g., JCTVC-W1005, 7.3.4 Scaling List Data Syntax)).
[0190] Of course, the information included in the transform information Tinfo is arbitrary and may include information other than the above information.
[0191] <Residual information Rinfo>
[0192] The residual information Rinfo (e.g., see 7.3.8.11 Residual Coding Syntax of JCTVC-W1005) may include, for example, a residual data presence / absence flag (cbf (coded_block_flag)). In addition, the residual information Rinfo may include the last non-zero coefficient X coordinate (last_sig_coeff_x_pos) and the last non-zero coefficient Y coordinate (last_sig_coeff_y_pos). In addition, the residual information Rinfo may include a sub-block non-zero coefficient presence / absence flag (coded_sub_block_flag) and a non-zero coefficient presence / absence flag (sig_coeff_flag).
[0193] In addition, the residual information Rinfo may include a GR1 flag (gr1_flag) and a GR2 flag (gr2_flag). The GR1 flag is a flag indicating whether the level of the non-zero coefficient is greater than 1, and the GR2 flag is a flag indicating whether the level of the non-zero coefficient is greater than 2. In addition, the residual information Rinfo may include a sign flag (sign_flag), which is a sign indicating the positive or negative of the non-zero coefficient. In addition, the residual information Rinfo may include a non-zero coefficient residual level (coeff_abs_level_remaining) that is the residual level of the non-zero coefficient.
[0194] Of course, the information included in the residual information Rinfo is arbitrary and may include information other than the above information.
[0195] <Filter information Finfo>
[0196] The filter information Finfo includes control information regarding filter processing. For example, the filter information Finfo may include control information regarding the deblocking filter (DBF). In addition, the filter information Finfo may include control information regarding sample adaptive offset (SAO). In addition, the filter information Finfo may include control information regarding the adaptive loop filter (ALF). In addition, the filter information Finfo may include control information regarding other linear and non-linear filters.
[0197] For example, the filter information Finfo may include information about the pictures to which each filter is applied and the regions in the specified pictures. In addition, the filter information Finfo may include filter enable control information or disable control information in units of CUs. In addition, the filter information Finfo may include filter enable control information or disable control information regarding the boundaries of slices or tiles.
[0198] Of course, the information included in the filter information Finfo is arbitrary and may include information other than the above information.
[0199] Returning to the description of the decoding unit 412. The decoding unit 412 refers to the residual information Rinfo and obtains the quantized transform coefficient level level at each coefficient position in each transform block. The decoding unit 412 provides the quantized transform coefficient level level to the inverse quantization unit 413.
[0200] In addition, the decoding unit 412 provides the parsed header information Hinfo, prediction mode information Pinfo, quantized transform coefficient level level, transform information Tinfo, and filter information Finfo to each block. As described in detail below.
[0201] The header information Hinfo is provided to the inverse quantization unit 413, inverse orthogonal transform unit 414, prediction unit 419, and in-loop filter unit 416. The prediction mode information Pinfo is provided to the inverse quantization unit 413 and the prediction unit 419. The transform information Tinfo is provided to the inverse quantization unit 413 and the inverse orthogonal transform unit 414. The filter information Finfo is provided to the in-loop filter unit 416.
[0202] Of course, the above examples are examples, and this embodiment is not limited to this example. For example, each coding parameter may be provided to any processing unit. In addition, other information may be provided to any processing unit.
[0203] <Inverse quantization unit>
[0204] The inverse quantization unit 413 performs processing related to inverse quantization. For example, the inverse quantization unit 413 obtains the transform information Tinfo and the quantized transform coefficient level level provided from the decoding unit 412. In addition, the inverse quantization unit 413 scales (inverse quantizes) the value of the quantized transform coefficient level level based on the transform information Tinfo to obtain the transform coefficient Coeff_IQ after inverse quantization.
[0205] Note that this inverse quantization is performed as the inverse process of the quantization performed by the quantization unit 314 of the image encoding device 300. In addition, the inverse quantization is a process similar to the inverse quantization performed by the inverse quantization unit 317 of the image encoding device 300. In other words, the inverse quantization unit 317 performs a process (inverse quantization) similar to that of the inverse quantization unit 413.
[0206] The inverse quantization unit 413 provides the obtained transform coefficient Coeff_IQ to the inverse orthogonal transform unit 414.
[0207] <Inverse orthogonal transform unit>
[0208] The inverse orthogonal transform unit 414 performs processing related to inverse orthogonal transform. For example, the inverse orthogonal transform unit 414 obtains the transform coefficient Coeff_IQ provided from the inverse quantization unit 413. In addition, the inverse orthogonal transform unit 414 obtains the transform information Tinfo provided from the decoding unit 412.
[0209] The inverse orthogonal transform unit 414 performs an inverse orthogonal transform process on the transform coefficient Coeff_IQ based on the transform information Tinfo to obtain the prediction residual D'. For example, the inverse orthogonal transform unit 414 performs an inverse secondary transform on the transform coefficient Coeff_IQ to generate a primary transform coefficient. In addition, the inverse orthogonal transform unit 414 performs an inverse primary transform on the primary transform coefficient to generate the prediction residual D'.
[0210] Note that this inverse orthogonal transform is performed as the inverse process of the orthogonal transform performed by the orthogonal transform unit 313 of the image encoding device 300. In addition, the inverse orthogonal transform is a process similar to the inverse orthogonal transform performed by the inverse orthogonal transform unit 318 of the image encoding device 300. That is, the inverse orthogonal transform unit 318 performs a process (inverse orthogonal transform) similar to that of the inverse orthogonal transform unit 414.
[0211] The inverse orthogonal transform unit 414 provides the obtained prediction residual D' to the calculation unit 415.
[0212] <Calculation unit>
[0213] The calculation unit 415 performs processing regarding adding information about an image. For example, the calculation unit 415 obtains the prediction residual D' provided from the inverse orthogonal transform unit 414. In addition, the calculation unit 415 obtains the predicted image P provided from the prediction unit 419. The calculation unit 415 adds the prediction residual D' and the predicted image P (prediction signal) corresponding to the prediction residual D' to obtain a locally decoded image Rlocal, as shown in the following expression.
[0214] Rlocal = D' + P
[0215] The calculation unit 415 provides the obtained locally decoded image Rlocal to the in-loop filter unit 416 and the frame memory 418.
[0216] <In-loop filter unit>
[0217] The in-loop filter unit 416 performs processing regarding in-loop filtering processing. For example, the in-loop filter unit 416 obtains the locally decoded image Rlocal provided from the calculation unit 415. In addition, the in-loop filter unit 416 obtains the filter information Finfo provided from the decoding unit 412. Note that the information input to the in-loop filter unit 416 is arbitrary, and information other than the above information can be input.
[0218] The in-loop filter unit 416 appropriately performs filtering processing on the locally decoded image Rlocal based on the filter information Finfo. For example, the in-loop filter unit 416 can apply a bilateral filter as the filtering processing. In addition, the in-loop filter unit 416 can apply a deblocking filter as the filtering processing. In addition, the in-loop filter unit 416 can apply an adaptive offset filter as the filtering processing. In addition, the in-loop filter unit 416 can apply an adaptive loop filter as the filtering processing. In addition, the in-loop filter unit 416 can combinatorially apply multiple of the above filters as the filtering processing. Note that which filter to apply and in what order to apply the filters is arbitrary and can be appropriately selected. For example, the in-loop filter unit 416 sequentially applies four in-loop filters, namely a bilateral filter, a deblocking filter, an adaptive offset filter, and an adaptive loop filter, as the filtering processing.
[0219] The in-loop filter unit 416 performs filtering processing corresponding to the filtering processing performed on the encoding side (e.g., by the in-loop filter unit 320 of the image encoding device 300). Of course, the filtering processing performed by the in-loop filter unit 416 is arbitrary and is not limited to the above example. For example, the in-loop filter unit 416 can apply a Wiener filter or the like.
[0220] The in-loop filter unit 416 supplies the filtered local decoded image Rlocal to the rearrangement buffer 417 and the frame memory 418.
[0221] <Rearrangement buffer>
[0222] The rearrangement buffer 417 receives the local decoded image Rlocal supplied from the in-loop filter unit 416 as an input, and stores (saves) the local decoded image Rlocal. The rearrangement buffer 417 uses the local decoded image Rlocal to reconstruct the decoded image R for each picture unit, and stores (saves) the decoded image R (in the buffer). The rearrangement buffer 417 rearranges the obtained decoded image R from the decoding order to the reproduction order. The rearrangement buffer 417 outputs the rearranged decoded image R group as moving image data to the outside of the image decoding device 400.
[0223] <Frame memory>
[0224] The frame memory 418 performs processing for storing data related to an image. For example, the frame memory 418 obtains the local decoded image Rlocal supplied from the calculation unit 415. Then, the frame memory 418 uses the local decoded image Rlocal to reconstruct the decoded image R for each picture unit. The frame memory 418 stores the reconstructed decoded image R in the buffer of the frame memory 418.
[0225] In addition, the frame memory 418 obtains the in-loop filtered local decoded image Rlocal supplied from the in-loop filter unit 416. Then, the frame memory 418 uses the in-loop filtered local decoded image Rlocal to reconstruct the decoded image R for each picture unit. The frame memory 418 stores the reconstructed decoded image R in the buffer of the frame memory 418.
[0226] In addition, the frame memory 418 appropriately supplies the stored decoded image R (or a part thereof) to the prediction unit 419 as a reference image.
[0227] Note that the frame memory 418 can store header information Hinfo, prediction mode information Pinfo, transform information Tinfo, filter information Finfo, etc. related to the generation of the decoded image.
[0228] <Prediction unit>
[0229] The prediction unit 419 performs processing related to the generation of a prediction image. For example, the prediction unit 419 acquires prediction mode information Pinfo provided from the decoding unit 412. In addition, the prediction unit 419 performs prediction processing by the prediction method specified by the prediction mode information Pinfo to obtain a prediction image P. When obtaining the prediction image P, the prediction unit 419 uses the decoded image R (or a part thereof) stored in the frame memory 418 as a reference image, and the decoded image R is specified by the prediction mode information Pinfo. The decoded image R can be an image before or after filtering. The prediction unit 419 provides the obtained prediction image P to the calculation unit 415.
[0230] <Configuration example>
[0231] Note that these processing units (accumulation buffer 411 to prediction unit 419) have arbitrary configurations. For example, each processing unit can be configured by a logic circuit that implements the above processing. In addition, each processing unit can include, for example, a CPU, ROM, RAM, etc., and implements the above processing by executing a program using the above resources. Of course, each processing unit can have both of these configurations, and implements a part of the above processing by a logic circuit and implements another part of the processing by executing a program. The configurations of the processing units can be independent of each other. For example, some of the processing units can implement a part of the above processing by a logic circuit, some of the other processing units can implement the above processing by executing a program, and some of the other processing units can implement the above processing by both a logic circuit and the execution of a program.
[0232] <Flow of image decoding processing>
[0233] Next, an example of the flow of image decoding processing performed by the image decoding apparatus 400 having the above configuration will be described with reference to Figure 8 the flowchart.
[0234] When the image decoding processing starts, in step S401, the accumulation buffer 411 acquires and saves (accumulates) encoded data (bitstream) provided from outside the image decoding apparatus 400.
[0235] In step S402, the decoding unit 412 decodes the encoded data (bitstream) to obtain the quantized transform coefficient level level. In addition, the decoding unit 412 parses (analyzes and acquires) various encoding parameters from the encoded data (bitstream) through this decoding.
[0236] In step S403, the inverse quantization unit 413 performs inverse quantization to obtain the transform coefficient Coeff_IQ, and the inverse quantization is the inverse process of the quantization performed on the quantized transform coefficient level level obtained through the processing in step S402 on the encoding side.
[0237] In step S404, the inverse orthogonal transform unit 414 performs an inverse orthogonal transform process to obtain a prediction residual D'. This inverse orthogonal transform process is the inverse process of the orthogonal transform process performed on the transform coefficients Coeff_IQ obtained in step S403 on the encoding side. For example, the inverse orthogonal transform unit 414 performs an inverse sub-transform on the transform coefficients Coeff_IQ (sub-transform coefficients) to generate primary transform coefficients. In addition, the inverse orthogonal transform unit 414 performs an inverse primary transform on the primary transform coefficients to generate the prediction residual D'.
[0238] In step S405, the prediction unit 419 performs a prediction process based on the information parsed in step S402 by the prediction method specified on the encoding side, and generates a prediction image P by referring to, for example, a reference image stored in the frame memory 418.
[0239] In step S406, the calculation unit 415 adds the prediction residual D' obtained in step S404 to the prediction image P obtained in step S405 to obtain a local decoded image Rlocal.
[0240] In step S407, the in-loop filter unit 416 performs an in-loop filtering process on the local decoded image Rlocal obtained through the process in step S406.
[0241] In step S408, the rearrangement buffer 417 uses the "filtered local decoded image Rlocal" obtained through the process in step S407 to obtain a decoded image R, and rearranges the group of decoded images R from the decoding order to the reproduction order. The group of decoded images R rearranged in the reproduction order is output as a moving image to the outside of the image decoding device 400.
[0242] In addition, in step S409, the frame memory 418 stores at least one of the local decoded image Rlocal obtained through the process in step S406 and the local decoded image Rlocal after the filtering process obtained through the process in step S407.
[0243] When the process in step S409 ends, the image decoding process ends.
[0244] <4. Supplementary>
[0245] <Computer>
[0246] The above series of processes can be executed by hardware or by software. In the case of executing a series of processes by software, a program for configuring the software is installed in a computer. Here, the computer includes a computer incorporated into dedicated hardware, a computer capable of executing various functions by installing various programs, such as a general-purpose personal computer, etc.
[0247] Figure 9 It is a block diagram showing a configuration example of the hardware of a computer that executes the above series of processes through a program.
[0248] In Figure 9 In the computer 800 shown, a central processing unit (CPU) 801, a read-only memory (ROM) 802, and a random access memory (RAM) 803 are interconnected via a bus 804.
[0249] An input / output interface 810 is also connected to the bus 804. An input unit 811, an output unit 812, a storage unit 813, a communication unit 814, and a drive 815 are connected to the input / output interface 810.
[0250] The input unit 811 includes, for example, a keyboard, a mouse, a microphone, a touchpad, an input terminal, etc. The output unit 812 includes, for example, a display, a speaker, an output terminal, etc. The storage unit 813 includes, for example, a hard disk, a RAM disk, a non-volatile memory, etc. The communication unit 814 includes, for example, a network interface. The drive 815 drives a removable medium 821 such as a magnetic disk, an optical disk, a magneto-optical disk, or a semiconductor memory.
[0251] In the computer configured as described above, the CPU 801 loads, for example, a program stored in the storage unit 813 into the RAM 803 via the input / output interface 810 and the bus 804, and executes the program so that the above series of processes are executed. In addition, the RAM 803 appropriately stores data and the like required for the CPU 801 to execute various types of processes.
[0252] For example, a program to be executed by the computer can be recorded and applied on a removable medium 821 such as a packaged medium, etc., and can be provided. In this case, by attaching the removable medium 821 to the drive 815, the program can be installed in the storage unit 813 via the input / output interface 810.
[0253] In addition, the program can be provided via a wired or wireless transmission medium such as a local area network, the Internet, or digital satellite broadcasting. In this case, the program can be received by the communication unit 814 and installed in the storage unit 813.
[0254] In addition to the above methods, the program can be pre-installed in the ROM 802 or the storage unit 813.
[0255] <Objects to which the present technology can be applied>
[0256] The present technology can be applied to any image encoding method. That is, the specifications of various types of processing for image encoding such as transformation (inverse transformation), quantization (inverse quantization), encoding, and prediction are arbitrary and are not limited to the above examples, as long as there is no contradiction with the present technology described above. In addition, as long as there is no contradiction with the present technology described above, some processing can be omitted.
[0257] In addition, the present technology can be applied to a multi-viewpoint image encoding system (or multi-viewpoint image decoding system) that performs encoding or decoding of a multi-viewpoint image including multiple viewpoints (views). In this case, the present technology only needs to be simply applied to the encoding and decoding of each viewpoint (view).
[0258] In addition, the present technology can be applied to a hierarchical image encoding (scalable encoding) system (or hierarchical image decoding system) that encodes or decodes a hierarchical image of multiple layers (hierarchical) to have a scalable function for a predetermined parameter. In this case, the present technology only needs to be simply applied to the encoding / decoding of each layer.
[0259] In addition, in the above description, the image encoding device 300 and the image decoding device 400 have been described as application examples of the present technology, but the present technology can be applied to any configuration.
[0260] The present technology can be applied to, for example, various electronic devices such as transmitters and receivers in satellite broadcasting (such as television receivers and mobile phones), cable broadcasting such as cable television, distribution on the Internet, distribution to terminals via cellular communication, or devices that record images on media such as optical discs, magnetic disks, and flash memories and reproduce images from these storage media (for example, hard disk recorders and imaging devices).
[0261] In addition, for example, the present technology can be implemented as a configuration of a part of a device, such as a processor (for example, a video processor) such as a system large-scale integration (LSI), a module (for example, a video module) using multiple processors, a unit (for example, a video unit) using multiple modules, or a kit (for example, a video kit) in which other functions are added to the unit (that is, a configuration of a part of a device).
[0262] In addition, for example, the present technology can also be applied to a network system including multiple devices. For example, the present technology can be implemented as cloud computing that is shared and collaboratively processed by multiple devices via a network. For example, the present technology can be implemented as a cloud service that provides services regarding images (moving images) to any terminal such as a computer, an audio-visual (AV) device, a portable information processing terminal, or an Internet of Things (IoT) device.
[0263] Note that in this specification, the term "system" means a collection of multiple configured elements (devices, modules (components), etc.), and it does not matter whether all the configured elements are in the same housing. Thus, both multiple devices housed in separate housings and connected via a network and a single device housing multiple modules in one housing are systems.
[0264] <Fields and Applications to which the Present Technology is Applicable>
[0265] Systems, devices, processing units, etc. to which the present technology is applied can be used in any fields such as transportation, medical care, crime prevention, agriculture, livestock farming, mining, beauty, factories, household appliances, weather and natural monitoring. In addition, the uses of the present technology are also arbitrary.
[0266] For example, the present technology can be applied to systems and devices provided for providing content for appreciation and the like. In addition, for example, the present technology can also be applied to systems and devices for transportation such as traffic condition detection and autonomous driving control. In addition, for example, the present technology can also be applied to systems and devices provided for security. In addition, for example, the present technology can be applied to systems and devices provided for automatic control of machines and the like. In addition, for example, the present technology can also be applied to systems and devices provided for agriculture or livestock farming. In addition, the present technology can also be applied to systems and devices for monitoring natural states such as volcanoes, forests and oceans, wild animals, etc. In addition, for example, the present technology can also be applied to systems and devices provided for sports.
[0267] <Others>
[0268] Note that the "flag" in this specification is information for identifying multiple states, and includes not only information for identifying two states of true (1) and false (0), but also information capable of identifying three or more states. Thus, the values that the "flag" can take can be, for example, binary values 1 / 0 or can be ternary values or more. That is, the number of bits constituting the "flag" is arbitrary and can be 1 bit or multiple bits. In addition, it is assumed that the identification information (including the flag) is not only in the form of including the identification information in the bit stream, but also in the form of including the difference information between the identification information and the specific reference information in the bit stream. Thus, in this specification, the "flag" and the "identification information" include not only the information itself, but also the difference information with respect to the reference information.
[0269] In addition, various types of information (such as metadata) regarding the encoded data (bitstream) can be sent or recorded in any form, as long as the various types of information are associated with the encoded data. Here, the term "associated" means that, for example, other data can be used (linked) when processing one data. That is, data that are associated with each other can be collected as one data or can be separate data. For example, information associated with the encoded data (image) can be sent on a transmission path different from the transmission path of the encoded data (image). In addition, for example, information associated with the encoded data (image) can be recorded on a recording medium different from the encoded data (image) (or another recording area of the same recording medium). Note that this "association" can be a part of the data, rather than the entire data. For example, an image and the information corresponding to the image can be associated with each other in any unit such as multiple frames, one frame, or a part of a frame.
[0270] Note that in this specification, terms such as "combine", "multiplex", "add", "integrate", "include", "store", and "insert" refer to putting multiple things into one, for example, putting the encoded data and metadata into one data, and refer to a method of the above-mentioned "association".
[0271] In addition, the embodiments of the present technology are not limited to the above embodiments, and various modifications can be made without departing from the gist of the present technology.
[0272] For example, the configuration described as one device (or processing unit) can be divided and configured into multiple devices (or processing units). Conversely, the configuration described as multiple devices (or processing units) can be configured together as one device (or processing unit). In addition, a configuration other than the above configuration can be added to the configuration of each device (or each processing unit). In addition, a part of the configuration of a specific device (or processing unit) can be included in the configuration of another device (or another processing unit), as long as the configuration and operation of the system as a whole are basically the same.
[0273] In addition, for example, the above program can be executed by any device. In this case, the device only needs to have the necessary functions (function blocks, etc.) and obtain the necessary information.
[0274] In addition, for example, each step of a flowchart can be executed by one device, or can be shared and executed by multiple devices. In addition, in the case where a step includes multiple processes, the multiple processes can be executed by one device, or can be shared and executed by multiple devices. In other words, the multiple processes included in one step can be executed as processes of multiple steps. Conversely, the processes described as multiple steps can be executed together as one step.
[0275] For example, a program executed by a computer may be configured such that the processes of the steps describing the program are executed in a time series in the order described in this specification. In addition, a program executed by a computer may be configured such that the processes of the steps describing the program are executed in parallel. In addition, a program executed by a computer may be configured such that the processes of the steps describing the program are executed individually at a necessary timing (e.g., when called). That is, as long as there is no contradiction, the processes of each step may be executed in an order different from the above order. In addition, the processes of the steps describing the program may be executed in parallel with the processes of other programs, or may be executed in combination with the processes of other programs.
[0276] In addition, for example, as long as there is no contradiction, multiple technologies related to the present technology may be independently implemented as a whole. Of course, any number of the present technologies may be implemented together. For example, part or all of the present technology described in any embodiment may be implemented in combination with part or all of the present technology described in other embodiments. In addition, part or all of any of the above-described present technologies may be implemented in combination with other technologies not described above.
[0277] Note that the present technology may also have the following configuration.
[0278] (1) An image processing apparatus, comprising:
[0279] An encoding mode setting unit configured to: set an encoding mode for encoding an image by setting a transform type having the minimum encoding cost in a non-united color difference encoding mode as the transform type in a united color difference encoding mode, and obtaining the encoding cost in the united color difference encoding mode.
[0280] (2) The image processing apparatus according to (1), wherein
[0281] the encoding mode setting unit performs the setting of the united color difference encoding mode and the setting of the transform type as the setting of the encoding mode.
[0282] (3) The image processing apparatus according to (2), wherein
[0283] the encoding mode setting unit sets whether to apply transform skip as the transform type in the united color difference encoding mode.
[0284] (4) The image processing apparatus according to (3), wherein
[0285] when not applying the transform skip, the encoding mode setting unit applies DCT2 as the transform type.
[0286] (5) The image processing apparatus according to (3) or (4), wherein
[0287] The encoding mode setting unit sets the value of the transform skip flag having the minimum encoding cost in the non-united chrominance encoding mode to the transform skip flag in the united chrominance encoding mode.
[0288] (6) The image processing apparatus according to any one of (2) to (5), wherein
[0289] The encoding mode setting unit sets the united chrominance encoding mode based on the chrominance encoding block flag in the non-united chrominance encoding mode.
[0290] (7) The image processing apparatus according to (6), wherein
[0291] The encoding mode setting unit sets an encoding component identifier based on the set united chrominance encoding mode.
[0292] (8) The image processing apparatus according to any one of (1) to (7), wherein
[0293] In the non-united chrominance encoding mode, the encoding mode setting unit obtains the encoding cost of each transform type, sets the transform type having the minimum encoding cost in the obtained encoding costs, and sets the set transform type as the transform type in the united chrominance encoding mode.
[0294] (9) The image processing apparatus according to any one of (1) to (8), wherein
[0295] The encoding mode setting unit compares the minimum encoding cost of the non-united chrominance encoding mode with the encoding cost of the united chrominance encoding mode, and selects the mode having the minimum encoding cost.
[0296] (10) The image processing apparatus according to any one of (1) to (9) further includes:
[0297] An orthogonal transform unit configured to perform an orthogonal transform on the coefficient data of the image according to the encoding mode set by the encoding mode setting unit.
[0298] (11) The image processing apparatus according to (10) further includes:
[0299] An encoding unit configured to encode the coefficient data orthogonally transformed by the orthogonal transform unit according to the encoding mode set by the encoding mode setting unit.
[0300] (12) The image processing apparatus according to (11), wherein
[0301] The coding mode setting unit sets a transform skip flag indicating whether to apply transform skip as the coding mode, and
[0302] the coding unit codes the transform skip flag set by the coding mode setting unit.
[0303] (13) The image processing apparatus according to (11) or (12) further includes:
[0304] a quantization unit configured to quantize the coefficient data orthogonally transformed by the orthogonal transform unit, where
[0305] the coding unit codes the coefficient data quantized by the quantization unit.
[0306] (14) The image processing apparatus according to any one of (10) to (13) further includes:
[0307] a calculation unit configured to generate a residual between the image and a predicted image, where
[0308] the orthogonal transform unit orthogonally transforms the coefficient data of the residual.
[0309] (15) An image processing method includes:
[0310] Setting a coding mode for coding an image by setting a transform type with the minimum coding cost in a non-united color difference coding mode as the transform type in a united color difference coding mode, and obtaining the coding cost in the united color difference coding mode.
[0311] Reference mark list
[0312] 300 Image coding apparatus
[0313] 301 Control unit
[0314] 312 Calculation unit
[0315] 313 Orthogonal transform unit
[0316] 314 Quantization unit
[0317] 315 Coding unit
Claims
1. An image processing apparatus, comprising: A coding mode setting unit, configured to: set a coding mode for encoding an image by setting a transform type with the minimum coding cost in a non-united color difference coding mode as the transform type in the united color difference coding mode, and obtaining the coding cost in the united color difference coding mode. Wherein, the coding mode setting unit is configured to compare the minimum coding cost of the non-united color difference coding mode with the coding cost of the united color difference coding mode, and select the mode with the minimum coding cost.
2. The image processing apparatus according to claim 1, wherein, The coding mode setting unit performs the setting of the united color difference coding mode and the setting of the transform type as the setting of the coding mode.
3. The image processing apparatus according to claim 2, wherein, The coding mode setting unit sets whether to apply transform skip as the transform type in the united color difference coding mode.
4. The image processing apparatus according to claim 3, wherein, In the case of not applying the transform skip, the coding mode setting unit applies DCT2 as the transform type.
5. The image processing apparatus according to claim 3, wherein, The coding mode setting unit sets the value of the transform skip flag with the minimum coding cost in the non-united color difference coding mode to the transform skip flag in the united color difference coding mode.
6. The image processing apparatus according to claim 2, wherein, The coding mode setting unit sets the united color difference coding mode based on the color difference coding block flag in the non-united color difference coding mode.
7. The image processing apparatus according to claim 6, wherein, The coding mode setting unit sets a coding component identifier based on the set united color difference coding mode.
8. The image processing apparatus according to claim 1, wherein, In the non-united color difference coding mode, the coding mode setting unit obtains the coding cost of each transform type, sets the transform type with the minimum coding cost in the obtained coding costs, and sets the set transform type as the transform type in the united color difference coding mode.
9. The image processing apparatus according to claim 1, further comprising: An orthogonal transform unit, configured to perform an orthogonal transform on the coefficient data of the image according to the coding mode set by the coding mode setting unit.
10. The image processing apparatus according to claim 9, further comprising: A coding unit, configured to encode the coefficient data orthogonally transformed by the orthogonal transform unit according to the coding mode set by the coding mode setting unit.
11. The image processing apparatus according to claim 10, wherein, The coding mode setting unit sets a transform skip flag indicating whether to apply transform skip as the coding mode, and The coding unit encodes the transform skip flag set by the coding mode setting unit.
12. The image processing apparatus according to claim 10, further comprising: A quantization unit, configured to quantize the coefficient data orthogonally transformed by the orthogonal transform unit, wherein The coding unit encodes the coefficient data quantized by the quantization unit.
13. The image processing apparatus according to claim 9, further comprising: A calculation unit, configured to generate a residual between the image and a predicted image, wherein The orthogonal transform unit performs an orthogonal transform on the coefficient data of the residual.
14. An image processing method, comprising: Set a coding mode for encoding an image by setting a transform type with the minimum coding cost in a non-united color difference coding mode as the transform type in the united color difference coding mode, and obtaining the coding cost in the united color difference coding mode. Wherein, compare the minimum coding cost of the non-united color difference coding mode with the coding cost of the united color difference coding mode, and select the mode with the minimum coding cost.
15. A computer-readable storage medium having stored thereon a computer-executable program, which when executed by a processor causes the processor to execute an image processing method, comprising: The encoding mode for encoding an image is set by setting the transform type with the minimum encoding cost in the non-united chromatic aberration encoding mode to the transform type in the united chromatic aberration encoding mode, and obtaining the encoding cost in the united chromatic aberration encoding mode. Among them, the minimum encoding cost of the non-united chromatic aberration encoding mode is compared with the encoding cost of the united chromatic aberration encoding mode, and the mode with the minimum encoding cost is selected.
16. A computer program product comprising a computer program / instructions which, when executed by a processor, implement the steps of the image processing method according to claim 14.
Citation Information
Patent Citations
Dynamic image encoding / decoding method and device
CN102057680A
Image encoding device, image encoding method, image decoding device, and image decoding method
WO2019188465A1