Image processing apparatus and image processing method

By parsing size identification data to control ACT and IACT processes based on block size, the image processing apparatus reduces memory requirements and costs, addressing the memory increase issue in image encoding methods like VVC.

JP7697571B2Active Publication Date: 2025-06-24SONY GROUP CORP
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
JP2024152944
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Priority Date
2019-12-27
Filing Date
2024-09-05
Publication Date
2025-06-24
Estimated Expiration
2040-12-18

AI Technical Summary

Technical Problem

The application of Adaptive Color Transform (ACT) processes in image encoding methods like VVC leads to an increase in memory size due to the need to temporarily accumulate YCgCo residual signals, which is particularly significant for larger orthogonal transform block sizes.

Method used

The implementation of an image processing apparatus and method that includes parsing size identification data to determine the maximum block size of an orthogonal transform block, enabling or disabling inverse adaptive color conversion processes only when the block size is 32, thereby reducing memory requirements by controlling the application of ACT and its inverse process (IACT) based on block size.

Benefits of technology

This approach effectively avoids an increase in memory size and associated costs by optimizing the use of ACT and IACT processes, ensuring efficient memory utilization and encoding efficiency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007697571000001
    Figure 0007697571000001
  • Figure 0007697571000002
    Figure 0007697571000002
  • Figure 0007697571000003
    Figure 0007697571000003
Patent Text Reader

Abstract

To avoid an increase in memory size.SOLUTION: An adaptive color conversion unit performs adaptive color conversion processing of adaptively performing conversion of a color space of an image to be encoded, on a residual signal of the image, and an orthogonal transform unit performs orthogonal transform processing for each of orthogonal transform blocks that are units of processing, on the residual signal of the image or on the residual signal of the image subjected to the adaptive color conversion processing. Then, control related to application of the adaptive color conversion processing is performed by a controller. The present technology can be applied to, for example, an image encoding device and an image decoding device that support ACT processing.SELECTED DRAWING: Figure 2
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to an image processing apparatus and an image processing method, and more particularly to an image processing apparatus and an image processing method capable of avoiding an increase in memory size.

Background Art

[0002] Conventionally, an apparatus for compressing and encoding an image by adopting an encoding method that digitally processes image information, aims at efficient information transmission and storage at that time, and utilizes redundancy peculiar to image information to perform orthogonal transformation such as discrete cosine transform and motion compensation for compression has been becoming popular.

[0003] Examples of this encoding method include MPEG (Moving Picture Experts Group), H.264 and MPEG-4 Part 10 (Advanced Video Coding, hereinafter referred to as H.264 / AVC), and H.265 and MPEG-H Part 2 (High Efficiency Video Coding, hereinafter referred to as H.265 / HEVC).

[0004] In addition, in order to further improve the encoding efficiency for AVC (Advanced Video Coding), HEVC (High Efficiency Video Coding), etc., the standardization of a coding method called VVC (Versatile Video Coding) is in progress (refer to the support of the embodiments described later).

[0005] As disclosed in Non-Patent Document 1, in VVC, a technique related to ACT (Adaptive Color Transform) for adaptively transforming the color space of an image is disclosed.

Prior Art Documents

Non-Patent Documents

[0006]

Non-Patent Document 1

Summary of the Invention

Problems to be Solved by the Invention

[0007] By the way, when applying the ACT process to convert, for example, the RGB color space to the YCgCo color space, it is necessary to temporarily accumulate the YCgCo residual signal output as the processing result. Therefore, depending on the block size of the orthogonal transform block in the orthogonal transform process performed after the ACT process, it is considered necessary to increase the memory size for accumulating the YCgCo residual signal.

[0008] The present disclosure has been made in view of such a situation, and enables avoidance of an increase in the memory size.

Means for Solving the Problems

[0009] The image processing apparatus according to the first aspect of the present disclosure includes a parsing unit that parses size identification data for identifying whether the maximum block size of an orthogonal transform block is 64 or 32 from a bit stream, and when the size identification data parsed by the parsing unit indicates that 32 is applied as the maximum block size of the orthogonal transform block, an inverse adaptive color conversion unit that applies the inverse adaptive color conversion process to the residual signal generated by applying the inverse orthogonal transform process to each orthogonal transform block that is a processing unit with respect to the conversion coefficients obtained by decoding the bit stream according to the identification data for identifying whether the adaptive color conversion process for adaptively converting the color space of the image or the inverse adaptive color conversion process for adaptively inverse-converting the color space of the image is Enable.

[0010] The image processing method according to the first aspect of the present disclosure includes parsing size identification data for identifying whether the maximum block size of an orthogonal transform block is 64 or 32 from a bit stream, and when the size identification data indicates that 32 is applied as the maximum block size of the orthogonal transform block, applying the inverse adaptive color conversion process to the residual signal generated by applying the inverse orthogonal transform process to each orthogonal transform block that is a processing unit with respect to the conversion coefficients obtained by decoding the bit stream according to the identification data for identifying whether the adaptive color conversion process for adaptively converting the color space of the image or the inverse adaptive color conversion process for adaptively inverse-converting the color space of the image is Enable.

[0011] In a first aspect of the present disclosure, size identification data for identifying whether the maximum block size of an orthogonal transform block is 64 or 32 is parsed from a bit stream, and only when the size identification data indicates that 32 is applied as the maximum block size of the orthogonal transform block, an inverse adaptive color conversion process that adaptively inverse-converts the color space of the image or an adaptive color conversion process that adaptively converts the color space of the image is enabled. An inverse adaptive color conversion process is applied to the residual signal generated by applying an inverse orthogonal transform process to each orthogonal transform block that is a processing unit with respect to the conversion coefficients obtained by decoding the bit stream according to the identification data.

[0012] An image processing apparatus according to a second aspect of the present disclosure includes a setting unit that sets identification data for identifying whether an adaptive color conversion process that adaptively converts the color space of the image or an inverse adaptive color conversion process that adaptively inverse-converts the color space of the image is enabled only when size identification data for identifying whether the maximum block size of an orthogonal transform block is 64 or 32 is set assuming that 32 is applied as the maximum block size of the orthogonal transform block, and an encoding unit that encodes the image and generates a bit stream including the identification data set by the setting unit.

[0013] An image processing method according to a second aspect of the present disclosure includes setting identification data for identifying whether an adaptive color conversion process that adaptively converts the color space of the image or an inverse adaptive color conversion process that adaptively inverse-converts the color space of the image is enabled only when size identification data for identifying whether the maximum block size of an orthogonal transform block is 64 or 32 is set assuming that 32 is applied as the maximum block size of the orthogonal transform block, and encoding the image to generate a bit stream including the identification data.

[0014] In a second aspect of the present disclosure, adaptive color conversion processing for adaptively converting the color space of an image or inverse adaptive color conversion processing for adaptively inverse-converting the color space of an image is enabled only when identification data for identifying whether the maximum block size of the orthogonal transform block is 64 or 32 is set assuming that 32 is applied as the maximum block size of the orthogonal transform block. The image is encoded to generate a bitstream including the identification data.

Brief Description of Drawings

[0015]

Figure 1

Figure 2

Figure 3

Figure 4

Figure 5

Figure 6

Figure 7

Figure 8

Figure 9

Figure 10

Figure 11

Figure 12

Embodiments for Carrying Out the Invention

[0016] <Literature etc. Supporting Technical Content and Technical Terms> The scope disclosed in this specification is not limited to the content of the embodiments. The content of the following reference documents REF1 to REF5, which were known at the time of filing, is also incorporated herein by reference. That is, the content described in reference documents REF1 to REF5 also serves as a basis for judging the support requirements. Furthermore, the documents cited in reference documents REF1 to REF5 also serve as a basis for judging the support requirements.

[0017] For example, even if structures such as Quad-Tre Block Structure, QTBT (Quad Tree Plus Binary Tree), Block Structure, and MTT (Multi-type Tree) Block Structure are not directly defined in the detailed description of the invention, they are within the scope of the present disclosure and are considered to meet the support requirements of the claims. Also, for example, technical terms such as Parsing, Syntax, and Semantics, even if not directly defined in the detailed description of the invention, are within the scope of the present disclosure and are considered to meet the support requirements of the claims. Also, for example, technical applications such as ACT (Adaptive Color Transform), even if not directly defined in the detailed description of the invention, are within the scope of the present disclosure and are considered to meet the support requirements of the claims.

[0018] REF1: Recommendation ITU-T H.264 (04 / 2017) “Advanced video coding for generic audiovisual services”, April 2017 REF2: Recommendation ITU-T H.265 (02 / 2018) “High efficiency video coding”, February 2018 REF3: Benjamin Bross, Jianle Chen, Shan Liu, Versatile Video Coding (Draft 7), JVET-P2001-v14 (version 14 - date 2019-11-14) REF4: Jianle Chen, Yan Ye, Seung Hwan Kim, Algorithm description for Versatile Video Coding and Test Model 7 (VTM 7), JVET-P2002-v1 (version 1 - date 2019-11-10) REF5: Xiaoyu Xiu, Yi-Wen Chen, Tsung-Chuan Ma, Hong-Jheng Jhu, Xianglin Wang, Support of adaptive color transform for 444 video coding in VVC, JVET-P0517_r1 (version 3 - date 2019-10-11)

[0019] <Terminology> In this application, the following terms are defined as follows.

[0020] <Block> The "block" (not the block indicating the processing unit) used in the description as a partial area or processing unit of an image (picture) indicates an arbitrary partial area in the picture, unless otherwise specified, and its size, shape, characteristics, etc. are not limited. For example, the "block" includes any partial area (processing unit) such as TB (Transform Block), TU (Transform Unit), PB (Prediction Block), PU (Prediction Unit), SCU (Smallest Coding Unit), CU (Coding Unit), LCU (Largest Coding Unit), CTB (Coding Tree Block), CTU (Coding Tree Unit), transform block, sub-block, macro-block, tile, or slice.

[0021] <Specification of Block Size> Also, when specifying the size of such a block, not only directly specify the block size, but also indirectly specify the block size. For example, the block size may be specified using identification information for identifying the size. Also, for example, the block size may be specified by a ratio or difference from the size of a reference block (e.g., LCU, SCU, etc.). For example, when transmitting information specifying the block size as a syntax element or the like, as that information, information indirectly specifying the size as described above may be used. By doing so, the amount of information of that information can be reduced, and in some cases, the coding efficiency can be improved. Also, the specification of this block size includes the specification of the range of the block size (e.g., the specification of the range of the allowable block size, etc.).

[0022] <Unit of Information / Processing> The data unit in which various information is set and the data unit to which various processes are applied are each arbitrary and not limited to the examples described above. For example, these information and processes may be set for each of TU (Transform Unit), TB (Transform Block), PU (Prediction Unit), PB (Prediction Block), CU (Coding Unit), LCU (Largest Coding Unit), sub-block, block, tile, slice, picture, sequence, or component, or the data of those data units may be targeted. Of course, this data unit can be set for each piece of information and process, and it is not necessary for the data units of all information and processes to be unified. Note that the storage location of this information is arbitrary and may be stored in the header or parameter set of the data unit described above. Also, it may be stored in multiple locations.

[0023] <Control Information> Control information related to this technology may be transmitted from the encoding side to the decoding side. For example, control information (such as enabled_flag) that controls whether to permit (or prohibit) the application of the above-described technology may be transmitted. Also, for example, control information indicating the target to which the above-described technology is applied (or the target to which it is not applied) may be transmitted. For example, control information specifying the block size (upper limit or lower limit, or both), frame, component, or layer, etc., to which this technology is applied (or for which application is permitted or prohibited) may be transmitted.

[0024] <Flag> Note that in this specification, a "flag" is information for identifying a plurality of states, and includes not only information used for identifying two states of true (1) or false (0), but also information capable of identifying three or more states. Therefore, the values that this "flag" can take may be, for example, two values of 1 / 0, or three or more values. That is, the number of bits constituting this "flag" is arbitrary and may be 1 bit or multiple bits. Also, identification information (including flags) is assumed to be included in the bit stream not only in the form of including the identification information itself, but also in the form of including the difference information of the identification information with respect to a certain reference information. Therefore, in this specification, "flag" and "identification information" include not only the information itself, but also the difference information with respect to the reference information.

[0025] <Associate metadata> Also, various types of information (such as metadata) related to the encoded data (bitstream) may be transmitted or recorded in any form as long as they are associated with the encoded data. Here, the term "associate" means, for example, enabling the use (linking) of one piece of data when processing the other piece of data. That is, the data associated with each other may be grouped as one piece of data or may be individual pieces of data. For example, the information associated with the encoded data (image) may be transmitted on a transmission path different from that of the encoded data (image). Also, for example, the information associated with the encoded data (image) may be recorded on a recording medium different from that of the encoded data (image) (or a different recording area of the same recording medium). Note that this "association" may be for a part of the data rather than the entire data. For example, an image and the information corresponding to the image may be associated with each other in any unit such as a plurality of frames, one frame, or a part within a frame.

[0026] Note that in this specification, terms such as "synthesize", "multiplex", "add", "integrate", "include", "store", "embed", "insert", "plug in", etc. mean, for example, grouping a plurality of things into one, such as grouping the encoded data and the metadata into one piece of data, and mean one method of the above-mentioned "associate". Also, in this specification, encoding includes not only the entire process of converting an image into a bitstream but also some processes. For example, it includes not only processes including prediction processing, orthogonal transformation, quantization, arithmetic coding, etc., but also processes collectively referred to as quantization and arithmetic coding, processes including prediction processing, quantization, and arithmetic coding, etc. Similarly, decoding includes not only the entire process of converting a bitstream into an image but also some processes. For example, it includes not only processes including inverse arithmetic decoding, inverse quantization, inverse orthogonal transformation, prediction processing, etc., but also processes including inverse arithmetic decoding and inverse quantization, processes including inverse arithmetic decoding, inverse quantization, and prediction processing, etc.

[0027] A prediction block means a block that is a processing unit when performing inter prediction, and includes sub-blocks within the prediction block. Also, when the orthogonal transformation block that is the processing unit when performing orthogonal transformation and the encoding block that is the processing unit when performing encoding processing have unified processing units, it means the same block as the prediction block, orthogonal transformation block, and encoding block.

[0028] Inter prediction is a general term for processing involving prediction between frames (prediction blocks) such as derivation of motion vectors by motion detection (Motion Prediction / Motion Estimation) and motion compensation using motion vectors (Motion Compensation), and includes some processing (for example, only motion compensation processing) or all processing (for example, motion detection processing + motion compensation processing) used when generating a predicted image. An inter prediction mode means including variables (parameters) referred to when deriving the inter prediction mode, such as the mode number when performing inter prediction, the index of the mode number, the block size of the prediction block, and the size of the sub-block that is the processing unit within the prediction block.

[0029] In the present disclosure, identification data for identifying a plurality of patterns can also be set as the syntax of the bitstream. In this case, the decoder can perform processing more efficiently by parsing and referring to the identification data. Methods (data) for identifying the block size include not only digitizing (bitizing) the block size itself but also methods (data) for identifying the difference value with respect to a reference block size (such as the maximum block size and the minimum block size).

[0030] Hereinafter, specific embodiments to which the present technology is applied will be described in detail with reference to the drawings.

[0031] <Configuration Example of Image Processing System> FIG. 1 is a block diagram showing a configuration example of an embodiment of an image processing system to which the present technology is applied.

[0032] As shown in FIG. 1, the image processing system 11 is configured to include an image encoding device 12 and an image decoding device 13. For example, in the image processing system 11, the image input to the image encoding device 12 is encoded, and the bitstream obtained by the encoding is transmitted to the image decoding device 13, and the decoded image decoded from the bitstream is output in the image decoding device 13.

[0033] As shown in FIG. 1, the image encoding device 12 has a prediction unit 21, an encoding unit 22, a storage unit 23, and a control unit 24, and the image decoding device 13 has a prediction unit 31, a decoding unit 32, a storage unit 33, and a control unit 34.

[0034] The prediction unit 21 performs inter prediction or intra prediction to generate a predicted image. For example, when the prediction unit 21 performs inter prediction, the prediction unit 21 generates a predicted image using a prediction block of a predetermined block size as a processing unit.

[0035] The encoding unit 22 encodes the image input to the image encoding device 12 according to a predetermined encoding method using an encoding block of a predetermined block size as a processing unit, and transmits the bitstream of the encoded data to the image decoding device 13. Further, this bitstream includes parameters related to blocks as described later with reference to FIGS. 4 to 6.

[0036] The storage unit 23 stores various data that needs to be stored when encoding an image in the image encoding device 12. For example, as will be described later with reference to FIG. 2, the storage unit 23 temporarily stores the YCgCo residual signal 1 output by the ACT process and the YCgCo residual signal 2 that is the target of the IACT process.

[0037] The control unit 24 performs control related to the application of the ACT process and the IACT process as described later with reference to FIG. 2.

[0038] The prediction unit 31 performs inter prediction or intra prediction to generate a predicted image. For example, when performing inter prediction, the prediction unit 21 generates a predicted image with a prediction block of a predetermined block size as a processing unit.

[0039] The decoding unit 32 decodes the bitstream transmitted from the image encoding device 12 in correspondence with the encoding method by the encoding unit 22, and outputs the decoded image.

[0040] The storage unit 33 stores various data that needs to be stored when decoding an image in the image decoding device 13. For example, as will be described later with reference to FIG. 3, the storage unit 33 temporarily stores the YCgCo residual signal 2 that is the target of the IACT process.

[0041] The control unit 34 performs control regarding the application of the IACT process as will be described later with reference to FIG. 3.

[0042] In the image processing system 11 configured in this way, by appropriately controlling the ACT process and the IACT process, it is possible to avoid an increase in the memory sizes of the storage unit 23 and the storage unit 33.

[0043] With reference to the block diagram shown in FIG. 2, the configuration of the image encoding device 12 will be further described.

[0044] As shown in FIG. 2, the image encoding device 12 includes an arithmetic unit 41, an adaptive color conversion unit 42, an orthogonal conversion unit 43, a quantization unit 44, an inverse quantization unit 45, an inverse orthogonal conversion unit 46, an inverse adaptive color conversion unit 47, an arithmetic unit 48, a prediction unit 49, and an encoding unit 50.

[0045] The arithmetic unit 41 performs an operation of subtracting the predicted image supplied from the prediction unit 49 from the image input to the image encoding device 12, and supplies the RGB residual signal 1, which is the difference information obtained by the operation, to the adaptive color conversion unit 42.

[0046] The adaptive color conversion unit 42 performs an ACT process that adaptively converts the color space of the image to be encoded on the RGB residual signal 1 supplied from the arithmetic unit 41. For example, the adaptive color conversion unit 42 performs an ACT process of converting the RGB color space to the YCgCo color space, thereby obtaining a YCgCo residual signal 1 from the RGB residual signal 1 and supplying it to the orthogonal conversion unit 43.

[0047] The orthogonal conversion unit 43 performs an orthogonal conversion process of performing orthogonal conversion for each orthogonal conversion block that is a processing unit on the YCgCo residual signal 1 supplied from the adaptive color conversion unit 42, thereby obtaining conversion coefficients and supplying them to the quantization unit 44. Also, when the orthogonal conversion unit 43 is controlled so that the ACT process is not performed in the adaptive color conversion unit 42, the orthogonal conversion unit 43 can perform an orthogonal conversion process on the RGB residual signal 1 supplied from the arithmetic unit 41.

[0048] The quantization unit 44 quantizes the conversion coefficients supplied from the orthogonal conversion unit 43 and supplies them to the inverse quantization unit 45 and the encoding unit 50. The inverse quantization unit 45 inverse quantizes the conversion coefficients quantized in the quantization unit 44 and supplies them to the inverse orthogonal conversion unit 46.

[0049] The inverse orthogonal conversion unit 46 performs an inverse orthogonal conversion process of performing inverse orthogonal conversion for each orthogonal conversion block that is a processing unit on the conversion coefficients supplied from the inverse quantization unit 45, thereby obtaining a YCgCo residual signal 2 and supplying it to the inverse adaptive color conversion unit 47. Also, when the inverse orthogonal conversion unit 46 is controlled so that the IACT process is not performed in the inverse adaptive color conversion unit 47, the inverse orthogonal conversion unit 46 can obtain an RGB residual signal 2 by the inverse orthogonal conversion process and supply it to the arithmetic unit 48.

[0050] The inverse adaptive color conversion unit 47 performs an IACT process that adaptively inverse-converts the color space of the image on the YCgCo residual signal 2 supplied from the inverse orthogonal conversion unit 46. For example, the inverse adaptive color conversion unit 47 performs an IACT process of inverse-converting the YCgCo color space to the RGB color space, thereby obtaining an RGB residual signal 2 from the YCgCo residual signal 2 and supplying it to the arithmetic unit 48.

[0051] The arithmetic unit 48 locally reconstructs (decodes) an image by performing an operation of adding the RGB residual signal 2 supplied from the inverse adaptive color conversion unit 47 to the predicted image supplied from the prediction unit 49, and outputs a reconstruction signal representing the reconstructed image. Further, when the arithmetic unit 48 is controlled so that the IACT process is not performed in the inverse adaptive color conversion unit 47, the arithmetic unit 48 can reconstruct (decode) an image from the RGB residual signal 2 supplied from the inverse orthogonal conversion unit 46.

[0052] The prediction unit 49 corresponds to the prediction unit 21 in FIG. 1, generates a predicted image predicted from the image reconstructed in the arithmetic unit 48, and supplies it to the arithmetic unit 41 and the arithmetic unit 48.

[0053] The encoding unit 50 corresponds to the encoding unit 22 in FIG. 1, and performs an encoding process on the transform coefficients quantized in the quantization unit 44 using, for example, CABAC (Context-based Adaptive Binary Arithmetic Coding), which is an encoding method with high encoding efficiency for continuous equivalents. As a result, the encoding unit 50 obtains a bitstream of encoded data and transmits it to the image decoding device 13.

[0054] In the image encoding device 12 configured as described above, by performing the ACT process of converting the RGB residual signal 1 into the YCgCo residual signal 1 in the adaptive color conversion unit 42, it is possible to improve the energy concentration degree of the signal. In this way, by improving the energy concentration degree of the signal, the image encoding device 12 can represent the image signal with a small amount of code, and an improvement in encoding efficiency is expected.

[0055] With reference to the block diagram shown in FIG. 3, the configuration of the image decoding device 13 will be further described.

[0056] As shown in FIG. 3, the image decoding device 13 includes a decoding unit 61, an inverse quantization unit 62, an inverse orthogonal conversion unit 63, an inverse adaptive color conversion unit 64, an arithmetic unit 65, and a prediction unit 66.

[0057] The decoding unit 61 corresponds to the decoding unit 32 in FIG. 1, and performs a decoding process using an encoding method (for example, CABAC) corresponding to the encoding by the encoding unit 50 in FIG. 2 on the bit stream of the encoded data transmitted from the image encoding device 12. As a result, the decoding unit 61 acquires the quantized transform coefficients from the bit stream of the encoded data and supplies them to the inverse quantization unit 62. At this time, the decoding unit 61 also acquires parameters regarding blocks as will be described later with reference to FIGS. 4 to 6 included in the bit stream of the encoded data.

[0058] The inverse quantization unit 62 inverse quantizes the quantized transform coefficients supplied from the decoding unit 61 and supplies them to the inverse orthogonal transform unit 63.

[0059] The inverse orthogonal transform unit 63 performs an inverse orthogonal transform process of performing an inverse orthogonal transform for each orthogonal transform block serving as a processing unit on the transform coefficients supplied from the inverse quantization unit 62 to obtain the YCgCo residual signal 2, and supplies it to the inverse adaptive color conversion unit 64. Further, when the inverse orthogonal transform unit 63 is controlled so that the IACT process is not performed in the inverse adaptive color conversion unit 64, the RGB residual signal 2 obtained by the inverse orthogonal transform process can be supplied to the arithmetic unit 65.

[0060] Similar to the inverse adaptive color conversion unit 47 in FIG. 2, the inverse adaptive color conversion unit 64 performs an IACT process of adaptively inverse-converting the color space of the image on the YCgCo residual signal 2 supplied from the inverse orthogonal transform unit 63. For example, the inverse adaptive color conversion unit 64 performs an IACT process of inverse-converting the YCgCo color space to the RGB color space to obtain the RGB residual signal 2 from the YCgCo residual signal 2 and supplies it to the arithmetic unit 65.

[0061] The arithmetic unit 65 performs an operation of adding the RGB residual signal 2 supplied from the inverse adaptive color conversion unit 64 to the predicted image supplied from the prediction unit 66, thereby locally reconstructing (decoding) the image and outputting a reconstruction signal representing the reconstructed image. Further, when the arithmetic unit 65 is controlled so that the IACT process is not performed in the inverse adaptive color conversion unit 64, the arithmetic unit 65 can reconstruct (decode) an image from the RGB residual signal 2 supplied from the inverse orthogonal conversion unit 63.

[0062] The prediction unit 66 corresponds to the prediction unit 31 in FIG. 1, and in the same manner as the prediction unit 49 in FIG. 2, generates a predicted image predicted from the image reconstructed in the arithmetic unit 65 and supplies it to the arithmetic unit 65.

[0063] In the image decoding apparatus 13 configured as described above, similarly to the image encoding apparatus 12, it can contribute to the improvement of the encoding efficiency.

[0064] As described above, the image processing system 11 is configured, and the ACT process for converting the RGB residual signal 1 into the YCgCo residual signal 1 is performed in the adaptive color conversion unit 42, and the IACT process for converting the YCgCo residual signal 2 into the RGB residual signal 2 is performed in the inverse adaptive color conversion units 47 and 64.

[0065] At this time, in the image encoding apparatus 12, since the ACT process is performed in units of three components in the adaptive color conversion unit 42, the YCgCo residual signal 1 for three components is temporarily stored in the storage unit 23 in FIG. 1. Therefore, the storage unit 23 stores the YCgCo residual signal 1 for three components corresponding to the block size (for example, 32×32) of the orthogonal conversion block in the orthogonal conversion unit 43. Further, the storage unit 23 stores the YCgCo residual signal 2 for three components corresponding to the block size (for example, 32×32) of the orthogonal conversion block in the inverse orthogonal conversion unit 46.

[0066] Similarly, in the image decoding device 13, the YCgCo residual signal 2 for three components corresponding to the block size of the orthogonal transformation block (for example, 32×32) in the inverse orthogonal transformation unit 63 is stored in the storage unit 33 in FIG. 1.

[0067] Thus, in the image processing system 11, when applying the ACT process and the IACT process, it is necessary to increase the memory size of the storage unit 23 so that the YCgCo residual signal 1 and the YCgCo residual signal 2 corresponding to the block size of the orthogonal transformation block can be stored. Similarly, when applying the ACT process and the IACT process, it is necessary to increase the memory size of the storage unit 33 so that the YCgCo residual signal 2 corresponding to the block size of the orthogonal transformation block can be stored. Therefore, as the memory size increases, the implementation cost of the image processing system 11 increases.

[0068] Therefore, in the image processing system 11, by appropriately controlling, by the control unit 24, the application of the ACT process by the adaptive color conversion unit 42 and the IACT process by the inverse adaptive color conversion unit 47, it is possible to avoid an increase in the memory size of the storage unit 23. Similarly, in the image processing system 11, by appropriately controlling, by the control unit 34, the application of the IACT process by the inverse adaptive color conversion unit 64, it is possible to avoid an increase in the memory size of the storage unit 33. Therefore, as the image processing system 11 avoids an increase in the memory size, it can suppress an increase in the implementation cost.

[0069] <The first concept regarding the application of the ACT process and the IACT process>

[0070] For example, in the image processing system 11, when a predetermined restriction (for example, restrictions on size, area, shape, etc.) is provided for the encoding block when encoding an image, control is performed so that the ACT process and the IACT process are applied.

[0071] For example, parameters of the encoding block for performing such a restriction include size, long side size, short side size, area, and shape. Examples of the size include 16×16, 16×8, 8×16, etc. The long side size for a 16×8 block is 16. The short side size for a 16×8 block is 8. Examples of the area include 16×16, 16×8, etc. Examples of the shape include square, rectangle, etc.

[0072] Conventionally, since a block size of 32×32 has been used for the orthogonal transform block in the orthogonal transform process and the inverse orthogonal transform process, a memory size capable of storing the YCgCo residual signal for three components with a block size of 32×32 has been required.

[0073] On the other hand, in the image processing system 11, a restriction is provided such that the block size of the encoding block when the ACT process and the IACT process are applied is a predetermined size (for example, 16×16) or less. Due to such a restriction, the storage unit 23 may have a memory size only for storing the YCgCo residual signal 1 for three components with a block size of 16×16. Similarly, for the YCgCo residual signal 2, the storage unit 23 and the storage unit 33 may have a memory size only for storing the YCgCo residual signal 2 for three components with a block size of 16×16.

[0074] For example, regarding the syntax of the bitstream, parameters (such as size, area, shape, etc.) of the encoding block are considered for the conditions under which the ACT process and the IACT process are applied. That is, in the image encoding apparatus 12, when transmitting a flag indicating that the ACT process is to be applied, it is confirmed that the block size of the encoding block is a predetermined size (for example, 16×16) or less. Then, only when the block size of the encoding block is a predetermined size (for example, 16×16) or less, a flag indicating that the ACT process is to be applied is transmitted, and when it is larger than the predetermined size (for example, 16×16), the syntax is determined so as not to transmit a flag indicating that the ACT process is to be applied.

[0075] As a result, in the image processing system 11, when the block size of the encoding block is larger than a predetermined size, it becomes unnecessary to transmit a flag indicating that the ACT process is to be applied, and since that flag can be removed from the bitstream, an improvement in encoding efficiency can be expected. Also, by not transmitting such a flag that does not need to be transmitted, an ambiguous signal can be removed from the syntax of the bitstream.

[0076] Next, the case where VVC (Versatile Video Coding) is used as the encoding method will be described. When VVC is used as the encoding method, 64×64 can be used as the block size of the orthogonal transform block in the orthogonal transform process and the inverse orthogonal transform process. Therefore, when VVC is used as the encoding method, a memory size capable of storing the YCgCo residual signal for three components with a block size of 64×64 was required.

[0077] Therefore, in the image processing system 11, when VVC is used as the encoding method, for example, the block size of the encoding block when the ACT process and the IACT process are applied is limited to 32×32 or less. Due to such a limitation, the storage unit 23 only needs a memory size capable of storing the YCgCo residual signal 1 for three components with a block size of 32×32. Similarly, for the YCgCo residual signal 2, the storage unit 23 and the storage unit 33 only need a memory size capable of storing the YCgCo residual signal 2 for three components with a block size of 32×32.

[0078] Even when HEVC (High Efficiency Video Coding) is used as the encoding method, a mechanism for performing ACT processing and IACT processing is provided. In HEVC, the maximum block size of the orthogonal transform block in the orthogonal transform process and the inverse orthogonal transform process was 32×32 according to the standard. In contrast, in VVC, the maximum block size of the orthogonal transform block can support 64×64, so the increase in memory size is larger compared to HEVC. Therefore, in the image processing system 11, when VVC is used as the encoding method, when ACT processing and IACT processing are applied, by restricting the orthogonal transform block in the orthogonal transform process and the inverse orthogonal transform process, it becomes possible to perform processing with the same memory size as HEVC.

[0079] FIG. 4 shows an example of a parameter set of the high-level syntax.

[0080] In the parameter set shown in FIG. 4, when sps_act_enabled_flag is 1, ACT processing and IACT processing can be applied, and it is specified that cu_act_enabled_flag may exist in the coding unit syntax. On the other hand, when sps_act_enabled_flag is 0, ACT processing and IACT processing are not applied, and it is specified that cu_act_enabled_flag does not exist in the coding unit syntax. Note that when sps_act_enabled_flag does not exist in the parameter set, sps_act_enabled_flag is presumed to be 0.

[0081] In the parameter set shown in FIG. 4, sps_log2_act_max_size_minus2 specifies the maximum block size used in ACT processing and IACT processing in the range from 0 to 7. Note that when sps_log2_act_max_size_minus2 does not exist in the parameter set, sps_log2_act_max_size_minus2 is presumed to be 0.

[0082] Also, when the MaxActSize variable is set to 1, it is much smaller than 1 << (sps_log2_act_max_size_minus2 + 2). For example, when sps_log2_act_max_size_minus2 is set to 2, the MaxActSize variable becomes 16 (= 1 << 4). This can prohibit ACT processing and IACT processing from being performed on sizes larger than MaxActSize.

[0083] FIG. 5 shows an example of a high-level syntax coding unit.

[0084] In the coding unit shown in FIG. 5, when cu_act_enabled_flag is 1, it specifies that the residual of the current coding unit is coded in the YCgCo color space. On the other hand, when cu_act_enabled_flag is 0, it specifies that the residual of the current coding unit is coded in the original color space. Note that when cu_act_enabled_flag does not exist, it is presumed that cu_act_enabled_flag is 0.

[0085] Here, by adding the condition that the width of the coding block is less than or equal to the maximum size of the block size in ACT processing and IACT processing, and the height of the coding block is less than or equal to the maximum size of the block size in ACT processing and IACT processing (&& cbWidth <= MaxActSize && cbHeight <= MaxActSize) to the condition for transmitting cu_act_enabled_flag, the memory size used in ACT processing and IACT processing can be restricted.

[0086] Also, by adding the condition that the width or height of the encoded block is less than or equal to the maximum size of the block size in the ACT process and the IACT process (&& (cbWidth * cbHeight) <= (MaxActSize * MaxActSize)) to the condition for transmitting the cu_act_enabled_flag, the memory size can be restricted. In this case, even if one side exceeds MaxActSize, if the other side is small enough to satisfy the above condition, the ACT process and the IACT process can be applied.

[0087] <Second Concept Regarding Application of ACT Process and IACT Process> In the second concept, even when the block size of the encoded block or the predicted block used in the image processing system 11 is large, control is performed so that orthogonal transformation processing and inverse orthogonal transformation processing are performed using orthogonal transformation blocks with a small size obtained by dividing the block size of the encoded block or the predicted block, and control is performed so that the ACT process and the IACT process are applied. That is, even when the block size of the encoded block or the predicted block used in the image processing system 11 is large, when the ACT process and the IACT process are applied, orthogonal transformation processing and inverse orthogonal transformation processing are performed using orthogonal transformation blocks smaller than the block size of the encoded block or the predicted block.

[0088] For example, when the block size of the encoded block is 64×64 and the block size of the prediction block for inter prediction is 64×64, usually, a 64×64 block size is also used in the orthogonal transformation block. In contrast, in the image processing system 11, even in the same case, when the ACT process and the IACT process are applied, orthogonal transformation blocks with a block size smaller than 64×64 (for example, four blocks with a block size of 32×32) are automatically used.

[0089] For example, by referring to a control signal indicating that the ACT process is applied, when the ACT process is applied, control is performed such that the orthogonal transformation process and the inverse orthogonal transformation process are performed using an orthogonal transformation block obtained by dividing the block size of the encoding block or prediction block at that time into a small block size.

[0090] Therefore, in the image processing system 11, by performing control according to such a second concept, the block size of the orthogonal transformation block in the orthogonal transformation process and the inverse orthogonal transformation process is reduced, and an increase in the memory size required when applying the ACT process and the IACT process can be avoided. For example, the storage unit 23 that stores the YCbCo residual signal 1 can have a memory size for three components with a block size of 32×32 instead of a memory size for three components with a block size of 64×64. As a result, in the image processing system 11, an increase in the implementation cost can be suppressed.

[0091] <Third Concept Regarding Application of ACT Process and IACT Process> In the third concept, when a small size is used as the maximum block size of the orthogonal transformation block in the orthogonal transformation process and the inverse orthogonal transformation process, control is performed so that the ACT process and the IACT process are applied.

[0092] For example, in the image processing system 11, 32 and 64 are defined as the maximum block sizes of the orthogonal transformation block. Then, the control unit 24 restricts the case where 32 is used as the maximum block size of the orthogonal transformation block, and causes the adaptive color conversion unit 42 to perform the ACT process and the inverse adaptive color conversion unit 47 to perform the IACT process. Similarly, the control unit 34 restricts the case where 32 is used as the maximum block size of the orthogonal transformation block, and causes the inverse adaptive color conversion unit 64 to perform the IACT process.

[0093] The control according to such a third concept can be realized by using the sps_max_luma_transform_size_64_flag included in the high-level syntax parameter set.

[0094] FIG. 6 shows an example of the high-level syntax parameter set used in the image processing system 11.

[0095] For example, when the sps_max_luma_transform_size_64_flag is 0, it is set to use 32 as the maximum block size of the orthogonal transform block. On the other hand, when the sps_max_luma_transform_size_64_flag is 1, it is set to use 64 as the maximum block size of the orthogonal transform block.

[0096] Therefore, in the image processing system 11, control is performed so that the ACT process and the IACT process are applied only when the sps_max_luma_transform_size_64_flag is 0. That is, in the image processing system 11, when the sps_max_luma_transform_size_64_flag is 1, the ACT process and the IACT process are not applied.

[0097] When sps_max_luma_transform_size_64_flag is 0, the control unit 24 sets 1, which indicates applying the adaptive color conversion process, to sps_act_enabled_flag that indicates whether to apply the adaptive color conversion process, and transmits it to the image decoding device 13. In this case, cu_act_enabled_flag may also be included in the coding unit syntax. When sps_act_enabled_flag is 0, it indicates not applying the adaptive color conversion process. In this case, cu_act_enabled_flag is not included in the coding unit syntax. Here, when sps_act_enabled_flag is not included in the parameter set, it is presumed that sps_act_enabled_flag is 0.

[0098] As described above, when sps_max_luma_transform_size_64_flag is 0, since the maximum block size of the orthogonal transform block is limited to 32, the image processing system 11 can avoid an increase in the memory size required when applying the ACT process and the IACT process.

[0099] For example, when such control is not performed, that is, when the maximum block size of the orthogonal transform block can be 64, the storage unit 23 and the storage unit 33 require a memory size capable of storing the YCgCo residual signals for three components with a block size of 64×64. On the other hand, only when the maximum block size of the orthogonal transform block is limited to 32, by performing control to apply the ACT process and the IACT process, the storage unit 23 and the storage unit 33 only need a memory size capable of storing the YCgCo residual signals for three components with a block size of 32×32.

[0100] Therefore, in the image processing system 11, by performing control according to such a third concept, an increase in the memory size can be avoided, and as a result, an increase in the implementation cost can be suppressed.

[0101] <Configuration Example of a Computer-Based System> FIG. 7 is a block diagram showing a configuration example of an embodiment of a computer-based system to which the present technology is applied.

[0102] FIG. 7 is a block diagram showing a configuration example of a network system in which one or more computers, servers, etc. are connected via a network. Note that the hardware and software environment shown in the embodiment of FIG. 7 is shown as an example that can provide a platform for implementing the software and / or method according to the present disclosure.

[0103] As shown in FIG. 7, the network system 101 includes a computer 102, a network 103, a remote computer 104, a web server 105, a cloud storage server 106, and a computer server 107. Here, in the present embodiment, a plurality of instances are executed by one or more of the functional blocks shown in FIG. 7.

[0104] Also, in FIG. 7, the detailed configuration of the computer 102 is illustrated. Note that the functional blocks shown in the computer 102 are illustrated to establish exemplary functions and are not limited to such a configuration. Also, the detailed configurations of the remote computer 104, the web server 105, the cloud storage server 106, and the computer server 107 are not illustrated, but these include configurations similar to the functional blocks shown in the computer 102.

[0105] As the computer 102, a personal computer, a desktop computer, a laptop computer, a tablet computer, a netbook computer, a portable information terminal, a smartphone, or another programmable electronic device capable of communicating with other devices on the network can be used.

[0106] And computer 102 is configured to include bus 111, processor 112, memory 113, non-volatile storage 114, network interface 115, peripheral device interface 116, and display interface 117. Each of these functions may be implemented in certain embodiments in individual electronic subsystems (integrated circuit chips or combinations of chips and associated devices), or in other embodiments, some of the functions may be combined and implemented in a single chip (system-on-chip or SoC (System on Chip)).

[0107] Bus 111 can adopt various proprietary or industry-standard high-speed parallel or serial peripheral interconnect buses.

[0108] Processor 112 can adopt one or more single or multi-chip microprocessors designed and / or manufactured as such.

[0109] Memory 113 and non-volatile storage 114 are storage media readable by computer 102. For example, memory 113 can adopt any suitable volatile storage device such as DRAM (Dynamic Random Access Memory) or SRAM (Static RAM). Non-volatile storage 114 can adopt at least one or more of flexible disks, hard disks, SSDs (Solid State Drives), ROMs (Read Only Memories), EPROMs (Erasable and Programmable Read Only Memories), flash memories, compact disks (CDs or CD-ROMs), DVDs (Digital Versatile Discs), card-type memories, or stick-type memories.

[0110] Also, the non-volatile storage 114 stores a program 121. The program 121 is, for example, a set of machine-readable instructions and / or data used to create, manage, and control certain software functions. In a configuration where the memory 113 is much faster than the non-volatile storage 114, the program 121 can be transferred from the non-volatile storage 114 to the memory 113 before being executed by the processor 112.

[0111] The computer 102 can communicate and interact with other computers via the network 103 through the network interface 115. The network 103 can adopt a configuration including, for example, a LAN (Local Area Network), a WAN (Wide Area Network) such as the Internet, or a combination of a LAN and a WAN, with wired, wireless, or optical fiber connections. Generally, the network 103 consists of any combination of connections and protocols that support communication between two or more computers and related devices.

[0112] The peripheral interface 116 can perform input and output of data with other devices that can be locally connected to the computer 102. For example, the peripheral interface 116 provides a connection to an external device 131. The external device 131 uses a keyboard, a mouse, a keypad, a touch screen, and / or other suitable input devices. The external device 131 can also include, for example, a portable computer-readable storage medium such as a thumb drive, a portable optical or magnetic disk, and a memory card.

[0113] In an embodiment of the present disclosure, for example, software and data used to implement program 121 may be stored in such a portable computer-readable storage medium. In such an embodiment, the software may be directly loaded into non-volatile storage 114 or into memory 113 via peripheral device interface 116. Peripheral device interface 116 may use industry standards such as RS-232 or USB (Universal Serial Bus) for connection to external device 131.

[0114] Display interface 117 can connect computer 102 to display 132 and use display 132 to present a command line or graphical user interface to the user of computer 102. For example, display interface 117 may adopt industry standards such as VGA (Video Graphics Array), DVI (Digital Visual Interface), DisplayPort, HDMI (High-Definition Multimedia Interface) (registered trademark).

[0115] <Configuration example of image encoding device> FIG. 8 shows the configuration of an embodiment of an image encoding device as an image processing device to which the present disclosure is applied.

[0116] The image encoding device 201 shown in FIG. 8 encodes image data using prediction processing. Here, as the encoding method, for example, a VVC (Versatile Video Coding) method, a HEVC (High Efficiency Video Coding) method, or the like is used.

[0117] The image encoding device 201 in FIG. 8 includes an A / D conversion unit 202, a screen rearrangement buffer 203, an arithmetic unit 204, an orthogonal conversion unit 205, a quantization unit 206, a reversible encoding unit 207, and an accumulation buffer 208. Further, the image encoding device 201 includes an inverse quantization unit 209, an inverse orthogonal conversion unit 210, an arithmetic unit 211, a deblocking filter 212, an adaptive offset filter 213, an adaptive loop filter 214, a frame memory 215, a selection unit 216, an intra prediction unit 217, a motion prediction / compensation unit 218, a predicted image selection unit 219, and a rate control unit 220.

[0118] The A / D conversion unit 202 performs A / D conversion on the input image data (Picture(s)) and supplies it to the screen rearrangement buffer 203. Note that, instead of providing the A / D conversion unit 202, a configuration may be adopted in which digital data images are input.

[0119] The screen rearrangement buffer 203 stores the image data supplied from the A / D conversion unit 202 and rearranges the images of the frames in the stored display order into the order of the frames for encoding according to the GOP (Group of Picture) structure. The screen rearrangement buffer 203 outputs the image with the frame order rearranged to the arithmetic unit 204, the intra prediction unit 217, and the motion prediction / compensation unit 218.

[0120] The arithmetic unit 204 subtracts the predicted image supplied from the intra prediction unit 217 or the motion prediction / compensation unit 218 via the predicted image selection unit 219 from the image output from the screen rearrangement buffer 203, and outputs the difference information to the orthogonal conversion unit 205.

[0121] For example, in the case of an image for which intra encoding is performed, the arithmetic unit 204 subtracts the predicted image supplied from the intra prediction unit 217 from the image output from the screen rearrangement buffer 203. Also, for example, in the case of an image for which inter encoding is performed, the arithmetic unit 204 subtracts the predicted image supplied from the motion prediction / compensation unit 218 from the image output from the screen rearrangement buffer 203.

[0122] The orthogonal transformation unit 205 performs an orthogonal transformation such as a discrete cosine transform or a Karhunen - Loeve transform on the differential information supplied from the arithmetic unit 204, and supplies the transformation coefficients to the quantization unit 206.

[0123] The quantization unit 206 quantizes the transformation coefficients output by the orthogonal transformation unit 205. The quantization unit 206 supplies the quantized transformation coefficients to the reversible coding unit 207.

[0124] The reversible coding unit 207 performs reversible coding such as variable - length coding or arithmetic coding on the quantized transformation coefficients.

[0125] The reversible coding unit 207 acquires parameters such as information indicating the intra - prediction mode from the intra - prediction unit 217, and acquires parameters such as information indicating the inter - prediction mode and motion vector information from the motion prediction / compensation unit 218.

[0126] The reversible coding unit 207 encodes the quantized transformation coefficients, encodes each acquired parameter (syntax element), and makes them part of the header information of the encoded data (multiplexes them). The reversible coding unit 207 supplies the encoded data obtained by encoding to the storage buffer 208 for storage.

[0127] For example, in the reversible coding unit 207, reversible coding processes such as variable - length coding or arithmetic coding are performed. Examples of variable - length coding include CAVLC (Context - Adaptive Variable Length Coding). Examples of arithmetic coding include CABAC (Context - Adaptive Binary Arithmetic Coding).

[0128] The storage buffer 208 temporarily holds the encoded stream (Encoded Data) supplied from the reversible coding unit 207, and outputs it as an encoded image at a predetermined timing to, for example, a recording device (not shown) or a transmission line in the subsequent stage. That is, the storage buffer 208 is also a transmission unit that transmits the encoded stream.

[0129] Further, the quantization coefficients quantized in the quantization unit 206 are also supplied to the inverse quantization unit 209. The inverse quantization unit 209 inverse-quantizes the quantized conversion coefficients in a method corresponding to the quantization by the quantization unit 206. The inverse quantization unit 209 supplies the obtained conversion coefficients to the inverse orthogonal transformation unit 210.

[0130] The inverse orthogonal transformation unit 210 inverse-orthogonally transforms the supplied conversion coefficients in a method corresponding to the orthogonal transformation process by the orthogonal transformation unit 205. The inverse-orthogonally transformed output (restored differential information) is supplied to the arithmetic unit 211.

[0131] The arithmetic unit 211 adds the inverse-orthogonal transformation result supplied from the inverse orthogonal transformation unit 210, that is, the restored differential information, to the prediction image supplied from the intra prediction unit 217 or the motion prediction / compensation unit 218 via the prediction image selection unit 219, to obtain a locally decoded image (decoded image).

[0132] For example, when the differential information corresponds to an image for which intra coding is performed, the arithmetic unit 211 adds the prediction image supplied from the intra prediction unit 217 to the differential information. Also, for example, when the differential information corresponds to an image for which inter coding is performed, the arithmetic unit 211 adds the prediction image supplied from the motion prediction / compensation unit 218 to the differential information.

[0133] The decoded image, which is the addition result, is supplied to the deblocking filter 212 and the frame memory 215.

[0134] The deblocking filter 212 suppresses the block distortion of the decoded image by appropriately performing deblocking filter processing on the image from the arithmetic unit 211, and supplies the filter processing result to the adaptive offset filter 213. The deblocking filter 212 has parameters β and Tc obtained based on the quantization parameter QP. The parameters β and Tc are thresholds (parameters) used for determination regarding the deblocking filter.

[0135] Note that the parameters β and Tc of the deblocking filter 212 are extended from β and Tc defined in the HEVC format. Each offset of the parameters β and Tc is encoded in the reversible encoding unit 207 as a parameter of the deblocking filter and transmitted to the image decoder 301 in FIG. 10 described later.

[0136] The adaptive offset filter 213 performs offset filter (SAO: Sample adaptive offset) processing for mainly suppressing ringing on the image after filtering by the deblocking filter 212.

[0137] There are a total of nine types of offset filters, including two types of band offsets, six types of edge offsets, and no offset. The adaptive offset filter 213 uses a quad-tree structure in which the type of offset filter is determined for each divided region and the offset value for each divided region to perform a filtering process on the image after filtering by the deblocking filter 212. The adaptive offset filter 213 supplies the image after the filtering process to the adaptive loop filter 214.

[0138] Note that in the image encoder 201, the quad-tree structure and the offset value for each divided region are calculated and used by the adaptive offset filter 213. The calculated quad-tree structure and the offset value for each divided region are encoded in the reversible encoding unit 207 as adaptive offset parameters and transmitted to the image decoder 301 in FIG. 10 described later.

[0139] The adaptive loop filter 214 performs an adaptive loop filter (ALF: Adaptive Loop Filter) process for each processing unit on the image after filtering by the adaptive offset filter 213 using filter coefficients. In the adaptive loop filter 214, for example, a two-dimensional Wiener filter is used as the filter. Of course, a filter other than the Wiener filter may be used. The adaptive loop filter 214 supplies the filter processing result to the frame memory 215.

[0140] Although not shown in the example of FIG. 8, in the image encoding apparatus 201, the filter coefficients are calculated and used by the adaptive loop filter 214 so as to minimize the residual from the original image from the screen rearrangement buffer 203 for each processing unit. The calculated filter coefficients are encoded in the reversible encoding unit 207 as adaptive loop filter parameters and transmitted to the image decoding apparatus 301 of FIG. 10 described later.

[0141] The frame memory 215 outputs the stored reference image to the intra prediction unit 217 or the motion prediction / compensation unit 218 via the selection unit 216 at a predetermined timing.

[0142] For example, in the case of an image for which intra encoding is performed, the frame memory 215 supplies the reference image to the intra prediction unit 217 via the selection unit 216. Also, for example, in the case of inter encoding, the frame memory 215 supplies the reference image to the motion prediction / compensation unit 218 via the selection unit 216.

[0143] The selection unit 216 supplies the reference image to the intra prediction unit 217 when the reference image supplied from the frame memory 215 is an image for which intra encoding is performed. Also, the selection unit 216 supplies the reference image to the motion prediction / compensation unit 218 when the reference image supplied from the frame memory 215 is an image for which inter encoding is performed.

[0144] The intra prediction unit 217 performs intra prediction (intra-frame prediction) that generates a prediction image using pixel values within the screen. The intra prediction unit 217 performs intra prediction in a plurality of modes (intra prediction modes).

[0145] The intra prediction unit 217 generates prediction images in all intra prediction modes, evaluates each prediction image, and selects an optimal mode. When the intra prediction unit 217 selects an optimal intra prediction mode, it supplies the prediction image generated in that optimal mode to the arithmetic unit 204 and the arithmetic unit 211 via the prediction image selection unit 219.

[0146] Also, as described above, the intra prediction unit 217 appropriately supplies parameters such as intra prediction mode information indicating the adopted intra prediction mode to the reversible coding unit 207.

[0147] The motion prediction / compensation unit 218 performs motion prediction on an image for which inter coding is to be performed, using the input image supplied from the screen rearrangement buffer 203 and the reference image supplied from the frame memory 215 via the selection unit 216. Also, the motion prediction / compensation unit 218 performs motion compensation processing according to the motion vector detected by the motion prediction, and generates a prediction image (inter prediction image information).

[0148] The motion prediction / compensation unit 218 performs inter prediction processing for all candidate inter prediction modes, and generates a prediction image. The motion prediction / compensation unit 218 supplies the generated prediction image to the arithmetic unit 204 and the arithmetic unit 211 via the prediction image selection unit 219. Also, the motion prediction / compensation unit 218 supplies parameters such as inter prediction mode information indicating the adopted inter prediction mode and motion vector information indicating the calculated motion vector to the reversible coding unit 207.

[0149] The prediction image selection unit 219 supplies the output of the intra prediction unit 217 to the arithmetic unit 204 and the arithmetic unit 211 in the case of an image to be intra-coded, and supplies the output of the motion prediction / compensation unit 218 to the arithmetic unit 204 and the arithmetic unit 211 in the case of an image to be inter-coded.

[0150] Based on the compressed image stored in the accumulation buffer 208, the rate control unit 220 controls the rate of the quantization operation of the quantization unit 206 so that overflow or underflow does not occur.

[0151] The image encoding device 201 is configured as described above, and an adaptive color conversion unit 42 (FIG. 2) is provided between the arithmetic unit 204 and the orthogonal conversion unit 205, and an inverse adaptive color conversion unit 47 (FIG. 2) is provided between the inverse orthogonal conversion unit 210 and the arithmetic unit 211. Then, in the image encoding device 201, by performing control according to the first to third concepts described above, an increase in the memory size can be avoided.

[0152] <Operation of Image Encoding Device> Referring to FIG. 9, the flow of the encoding process executed by the image encoding device 201 as described above will be described.

[0153] In step S101, the A / D conversion unit 202 performs A / D conversion on the input image.

[0154] In step S102, the screen rearrangement buffer 203 stores the image A / D-converted by the A / D conversion unit 202, and rearranges the order of display of each picture to the order of encoding.

[0155] When the image to be processed supplied from the screen rearrangement buffer 203 is an image of a block to be intra-processed, the decoded image to be referred to is read from the frame memory 215 and supplied to the intra prediction unit 217 via the selection unit 216.

[0156] Based on these images, in step S103, the intra prediction unit 217 performs intra prediction on the pixels of the block to be processed in all candidate intra prediction modes. Note that, as the decoded pixels to be referred to, pixels that have not been filtered by the deblocking filter 212 are used.

[0157] By this process, intra prediction is performed in all candidate intra prediction modes, and cost function values are calculated for all candidate intra prediction modes. Then, based on the calculated cost function values, the optimal intra prediction mode is selected, and the predicted image generated by the intra prediction of the optimal intra prediction mode and its cost function value are supplied to the predicted image selection unit 219.

[0158] When the image to be processed supplied from the screen rearrangement buffer 203 is an image to be inter-processed, the reference image is read from the frame memory 215 and supplied to the motion prediction / compensation unit 218 via the selection unit 216. Based on these images, in step S104, the motion prediction / compensation unit 218 performs motion prediction / compensation processing.

[0159] By this process, motion prediction processing is performed in all candidate inter prediction modes, cost function values are calculated for all candidate inter prediction modes, and based on the calculated cost function values, the optimal inter prediction mode is determined. Then, the predicted image generated by the optimal inter prediction mode and its cost function value are supplied to the predicted image selection unit 219.

[0160] In step S105, the predicted image selection unit 219 determines one of the optimal intra prediction mode and the optimal inter prediction mode as the optimal prediction mode based on each cost function value output from the intra prediction unit 217 and the motion prediction / compensation unit 218. Then, the predicted image selection unit 219 selects the predicted image of the determined optimal prediction mode and supplies it to the arithmetic units 204 and 211. This predicted image is used in the arithmetic operations of steps S106 and S111 described later.

[0161] Note that the selection information of this predicted image is supplied to the intra prediction unit 217 or the motion prediction / compensation unit 218. When the predicted image of the optimal intra prediction mode is selected, the intra prediction unit 217 supplies information indicating the optimal intra prediction mode (that is, parameters related to intra prediction) to the reversible coding unit 207.

[0162] When the prediction image of the optimal inter prediction mode is selected, the motion prediction / compensation unit 218 outputs information indicating the optimal inter prediction mode and information corresponding to the optimal inter prediction mode (i.e., parameters related to motion prediction) to the reversible coding unit 207. Examples of the information corresponding to the optimal inter prediction mode include motion vector information and reference frame information.

[0163] In step S106, the arithmetic unit 204 calculates the difference between the image rearranged in step S102 and the prediction image selected in step S105. The prediction image is supplied to the arithmetic unit 204 via the prediction image selection unit 219 from the motion prediction / compensation unit 218 in the case of inter prediction and from the intra prediction unit 217 in the case of intra prediction.

[0164] The amount of difference data is smaller than that of the original image data. Therefore, the amount of data can be compressed compared with the case of directly coding the image.

[0165] In step S107, the orthogonal transformation unit 205 orthogonally transforms the difference information supplied from the arithmetic unit 204. Specifically, orthogonal transformations such as discrete cosine transform and Karhunen - Loeve transform are performed, and transformation coefficients are output.

[0166] In step S108, the quantization unit 206 quantizes the transformation coefficients. During this quantization, the rate is controlled as will be described in the process of step S118.

[0167] The difference information quantized as described above is locally decoded as follows. That is, in step S109, the inverse quantization unit 209 inverse - quantizes the transformation coefficients quantized by the quantization unit 206 with characteristics corresponding to the characteristics of the quantization unit 206. In step S110, the inverse orthogonal transformation unit 210 inverse - orthogonally transforms the transformation coefficients inverse - quantized by the inverse quantization unit 209 with characteristics corresponding to the characteristics of the orthogonal transformation unit 205.

[0168] In step S111, the arithmetic unit 211 adds the predicted image input via the predicted image selection unit 219 to the locally decoded differential information to generate a locally decoded (i.e., locally decoded) image (the image corresponding to the input to the arithmetic unit 204).

[0169] In step S112, the deblocking filter 212 performs deblocking filter processing on the image output from the arithmetic unit 211. At this time, as the determination threshold values for the deblocking filter, the parameters β and Tc extended from β and Tc defined in the HEVC method are used. The filtered image from the deblocking filter 212 is output to the adaptive offset filter 213.

[0170] Note that each offset of the parameters β and Tc input by the user operating an operation unit or the like and used in the deblocking filter 212 is supplied to the reversible encoding unit 207 as a parameter of the deblocking filter.

[0171] In step S113, the adaptive offset filter 213 performs adaptive offset filter processing. By this processing, using the quad-tree structure in which the type of offset filter is determined for each divided region and the offset value for each divided region, filter processing is performed on the image filtered by the deblocking filter 212. The filtered image is supplied to the adaptive loop filter 214.

[0172] Note that the determined quad-tree structure and the offset value for each divided region are supplied to the reversible encoding unit 207 as adaptive offset parameters.

[0173] In step S114, the adaptive loop filter 214 performs adaptive loop filtering on the image after filtering by the adaptive offset filter 213. For example, for the image after filtering by the adaptive offset filter 213, filtering is performed on the image for each processing unit using filter coefficients, and the filtering result is supplied to the frame memory 215.

[0174] In step S115, the frame memory 215 stores the filtered image. Note that the frame memory 215 is also supplied with and stores the unfiltered image from the arithmetic unit 211 by the deblocking filter 212, the adaptive offset filter 213, and the adaptive loop filter 214.

[0175] On the other hand, the quantization conversion coefficients in step S108 described above are also supplied to the reversible encoding unit 207. In step S116, the reversible encoding unit 207 encodes the quantization conversion coefficients output from the quantization unit 206 and the supplied respective parameters. That is, the differential image is reversibly encoded and compressed by variable length encoding, arithmetic encoding, or the like. Here, examples of the respective parameters to be encoded include parameters of the deblocking filter, parameters of the adaptive offset filter, parameters of the adaptive loop filter, quantization parameters, motion vector information, reference frame information, prediction mode information, and the like.

[0176] In step S117, the accumulation buffer 208 accumulates the encoded differential image (i.e., the encoded stream) as a compressed image. The compressed image accumulated in the accumulation buffer 208 is appropriately read out and transmitted to the decoding side via the transmission path.

[0177] In step S118, the rate control unit 220 controls the rate of the quantization operation of the quantization unit 206 so that overflow or underflow does not occur based on the compressed image accumulated in the accumulation buffer 208.

[0178] When the process of step S118 ends, the encoding process ends.

[0179] In the encoding process as described above, the ACT process by the adaptive color conversion unit 42 (Fig. 2) is performed between step S106 and step S107, and the IACT process by the inverse adaptive color conversion unit 47 (Fig. 2) is performed between step S110 and step S111. Then, in the encoding process, control regarding the application of the ACT process and the IACT process is performed according to the first to third concepts described above.

[0180] <Configuration example of image decoding device> Fig. 10 shows the configuration of an embodiment of an image decoding device as an image processing device to which the present disclosure is applied. The image decoding device 301 shown in Fig. 10 is a decoding device corresponding to the image encoding device 201 in Fig. 8.

[0181] The encoded stream (Encoded Data) encoded by the image encoding device 201 is transmitted via a predetermined transmission path to the image decoding device 301 corresponding to this image encoding device 201 and is decoded.

[0182] As shown in Fig. 10, the image decoding device 301 includes an accumulation buffer 302, a reversible decoding unit 303, an inverse quantization unit 304, an inverse orthogonal transformation unit 305, an arithmetic unit 306, a deblocking filter 307, an adaptive offset filter 308, an adaptive loop filter 309, a screen rearrangement buffer 310, a D / A conversion unit 311, a frame memory 312, a selection unit 313, an intra prediction unit 314, a motion prediction / compensation unit 315, and a selection unit 316.

[0183] The accumulation buffer 302 is also a receiving unit that receives the transmitted encoded data. The accumulation buffer 302 receives and accumulates the transmitted encoded data. This encoded data is encoded by the image encoding device 201. The reversible decoding unit 303 decodes the encoded data read from the accumulation buffer 302 at a predetermined timing in a method corresponding to the encoding method of the reversible encoding unit 207 in Fig. 8.

[0184] The inverse decoding unit 303 supplies parameters such as information indicating the decoded intra prediction mode to the intra prediction unit 314, and supplies parameters such as information indicating the inter prediction mode and motion vector information to the motion prediction / compensation unit 315. Further, the inverse decoding unit 303 supplies the parameters of the decoded deblocking filter to the deblocking filter 307, and supplies the decoded adaptive offset parameters to the adaptive offset filter 308.

[0185] The inverse quantization unit 304 inverse-quantizes the coefficient data (quantized coefficients) obtained by decoding by the inverse decoding unit 303 in a method corresponding to the quantization method of the quantization unit 206 in FIG. 8. That is, the inverse quantization unit 304 performs inverse quantization of the quantized coefficients in the same method as the inverse quantization unit 209 in FIG. 8 using the quantization parameters supplied from the image encoding device 201.

[0186] The inverse quantization unit 304 supplies the inverse-quantized coefficient data, that is, the orthogonal transform coefficients, to the inverse orthogonal transform unit 305. The inverse orthogonal transform unit 305 inverse-orthogonally transforms the orthogonal transform coefficients in a method corresponding to the orthogonal transform method of the orthogonal transform unit 205 in FIG. 8, and obtains decoded residual data corresponding to the residual data before orthogonal transformation in the image encoding device 201.

[0187] The decoded residual data obtained by inverse-orthogonal transformation is supplied to the arithmetic unit 306. Further, a predicted image is supplied to the arithmetic unit 306 from the intra prediction unit 314 or the motion prediction / compensation unit 315 via the selection unit 316.

[0188] The arithmetic unit 306 adds the decoded residual data and the predicted image to obtain decoded image data corresponding to the image data before the predicted image is subtracted by the arithmetic unit 204 of the image encoding device 201. The arithmetic unit 306 supplies the decoded image data to the deblocking filter 307.

[0189] The deblocking filter 307 performs deblocking filter processing on the image from the arithmetic unit 306 as appropriate to suppress block distortion of the decoded image, and supplies the filter processing result to the adaptive offset filter 308. The deblocking filter 307 is basically configured in the same manner as the deblocking filter 212 in FIG. 8. That is, the deblocking filter 307 has parameters β and Tc obtained based on the quantization parameter. The parameters β and Tc are threshold values used for determination regarding the deblocking filter.

[0190] Note that β and Tc, which are parameters of the deblocking filter 307, are extended from β and Tc defined in the HEVC format. Each offset of the parameters β and Tc of the deblocking filter encoded by the image encoding device 201 is received as a parameter of the deblocking filter in the image decoding device 301, decoded by the reversible decoding unit 303, and used by the deblocking filter 307.

[0191] The adaptive offset filter 308 performs offset filter (SAO) processing mainly for suppressing ringing on the image after filtering by the deblocking filter 307.

[0192] The adaptive offset filter 308 performs filter processing on the image after filtering by the deblocking filter 307 using a quad-tree structure in which the type of offset filter is determined for each divided region and the offset value for each divided region. The adaptive offset filter 308 supplies the image after the filter processing to the adaptive loop filter 309.

[0193] Note that this quad-tree structure and the offset value for each divided area are calculated by the adaptive offset filter 213 of the image encoding device 201, and are encoded and sent as adaptive offset parameters. Then, the quad-tree structure and the offset value for each divided area encoded by the image encoding device 201 are received as adaptive offset parameters in the image decoding device 301, decoded by the reversible decoding unit 303, and used by the adaptive offset filter 308.

[0194] The adaptive loop filter 309 performs filter processing for each processing unit using filter coefficients on the image after filtering by the adaptive offset filter 308, and supplies the filter processing result to the frame memory 312 and the screen rearrangement buffer 310.

[0195] Although not shown in the example of FIG. 10, in the image decoding device 301, the filter coefficients are calculated for each LUC by the adaptive loop filter 214 of the image encoding device 201, and those encoded and sent as adaptive loop filter parameters are decoded and used by the reversible decoding unit 303.

[0196] The screen rearrangement buffer 310 rearranges the image and supplies it to the D / A conversion unit 311. That is, the order of the frames rearranged for the encoding order by the screen rearrangement buffer 203 in FIG. 8 is rearranged to the original display order.

[0197] The D / A conversion unit 311 performs D / A conversion on the image (Decoded Picture(s)) supplied from the screen rearrangement buffer 310 and outputs it to a display (not shown) for display. Note that a configuration may be adopted in which the image is output as digital data without providing the D / A conversion unit 311.

[0198] The output of the adaptive loop filter 309 is further supplied to the frame memory 312.

[0199] The frame memory 312, selection unit 313, intra prediction unit 314, motion prediction / compensation unit 315, and selection unit 316 respectively correspond to the frame memory 215, selection unit 216, intra prediction unit 217, motion prediction / compensation unit 218, and predicted image selection unit 219 of the image encoding device 201.

[0200] The selection unit 313 reads out the inter-processed image and the reference image from the frame memory 312 and supplies them to the motion prediction / compensation unit 315. Also, the selection unit 313 reads out the image used for intra prediction from the frame memory 312 and supplies it to the intra prediction unit 314.

[0201] To the intra prediction unit 314, information indicating the intra prediction mode obtained by decoding the header information and the like is appropriately supplied from the reversible decoding unit 303. Based on this information, the intra prediction unit 314 generates a predicted image from the reference image acquired from the frame memory 312 and supplies the generated predicted image to the selection unit 316.

[0202] To the motion prediction / compensation unit 315, information (prediction mode information, motion vector information, reference frame information, flag, and various parameters, etc.) obtained by decoding the header information is supplied from the reversible decoding unit 303.

[0203] Based on the information supplied from the reversible decoding unit 303, the motion prediction / compensation unit 315 generates a predicted image from the reference image acquired from the frame memory 312 and supplies the generated predicted image to the selection unit 316.

[0204] The selection unit 316 selects the predicted image generated by the motion prediction / compensation unit 315 or the intra prediction unit 314 and supplies it to the arithmetic unit 306.

[0205] The image decoding device 301 is configured in this way, and an inverse adaptive color conversion unit 64 (FIG. 3) is provided between the inverse orthogonal transformation unit 305 and the arithmetic unit 306. Then, in the image decoding device 301, by performing control according to the above-described first to third concepts, an increase in the memory size can be avoided.

[0206] <Operation of Image Decoding Device> Referring to FIG. 11, an example of the flow of the decoding process executed by the image decoding device 301 as described above will be described.

[0207] When the decoding process is started, in step S201, the accumulation buffer 302 receives and accumulates the transmitted encoded stream (data). In step S202, the reversible decoding unit 303 decodes the encoded data supplied from the accumulation buffer 302. The I picture, P picture, and B picture encoded by the reversible encoding unit 207 in FIG. 8 are decoded.

[0208] Prior to the decoding of the picture, information on parameters such as motion vector information, reference frame information, and prediction mode information (intra prediction mode or inter prediction mode) is also decoded.

[0209] When the prediction mode information is intra prediction mode information, the prediction mode information is supplied to the intra prediction unit 314. When the prediction mode information is inter prediction mode information, the prediction mode information and the corresponding motion vector information, etc. are supplied to the motion prediction / compensation unit 315. Also, the parameters of the deblocking filter and the adaptive offset parameters are decoded and supplied to the deblocking filter 307 and the adaptive offset filter 308, respectively.

[0210] In step S203, the intra prediction unit 314 or the motion prediction / compensation unit 315 performs prediction image generation processing corresponding to the prediction mode information supplied from the reversible decoding unit 303, respectively.

[0211] That is, when intra prediction mode information is supplied from the reversible decoding unit 303, the intra prediction unit 314 generates an intra prediction image in the intra prediction mode. When inter prediction mode information is supplied from the reversible decoding unit 303, the motion prediction / compensation unit 315 performs motion prediction / compensation processing in the inter prediction mode and generates an inter prediction image.

[0212] Through this process, the predicted image (intra-predicted image) generated by the intra prediction unit 314 or the predicted image (inter-predicted image) generated by the motion prediction / compensation unit 315 is supplied to the selection unit 316.

[0213] In step S204, the selection unit 316 selects a predicted image. That is, the predicted image generated by the intra prediction unit 314 or the predicted image generated by the motion prediction / compensation unit 315 is supplied. Therefore, the supplied predicted image is selected and supplied to the arithmetic unit 306, and is added to the output of the inverse orthogonal transformation unit 305 in step S207 described later.

[0214] In the above-described step S202, the transformation coefficients decoded by the reversible decoding unit 303 are also supplied to the inverse quantization unit 304. In step S205, the inverse quantization unit 304 inverse-quantizes the transformation coefficients decoded by the reversible decoding unit 303 with characteristics corresponding to the characteristics of the quantization unit 206 in FIG. 8.

[0215] In step S206, the inverse orthogonal transformation unit 305 inverse-orthogonally transforms the transformation coefficients inverse-quantized by the inverse quantization unit 304 with characteristics corresponding to the characteristics of the orthogonal transformation unit 205 in FIG. 8. As a result, the differential information corresponding to the input of the orthogonal transformation unit 205 (output of the arithmetic unit 204) in FIG. 8 is decoded.

[0216] In step S207, the arithmetic unit 306 adds the predicted image selected in the process of step S204 described above and input via the selection unit 316 to the differential information. Thereby, the original image is decoded.

[0217] In step S208, the deblocking filter 307 performs deblocking filter processing on the image output from the arithmetic unit 306. At this time, as the determination threshold values for the deblocking filter, the parameters β and Tc extended from β and Tc defined in the HEVC format are used. The image after filtering by the deblocking filter 307 is output to the adaptive offset filter 308. In addition, in the deblocking filter processing, the offset values of the parameters β and Tc of the deblocking filter supplied from the reversible decoding unit 303 are also used.

[0218] In step S209, the adaptive offset filter 308 performs adaptive offset filter processing. By this processing, using the quad-tree structure in which the type of offset filter is determined for each divided region and the offset value for each divided region, filter processing is performed on the image after filtering by the deblocking filter 307. The image after filtering is supplied to the adaptive loop filter 309.

[0219] In step S210, the adaptive loop filter 309 performs adaptive loop filter processing on the image after filtering by the adaptive offset filter 308. The adaptive loop filter 309 performs filter processing on the input image for each processing unit using the filter coefficients calculated for each processing unit, and supplies the filter processing result to the screen rearrangement buffer 310 and the frame memory 312.

[0220] In step S211, the frame memory 312 stores the filtered image.

[0221] In step S212, after rearranging the image after the adaptive loop filter 309, the screen rearrangement buffer 310 supplies it to the D / A conversion unit 311. That is, the order of the frames rearranged for encoding by the screen rearrangement buffer 203 of the image encoding device 201 is rearranged to the original display order.

[0222] In step S213, the D / A conversion unit 311 D / A-converts the image rearranged in the screen rearrangement buffer 310 and outputs it to a display (not shown), and the image is displayed.

[0223] When the process of step S213 ends, the decoding process ends.

[0224] In the decoding process as described above, IACT processing by the inverse adaptive color conversion unit 64 (FIG. 3) is performed between step S206 and step S207. And in the decoding process, control regarding the application of ACT processing and IACT processing is performed according to the first to third concepts described above.

[0225] <Configuration example of computer> Next, the above-described series of processes (image processing method) can be performed by hardware or by software. When the series of processes are performed by software, the program constituting the software is installed in a general-purpose computer or the like.

[0226] FIG. 12 is a block diagram showing a configuration example of an embodiment of a computer in which a program for executing the above-described series of processes is installed.

[0227] The program can be pre-recorded in a hard disk 1005 or a ROM 1003 as a recording medium built in the computer.

[0228] Alternatively, the program can be stored (recorded) in a removable recording medium 1011 driven by a drive 1009. Such a removable recording medium 1011 can be provided as so-called package software. Here, examples of the removable recording medium 1011 include a flexible disk, a CD-ROM (Compact Disc Read Only Memory), an MO (Magneto Optical) disk, a DVD (Digital Versatile Disc), a magnetic disk, a semiconductor memory, and the like.

[0229] In addition to being installed on the computer from the removable recording medium 1011 as described above, the program can also be downloaded to the computer via a communication network or a broadcast network and installed on the built-in hard disk 1005. That is, the program can be wirelessly transferred to the computer from, for example, a download site via an artificial satellite for digital satellite broadcasting, or can be wiredly transferred to the computer via a network such as a LAN (Local Area Network) or the Internet.

[0230] The computer incorporates a CPU (Central Processing Unit) 1002, and an input / output interface 1010 is connected to the CPU 1002 via a bus 1001.

[0231] When a command is input by the user operating the input unit 1007 etc. via the input / output interface 1010, the CPU 1002 executes the program stored in the ROM (Read Only Memory) 1003 accordingly. Alternatively, the CPU 1002 loads and executes the program stored in the hard disk 1005 into the RAM (Random Access Memory) 1004.

[0232] Thereby, the CPU 1002 performs the processing according to the above-described flowchart or the processing performed according to the configuration of the above-described block diagram. And the CPU 1002 outputs the processing result from the output unit 1006 via, for example, the input / output interface 1010 as needed, or transmits it from the communication unit 1008, and further records it on the hard disk 1005 etc.

[0233] Note that the input unit 1007 is composed of a keyboard, a mouse, a microphone, etc. Also, the output unit 1006 is composed of an LCD (Liquid Crystal Display), a speaker, etc.

[0234] Here, in this specification, the processing performed by a computer according to a program does not necessarily have to be performed in time series in the order described as a flowchart. That is, the processing performed by a computer according to a program also includes processing that is executed in parallel or individually (for example, parallel processing or object-based processing).

[0235] Also, the program may be processed by one computer (processor) or may be distributedly processed by a plurality of computers. Furthermore, the program may be transferred to a remote computer for execution.

[0236] Furthermore, in this specification, a system means a collection of a plurality of components (devices, modules (parts), etc.), and it does not matter whether all the components are in the same housing. Therefore, a plurality of devices housed in separate housings and connected via a network, and one device in which a plurality of modules are housed in one housing are both systems.

[0237] Also, for example, the configuration described as one device (or processing unit) may be divided and configured as a plurality of devices (or processing units). Conversely, the configuration described as a plurality of devices (or processing units) above may be combined and configured as one device (or processing unit). Of course, a configuration other than those described above may be added to the configuration of each device (or each processing unit). Furthermore, if the configuration and operation of the entire system are substantially the same, a part of the configuration of a certain device (or processing unit) may be included in the configuration of another device (or another processing unit).

[0238] Also, for example, this technology can take a configuration of cloud computing in which one function is shared and jointly processed by a plurality of devices via a network.

[0239] Also, for example, the above-described program can be executed on any device. In that case, it suffices if the device has the necessary functions (such as function blocks) and can obtain the necessary information.

[0240] Also, for example, each step described in the above flowchart can be executed by one device or can be shared and executed by a plurality of devices. Further, when a plurality of processes are included in one step, the plurality of processes included in that one step can be executed by one device or can be shared and executed by a plurality of devices. In other words, the plurality of processes included in one step can also be executed as the processes of a plurality of steps. Conversely, the processes described as a plurality of steps can also be executed together as one step.

[0241] Note that the program executed by the computer may be such that the processing of the steps of describing the program is executed in time series along the order described in this specification, or may be executed in parallel, or may be executed individually at a necessary timing such as when a call is made. That is, as long as there is no contradiction, the processing of each step may be executed in an order different from the above-described order. Further, the processing of the steps of describing this program may be executed in parallel with the processing of another program, or may be executed in combination with the processing of another program.

[0242] Note that the present technology described a plurality of times in this specification can be implemented independently and individually as long as there is no contradiction. Of course, any plurality of the present technologies can also be implemented in combination. For example, a part or all of the present technology described in any one of the embodiments can also be implemented in combination with a part or all of the present technology described in other embodiments. Also, a part or all of any of the above-described present technologies can also be implemented in combination with other technologies not described above.

[0243] <Example of configuration combination> Note that the present technology can also have the following configuration. (1) An adaptive color conversion unit that performs an adaptive color conversion process for adaptively converting the color space of the image to be symbolized on the residual signal of the image, and a orthogonal conversion unit that performs an orthogonal conversion process for each orthogonal conversion block that is a processing unit on the residual signal of the image or on the residual signal of the image on which the adaptive color conversion process has been performed, and a control unit that controls the application of the adaptive color conversion process An image processing apparatus comprising the above. (2) As the maximum block size of the orthogonal conversion block, a first block size and a second block size larger than the first block size are defined, When the control unit uses the first block size as the maximum block size of the orthogonal conversion block, the control unit controls the adaptive color conversion unit to apply the adaptive color conversion process The image processing apparatus according to (1) above. (3) The first block size is 32, and the second block size is 64, The control unit restricts the case where 32 is used as the maximum block size of the orthogonal conversion block and causes the adaptive color conversion process to be applied The image processing apparatus according to (2) above. (4) When sps_max_luma_transform_size_64_flag included in the high-level syntax parameter set is 0, the control unit transmits sps_act_enabled_flag indicating that the adaptive color conversion process is to be applied to the decoding side The image processing apparatus according to (3) above. (5) When a predetermined restriction is provided for the encoding block when encoding the image, the control unit controls the adaptive color conversion unit to apply the adaptive color conversion process The image processing apparatus according to any one of (1) to (4) above. (6) When the block size of an encoding block, which is a processing unit when encoding the image, is restricted to a predetermined size or less, the control unit causes the adaptive color conversion process to be applied. The image processing apparatus according to (5) above. (7) When the block size of the encoding block is restricted to 16×16 or less, the control unit causes the adaptive color conversion process to be applied. The image processing apparatus according to (6) above. (8) When the block size of an encoding block, which is a processing unit when encoding the image, is larger than a predetermined size, the control unit controls the orthogonal transformation unit to perform the orthogonal transformation process using the orthogonal transformation block having a small block size obtained by dividing the encoding block, and controls the adaptive color conversion unit to apply the adaptive color conversion process. The image processing apparatus according to any one of (1) to (7) above. (9) When the block size of the encoding block is 64×64, the control unit causes the orthogonal transformation process to be performed with the orthogonal transformation block being 32×32, and causes the adaptive color conversion process to be applied. The image processing apparatus according to (8) above. (10) An inverse orthogonal transformation unit that acquires the residual signal by performing an inverse orthogonal transformation process for each orthogonal transformation block on the transformation coefficients obtained by performing the orthogonal transformation process, and An inverse adaptive color conversion unit that performs an inverse adaptive color conversion process for adaptively inverse-converting the color space of the image on the residual signal acquired by the inverse orthogonal transformation unit further comprising The control unit performs control regarding the application of the inverse adaptive color conversion process in correspondence with the adaptive color conversion process. The image processing apparatus according to any one of (1) to (9) above. (11) Performing an adaptive color conversion process for adaptively converting the color space of the image to be encoded on the residual signal of the image. Performing orthogonal transformation processing for each orthogonal transformation block serving as a processing unit on the residual signal of the image or on the residual signal of the image subjected to the adaptive color conversion processing; Performing control regarding application of the adaptive color conversion processing; An image processing method including the above. (12) An inverse orthogonal transformation unit that obtains the residual signal by performing inverse orthogonal transformation processing for each orthogonal transformation block serving as a processing unit on the transformation coefficients obtained by performing orthogonal transformation processing on the residual signal of the image to be decoded on the encoding side; An inverse adaptive color conversion unit that performs inverse adaptive color conversion processing for adaptively inverse-converting the color space of the image on the residual signal; A control unit that performs control regarding application of the inverse adaptive color conversion processing; An image processing apparatus including the above. (13) Obtaining the residual signal by performing inverse orthogonal transformation processing for each orthogonal transformation block serving as a processing unit on the transformation coefficients obtained by performing orthogonal transformation processing on the residual signal of the image to be decoded on the encoding side; Performing inverse adaptive color conversion processing for adaptively inverse-converting the color space of the image on the residual signal; Performing control regarding application of the inverse adaptive color conversion processing; An image processing method including the above.

[0244] Note that the present embodiment is not limited to the above-described embodiment, and various modifications are possible without departing from the gist of the present disclosure. Also, the effects described in this specification are merely examples and are not limiting, and there may be other effects.

Explanation of Signs

[0245] 11 Image processing system, 12 Image encoding device, 13 Image decoding device, 21 Prediction unit, 22 Encoding unit, 23 Memory unit, 24 Control unit, 31 Prediction unit, 32 Decoding unit, 33 Memory unit, 34 Control unit, 41 Arithmetic unit, 42 Adaptive color conversion unit, 43 Orthogonal transformation unit, 44 Quantization unit, 45 Inverse quantization unit, 46 Inverse orthogonal transformation unit, 47 Inverse adaptive color conversion unit, 48 Arithmetic unit, 49 Prediction unit, 50 Encoding unit, 61 Decoding unit, 62 Inverse quantization unit, 63 Inverse orthogonal transformation unit, 64 Inverse adaptive color conversion unit, 65 Arithmetic unit, 66 Prediction unit

Claims

1. a parsing unit for parsing size identification data from the bit stream, the size identification data identifying whether a maximum block size of an orthogonal transform block is 64 or 32; an inverse adaptive color transform unit that applies the inverse adaptive color transform processing to a residual signal generated by applying an inverse orthogonal transform processing to transform coefficients obtained by decoding the bitstream for each orthogonal transform block serving as a processing unit, in accordance with identification data that identifies whether or not an adaptive color transform processing that adaptively transforms a color space of an image or an inverse adaptive color transform processing that adaptively inverse transforms a color space of an image is enabled only when the size identification data parsed by the parser indicates that 32 is applied as a maximum block size of an orthogonal transform block; An image processing device comprising:

2. The size identification data is set as a sequence parameter set of the bitstream. The image processing device according to claim 1 .

3. The size identification data is set as the syntax of sps_max_luma_transform_size_64_flag. The image processing device according to claim 2 .

4. The identification data is referenced only when the value of the sps_max_luma_transform_size_64_flag is set to 0. The image processing device according to claim 3 .

5. an arithmetic decoding unit that arithmetically decodes the bitstream to generate quantized transform coefficients; an inverse quantization unit that inverse quantizes the quantized transform coefficients generated by the arithmetic decoding unit to generate the transform coefficients; The image processing device according to claim 1 , further comprising:

6. parsing size identification data from the bitstream that identifies whether a maximum block size of the orthogonal transform block is 64 or 32; only when the size identification data indicates that 32 is to be applied as the maximum block size of the orthogonal transform block, applying the inverse adaptive color transform processing to a residual signal generated by applying an inverse orthogonal transform processing to transform coefficients obtained by decoding the bit stream for each orthogonal transform block serving as a processing unit, in accordance with identification data for identifying whether or not an adaptive color transform processing for adaptively transforming a color space of an image or an inverse adaptive color transform processing for adaptively inverse transforming the color space of an image is Enable; An image processing method comprising:

7. a setting unit which sets identification data for identifying whether an adaptive color conversion process for adaptively converting a color space of an image or an inverse adaptive color conversion process for adaptively inversely converting a color space of an image is enabled only when size identification data for identifying whether a maximum block size of an orthogonal transformation block is 64 or 32 is set to be 32 as the maximum block size of the orthogonal transformation block; an encoding unit that encodes the image to generate a bit stream including the identification data set by the setting unit; An image processing device comprising:

8. The setting unit sets the size identification data as a sequence parameter set of the bit stream. The image processing device according to claim 7.

9. The setting unit sets the size identification data as a syntax of sps_max_luma_transform_size_64_flag. The image processing device according to claim 8.

10. The setting unit sets the identification data only when the value of the sps_max_luma_transform_size_64_flag is set to 0. The image processing device according to claim 9 .

11. an orthogonal transform unit that applies an orthogonal transform process to a residual signal of the image for each orthogonal transform block that is a processing unit; a quantization unit that quantizes the transform coefficients generated by the orthogonal transform unit to generate quantized transform coefficients; an arithmetic coding unit that arithmetically codes the quantized transform coefficients generated by the quantization unit to generate the bit stream; The image processing device according to claim 7 , further comprising:

12. Only when size identification data for identifying whether the maximum block size of the orthogonal transformation block is 64 or 32 is set to be 32 as the maximum block size of the orthogonal transformation block, identification data for identifying whether an adaptive color transformation process for adaptively transforming a color space of an image or an inverse adaptive color transformation process for adaptively inverse transforming a color space of an image is enabled is set; encoding said image to generate a bitstream including said identification data; An image processing method comprising: