Image processing device and image processing method
By controlling the adaptive color conversion unit and the inverse adaptive color conversion unit, the problem of increased storage capacity in the adaptive color conversion process is solved, and the storage and installation costs of the image processing device are reduced.
Patent Information
- Application Number
- CN202080086941.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Priority Date
- 2019-12-27
- Filing Date
- 2020-12-18
- Publication Date
- 2025-09-23
- Estimated Expiration
- 2040-12-18
AI Technical Summary
When the adaptive color conversion process is applied, the YCgCo residual signal needs to be temporarily accumulated, resulting in an increase in the amount of storage, which increases the storage requirements and installation costs of the image processing device.
Color space conversion and inverse conversion of an image are adaptively performed by an adaptive color conversion unit and an inverse adaptive color conversion unit, and a controller controls the application of these processes to avoid an increase in storage capacity.
By controlling the storage capacity of the storage unit, an increase in storage capacity is avoided and the installation cost of the image processing device is reduced.
Smart Images

Figure CN114830671B_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to an image processing apparatus and an image processing method, and more particularly, to an image processing apparatus and an image processing method that can avoid an increase in memory amount. Background Art
[0002] Devices are being consistently developed that digitally process image information and, at the same time, compress and encode images for efficient transmission and accumulation of information by using redundancy specific to the image information by employing coding methods that perform coding by orthogonal transforms such as discrete cosine transforms and motion compensation.
[0003] Examples of encoding methods include Moving Picture Experts Group (MPEG), H.264 and MPEG-4 Part 10 (Advanced Video Coding, hereinafter referred to as H.264 / AVC), and H.265 and MPEG-H Part 2 (High Efficiency Video Coding, hereinafter referred to as H.265 / HEVC).
[0004] Furthermore, in order to further improve the coding efficiency of Advanced Video Coding (AVC), High Efficiency Video Coding (HEVC), and the like, standardization of a coding scheme called Versatile Video Coding (VVC) is underway (see support of embodiments described later).
[0005] As disclosed in Non-Patent Document 1, in VVC, a technology related to Adaptive Color Transform (ACT) that adaptively performs conversion of the color space of an image is disclosed.
[0006] Reference List
[0007] Non-patent literature
[0008] Non-patent literature 1: Xiaoyu Xiu, Yi-Wen Chen, Tsung-Chuan Ma, Hong-Jheng Jhu, Xianglin Wang, Support of adaptive color transform for 444 video coding in VVC, JVET-P0517_r1 (3rd edition - dated October 11, 2019). Summary of the Invention
[0009] Problems to be solved by the present invention
[0010] In addition, when applying the ACT process to convert the RGB color space into the YCgCo color space, for example, it is necessary to temporarily accumulate the YCgCo residual signal output as a result of the process. Therefore, it is considered necessary to increase the amount of memory used to accumulate the YCgCo residual signal according to the block size of the orthogonal transform block in the orthogonal transform process performed after the ACT process.
[0011] The present disclosure has been made in view of such circumstances, and can avoid an increase in storage capacity.
[0012] Solution to the problem
[0013] An image processing device according to a first aspect of the present disclosure includes: an adaptive color conversion unit that performs adaptive color conversion processing on a residual signal of an image to be encoded, the adaptive color conversion processing adaptively performing conversion of a color space of the image; an orthogonal transform unit that performs orthogonal transform processing on the residual signal of the image or on the residual signal of the image subjected to the adaptive color conversion processing for each orthogonal transform block in an orthogonal transform block serving as a processing unit; and a controller that performs control related to application of the adaptive color conversion processing.
[0014] An image processing method according to a first aspect of the present disclosure includes: performing adaptive color conversion processing on a residual signal of an image to be encoded, the adaptive color conversion processing adaptively performing conversion of a color space of the image; performing orthogonal transform processing on the residual signal of the image or on the residual signal of the image subjected to the adaptive color conversion processing for each orthogonal transform block in an orthogonal transform block as a processing unit; and performing control related to application of the adaptive color conversion processing.
[0015] In a first aspect of the present disclosure, adaptive color conversion processing is performed on a residual signal of an image to be encoded, the adaptive color conversion processing adaptively performing conversion of a color space of the image; orthogonal transform processing is performed on the residual signal of the image or on the residual signal of the image subjected to the adaptive color conversion processing for each of the orthogonal transform blocks serving as processing units; and control related to application of the adaptive color conversion processing is performed.
[0016] An image processing device according to a second aspect of the present disclosure includes: an inverse orthogonal transform unit that obtains a residual signal of an image to be decoded by performing, for each of orthogonal transform blocks serving as a processing unit, an inverse orthogonal transform process on a transform coefficient obtained when an orthogonal transform process is performed on the residual signal on the encoding side; an inverse adaptive color conversion unit that performs, on the residual signal, an inverse adaptive color conversion process that adaptively performs inverse conversion of a color space of the image; and a controller that performs control related to application of the inverse adaptive color conversion process.
[0017] The image processing method of the second aspect of the present disclosure includes: obtaining a residual signal of an image to be decoded by: performing inverse orthogonal transform processing on transform coefficients obtained when orthogonal transform processing is performed on the residual signal on the encoding side for each orthogonal transform block in the orthogonal transform blocks serving as processing units; performing inverse adaptive color conversion processing on the residual signal, the inverse adaptive color conversion processing adaptively performing inverse conversion of the color space of the image; and performing control related to application of the inverse adaptive color conversion processing.
[0018] In a second aspect of the present disclosure, for each of the orthogonal transform blocks serving as processing units, inverse orthogonal transform processing is performed on transform coefficients obtained when orthogonal transform processing is performed on a residual signal of an image to be decoded on the encoding side, thereby obtaining a residual signal; inverse adaptive color conversion processing is performed on the residual signal, the inverse adaptive color conversion processing adaptively performing inverse conversion of a color space of the image; and control related to application of the inverse adaptive color conversion processing is performed. BRIEF DESCRIPTION OF THE DRAWINGS
[0019] Figure 1 is a block diagram showing a configuration example of an embodiment of an image processing system to which the present technology is applied.
[0020] Figure 2 is a diagram showing a configuration example of an image encoding device.
[0021] Figure 3 is a diagram showing a configuration example of an image decoding device.
[0022] Figure 4 is a diagram illustrating an example of a parameter set for a high-level syntax.
[0023] Figure 5 is a diagram illustrating an example of an encoding unit of a high-level syntax.
[0024] Figure 6 is a diagram illustrating an example of a parameter set for a high-level syntax.
[0025] Figure 7is a block diagram illustrating a configuration example of an embodiment of a computer-based system to which the present technology is applied.
[0026] Figure 8 is a block diagram showing a configuration example of an embodiment of an image encoding device.
[0027] Figure 9 is a flowchart describing the encoding process.
[0028] Figure 10 is a block diagram showing a configuration example of an embodiment of an image decoding device.
[0029] Figure 11 is a flowchart describing the decoding process.
[0030] Figure 12 is a block diagram illustrating a configuration example of an embodiment of a computer to which the present technology is applied. DETAILED DESCRIPTION
[0031] <Documents supporting technical content and technical terminology, etc.>
[0032] The scope of the disclosure in this specification is not limited to the contents of the embodiments, and the contents of the following references REF1 to REF5, which were known at the time of filing, are also incorporated herein by reference. That is, the contents described in references REF1 to REF5 are also the basis for determining supporting requirements. In addition, the documents mentioned in references REF1 to REF5 are also the basis for determining supporting requirements.
[0033] For example, even if a quadtree block structure, a quadtree plus binary tree (QTBT), a block structure, a multi-type tree (MTT) block structure, etc. are not directly defined in the detailed description of the present invention, the above block structures are also within the scope of the present disclosure and should meet the support requirements of the claims. In addition, similarly, even if technical terms such as parsing, grammar, semantics, etc. are not directly defined in the detailed description of the present invention, the above technical terms are also within the scope of the present disclosure and should meet the support requirements of the claims. In addition, equivalently, even if technical applications such as adaptive color transform (ACT) are not directly defined in the detailed description of the present invention, the above technical applications are also within the scope of the present disclosure and should meet the support requirements of the claims.
[0034] REF1: Recommendation ITU-T H.264 (04 / 2017) “Advanced video coding for generic audiovisual services”, April 2017;
[0035] REF2: Recommendation ITU-T H.265 (02 / 2018) “High efficiency video coding”, February 2018;
[0036] REF3: Benjamin Bross, Jianle Chen, Shan Liu, Versatile Video Coding (Draft7), JVET-P2001-v14 (Version 14 - Date November 14, 2019);
[0037] REF4: Jianle Chen, Yan Ye, Seung Hwan Kim, Algorithm description for Versatile Video Coding and Test Model 7 (VTM 7), JVET-P2002-v1 (Version 1 - Date November 10, 2019);
[0038] REF5: Xiaoyu Xiu, Yi-Wen Chen, Tsung-Chuan Ma, Hong-Jheng Jhu, Xianglin Wang, Support of adaptive color transform for 444 video coding in VVC, JVET-P0517_r1 (version 3 - dated October 11, 2019).
[0039] <Term>
[0040] In this application, the following terms are defined as follows.
[0041] <block>
[0042] Unless otherwise specified, a “block” (not a block indicating a processing unit) used to describe a partial region or unit of processing as an image (picture) indicates any partial region in a picture and is not limited in size, shape, characteristics, etc. For example, a “block” includes any partial region (processing unit) such as a transform block (TB), a transform unit (TU), a prediction block (PB), a prediction unit (PU), a smallest coding unit (SCU), a coding unit (CU), a largest coding unit (LCU), a coding tree block (CTB), a coding tree unit (CTU), a transform block, a subblock, a macroblock, a tile, or a slice.
[0043] <Description of Block Size>
[0044] In addition, in the description of such a block size, the block size can be specified not only directly but also indirectly. For example, the block size can be specified by using identification information for identifying the size. In addition, for example, the block size can be specified by a ratio or difference with the size of a reference block (e.g., LCU, SCU, etc.). For example, in the case where information for specifying the block size is sent as a syntax element, etc., information for indirectly specifying the size as described above can be used as the information. In this way, the amount of information can be reduced, and in some cases, the coding efficiency can be improved. In addition, the description of the block size also includes a description of the block size range (e.g., a description of the allowed block size range, etc.).
[0045] <Unit of Information and Processing>
[0046] The data units for which various types of information are set and the data units for which various types of processing are targeted are arbitrary and are not limited to the examples above. For example, these information and processing can each be set for each transform unit (TU), transform block (TB), prediction unit (PU), prediction block (PB), coding unit (CU), maximum coding unit (LCU), sub-block, block, tile, slice, picture, sequence or component, or can be set for the data in these data units. Of course, a data unit can be set for each information or processing, and it is not necessary to unify the data units for all information and processing. Note that the storage location of this information is arbitrary and can be stored in the headers, parameter sets, etc. of the above-mentioned data units. In addition, these can be stored in multiple locations.
[0047] <Control Information>
[0048] Control information related to this technology can be sent from the encoding side to the decoding side. For example, the following control information (e.g., enabled_flag) can be sent, which controls whether to enable (or disable) the application of the above-mentioned technology. In addition, for example, the following control information can be sent, which indicates the objects to which this technology is applied (or the objects to which this technology is not applied). For example, the following control information can be sent, which specifies the block size (upper limit, lower limit, or both), frame, component, layer, etc. to which this technology is applied (or to which application is enabled or disabled).
[0049] <Logo>
[0050] Note that in this specification, a "flag" is information for identifying a plurality of states, and includes not only information for identifying two states of true (1) or false (0), but also information capable of identifying three or more states. Therefore, the value that can be adopted by the "flag" may be two values such as 1 / 0, or three or more values. That is, the number of bits constituting the "flag" is arbitrary, and may be 1 bit or more bits. Furthermore, it is assumed that identification information (including a flag) includes not only identification information in a bit stream, but also different information of identification information relative to a certain reference information in the bit stream, so that in this specification, "flag" and "identification information" include not only information but also different information relative to the reference information.
[0051] <associated metadata>
[0052] In addition, various types of information (metadata, etc.) about the encoded data (bitstream) can be transmitted or recorded in any form as long as the information is associated with the encoded data. Here, the term "associated" means, for example, that when one data is processed, another data is made available (linkable). That is, data associated with each other can be collected as one data, or can be a single data. For example, information associated with the encoded data (image) can be transmitted on a transmission path different from the transmission path of the encoded data (image). In addition, for example, information associated with the encoded data (image) can be recorded in a recording medium different from the recording medium of the encoded data (image) (or in a different recording area of the same recording medium). Note that this "association" can be a part of the data, not the entire data. For example, an image and information corresponding to the image can be associated with each other in arbitrary units such as multiple frames, one frame, or a part within a frame.
[0053] Note that in this specification, the terms "combine", "multiplex", "add", "integrate", "include", "store", "put in", "surround", "insert", etc. refer to combining multiple objects into one, for example, combining coded data and metadata into one, and the term refers to a method of "associating" described above. In addition, in this specification, encoding includes not only the entire process of converting an image into a bit stream, but also a part of the process. For example, it includes not only a process including prediction processing, orthogonal transform, quantization, arithmetic coding, etc., but also a process collectively referring to quantization and arithmetic coding, a process including prediction processing, quantization and arithmetic coding, etc. Similarly, decoding includes not only the entire process of converting a bit stream into an image, but also a part of the process. For example, it includes not only a process including inverse arithmetic decoding, inverse quantization, inverse orthogonal transform, prediction processing, etc., but also a process including inverse arithmetic decoding and inverse quantization, a process including inverse arithmetic decoding, inverse quantization and prediction processing, etc.
[0054] The prediction block refers to a block that serves as a processing unit when performing inter-frame prediction, and also includes subblocks in the prediction block. In addition, when the processing unit coincides with the orthogonal transform block that serves as the processing unit when performing orthogonal transform and the coding block that serves as the processing unit when performing encoding, the prediction block, the orthogonal transform block, and the coding block refer to the same block.
[0055] Inter-frame prediction is a general term for processes involving prediction between frames (prediction blocks), such as deriving motion vectors through motion detection (motion prediction / motion estimation) and motion compensation using motion vectors, and includes some processes (e.g., motion compensation only) or all types of processes for generating predicted images (e.g., motion detection + motion compensation). Inter-frame prediction mode refers inclusively to variables (parameters) mentioned when deriving an inter-frame prediction mode, such as a mode number, an index of the mode number, the block size of a prediction block, and the size of a subblock as a processing unit in a prediction block when performing inter-frame prediction.
[0056] In the present disclosure, identification data for identifying multiple patterns may also be provided as syntax for a bitstream. In this case, the decoder can perform processing more efficiently by parsing and referring to the identification data. Methods (data) for identifying block sizes include not only methods for digitizing the block size itself (bit conversion), but also methods (data) for identifying different values (maximum block size, minimum block size, etc.) relative to a block size serving as a reference.
[0057] Hereinafter, specific embodiments to which the present technology is applied will be described in detail with reference to the accompanying drawings.
[0058] <Configuration Example of Image Processing System>
[0059] Figure 1 is a block diagram showing a configuration example of an embodiment of an image processing system to which the present technology is applied.
[0060] like Figure 1 As shown in FIG, the image processing system 11 includes an image encoding device 12 and an image decoding device 13. For example, in the image processing system 11, an image input to the image encoding device 12 is encoded, a bit stream obtained by the encoding is sent to the image decoding device 13, and a decoded image decoded from the bit stream in the image decoding device 13 is output.
[0061] like Figure 1 As shown in , the image encoding device 12 includes a predictor 21 , an encoder 22 , a storage unit 23 , and a controller 24 , and the image decoding device 13 includes a predictor 31 , a decoder 32 , a storage unit 33 , and a controller 34 .
[0062] The predictor 21 performs inter prediction or intra prediction to generate a predicted image. For example, when performing inter prediction, the predictor 21 generates a predicted image using a prediction block (prediction unit) having a predetermined block size as a processing unit.
[0063] The encoder 22 encodes the image input to the image encoding device 12 according to a predetermined encoding method using an encoding block (encoding unit) of a predetermined block size as a processing unit, and sends a bit stream of the encoded data to the image decoding device 13. In addition, the bit stream includes Figures 4 to 6 Describes the parameters related to the block, etc.
[0064] The storage unit 23 stores various types of data required to be stored when encoding an image in the image encoding device 12. For example, as will be described later, Figure 2 As described above, the storage unit 23 temporarily accumulates the YCgCo residual signal 1 output through the ACT process and the YCgCo residual signal 2 to be subjected to the IACT process.
[0065] The controller 24 executes the same Figure 2 Control related to the application of the ACT process and IACT process.
[0066] The predictor 31 performs inter prediction or intra prediction to generate a predicted image. For example, in the case of performing inter prediction, the predictor 21 generates a predicted image using a prediction block having a predetermined block size as a processing unit.
[0067] The decoder 32 decodes the bit stream transmitted from the image encoding device 12 according to the encoding method of the encoder 22 and outputs a decoded image.
[0068] The storage unit 33 stores various types of data required to be stored when decoding an image in the image decoding device 13. For example, as will be described later, Figure 3 As described above, the storage unit 33 temporarily accumulates the YCgCo residual signal 2 to be subjected to the IACT process.
[0069] The controller 34 performs the same operations as those described later. Figure 3 The IACT process applies relevant controls.
[0070] In the image processing system 11 configured as described above, control related to the ACT process and the IACT process is appropriately performed, whereby an increase in the storage amount of the storage unit 23 and the storage unit 33 can be avoided.
[0071] Will refer to Figure 2 The block diagram shown in further describes the configuration of the image encoding device 12.
[0072] like Figure 2As shown in , the image encoding device 12 includes a calculation unit 41, an adaptive color conversion unit 42, an orthogonal transformation unit 43, a quantization unit 44, an inverse quantization unit 45, an inverse orthogonal transformation unit 46, an inverse adaptive color conversion unit 47, a calculation unit 48, a predictor 49 and an encoder 50.
[0073] The calculation unit 41 performs calculation of subtracting the predicted image supplied from the predictor 49 from the image input to the image encoding device 12 , and supplies the RGB residual signal 1 as difference information obtained by the calculation to the adaptive color conversion unit 42 .
[0074] The adaptive color conversion unit 42 performs ACT processing on the RGB residual signal 1 supplied from the calculation unit 41. The ACT processing adaptively converts the color space of the image to be encoded. For example, the adaptive color conversion unit 42 performs ACT processing to convert the RGB color space into the YCgCo color space, thereby obtaining the YCgCo residual signal 1 from the RGB residual signal 1 and supplies the YCgCo residual signal 1 to the orthogonal transformation unit 43.
[0075] The orthogonal transform unit 43 obtains a transform coefficient by performing an orthogonal transform process on the YCgCo residual signal 1 supplied from the adaptive color conversion unit 42, which performs an orthogonal transform for each orthogonal transform block in the orthogonal transform block serving as the processing unit, and supplies the transform coefficient to the quantization unit 44. Furthermore, in the case where control is performed so that the ACT process is not performed in the adaptive color conversion unit 42, the orthogonal transform unit 43 may perform the orthogonal transform process on the RGB residual signal 1 supplied from the calculation unit 41.
[0076] The quantization unit 44 quantizes the transform coefficient supplied from the orthogonal transform unit 43 and supplies the transform coefficient to the inverse quantization unit 45 and the encoder 50. The inverse quantization unit 45 performs inverse quantization on the transform coefficient quantized in the quantization unit 44 and supplies the transform coefficient to the inverse orthogonal transform unit 46.
[0077] The inverse orthogonal transform unit 46 obtains a YCgCo residual signal 2 by performing an inverse orthogonal transform process on the transform coefficient supplied from the inverse quantization unit 45, and supplies the YCgCo residual signal 2 to the inverse adaptive color conversion unit 47. In addition, when control is performed so that the IACT process is not performed in the inverse adaptive color conversion unit 47, the inverse orthogonal transform unit 46 can obtain an RGB residual signal 2 through the inverse orthogonal transform process and supply the RGB residual signal 2 to the calculation unit 48.
[0078] The inverse adaptive color conversion unit 47 performs an IACT process that adaptively performs inverse conversion of the color space of the image on the YCgCo residual signal 2 supplied from the inverse orthogonal transform unit 46. For example, the inverse adaptive color conversion unit 47 performs an IACT process that inversely converts the YCgCo color space into the RGB color space, thereby obtaining an RGB residual signal 2 from the YCgCo residual signal 2 to supply the RGB residual signal 2 to the calculation unit 48.
[0079] The calculation unit 48 locally reconfigures (decodes) the image by performing a calculation of adding the RGB residual signal 2 supplied from the inverse adaptive color conversion unit 47 to the predicted image supplied from the predictor 49, and outputs a reconfigured signal representing the reconfigured image. In addition, when control is performed so that the IACT process is not performed in the inverse adaptive color conversion unit 47, the calculation unit 48 can reconfigure (decode) the image based on the RGB residual signal 2 supplied from the inverse orthogonal transform unit 46.
[0080] Predictor 49 corresponds to Figure 1 The predictor 21 in generates a predicted image predicted from the image reconfigured in the computing unit 48, and supplies the predicted image to the computing unit 41 and the computing unit 48.
[0081] The encoder 50 corresponds to Figure 1 The encoder 22 in the image decoding unit 50 performs encoding processing using, for example, context-based adaptive binary arithmetic coding (CABAC), which is an encoding method with high encoding efficiency for consecutive equal values, on the transform coefficients quantized in the quantization unit 44. Thus, the encoder 50 obtains a bit stream of encoded data and transmits the bit stream to the image decoding device 13.
[0082] In the image encoding device 12 configured as described above, the adaptive color conversion unit 42 performs ACT processing to convert the RGB residual signal 1 into the YCgCo residual signal 1, thereby improving the energy concentration of the signal. As described above, by improving the energy concentration of the signal, the image encoding device 12 can express the image signal with a small amount of code, and it is expected that the encoding efficiency will be improved.
[0083] Will refer to Figure 3 The block diagram shown in further describes the configuration of the image decoding device 13.
[0084] like Figure 3 As shown in , the image decoding device 13 includes a decoder 61, an inverse quantization unit 62, an inverse orthogonal transform unit 63, an inverse adaptive color conversion unit 64, a calculation unit 65, and a predictor 66.
[0085] Decoder 61 corresponds to Figure 1 The decoder 32 in the image encoding device 12 performs the decoding of the bit stream of the encoded data sent from the image encoding device 12 using the same Figure 2 Therefore, the decoder 61 obtains the quantized transform coefficients from the bit stream of the encoded data and supplies the transform coefficients to the inverse quantization unit 62. At this time, the decoder 61 also obtains the quantized transform coefficients included in the bit stream of the encoded data and referred to later. Figures 4 to 6 Describes the parameters related to the block, etc.
[0086] The inverse quantization unit 62 performs inverse quantization on the quantized transform coefficient supplied from the decoder 61 , and supplies the transform coefficient to the inverse orthogonal transform unit 63 .
[0087] The inverse orthogonal transform unit 63 obtains a YCgCo residual signal 2 by performing an inverse orthogonal transform on the transform coefficients supplied from the inverse quantization unit 62, and supplies the YCgCo residual signal 2 to the inverse adaptive color conversion unit 64. Furthermore, when control is performed so that the IACT process is not performed in the inverse adaptive color conversion unit 64, the inverse orthogonal transform unit 63 may supply the RGB residual signal 2 obtained by the inverse orthogonal transform process to the calculation unit 65.
[0088] Similar to Figure 2 The inverse adaptive color conversion unit 47 in the inverse adaptive color conversion unit 64 performs an IACT process of adaptively performing inverse conversion of the color space of the image on the YCgCo residual signal 2 supplied from the inverse orthogonal transform unit 63. For example, the inverse adaptive color conversion unit 64 performs an IACT process of inversely converting the YCgCo color space into the RGB color space, thereby obtaining an RGB residual signal 2 from the YCgCo residual signal 2 to supply the RGB residual signal 2 to the calculation unit 65.
[0089] The calculation unit 65 locally reconfigures (decodes) the image by performing a calculation of adding the RGB residual signal 2 supplied from the inverse adaptive color conversion unit 64 to the predicted image supplied from the predictor 66, and outputs a reconfigured signal representing the reconfigured image. In addition, when control is performed so that the IACT process is not performed in the inverse adaptive color conversion unit 64, the calculation unit 65 can reconfigure (decode) the image based on the RGB residual signal 2 supplied from the inverse orthogonal transform unit 63.
[0090] Similar to Figure 2 Predictor 49 in , predictor 66 corresponds to Figure 1The predictor 31 in generates a predicted image predicted from the image reconfigured in the computing unit 65 and supplies the predicted image to the computing unit 65.
[0091] The image decoding device 13 configured as described above can contribute to improvement in encoding efficiency similarly to the image encoding device 12 .
[0092] The image processing system 11 is configured as described above, and performs ACT processing of converting the RGB residual signal 1 into the YCgCo residual signal 1 in the adaptive color conversion unit 42, and performs IACT processing of converting the YCgCo residual signal 2 into the RGB residual signal 2 in the inverse adaptive color conversion units 47 and 64.
[0093] At this time, in the image encoding device 12, the ACT process is performed in units of three components in the adaptive color conversion unit 42, so that the YCgCo residual signal 1 of the three components is temporarily stored in Figure 1 . Therefore, in the storage unit 23, the YCgCo residual signal 1 is stored as three components corresponding to the block size (for example, 32×32) of the orthogonal transformation block in the orthogonal transformation unit 43. In addition, in the storage unit 23, the YCgCo residual signal 2 is stored as three components corresponding to the block size (for example, 32×32) of the orthogonal transformation block in the inverse orthogonal transformation unit 46.
[0094] Similarly, in the image decoding device 13, the YCgCo residual signal 2 of three components corresponding to the block size (for example, 32×32) of the orthogonal transformation block in the inverse orthogonal transformation unit 63 is stored in Figure 1 In the storage unit 33.
[0095] As described above, in the image processing system 11, when the ACT process and the IACT process are applied, the memory capacity of the storage unit 23 needs to be increased so that the YCgCo residual signal 1 and the YCgCo residual signal 2 can be stored according to the block size of the orthogonal transform block. Similarly, when the ACT process and the IACT process are applied, the memory capacity of the storage unit 33 needs to be increased so that the YCgCo residual signal 2 can be stored according to the block size of the orthogonal transform block. Therefore, as the memory capacity increases, the installation cost of the image processing system 11 increases.
[0096] Therefore, in the image processing system 11, the controller 24 appropriately executes control related to the application of the ACT process performed by the adaptive color conversion unit 42 and the IACT process performed by the inverse adaptive color conversion unit 47, thereby avoiding an increase in the storage amount of the storage unit 23. Similarly, in the image processing system 11, the controller 34 appropriately executes control related to the application of the IACT process performed by the inverse adaptive color conversion unit 64, thereby avoiding an increase in the storage amount of the storage unit 33. Therefore, the image processing system 11 can suppress an increase in installation cost because an increase in storage amount is avoided.
[0097] <First Concept Related to Application of ACT Processing and IACT Processing>
[0098] For example, in the image processing system 11, control is performed so that the ACT process and the IACT process are applied while providing predetermined restrictions (for example, restrictions on size, area, shape, etc.) for encoding blocks when an image is encoded.
[0099] For example, the coding block parameters used to implement this restriction include size, long side size, short side size, area, and shape. Sizes include 16×16, 16×8, 8×16, and others. The long side size includes 16 for a 16×8 block. The short side size includes 8 for a 16×8 block. Areas include 16×16, 16×8, and others. Shapes include square and rectangular shapes.
[0100] Since 32×32 is generally used as the block size of an orthogonal transform block in an orthogonal transform process and an inverse orthogonal transform process, a memory amount capable of storing the YCgCo residual signal of three components of the 32×32 block size is required.
[0101] On the other hand, in the image processing system 11, when the ACT process and the IACT process are applied, the block size of the coding block is restricted to be less than or equal to a predetermined size (e.g., 16×16). Due to this restriction, the storage unit 23 only needs to have a storage capacity sufficient to store the three components of the YCgCo residual signal 1 with a block size of 16×16. Similarly, for the YCgCo residual signal 2, the storage unit 23 and the storage unit 33 only need to have a storage capacity sufficient to store the three components of the YCgCo residual signal 2 with a block size of 16×16.
[0102] For example, regarding the syntax of the bitstream, parameters (size, area, shape, etc.) of the coding block are considered as conditions for applying the ACT process and the IACT process. That is, when sending a flag indicating the application of the ACT process, the image coding device 12 confirms that the block size of the coding block is less than or equal to a predetermined size (e.g., 16×16). Then, the syntax is determined so that the flag indicating the application of the ACT process is sent only when the block size of the coding block is less than or equal to the predetermined size (e.g., 16×16), and the flag indicating the application of the ACT process is not sent when the block size is greater than the predetermined size (e.g., 16×16).
[0103] Therefore, in the image processing system 11, when the block size of the coding block is larger than the predetermined size, it is not necessary to transmit a flag indicating the application of the ACT process, and the flag can be removed from the bitstream, so that it can be expected that the coding efficiency will be improved. In addition, by not transmitting such a flag that does not need to be transmitted, it is possible to remove ambiguous signals in the syntax of the bitstream.
[0104] Furthermore, the case of using Versatile Video Coding (VVC) as the encoding method will be described. When VVC is used as the encoding method, a 64×64 block size can be used as the orthogonal transform block in the orthogonal transform process and the inverse orthogonal transform process. Therefore, when VVC is used as the encoding method, a memory capacity capable of storing the YCgCo residual signal of the three components in a 64×64 block size is required.
[0105] Therefore, in the image processing system 11, when VVC is used as the encoding method, for example, when the ACT process and the IACT process are applied, the block size of the encoding block is limited to 32×32 or less. Due to this limitation, the storage unit 23 only needs to have a storage capacity sufficient to store the three components of the YCgCo residual signal 1 with a block size of 32×32. Similarly, for the YCgCo residual signal 2, the storage unit 23 and the storage unit 33 only need to have a storage capacity sufficient to store the three components of the YCgCo residual signal 2 with a block size of 32×32.
[0106] Note that, in addition, in the case of using High Efficiency Video Coding (HEVC) as the encoding method, a mechanism for performing ACT processing and IACT processing is provided, and in HEVC, the standard for the maximum block size of the orthogonal transform block in the orthogonal transform process and the inverse orthogonal transform process is 32×32. On the other hand, in VVC, 64×64 can already be supported as the maximum block size of the orthogonal transform block, and the increase in storage amount is greater than the increase in storage amount in HEVC. Therefore, in the image processing system 11, in the case of VVC being used as the encoding method, when ACT processing and IACT processing are applied, the orthogonal transform blocks in the orthogonal transform process and the inverse orthogonal transform process are limited, thereby making it possible to perform processing with the same amount of storage as in HEVC.
[0107] Figure 4 An example of a parameter set for a high-level syntax is shown.
[0108] exist Figure 4 In the parameter set shown in , when sps_act_enabled_flag is 1, it is specified that ACT processing and IACT processing can be applied, and cu_act_enabled_flag can be present in the coding unit syntax. On the other hand, when sps_act_enabled_flag is 0, it is specified that ACT processing and IACT processing are not applied, and cu_act_enabled_flag is not present in the coding unit syntax. Note that when sps_act_enabled_flag is not present in the parameter set, sps_act_enabled_flag is evaluated as 0.
[0109] exist Figure 4 In the parameter set shown in , sps_log2_act_max_size_minus2 specifies the maximum block size used in ACT processing and IACT processing in the range of 0 to 7. Note that in the case where sps_log2_act_max_size_minus2 is not present in the parameter set, sps_log2_act_max_size_minus2 is estimated to be 0.
[0110] Furthermore, when set to 1, the MaxActSize variable is much smaller than sps_log2_act_max_size_minus2+2 (the variable MaxActSize is set to 1<<(sps_log2_act_max_size_minus2+2)). For example, if sps_log2_act_max_size_minus2 is set to 2, the MaxActSize variable becomes 16 (=1<<4). Therefore, it is possible to prohibit the execution of ACT processing and IACT processing with a size larger than MaxActSize.
[0111] Figure 5 An example of an encoding unit of a high-level syntax is shown.
[0112] exist Figure 5 In the coding unit shown in , if cu_act_enabled_flag is 1, it specifies that the residual of the current coding unit is encoded in the YCgCo color space. On the other hand, if cu_act_enabled_flag is 0, it specifies that the residual of the current coding unit is encoded in the original color space. Note that if cu_act_enabled_flag is not present, cu_act_enabled_flag is estimated to be 0.
[0113] Here, by adding the conditions for sending cu_act_enabled_flag that the width of the coding block is less than or equal to the maximum size of the block size in ACT processing and IACT processing, and the height of the coding block is less than or equal to the maximum size of the block size in ACT processing and IACT processing (&&cbWidth<=MaxActSize&&cbHeight<=MaxActSize), the amount of storage used in ACT processing and IACT processing can be limited.
[0114] Furthermore, the amount of memory can be limited by adding the condition that the width or height of the coding block is less than or equal to the maximum size of the block size in the ACT process and the IACT process (&&(cbWidth*cbHeight)<=(MaxActsize*MaxActSize)) to the conditions for transmitting the cu_act_enabled_flag. In this case, even if one side exceeds MaxActSize, if the other side is smaller and satisfies the above conditions, the ACT process and the IACT process can be applied.
[0115] <Second Concept Related to Application of ACT Processing and IACT Processing>
[0116] In the second concept, even in the case where the block size of the coding block or the expected block used in the image processing system 11 is large, control is performed so that the orthogonal transform process and the inverse orthogonal transform process are performed by using orthogonal transform blocks having a small size obtained by dividing the block size of the coding block or the expected block, and control is performed so that the ACT process and the IACT process are applied. That is, even in the case where the block size of the coding block or the expected block used in the image processing system 11 is large, when the ACT process and the IACT process are applied, the orthogonal transform process and the inverse orthogonal transform process are performed in an orthogonal transform block having a block size smaller than that of the coding block or the expected block.
[0117] For example, when the block size of the coding block is 64×64 and the block size of the prediction block of the inter-frame prediction is 64×64, a block size of 64×64 is generally used in the orthogonal transform block. On the other hand, in the image processing system 11, even in similar cases, when the ACT process and the IACT process are applied, an orthogonal transform block having a block size smaller than 64×64 (for example, four blocks having a block size of 32×32) is automatically used.
[0118] For example, by referring to a control signal indicating application of ACT processing, in the case of applying ACT processing, control is performed so that orthogonal transform processing and inverse orthogonal transform processing are performed using an orthogonal transform block with a small block size obtained by dividing the block size of the coding block or prediction block at this time.
[0119] Therefore, in the image processing system 11, control is performed according to this second concept, thereby reducing the block size of the orthogonal transform block in the orthogonal transform process and the inverse orthogonal transform process, and it is possible to avoid an increase in the required memory capacity when applying the ACT process and the IACT process. For example, the storage unit 23 storing the YCbCo residual signal 1 can have a memory capacity for three components of a 32×32 block size, rather than a memory capacity for three components of a 64×64 block size. Therefore, in the image processing system 11, an increase in installation cost can be suppressed.
[0120] <Third Concept Related to Application of ACT Processing and IACT Processing>
[0121] In the third concept, control is performed such that the ACT process and the IACT process are applied in the case of using a small size as the maximum block size of the orthogonal transform block in the orthogonal transform process and the inverse orthogonal transform process.
[0122] For example, in the image processing system 11, 32 and 64 are defined as the maximum block sizes of the orthogonal transform block. Then, only in the case where 32 is used as the maximum block size of the orthogonal transform block, the controller 24 causes the adaptive color conversion unit 42 to perform the ACT process and the inverse adaptive color conversion unit 47 to perform the IACT process. Similarly, only in the case where 32 is used as the maximum block size of the orthogonal transform block, the controller 34 causes the inverse adaptive color conversion unit 64 to perform the IACT process.
[0123] Such control according to the third concept can be achieved by using sps_max_luma_transform_size_64_flag included in the parameter set of the high-level syntax.
[0124] Figure 6 An example of a parameter set of a high-level syntax used in the image processing system 11 is shown.
[0125] For example, when sps_max_luma_transform_size_64_flag is 0, 32 is set as the maximum block size used as the orthogonal transform block. On the other hand, when sps_max_luma_transform_size_64_flag is 1, 64 is set as the maximum block size used as the orthogonal transform block.
[0126] Therefore, in the image processing system 11 , control is performed so that the ACT process and the IACT process are applied only when sps_max_luma_transform_size_64_flag is 0. That is, in the image processing system 11 , when sps_max_luma_transform_size_64_flag is 1, the ACT process and the IACT process are not applied.
[0127] Then, if sps_max_luma_transform_size_64_flag is 0, the controller 24 sets 1 to sps_act_enabled_flag, which indicates whether adaptive color conversion processing is applied, indicating that adaptive color conversion processing is applied, and transmits sps_max_luma_transform_size_64_flag to the image decoding device 13. In this case, cu_act_enabled_flag may also be included in the coding unit syntax. Note that if sps_act_enabled_flag is 0, it indicates that adaptive color conversion processing is not applied, and in this case, cu_act_enabled_flag is not included in the coding unit syntax. Here, if sps_act_enabled_flag is not included in the parameter set, sps_act_enabled_flag is estimated to be 0.
[0128] As described above, when sps_max_luma_transform_size_64_flag is 0, the maximum block size of the orthogonal transform block is limited to 32, and therefore, in the image processing system 11 , an increase in the amount of memory required when applying the ACT process and the IACT process can be avoided.
[0129] For example, when such control is not performed, that is, when the maximum block size of the orthogonal transform block can be 64, the storage unit 23 and the storage unit 33 require a storage capacity capable of storing the YCgCo residual signal of the three components of the block size of 64×64. On the other hand, control is performed so that the ACT process and the IACT process are applied only when the maximum block size of the orthogonal transform block is limited to 32, thereby requiring the storage unit 23 and the storage unit 33 to have a storage capacity capable of storing only the YCgCo residual signal of the three components of the block size of 32×32.
[0130] Therefore, in the image processing system 11, control is performed according to such a third concept, whereby an increase in the storage amount can be avoided, and therefore, an increase in installation cost can be suppressed.
[0131] <Configuration Example for Computer-Based Systems>
[0132] Figure 7 is a block diagram illustrating a configuration example of an embodiment of a computer-based system to which the present technology is applied.
[0133] Figure 7 1 is a block diagram showing a configuration example of a network system that connects one or more computers, servers, etc. to each other via a network. Figure 7The hardware and software environments shown in the embodiments are shown as examples of platforms that can provide for implementing software and / or methods according to the present disclosure.
[0134] like Figure 7 As shown in FIG, the network system 101 includes a computer 102, a network 103, a remote computer 104, a web server 105, a cloud storage server 106, and a computer server 107. Figure 7 One or more functional blocks shown in are executed.
[0135] In addition, Figure 7 , a detailed configuration of the computer 102 is shown. Note that the functional blocks shown in the computer 102 are shown to establish exemplary functions and are not limited to this configuration. In addition, although the detailed configurations of the remote computer 104, web server 105, cloud storage server 106, and computer server 107 are not shown, the remote computer 104, web server 105, cloud storage server 106, and computer server 107 include components similar to the functional blocks shown in the computer 102.
[0136] As the computer 102 , a personal computer, a desktop computer, a laptop computer, a tablet computer, a netbook computer, a personal digital assistant, a smart phone, or other programmable electronic device capable of communicating with other devices on a network may be used.
[0137] Computer 102 then includes a bus 111, a processor 112, a memory 113, a non-volatile storage device 114, a network interface 115, a peripheral device interface 116, and a display interface 117. In some embodiments, each of these functions may be implemented in a single electronic subsystem (an integrated circuit chip or a combination of a chip and associated devices), or in other embodiments, some of the functions may be combined and implemented in a single chip (a system on a chip or system on a chip (SoC)).
[0138] As the bus 111 , various proprietary or industry-standard high-speed parallel or serial peripheral interconnect buses may be used.
[0139] As the processor 112 , a processor designed and / or manufactured as one or more single-chip or multi-chip microprocessors may be used.
[0140] The memory 113 and the non-volatile storage device 114 are storage media that can be read by the computer 102. For example, any suitable volatile storage device, such as a dynamic random access memory (DRAM) or a static RAM (SRAM), can be used as the memory 113. The non-volatile storage device 114 can include at least one or more of a flexible disk, a hard disk, a solid-state drive (SSD), a read-only memory (ROM), an erasable and programmable read-only memory (EPROM), a flash memory, a compact disk (CD or CD-ROM), a digital versatile disk (DVD), a card memory, or a sticky memory.
[0141] In addition, a program 121 is stored in the non-volatile storage device 114. The program 121 is, for example, a set of machine-readable instructions and / or data used to create, manage, and control specific software functions. Note that in configurations where the memory 113 is much faster than the non-volatile storage device 114, the program 121 can be transferred from the non-volatile storage device 114 to the memory 113 before being executed by the processor 112.
[0142] Computer 102 can communicate and interact with other computers over network 103 via network interface 115. Network 103 can employ a configuration including wired, wireless, or fiber optic connections through, for example, a local area network (LAN), a wide area network (WAN) such as the Internet, or a combination of a LAN and a WAN. In general, network 103 includes any combination of connections and protocols that support communication between two or more computers and associated devices.
[0143] The peripheral device interface 116 can input data to and output data from devices that can be locally connected to the computer 102. For example, the peripheral device interface 116 provides a connection to an external device 131. A keyboard, a mouse, a keypad, a touch screen, and / or other suitable input devices can be used as the external device 131. The external device 131 can also include a portable computer-readable storage medium, such as a thumb drive, a portable optical or magnetic disk, a memory card, and the like.
[0144] In an embodiment of the present disclosure, for example, software and data for implementing the program 121 may be stored in such a portable computer-readable storage medium. In such an embodiment, the software may be loaded into the non-volatile storage device 114 via the peripheral device interface 116, or directly loaded into the memory 113. The peripheral device interface 116 may use industry standards such as RS-232, a universal serial bus (USB), etc. for connection to the external device 131.
[0145] The display interface 117 can connect the computer 102 to the display 132 and can present a command line or a graphical user interface to the user of the computer 102 by using the display 132. For example, as the display interface 117, an industry standard such as Video Graphics Array (VGA), Digital Visual Interface (DVI), DisplayPort, or High-Definition Multimedia Interface (HDMI) (registered trademark) can be adopted.
[0146] <Configuration Example of Image Encoding Device>
[0147] Figure 8 : is a block diagram showing a configuration example of an embodiment of an image encoding device as an image processing device to which the present disclosure is applied.
[0148] Figure 8 The image encoding device 201 shown in FIG encodes image data by using prediction processing. Here, as an encoding method, for example, a Versatile Video Coding (VVC) method, a High Efficiency Video Coding (HEVC) method, or the like is used.
[0149] Figure 8 The image encoding device 201 in FIG. 1 includes an A / D conversion unit 202, a screen rearrangement buffer 203, a calculation unit 204, an orthogonal transform unit 205, a quantization unit 206, a lossless encoder 207, and an accumulation buffer 208. In addition, the image encoding device 201 includes an inverse quantization unit 209, an inverse orthogonal transform unit 210, a calculation unit 211, a deblocking filter 212, an adaptive offset filter 213, an adaptive loop filter 214, a frame memory 215, a selection unit 216, an intra predictor 217, a motion prediction and compensation unit 218, a predicted image selection unit 219, and a rate controller 220.
[0150] The A / D conversion unit 202 performs A / D conversion on input image data (picture), and supplies the converted image data to the screen rearrangement buffer 203. Note that an image of digital data may be input without providing the A / D conversion unit 202.
[0151] The screen rearrangement buffer 203 stores the image data supplied from the A / D conversion unit 202 and rearranges the images of the frames stored in the display order in the order of frames for encoding according to the group of pictures (GOP) structure. The screen rearrangement buffer 203 outputs the images whose order of frames has been rearranged to the calculation unit 204, the intra predictor 217, and the motion prediction and compensation unit 218.
[0152] The calculation unit 204 subtracts the predicted image supplied from the intra predictor 217 or the motion prediction and compensation unit 218 via the predicted image selection unit 219 from the image output by the screen rearrangement buffer 203 , and outputs the difference information to the orthogonal transformation unit 205 .
[0153] For example, in the case of an image on which intra encoding is performed, the calculation unit 204 subtracts the predicted image supplied from the intra predictor 217 from the image output from the screen rearrangement buffer 203. Furthermore, for example, in the case of an image on which inter encoding is performed, the calculation unit 204 subtracts the predicted image supplied from the motion prediction and compensation unit 218 from the image output from the screen rearrangement buffer 203.
[0154] The orthogonal transform unit 205 performs orthogonal transform, such as discrete cosine transform or Karhunen-Loeve transform, on the difference information supplied from the calculation unit 204 , and supplies the transform coefficient to the quantization unit 206 .
[0155] The quantization unit 206 quantizes the transform coefficient output from the orthogonal transform unit 205. The quantization unit 206 supplies the quantized transform coefficient to the lossless encoder 207.
[0156] The lossless encoder 207 performs lossless encoding, such as variable length encoding and arithmetic encoding, on the quantized transform coefficients.
[0157] The lossless encoder 207 acquires parameters such as information indicating the intra prediction mode from the intra predictor 217 , and acquires parameters such as information indicating the inter prediction mode and motion vector information from the motion prediction and compensation unit 218 .
[0158] The lossless encoder 207 encodes the quantized transform coefficients and encodes each acquired parameter (syntax element) as part of the header information of the encoded data (multiplexed into the header information of the encoded data). The lossless encoder 207 supplies the encoded data obtained by encoding to the accumulation buffer 208 for accumulation.
[0159] For example, in the lossless encoder 207, lossless encoding processing such as variable length coding or arithmetic coding is performed. Examples of variable length coding include context-adaptive variable length coding (CAVLC). Examples of arithmetic coding include context-adaptive binary arithmetic coding (CABAC).
[0160] The accumulation buffer 208 temporarily stores the encoded stream (encoded data) supplied from the lossless encoder 207, and outputs the encoded stream, which is an encoded image subjected to encoding, to, for example, a recording device or a transmission path (not shown) in a subsequent stage at a predetermined timing. That is, the accumulation buffer 208 also serves as a transmission unit that transmits the encoded stream.
[0161] Furthermore, the transform coefficient quantized in the quantization unit 206 is also supplied to the inverse quantization unit 209. The inverse quantization unit 209 inversely quantizes the quantized transform coefficient by a method corresponding to the quantization performed by the quantization unit 206. The inverse quantization unit 209 supplies the obtained transform coefficient to the inverse orthogonal transform unit 210.
[0162] The inverse orthogonal transform unit 210 performs inverse orthogonal transform on the supplied transform coefficient by a method corresponding to the orthogonal transform process performed by the orthogonal transform unit 205. The output subjected to the inverse orthogonal transform (restored difference information) is supplied to the calculation unit 211.
[0163] The calculation unit 211 adds the predicted image supplied from the intra predictor 217 or the motion prediction and compensation unit 218 via the predicted image selection unit 219 and the inverse orthogonal transformation result supplied from the inverse orthogonal transformation unit 210, i.e., the restored difference information, to obtain a local decoded image (decoded image).
[0164] For example, in the case where the difference information corresponds to an image on which intra encoding is performed, the calculation unit 211 adds the predicted image supplied from the intra predictor 217 to the difference information. Furthermore, for example, in the case where the difference information corresponds to an image on which inter encoding is performed, the calculation unit 211 adds the predicted image supplied from the motion prediction and compensation unit 218 to the difference information.
[0165] The decoded image as a result of the addition is supplied to the deblocking filter 212 and the frame memory 215 .
[0166] The deblocking filter 212 suppresses block distortion in the decoded image by appropriately performing deblocking filtering on the image from the calculation unit 211, and supplies the filter processing result to the adaptive offset filter 213. The deblocking filter 212 has parameters β and Tc obtained based on the quantization parameter QP. The parameters β and Tc are thresholds (parameters) used for determination related to the deblocking filter.
[0167] Note that β and Tc as parameters of the deblocking filter 212 are extended from β and Tc defined in the HEVC scheme. The offset of the parameters β and Tc is encoded as a parameter of the deblocking filter in the lossless encoder 207 and sent to the . Figure 10 The image decoding device 301 in.
[0168] The adaptive offset filter 213 performs an offset filter (sample adaptive offset (SAO)) process for mainly suppressing ringing on the image filtered by the deblocking filter 212 .
[0169] There are nine types of offset filters, including two types of band offset, six types of edge offset, and no offset. The adaptive offset filter 213 performs filtering processing on the image filtered by the deblocking filter 212 by using a quadtree structure that determines the type of offset filter for each divided area and the offset value for each divided area. The adaptive offset filter 213 supplies the filtered image to the adaptive loop filter 214.
[0170] Note that in the image encoding device 201, the quadtree structure and the offset value of each divided region are calculated and used by the adaptive offset filter 213. The calculated quadtree structure and the offset value of each divided region are encoded as adaptive offset parameters in the lossless encoder 207 and sent to the adaptive offset filter described later. Figure 10 The image decoding device 301 in.
[0171] The adaptive loop filter 214 performs adaptive loop filter (ALF) processing on the image filtered by the adaptive offset filter 213 for each processing unit using the filter coefficients. In the adaptive loop filter 214, for example, a two-dimensional Wiener filter is used as a filter. Of course, a filter other than the Wiener filter can be used. The adaptive loop filter 214 supplies the filtering processing results to the frame memory 215.
[0172] Note that although Figure 8 Although not shown in the example of , in the image encoding device 201, for each processing unit, a filter coefficient is calculated by the adaptive loop filter 214 so as to minimize the residual of the original image from the screen rearrangement buffer 203, and the filter coefficient is used. The calculated filter coefficient is encoded as an adaptive loop filter parameter in the lossless encoder 207 and sent to the CMOS controller described later. Figure 10 The image decoding device 301 in.
[0173] The frame memory 215 outputs the accumulated reference image to the intra predictor 217 or the motion prediction and compensation unit 218 via the selection unit 216 at predetermined timing.
[0174] For example, in the case of an intra-coded image, the frame memory 215 supplies the reference image to the intra predictor 217 via the selection unit 216. Furthermore, for example, in the case of inter-coded images, the frame memory 215 supplies the reference image to the motion prediction and compensation unit 218 via the selection unit 216.
[0175] In the case where the reference image supplied from the frame memory 215 is an image to be subjected to intra-frame encoding, the selection unit 216 supplies the reference image to the intra-frame predictor 217. In addition, in the case where the reference image supplied from the frame memory 215 is an image to be subjected to inter-frame encoding, the selection unit 216 supplies the reference image to the motion prediction and compensation unit 218.
[0176] The intra predictor 217 performs intra prediction (intra screen prediction) of generating a predicted image by using pixel values in a screen. The intra predictor 217 performs intra prediction in a plurality of modes (intra prediction modes).
[0177] The intra-frame predictor 217 generates prediction images in all intra-frame prediction modes, evaluates each prediction image, and selects the optimal mode. When the optimal intra-frame prediction mode is selected, the intra-frame predictor 217 supplies the prediction image generated in the optimal mode to the calculation unit 204 and the calculation unit 211 via the prediction image selection unit 219.
[0178] Furthermore, as described above, the intra predictor 217 appropriately supplies parameters such as intra prediction mode information indicating the employed intra prediction mode to the lossless encoder 207 .
[0179] The motion prediction and compensation unit 218 performs motion prediction on the image on which inter encoding is performed by using the input image supplied from the screen rearrangement buffer 203 and the reference image supplied from the frame memory 215 via the selection unit 216. Furthermore, the motion prediction and compensation unit 218 performs motion compensation processing based on the motion vector detected by the motion prediction, and generates a predicted image (inter-prediction image information).
[0180] The motion prediction and compensation unit 218 performs inter-frame prediction processing in all candidate inter-frame prediction modes and generates a predicted image. The motion prediction and compensation unit 218 supplies the generated predicted image to the calculation unit 204 and the calculation unit 211 via the predicted image selection unit 219. In addition, the motion prediction and compensation unit 218 supplies parameters such as inter-frame prediction mode information indicating the adopted inter-frame prediction mode and motion vector information indicating the calculated motion vector to the lossless encoder 207.
[0181] The predicted image selection unit 219 supplies the output of the intra-frame predictor 217 to the calculation unit 204 and the calculation unit 211 in the case of an image to be subjected to intra-frame encoding, and supplies the output of the motion prediction and compensation unit 218 to the calculation unit 204 and the calculation unit 211 in the case of an image to be subjected to inter-frame encoding.
[0182] The rate controller 220 controls the rate of the quantization operation of the quantization unit 206 based on the compressed image accumulated in the accumulation buffer 208 so that overflow or underflow does not occur.
[0183] The image encoding device 201 is configured as described above, and the adaptive color conversion unit 42 ( Figure 2 ) is provided between the calculation unit 204 and the orthogonal transformation unit 205, and the inverse adaptive color conversion unit 47 ( Figure 2 ) is provided between the inverse orthogonal transform unit 210 and the calculation unit 211. Then, in the image encoding device 201, control is performed according to the first concept to the third concept described above, whereby an increase in the memory amount can be avoided.
[0184] <Operation of Image Coding Apparatus>
[0185] Will refer to Figure 9 The flow of the encoding process performed by the image encoding device 201 described above is described.
[0186] In step S101 , the A / D conversion unit 202 performs A / D conversion on an input image.
[0187] In step S102 , the screen rearrangement buffer 203 stores the image subjected to A / D conversion by the A / D conversion unit 202 , and performs rearrangement from the display order to the encoding order of each picture.
[0188] In the case where the image to be processed supplied from the screen rearrangement buffer 203 is an image of a block to be subjected to intra processing, the decoded image to be referenced is read from the frame memory 215 and supplied to the intra predictor 217 via the selection unit 216 .
[0189] Based on these images, in step S103, the intra predictor 217 performs intra prediction on the pixels of the block to be processed in all candidate intra prediction modes. Note that as decoded pixels to be referenced, pixels not filtered by the deblocking filter 212 are used.
[0190] This process performs intra prediction in all candidate intra prediction modes, and calculates cost function values for all candidate intra prediction modes. An optimal intra prediction mode is then selected based on the calculated cost function values, and a predicted image generated by intra prediction in the optimal intra prediction mode and the cost function value of the optimal intra prediction mode are supplied to the predicted image selection unit 219.
[0191] In the case where the image to be processed supplied from the screen rearrangement buffer 203 is an image to be subjected to inter-frame processing, the image to be referenced is read from the frame memory 215, and the image to be referenced is supplied to the motion prediction and compensation unit 218 via the selection unit 216. Based on these images, in step S104, the motion prediction and compensation unit 218 performs motion prediction and compensation processing.
[0192] This process performs motion prediction processing on all candidate inter prediction modes, calculates cost function values for all candidate inter prediction modes, and determines the optimal inter prediction mode based on the calculated cost function values. The predicted image generated by the optimal inter prediction mode and the cost function value of the optimal inter prediction mode are then supplied to the predicted image selection unit 219.
[0193] In step S105, the predicted image selection unit 219 determines one of the optimal intra prediction mode or the optimal inter prediction mode as the optimal prediction mode based on the cost function values output from the intra predictor 217 and the motion prediction and compensation unit 218. Then, the predicted image selection unit 219 selects a predicted image in the determined optimal prediction mode and supplies the predicted image to the calculation units 204 and 211. The predicted image is used for calculations in steps S106 and S111 described later.
[0194] Note that the selection information of the predicted image is supplied to the intra predictor 217 or the motion prediction and compensation unit 218. In the case of selecting the predicted image in the optimal intra prediction mode, the intra predictor 217 supplies information indicating the optimal intra prediction mode (i.e., parameters related to intra prediction) to the lossless encoder 207.
[0195] In the case where the predicted image in the optimal inter prediction mode is selected, the motion prediction and compensation unit 218 outputs information indicating the optimal inter prediction mode and information corresponding to the optimal inter prediction mode (i.e., parameters related to motion prediction) to the lossless encoder 207. Examples of the information corresponding to the optimal inter prediction mode include motion vector information and reference frame information.
[0196] In step S106, the calculation unit 204 calculates the difference between the image rearranged in step S102 and the predicted image selected in step S105. In the case of performing inter prediction, the predicted image is supplied to the calculation unit 204 from the motion prediction and compensation unit 218 via the predicted image selection unit 219, and in the case of performing intra prediction, the predicted image is supplied to the calculation unit 204 from the intra predictor 217 via the predicted image selection unit 219.
[0197] The amount of difference data is smaller than that of the original image data, so the amount of data can be compressed compared to the case of directly encoding the image.
[0198] In step S107, the orthogonal transform unit 205 performs orthogonal transform on the difference information supplied from the calculation unit 204. Specifically, orthogonal transform such as discrete cosine transform or Karhunen-Loeve transform is performed, and a transform coefficient is output.
[0199] In step S108, the transform coefficient is quantized by the quantization unit 206. At the time of this quantization, the rate is controlled as described in the process of step S118 described later.
[0200] The difference information quantized as described above is locally decoded as follows. That is, in step S109, the inverse quantization unit 209 inversely quantizes the transform coefficients quantized by the quantization unit 206 with characteristics corresponding to the characteristics of the quantization unit 206. In step S110, the inverse orthogonal transform unit 210 performs inverse orthogonal transform on the transform coefficients inversely quantized by the inverse quantization unit 209 with characteristics corresponding to the characteristics of the orthogonal transform unit 205.
[0201] In step S111 , the calculation unit 211 adds the predicted image input via the predicted image selection unit 219 and the locally decoded difference information to generate a locally decoded image (ie, an image subjected to local decoding) (an image corresponding to the input to the calculation unit 204 ).
[0202] In step S112, the deblocking filter 212 performs a deblocking filter process on the image output from the calculation unit 211. At this time, the parameters β and Tc extended from the β and Tc defined in the HEVC scheme are used as thresholds for determination related to the deblocking filter. The filtered image from the deblocking filter 212 is output to the adaptive offset filter 213.
[0203] Note that the parameter β and the offset amount of Tc, which are input by the user who operates the operation unit or the like and used in the deblocking filter 212 , are supplied to the lossless encoder 207 as parameters of the deblocking filter.
[0204] In step S113, the adaptive offset filter 213 performs adaptive offset filtering. This process performs filtering on the image filtered by the deblocking filter 212 using a quadtree structure that determines the type of offset filtering for each divided region and the offset value for each divided region. The filtered image is supplied to the adaptive loop filter 214.
[0205] Note that the determined quadtree structure and the offset value of each divided area are supplied to the lossless encoder 207 as adaptive offset parameters.
[0206] In step S114, the adaptive loop filter 214 performs adaptive loop filtering processing on the image filtered by the adaptive offset filter 213. For example, for the image filtered by the adaptive offset filter 213, filtering processing is performed on the image for each processing unit by using a filter coefficient, and the filtering processing result is supplied to the frame memory 215.
[0207] In step S115 , the frame memory 215 stores the filtered image. Note that an image not filtered by the deblocking filter 212 , the adaptive offset filter 213 , and the adaptive loop filter 214 is also supplied from the calculation unit 211 and stored in the frame memory 215 .
[0208] On the other hand, the transform coefficients quantized in step S108 are also supplied to the lossless encoder 207. In step S116, the lossless encoder 207 encodes the quantized transform coefficients output from the quantization unit 206 and the supplied parameters. That is, the difference image is subjected to lossless encoding such as variable length encoding or arithmetic encoding and compressed. Examples of parameters to be encoded include deblocking filter parameters, adaptive offset filter parameters, adaptive loop filter parameters, quantization parameters, motion vector information, reference frame information, prediction mode information, and the like.
[0209] In step S117, the accumulation buffer 208 accumulates the encoded difference image (ie, the encoded stream) as a compressed image. The compressed image accumulated in the accumulation buffer 208 is appropriately read and transmitted to the decoding side via the transmission path.
[0210] In step S118 , the rate controller 220 controls the rate of the quantization operation of the quantization unit 206 based on the compressed image accumulated in the accumulation buffer 208 so that overflow or underflow does not occur.
[0211] When the process of step S118 ends, the encoding process ends.
[0212] In the encoding process described above, the adaptive color conversion unit 42 ( Figure 2 ) and is processed by the inverse adaptive color conversion unit 47 ( Figure 2 ) is performed between step S110 and step S111. Then, in the encoding process, control related to the application of the ACT process and the IACT process is performed according to the first to third concepts described above.
[0213] <Configuration Example of Image Decoding Device>
[0214] Figure 10 A configuration of an embodiment of an image decoding device as an image processing device to which the present disclosure is applied is shown. Figure 10 The image decoding apparatus 301 shown in FIG. Figure 8 The decoding device corresponding to the image encoding device 201 in.
[0215] The encoded stream (encoded data) encoded by the image encoding device 201 is transmitted to the image decoding device 301 corresponding to the image encoding device 201 via a predetermined transmission path, and is decoded.
[0216] like Figure 10 As shown in the figure, the image decoding device 301 includes an accumulation buffer 302, a lossless decoder 303, an inverse quantization unit 304, an inverse orthogonal transform unit 305, a calculation unit 306, a deblocking filter 307, an adaptive offset filter 308, an adaptive loop filter 309, a picture rearrangement buffer 310, a D / A conversion unit 311, a frame memory 312, a selection unit 313, an intra-frame predictor 314, a motion prediction and compensation unit 315, and a selection unit 316.
[0217] The accumulation buffer 302 is also a receiving unit for receiving the transmitted coded data. The accumulation buffer 302 receives and accumulates the transmitted coded data. The coded data is coded by the image coding device 201. The lossless decoder 303 Figure 8 The encoded data read from the accumulation buffer 302 is decoded at a predetermined timing by a method corresponding to the encoding method of the lossless encoder 207 in FIG.
[0218] The lossless decoder 303 supplies parameters such as information indicating the decoded intra prediction mode to the intra predictor 314, and supplies parameters such as information indicating the inter prediction mode and motion vector information to the motion prediction and compensation unit 315. In addition, the lossless decoder 303 supplies the decoded parameters of the deblocking filter to the deblocking filter 307, and supplies the decoded adaptive offset parameters to the adaptive offset filter 308.
[0219] The inverse quantization unit 304 is Figure 8 The coefficient data (quantized coefficient) obtained by decoding by the lossless decoder 303 is inversely quantized by a method corresponding to the quantization method of the quantization unit 206 in FIG. Figure 8 The quantization coefficient is inversely quantized by using the quantization parameter supplied from the image encoding device 201 in a method similar to that of the inverse quantization unit 209 in FIG.
[0220] The inverse quantization unit 304 supplies the inverse quantized coefficient data, that is, the orthogonal transformation coefficient, to the inverse orthogonal transformation unit 305. The inverse orthogonal transformation unit 305 Figure 8 A method corresponding to the orthogonal transform method of the orthogonal transform unit 205 in performs inverse orthogonal transform on the orthogonal transform coefficients, and obtains decoded residual data corresponding to the residual data before being subjected to orthogonal transform in the image encoding device 201.
[0221] The decoded residual data obtained by being subjected to the inverse orthogonal transform is supplied to the calculation unit 306. Furthermore, the predicted image is supplied to the calculation unit 306 from the intra predictor 314 or the motion prediction and compensation unit 315 via the selection unit 316.
[0222] The calculation unit 306 adds the decoded residual data and the predicted image together to obtain decoded image data corresponding to the image data before the predicted image is subtracted by the calculation unit 204 of the image encoding device 201. The calculation unit 306 supplies the decoded image data to the deblocking filter 307.
[0223] The deblocking filter 307 suppresses block distortion of the decoded image by appropriately performing deblocking filtering processing on the image from the calculation unit 306, and supplies the result of the filtering processing to the adaptive offset filter 308. The deblocking filter 307 is basically similar to Figure 8 The deblocking filter 212 in FIG. 307 is configured. That is, the deblocking filter 307 has parameters β and Tc obtained based on the quantization parameter. The parameters β and Tc are thresholds used for determination related to the deblocking filter.
[0224] Note that β and Tc, which are parameters of the deblocking filter 307, are extended from β and Tc defined in the HEVC scheme. The offset of the parameters β and Tc of the deblocking filter encoded by the image encoding device 201 is received by the image decoding device 301 as the parameters of the deblocking filter, decoded by the lossless decoder 303, and used by the deblocking filter 307.
[0225] The adaptive offset filter 308 performs offset filtering (SAO) processing for mainly suppressing ringing on the image filtered by the deblocking filter 307 .
[0226] The adaptive offset filter 308 performs filtering processing on the image filtered by the deblocking filter 307 by using a quadtree structure that determines the type of offset filter for each divided area and an offset value for each divided area. The adaptive offset filter 308 supplies the image after filtering processing to the adaptive loop filter 309.
[0227] Note that the quadtree structure and the offset value for each divided region are calculated by the adaptive offset filter 213 of the image encoding device 201, encoded as adaptive offset parameters, and transmitted. The quadtree structure and the offset value for each divided region encoded by the image encoding device 201 are then received as adaptive offset parameters by the image decoding device 301, decoded by the lossless decoder 303, and used by the adaptive offset filter 308.
[0228] The adaptive loop filter 309 performs filtering processing for each processing unit by using a filter coefficient on the image filtered by the adaptive offset filter 308 , and supplies the filtering processing result to the frame memory 312 and the screen rearrangement buffer 310 .
[0229] Note that although Figure 10 Although not shown in the example of , in the image decoding device 301 , the filter coefficients calculated for each LUC by the adaptive loop filter 214 of the image encoding device 201 and encoded and transmitted as adaptive loop filter parameters are decoded by the lossless decoder 303 and used.
[0230] The screen rearrangement buffer 310 rearranges the image and supplies the rearranged image to the D / A conversion unit 311. Figure 8 The picture rearrangement buffer 203 rearranges the order of frames that were rearranged in the encoding order to the original display order.
[0231] The D / A conversion unit 311 performs D / A conversion on the image (decoded picture) supplied from the screen rearrangement buffer 310, outputs the image to a display (not shown), and displays the image. Note that the image can be output as digital data without providing the D / A conversion unit 311.
[0232] The output of the adaptive loop filter 309 is also supplied to the frame memory 312 .
[0233] The frame memory 312, the selection unit 313, the intra-frame predictor 314, the motion prediction and compensation unit 315, and the selection unit 316 correspond to the frame memory 215, the selection unit 216, the intra-frame predictor 217, the motion prediction and compensation unit 218, and the predicted image selection unit 219 of the image encoding device 201, respectively.
[0234] The selection unit 313 reads an image to be subjected to inter processing and an image to be referenced from the frame memory 312, and supplies the images to the motion prediction and compensation unit 315. In addition, the selection unit 313 reads an image to be used for intra prediction from the frame memory 312, and supplies the image to the intra predictor 314.
[0235] Information indicating the intra prediction mode obtained by decoding the header information and the like are appropriately supplied from the lossless decoder 303 to the intra predictor 314. Based on this information, the intra predictor 314 generates a predicted image from the reference image acquired from the frame memory 312 and supplies the generated predicted image to the selection unit 316.
[0236] Information obtained by decoding the header information (prediction mode information, motion vector information, reference frame information, flags, various parameters, and the like) is supplied from the lossless decoder 303 to the motion prediction and compensation unit 315 .
[0237] The motion prediction and compensation unit 315 generates a predicted image from the reference image acquired from the frame memory 312 based on the information supplied from the lossless decoder 303 , and supplies the generated predicted image to the selection unit 316 .
[0238] The selection unit 316 selects the predicted image generated by the motion prediction and compensation unit 315 or the intra predictor 314 , and supplies the predicted image to the calculation unit 306 .
[0239] The image decoding device 301 is configured as described above, and the inverse adaptive color conversion unit 64 ( Figure 3 ) is provided between the inverse orthogonal transform unit 305 and the calculation unit 306. Then, in the image decoding device 301, control is performed according to the first concept to the third concept described above, whereby an increase in the memory amount can be avoided.
[0240] <Operation of Image Decoding Device>
[0241] Will refer to Figure 11 An example of the flow of a decoding process performed by the image decoding device 301 as described above is described.
[0242] When the decoding process is started, in step S201, the accumulation buffer 302 receives and accumulates the transmitted coded stream (data). In step S202, the lossless decoder 303 decodes the coded data supplied from the accumulation buffer 302. Figure 8 The I picture, P picture, and B picture encoded by the lossless encoder 207 in are decoded.
[0243] Before decoding a picture, parameter information such as motion vector information, reference frame information, and prediction mode information (intra-frame prediction mode or inter-frame prediction mode) is also decoded.
[0244] When the prediction mode information is intra-frame prediction mode information, the prediction mode information is supplied to the intra-frame predictor 314. When the prediction mode information is inter-frame prediction mode information, the prediction mode information and corresponding motion vector information and the like are supplied to the motion prediction and compensation unit 315. In addition, the parameters of the deblocking filter and the adaptive offset parameters are also decoded and supplied to the deblocking filter 307 and the adaptive offset filter 308, respectively.
[0245] In step S203 , the intra predictor 314 or the motion prediction and compensation unit 315 performs a predicted image generation process corresponding to the prediction mode information supplied from the lossless decoder 303 .
[0246] That is, when the intra prediction mode information is supplied from the lossless decoder 303, the intra predictor 314 generates an intra prediction image in the intra prediction mode. When the inter prediction mode information is supplied from the lossless decoder 303, the motion prediction and compensation unit 315 performs motion prediction and compensation processing in the inter prediction mode to generate an inter prediction image.
[0247] With this processing, the predicted image (intra-predicted image) generated by the intra predictor 314 or the predicted image (inter-predicted image) generated by the motion prediction and compensation unit 315 is supplied to the selection unit 316 .
[0248] In step S204, the selection unit 316 selects a predicted image. That is, the predicted image generated by the intra predictor 314 or the predicted image generated by the motion prediction and compensation unit 315 is supplied. Thus, the supplied predicted image is selected and supplied to the calculation unit 306, and is added to the output of the inverse orthogonal transform unit 305 in step S207 described later.
[0249] In the above step S202, the transform coefficients decoded by the lossless decoder 303 are also supplied to the inverse quantization unit 304. In step S205, the inverse quantization unit 304 uses Figure 8 The transform coefficients decoded by the lossless decoder 303 are inversely quantized according to characteristics corresponding to the characteristics of the quantization unit 206 in FIG.
[0250] In step S206, the inverse orthogonal transform unit 305 uses Figure 8 The inverse orthogonal transform is performed on the transform coefficients inversely quantized by the inverse quantization unit 304 according to the characteristics corresponding to the characteristics of the orthogonal transform unit 205 in FIG. Figure 8 The difference information corresponding to the input of the orthogonal transform unit 205 (the output of the calculation unit 204) is decoded.
[0251] In step S207, the calculation unit 306 adds the predicted image selected in the process in step S204 described above and input via the selection unit 316 to the difference information. Thus, the original image is decoded.
[0252] In step S208, the deblocking filter 307 performs deblocking filtering on the image output from the calculation unit 306. At this time, the parameters β and Tc, which are extended from the β and Tc defined in the HEVC scheme, are used as thresholds for determination related to the deblocking filter. The filtered image from the deblocking filter 307 is output to the adaptive offset filter 308. Note that the deblocking filter also uses the offset of the deblocking filter parameters β and Tc supplied from the lossless decoder 303.
[0253] In step S209, the adaptive offset filter 308 performs adaptive offset filtering. This process filters the image filtered by the deblocking filter 307 using a quadtree structure that determines the type of offset filtering for each divided region and the offset value for each divided region. The filtered image is then supplied to the adaptive loop filter 309.
[0254] In step S210, the adaptive loop filter 309 performs adaptive loop filtering processing on the image filtered by the adaptive offset filter 308. The adaptive loop filter 309 performs filtering processing on the input image for each processing unit by using the filter coefficient calculated for each processing unit, and supplies the filtering processing result to the screen rearrangement buffer 310 and the frame memory 312.
[0255] In step S211 , the frame memory 312 stores the filtered image.
[0256] In step S212, the screen rearrangement buffer 310 rearranges the image after the adaptive loop filter 309, and then supplies the rearranged image to the D / A conversion unit 311. That is, the order of the frames rearranged for encoding by the screen rearrangement buffer 203 of the image encoding device 201 is rearranged in the original display order.
[0257] In step S213 , the D / A conversion unit 311 performs D / A conversion on the image rearranged by the screen rearrangement buffer 310 , and outputs the converted image to a display (not shown) to display the image.
[0258] When the process of step S213 ends, the decoding process ends.
[0259] In the decoding process described above, the inverse adaptive color conversion unit 64 ( Figure 3 Then, in the decoding process, control related to the application of the ACT process and the IACT process is performed according to the first to third concepts described above.
[0260] <Computer Configuration Example>
[0261] Next, the above-mentioned series of processing (image processing method) can be executed by hardware or software. In the case of executing the series of processing by software, a program constituting the software is installed in a general-purpose computer or the like.
[0262] Figure 12 : is a block diagram showing a configuration example of an embodiment of a computer in which a program for executing the above-described series of processes is installed.
[0263] The program may be recorded in advance on the hard disk 1005 or the ROM 1003 as a recording medium incorporated in the computer.
[0264] Alternatively, the program may be stored (recorded) in a removable recording medium 1011 driven by the drive 1009. Such a removable recording medium 1011 may be provided as so-called packaged software. Examples of the removable recording medium 1011 include a flexible disk, a compact disk read-only memory (CD-ROM), a magneto-optical (MO) disk, a digital versatile disk (DVD), a magnetic disk, a semiconductor memory, and the like.
[0265] Note that the program can be installed on the computer from the removable recording medium 1011 as described above, or can be downloaded to the computer via a communication network or a broadcast network and installed on the incorporated hard disk 1005. In other words, for example, the program can be wirelessly transferred from a download site to the computer via an artificial satellite used for digital satellite broadcasting, or can be sent to the computer via a network such as a local area network (LAN) or the Internet over a wire.
[0266] The computer includes a central processing unit (CPU) 1002 , and an input / output interface 1010 is connected to the CPU 1002 via a bus 1001 .
[0267] When a command is input by a user operating the input unit 1007 or the like via the input / output interface 1010, the CPU 1002 executes a program stored in the read-only memory (ROM) 1003 according to the command. Alternatively, the CPU 1002 loads a program stored in the hard disk 1005 into the random access memory (RAM) 1004 and executes the program.
[0268] Thus, the CPU 1002 executes the processing according to the above flowchart or the processing executed by the configuration of the above block diagram. Then, the CPU 1002 causes the processing result to be output from the output unit 1006 via the input / output interface 1010 or transmitted from the communication unit 1008 as needed, and also causes the processing result to be recorded on, for example, the hard disk 1005.
[0269] Note that the input unit 1007 includes a keyboard, a mouse, a microphone, etc. Furthermore, the output unit 1006 includes a liquid crystal display (LCD), a speaker, and the like.
[0270] Here, in this specification, the processing executed by a computer according to a program does not necessarily have to be performed in chronological order according to the order described in the flowchart. That is, the processing executed by a computer according to a program also includes processing executed in parallel or individually (for example, parallel processing or processing by an object).
[0271] In addition, the program may be processed by one computer (processor), or may be distributed and processed by a plurality of computers. In addition, the program may be transmitted to a remote computer and executed.
[0272] In this specification, a system refers to a group of multiple components (devices, modules (components), etc.), and it does not matter whether all components are in the same cabinet. Therefore, multiple devices housed in separate cabinets and connected to each other via a network, as well as a single device housing multiple modules in a single cabinet, are both systems.
[0273] Furthermore, for example, a configuration described as one device (or processing unit) may be divided and configured as a plurality of devices (or processing units). Conversely, a configuration described as a plurality of devices (or processing units) may be collectively configured as one device (or processing unit). Furthermore, of course, configurations other than those described above may be added to the configuration of each device (or each processing unit). Furthermore, as long as the overall configuration and operation of the system are substantially the same, a portion of the configuration of a certain device (or processing unit) may be included in the configuration of another device (or another processing unit).
[0274] Furthermore, for example, the present technology can employ a configuration of cloud computing in which one function is shared among a plurality of devices via a network to perform collaborative processing.
[0275] In addition, for example, the above-mentioned program can be executed in any device. In this case, it is sufficient that the device has necessary functions (functional blocks, etc.) and can obtain necessary information.
[0276] Furthermore, for example, each step described in the flowchart above can be shared among multiple devices rather than being executed by a single device. Furthermore, when multiple processes are included in a single step, the multiple processes included in a single step can be shared among multiple devices rather than being executed by a single device. In other words, the multiple processes included in a single step can be executed as a single process. Conversely, processes described as multiple steps can be collectively executed as a single step.
[0277] Note that in a program executed by a computer, the processing of the steps describing the program can be performed in chronological order and in the order described in this specification, or in parallel, or can be performed individually at necessary timings, such as when each step is called. In other words, the processing of each step can be performed in an order different from the order described above, as long as no inconsistency occurs. Furthermore, the processing of the steps describing the program can be performed in parallel with the processing of another program, or can be performed in combination with the processing of another program.
[0278] Note that, as long as there are no inconsistencies, each of the multiple present technologies described in this specification can be implemented independently and individually. Of course, it is also possible to implement it by combining any of the multiple present technologies. For example, part or all of the present technology described in any one of the embodiments can be implemented in combination with part or all of the present technology described in other embodiments. In addition, part or all of the above-mentioned present technology can be implemented in combination with another technology not described above.
[0279] <Configuration combination example>
[0280] Note that the present technology can also be configured as described below. (1)
[0282] An image processing device, comprising:
[0283] an adaptive color conversion unit that performs adaptive color conversion processing on a residual signal of an image to be encoded, the adaptive color conversion processing adaptively performing conversion of a color space of the image;
[0284] an orthogonal transform unit that performs an orthogonal transform process on a residual signal of the image or on a residual signal of the image subjected to the adaptive color conversion process, for each of the orthogonal transform blocks serving as a processing unit; and
[0285] A controller performs control related to application of the adaptive color conversion process. (2)
[0287] The image processing device according to (1), wherein
[0288] A first block size and a second block size larger than the first block size are defined as maximum block sizes of the orthogonal transform block, and
[0289] The controller performs control so that the adaptive color conversion unit applies the adaptive color conversion process if the first block size is used as the maximum block size of the orthogonal transform block. (3)
[0291] The image processing device according to (2), wherein
[0292] The first block size is 32, and the second block size is 64, and
[0293] The controller causes the adaptive color conversion process to be applied only in a case where 32 is used as the maximum block size of the orthogonal transform block. (4)
[0295] The image processing device according to (3), wherein
[0296] The controller transmits sps_act_enabled_flag to the decoding side, the sps_act_enabled_flag indicating that the adaptive color conversion process is applied if sps_max_luma_transform_size_64_flag included in a parameter set of a high-level syntax is 0. (5)
[0298] The image processing device according to any one of (1) to (4), wherein
[0299] The controller performs control to cause the adaptive color conversion unit to apply the adaptive color conversion process while providing a predetermined restriction for an encoding block when the image is encoded. (6)
[0301] The image processing device according to (5), wherein
[0302] The controller causes the adaptive color conversion process to be applied with a block size of an encoding block as a processing unit being restricted to be smaller than or equal to a predetermined size when the image is encoded. (7)
[0304] The image processing device according to (6), wherein
[0305] The controller causes the adaptive color conversion process to be applied if a block size of the encoding block is restricted to be less than or equal to 16×16. (8)
[0307] The image processing device according to any one of (1) to (7), wherein
[0308] In a case where a block size of a coding block as a processing unit when the image is encoded is larger than a predetermined size, the controller performs control so that the orthogonal transform unit performs the orthogonal transform process using the orthogonal transform block having a small block size obtained by dividing the coding block, and performs control so that the adaptive color conversion unit applies the adaptive color conversion process. (9)
[0310] The image processing device according to (8), wherein
[0311] In a case where the block size of the encoding block is 64×64, the controller sets the orthogonal transform block to 32×32 to cause the orthogonal transform process to be performed, and causes the adaptive color conversion process to be applied. (10)
[0313] The image processing apparatus according to any one of (1) to (9), further comprising:
[0314] an inverse orthogonal transform unit that acquires the residual signal by performing an inverse orthogonal transform process on a transform coefficient obtained when the orthogonal transform process is performed for each of the orthogonal transform blocks; and
[0315] an inverse adaptive color conversion unit that performs an inverse adaptive color conversion process on the residual signal obtained by the inverse orthogonal transform unit, the inverse adaptive color conversion process adaptively performing inverse conversion of the color space of the image,
[0316] in,
[0317] The controller performs control related to application of the inverse adaptive color conversion process corresponding to the adaptive color conversion process. (11)
[0319] An image processing method, comprising:
[0320] performing an adaptive color conversion process on a residual signal of an image to be encoded, wherein the adaptive color conversion process adaptively performs conversion of a color space of the image;
[0321] performing, for each of the orthogonal transform blocks as processing units, an orthogonal transform process on a residual signal of the image or on a residual signal of the image subjected to the adaptive color conversion process; and
[0322] Control related to application of the adaptive color conversion process is performed. (12)
[0324] An image processing device, comprising:
[0325] an inverse orthogonal transform unit that obtains a residual signal of an image to be decoded by performing an inverse orthogonal transform process on a transform coefficient obtained when an orthogonal transform process is performed on the residual signal on the encoding side, for each of the orthogonal transform blocks serving as the processing unit;
[0326] an inverse adaptive color conversion unit that performs an inverse adaptive color conversion process on the residual signal, the inverse adaptive color conversion process adaptively performing an inverse conversion of a color space of an image; and
[0327] A controller performs control related to application of the inverse adaptive color conversion process. (13)
[0329] An image processing method, comprising:
[0330] Acquire a residual signal of an image to be decoded by: performing, for each of the orthogonal transform blocks serving as a processing unit, an inverse orthogonal transform process on a transform coefficient obtained when an orthogonal transform process is performed on the residual signal on the encoding side;
[0331] performing an inverse adaptive color conversion process on the residual signal, the inverse adaptive color conversion process adaptively performing an inverse conversion of a color space of an image; and
[0332] Control related to application of the inverse adaptive color conversion process is performed.
[0333] Note that this embodiment is not limited to the above-described embodiment, and various modifications can be made without departing from the scope of the present disclosure. In addition, the advantageous effects described in this specification are merely examples, and are not limited to these advantageous effects, and may include other effects.
[0334] Reference Signs List
[0335] 11 Image Processing System
[0336] 12 Image Coding Device
[0337] 13 Image Decoding Device
[0338] 21 Predictor
[0339] 22 Encoder
[0340] 23 storage units
[0341] 24 Controller
[0342] 31 Predictor
[0343] 32 decoder
[0344] 33 storage units
[0345] 34 Controller
[0346] 41 computing units
[0347] 42 adaptive color conversion units
[0348] 43 Orthogonal Transformation Unit
[0349] 44 Quantitative units
[0350] 45 Inverse Quantization Unit
[0351] 46 Inverse Orthogonal Transformation Unit
[0352] 47 Inverse Adaptive Color Conversion Unit
[0353] 48 computing units
[0354] 49 Predictor
[0355] 50 encoder
[0356] 61 Decoder
[0357] 62 Inverse Quantization Units
[0358] 63 Inverse Orthogonal Transformation Unit
[0359] 64 inverse adaptive color conversion units
[0360] 65 computing units
[0361] 66 Predictor
Claims
1. An image processing device, comprising: an adaptive color conversion unit that performs an adaptive color conversion process on a residual signal of an image to be encoded, wherein the adaptive color conversion process adaptively performs conversion of a color space of the image; an orthogonal transform unit that performs an orthogonal transform process on a residual signal of the image or on a residual signal of the image subjected to the adaptive color conversion process, for each of the orthogonal transform blocks serving as a processing unit; and a controller that performs control related to application of the adaptive color conversion process, wherein a first block size and a second block size larger than the first block size are defined as the maximum block size of the orthogonal transform block, and Here, the controller performs control so that the adaptive color conversion unit applies the adaptive color conversion process in a case where the first block size is used as a maximum block size of the orthogonal transform block.
2. The image processing apparatus according to claim 1, wherein: The first block size is 32, and the second block size is 64, and The controller causes the adaptive color conversion process to be applied only in a case where 32 is used as the maximum block size of the orthogonal transform block.
3. The image processing apparatus according to claim 2, wherein: The controller transmits sps_act_enabled_flag to the decoding side, the sps_act_enabled_flag indicating that the adaptive color conversion process is applied if sps_max_luma_transform_size_64_flag included in a parameter set of a high-level syntax is 0.
4. The image processing apparatus according to claim 1, wherein: The controller performs control to cause the adaptive color conversion unit to apply the adaptive color conversion process while providing a predetermined restriction for an encoding block when the image is encoded.
5. The image processing apparatus according to claim 4, wherein: The controller causes the adaptive color conversion process to be applied with a block size of an encoding block as a processing unit being restricted to be smaller than or equal to a predetermined size when the image is encoded. The image processing apparatus according to claim 5 , wherein: The controller causes the adaptive color conversion process to be applied if a block size of the encoding block is restricted to be less than or equal to 16×16.
7. The image processing apparatus according to claim 1, wherein: In a case where a block size of a coding block as a processing unit when the image is encoded is larger than a predetermined size, the controller performs control so that the orthogonal transform unit performs the orthogonal transform process using the orthogonal transform block having a small block size obtained by dividing the coding block, and performs control so that the adaptive color conversion unit applies the adaptive color conversion process.
8. The image processing apparatus according to claim 7, wherein: In a case where the block size of the encoding block is 64×64, the controller sets the orthogonal transform block to 32×32 to cause the orthogonal transform process to be performed, and causes the adaptive color conversion process to be applied.
9. The image processing apparatus according to claim 1, further comprising: an inverse orthogonal transform unit that acquires the residual signal by performing an inverse orthogonal transform process on a transform coefficient obtained when the orthogonal transform process is performed for each of the orthogonal transform blocks; and an inverse adaptive color conversion unit that performs an inverse adaptive color conversion process on the residual signal obtained by the inverse orthogonal transform unit, the inverse adaptive color conversion process adaptively performing an inverse conversion of the color space of the image, in, The controller performs control related to application of the inverse adaptive color conversion process corresponding to the adaptive color conversion process.
10. An image processing method, comprising: performing an adaptive color conversion process on a residual signal of an image to be encoded, wherein the adaptive color conversion process adaptively performs conversion of a color space of the image; performing, for each of the orthogonal transform blocks serving as a processing unit, an orthogonal transform process on a residual signal of the image or on a residual signal of the image subjected to the adaptive color conversion process; and performing control related to application of said adaptive color conversion process, wherein a first block size and a second block size larger than the first block size are defined as the maximum block size of the orthogonal transform block, and wherein performing the control includes applying the adaptive color conversion process when the first block size is used as a maximum block size of the orthogonal transform block.
11. An image processing device, comprising: an inverse orthogonal transform unit that obtains a residual signal of an image to be decoded by performing an inverse orthogonal transform process on a transform coefficient obtained when an orthogonal transform process is performed on the residual signal on the encoding side, for each of the orthogonal transform blocks serving as the processing unit; an inverse adaptive color conversion unit that performs an inverse adaptive color conversion process on the residual signal, wherein the inverse adaptive color conversion process adaptively performs an inverse conversion of the color space of the image; as well as a controller that performs control related to application of the inverse adaptive color conversion process, wherein a first block size and a second block size larger than the first block size are defined as a maximum block size of the orthogonal transform block; and Here, the controller performs control so that the inverse adaptive color conversion unit applies the inverse adaptive color conversion process in a case where the first block size is used as a maximum block size of the orthogonal transform block.
12. An image processing method, comprising: Acquire a residual signal of an image to be decoded by: performing, for each of the orthogonal transform blocks serving as a processing unit, an inverse orthogonal transform process on a transform coefficient obtained when an orthogonal transform process is performed on the residual signal on the encoding side; performing an inverse adaptive color conversion process on the residual signal, wherein the inverse adaptive color conversion process adaptively performs an inverse conversion of a color space of an image; as well as performing control related to application of said inverse adaptive color conversion process, wherein a first block size and a second block size larger than the first block size are defined as a maximum block size of the orthogonal transform block; and wherein performing the control includes applying the inverse adaptive color conversion process when the first block size is used as a maximum block size of the orthogonal transform block. 13 . A non-transitory computer-readable medium having a program embodied thereon, the program, when executed by a computer, causing the computer to perform the image processing method according to claim 10 or 12 . 14 . A computer program product comprising a computer program, which, when executed by a computer, causes the computer to execute the image processing method according to claim 10 .
Citation Information
Patent Citations
Video encoding methods and systems using adaptive color transform
EP3104606A1