Image encoding device, control method thereof, and program

The image encoding device efficiently encodes RAW data from sensors with luminance pixels by converting it into multiple planes with uniform pixel spacing, addressing compatibility issues and improving encoding efficiency.

JP7723480B2Active Publication Date: 2025-08-14CANON KK
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
JP2021020069
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Filing Date
2021-02-10
Publication Date
2025-08-14
Estimated Expiration
2041-02-10

AI Technical Summary

Technical Problem

Existing methods for encoding RAW image data from imaging devices with color filter arrays that include luminance pixels, such as white pixels, are not compatible with CFAs like the Dual Bayer + HDR array, leading to inefficiencies in data encoding.

Method used

An image encoding device that converts RAW image data from a color filter array with luminance pixels into multiple planes of low-frequency and high-frequency component data, using a plane conversion process that generates planes with uniform pixel spacing for efficient encoding.

Benefits of technology

Enables efficient encoding of RAW image data from sensors with luminance pixels, maintaining image quality while reducing data volume, and allowing for standardized circuit design and processing times.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007723480000001
    Figure 0007723480000001
  • Figure 0007723480000002
    Figure 0007723480000002
  • Figure 0007723480000003
    Figure 0007723480000003
Patent Text Reader

Abstract

To efficiently encode RAW image data from an imaging element having a color filter array that generates a luminance pixel such as a white pixel in addition to three primary color components.SOLUTION: In an image encoding device that encodes RAW image data obtained by an imaging element having a color filter array in which a plurality of filters of the respective three primary colors and a plurality of filters of a specific color for luminance are arranged in an N×N pixel area, and the filter for the N×N pixel area is repeated includes a conversion unit that converts RAW image data into a plurality of planes, each of which is composed of a single color component, and an encoding unit that encodes each plane obtained by the conversion unit. Here, for each component representing the three primary colors, the conversion unit refers to the pixel value of the same component in the N×N pixel area, thereby generating a plane composed of low-frequency component data and a plane composed of high-frequency component data.SELECTED DRAWING: Figure 1
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] The present invention relates to a technique for encoding RAW image data. Regarding. [Background technology]

[0002] Nowadays, imaging devices such as digital cameras and digital camcorders use CCD sensors or CMOS sensors as their imaging elements. These sensors have a color filter array (hereafter referred to as CFA) arranged on their surface, with each color filter corresponding to one pixel and one color component. A typical example of a CFA is a periodic pattern arrangement of R (red), G0 (green), B (blue), and G1 (green), as shown in Figure 3. This pattern arrangement is generally called a Bayer array. Image data (hereafter referred to as RAW image data) is obtained through this CFA.

[0003] Human vision is known to be highly sensitive to luminance components. For this reason, in a typical Bayer array, as shown in Figure 3, the green (G) component, which contains a large amount of luminance, is assigned twice as many pixels as the red and blue components. In RAW image data, each pixel contains only one color component. Therefore, a process called demosaicing is required to generate red (R), blue (B), and green (G) information for each pixel. The RGB signal obtained by demosaicing, or the YUV signal obtained by further converting the RGB signal, is generally encoded before being recorded on a recording medium. However, because demosaicing generates an image with three color components per pixel, the data volume is three times that of RAW image data. Therefore, several methods have been proposed for directly encoding and recording RAW data before demosaicing.

[0004] For example, Patent Document 1 discloses a method of separating RAW data into R, G0, B, and G1 planes (that is, four planes) and then encoding each plane.

[0005] Patent document 2 also shows a method of separating RAW data into four planes, R, G0, B, and G1, as in Patent document 1, and then converting and encoding the data approximately into luminance (Y) and color difference (Co, Cg, Dg). [Prior art documents] [Patent documents]

[0006] [Patent Document 1] Japanese Patent Application Laid-Open No. 2003-125209 [Patent Document 2] Japanese Patent Application Laid-Open No. 2006-121669 Summary of the Invention [Problem to be solved by the invention]

[0007] Meanwhile, sensors with color filter arrays different from those shown in Figure 3 have been developed to improve sensitivity in low-light conditions. Figure 4 shows a CFA with a so-called dual Bayer + HDR (High Dynamic Range) array structure, which generates image data for red (R), green (G), blue (B), and white (W) pixels. A 4x4 pixel array is repeated, with the ratio of the number of pixels for each color being R:G:B:W = 2:4:2:8. The W pixels transmit the entire visible light range without using a color filter. Therefore, the W pixels have higher sensitivity than the RGB pixels. Therefore, the CFA shown in Figure 4 can improve luminance sensitivity compared to a CFA consisting only of RGB. The plane conversion methods shown in Patent Documents 1 and 2 are based on the Bayer array shown in Figure 3 and are not compatible with CFAs that have W pixels in addition to RGB.

[0008] The present invention aims to provide a technique for efficiently encoding RAW data obtained by an imaging device with a color filter array (CGA) that generates luminance pixels such as white pixels in addition to RGB. [Means for solving the problem]

[0009] In order to solve this problem, for example, an image coding device of the present invention has the following arrangement: An image encoding device that encodes RAW image data obtained from an image sensor having a color filter array in which a plurality of filters for each of three primary colors and a plurality of filters of a specific color for luminance are arranged within an N×N pixel area, and the filters of the N×N pixel area are repeated, a conversion means for converting the RAW image data into a plurality of planes each composed of a single color component; encoding means for encoding each plane obtained by said converting means; The conversion means For each component representing the three primary colors, a plane consisting of low-frequency component data and a plane consisting of high-frequency component data are generated by referencing the pixel value of the same component within the NxN pixel area. death, For each component representing the three primary colors, a plane consisting of low-frequency component data and a plane consisting of high-frequency component data are generated by referencing the pixel values of the same component that are diagonally adjacent within the N×N pixel area. It is characterized by: [Effects of the Invention]

[0010] According to the present invention, it is possible to efficiently encode raw image data from an image sensor having a color filter array that generates luminance pixels such as white pixels in addition to the three primary color components. [Brief explanation of the drawings]

[0011] [Figure 1] FIG. 1 is a block diagram showing the configuration of an encoding device according to a first embodiment. [Figure 2] FIG. 1 is a block diagram showing the configuration of a decoding device according to a first embodiment. [Figure 3] FIG. 1 is a diagram for explaining a Bayer array. [Figure 4] FIG. 2 is a diagram showing an example of a Dual Bayer+HDR arrangement according to the first embodiment. [Figure 5] FIG. 10 is a diagram for explaining a plane conversion method for a G component. [Figure 6] FIG. 1 is a diagram illustrating a wavelet transform. [Figure 7] FIG. 10 is a diagram showing an example of a Dual Bayer+HDR arrangement according to the second embodiment. [Figure 8] 6 is a flowchart showing the processing procedure of a plane conversion unit in the first embodiment. DETAILED DESCRIPTION OF THE INVENTION

[0012] Hereinafter, embodiments will be described in detail with reference to the accompanying drawings. Note that the following embodiments do not limit the scope of the claimed invention. Although multiple features are described in the embodiments, not all of these multiple features are necessarily essential to the invention, and multiple features may be combined arbitrarily. Furthermore, in the accompanying drawings, the same reference numerals are used to designate the same or similar components, and redundant explanations will be omitted.

[0013] [First embodiment] Hereinafter, embodiments of the present invention will be described in detail with reference to the drawings.

[0014] FIG. 1 shows a block diagram of an image encoding device according to an embodiment when applied to an imaging device.

[0015] As shown in Fig. 1, the image coding device includes an imaging unit 100, a plane conversion unit 101, and a RAW coding unit 110. The RAW coding unit 110 includes a frequency conversion unit 102, a code amount control unit 103, a quantization unit 104, and an entropy coding unit 105. As will be described in detail later, the RAW coding Department The plane converter 110 encodes the encoding target plane input from the plane converter 101 to a target amount set based on its color and type (whether it is a low-frequency plane or a high-frequency plane, etc.). Note that the imaging device also has components such as an operation unit for user operation and a recording unit for recording captured images, but these are not the focus of the present invention and are not shown. Also, in this embodiment, JPEG2000 will be used as an example of an encoding method, but the type of encoding method is not particularly important.

[0016] First, a method for encoding an input image will be described with reference to FIG.

[0017] The imaging unit 100 has a general imaging optical section consisting of an optical lens, an aperture mechanism, a shutter mechanism, an imager (imaging element), etc. The imager is a type that captures images by color separation, such as a CCD or MOS type with a color filter array (CFA) for color separation on its surface. This imager separates the formed optical image into colors and converts them into electrical signals according to the amount of light.

[0018] In this embodiment, the color separation color filters of the imager of the imaging unit 100 are described as being filters in a Dual Bayer+HDR array arrangement shown in FIG. 4. As shown in FIG. 4, W (white) filters are arranged in a checkerboard pattern in a 4×4 pixel area, with eight filters, half the area. Four G filters are arranged consecutively in a diagonal direction. Two R filters and two B filters are arranged adjacent to each other in a diagonal direction. Therefore, each pixel of the RAW image data output by the imaging unit 100 has a corresponding color (component) value shown in the Dual Bayer+HDR array arrangement shown in FIG. 4.

[0019] The plane conversion unit 101 receives RAW image data of a 4×4 pixel array as shown in FIG. 4 and performs processing (plane conversion) to generate a plurality of plane data, each of which is a single color component.

[0020] First, the plane conversion of the G component will be described. In the plane conversion of the G component, the plane conversion unit 101 first performs the following linear conversion using G0, G1, G2, and G3. GL01 = G0 + G1 GH01 = G0 - G1 GL23 = G2 + G3 GH23 = G2 - G3

[0021] Next, the plane conversion unit 101 executes the following secondary conversion using the low frequency values GL01 and GL23 obtained in the above primary conversion. GL = GL01 + GL23 GH = GL01 - GL23 The plane conversion unit 101 then transmits a total of four single-component planes, the component values GL, GH, GH01, and GH23 obtained by the conversion, to the frequency conversion unit 102. Here, it can be said that the component value GL represents the low-frequency value (low-frequency plane) of the G component in a 4×4 pixel region in the RAW image, and the other component values GH, GH01, and GH23 represent the high-frequency values (high-frequency planes) of the G component in the 4×4 pixel region. Incidentally, the frequency components can be said to have the relationship "GH01 = GH23 > GH01 > GL" in descending order.

[0022] Next, the plane conversion of the R component will be described. In the plane conversion of the R component, the plane conversion unit 101 performs the following conversion using R0 and R1. RL = R0 + R1 RH = R0 - R1 The plane conversion unit 101 then transmits the component values RL and RH obtained by the conversion, a total of two planes, to the frequency conversion unit 102. Here, it can be said that the component value RL represents the low-frequency value of the R component in the 4×4 pixel region, and the component value RH represents the high-frequency value of the R component in the 4×4 pixel region.

[0023] Next, the plane conversion of the B component will be described. In the plane conversion of the B component, the plane conversion unit 101 performs the following conversion using B0 and B1. BL = B0 + B1 BH = B0 - B1 The plane conversion unit 101 then transmits the component values BL and BH obtained by the conversion, a total of two planes, to the frequency conversion unit 102. Here, it can be said that the component value BL represents the low-frequency value of the B component in the 4×4 pixel region, and the component value BH represents the high-frequency value of the B component in the 4×4 pixel region.

[0024] The plane conversion unit 101 does not perform plane conversion on the W component within the 4×4 pixel region, but separates W0 to W7 and transmits each to the frequency conversion unit 102. In other words, the plane conversion unit 101 generates and outputs a total of eight planes for the W component, including a plane made up of only W0, a plane made up of only W1, ..., a plane made up of only W7, all at the same relative position within the 4×4 pixel region.

[0025] To summarize the above, the plane conversion unit 101 in this embodiment performs plane conversion on a 4×4 pixel area, For G components, GL, GH, GH01, GH23 For the R component, RL and RH For B component, BL and BH For the W component, W1 to W7 A total of 16 component values are generated and transmitted to the frequency transform unit 102. When the number of horizontal pixels in the RAW image is W and the number of vertical pixels is H, for example, the size of a GL plane consisting only of GL component values among the 16 component values is W / 4 × H / 4. Planes of component values other than the GL component also have the same size, W / 4 × H / 4.

[0026] FIG. 5(a) is a diagram showing only the G component of the Dual Bayer+HDR array arrangement in FIG.

[0027] Looking at the arrangement of G component pixels, we see that while there are three pixels between them in the vertical and horizontal directions, they are adjacent in the right diagonal direction. Because the pixel spacing is not constant, this arrangement is not suitable for separating high-frequency and low-frequency components in the frequency conversion unit 102 at the subsequent stage.

[0028] Here, the effect of plane transformation on coding efficiency will be described using plane transformation of the G component as an example.

[0029] GL01, generated in the first plane transformation of the G component, is the sum of G0 and G1. Because it is equivalent to taking the average of G0 and G1, it is equivalent to constructing the low-frequency component of frequency transformation. In other words, generating GL01 is equivalent to generating an intermediate pixel value between G0 and G1, as shown in Figure 5(b). Similarly, generating GL23 is equivalent to generating an intermediate pixel value between G2 and G3, as shown in Figure 5(b). On the other hand, GH01 and GH23 are the difference, so they correspond to the high-frequency component of frequency transformation. Because GL is generated by the sum of GL01 and GL23, it corresponds to the low-frequency component of the second-stage frequency transformation, and is equivalent to generating an intermediate pixel value between GL01 and GL23, as shown in Figure 5(c).

[0030] Furthermore, when we look at the pixel arrangement in the finally generated GL plane, we see that the pixels are evenly spaced horizontally and vertically every three pixels. Because the pixels are evenly spaced horizontally and vertically, this arrangement is appropriate for separating high-frequency components and low-frequency components in the subsequent frequency conversion unit 102. This makes it possible to appropriately reduce high-frequency components while leaving low-frequency components, which have a large impact on image quality, when quantization is performed in the quantization unit 104.

[0031] On the other hand, the GH plane corresponds to the plane of the second-stage high frequency components.

[0032] In the plane conversion of this embodiment, RL, BL, and W0 to W7 also have equal horizontal and vertical pixel arrangements every three pixels, enabling efficient encoding processing. Also, RH and BH correspond to high-frequency components of frequency conversion similar to GH01 and GH23.

[0033] The frequency transform unit 102 performs wavelet transform on each of the 16 pieces of plane data GL, GH, GH01, GH23, RL, RH, BL, BH, and W0 to W7 input from the plane transform unit 101. As a result, multiple subbands (each subband includes multiple transform coefficients) are generated from one piece of plane data. The frequency transform unit 102 outputs the transform coefficients of the multiple subbands obtained from each piece of plane data to the quantization unit 104.

[0034] Here, the wavelet transform will be explained using the example of the configuration of the wavelet transform unit in FIG.

[0035] FIG. 6 shows an example of a wavelet transform unit that performs subband decomposition at only one level (one time), a method also adopted in JPEG2000. In the figure, "LPF" stands for low-pass filter, and "HPF" stands for high-pass filter. In FIG. 6, a vertical low-pass filter 401 performs vertical low-frequency filtering on an input plane 400 to generate vertical low-frequency component data, which it outputs to a downsampling circuit 403. A vertical high-pass filter 402 performs vertical high-frequency filtering on the input plane 400 to generate vertical high-frequency component data, which it outputs to a downsampling circuit 404. Each of the downsampling circuits 403 and 404 downsamples the input data by 2:1. Specifically, the downsampling circuit 403 outputs low-frequency component data with half the vertical resolution of the original input plane 400, and the downsampling circuit 404 outputs high-frequency component data with half the vertical resolution.

[0036] The downsampling circuit 403 supplies the vertical low-frequency component data to a horizontal low-pass filter 405 and a horizontal high-pass filter 406. The horizontal low-pass filter 405 performs horizontal low-frequency filtering and outputs the result to a downsampling circuit 409. The horizontal high-pass filter 406 performs horizontal high-frequency filtering and outputs the result to a downsampling circuit 410. Each of the downsampling circuits 409 and 410 downsamples the input data at a ratio of 2:1.

[0037] Meanwhile, the downsampling circuit 404 supplies the vertical high-frequency component data to a horizontal low-pass filter 407 and a horizontal high-pass filter 408. The horizontal low-pass filter 407 performs horizontal low-frequency filtering and outputs the result to a downsampling circuit 411. The horizontal high-pass filter 406 also performs horizontal high-frequency filtering and outputs the result to a downsampling circuit 412. Each of the downsampling circuits 411 and 412 downsamples the input data at a ratio of 2:1.

[0038] As a result of the above, subband 413 can be obtained. Subband 413 is composed of LL blocks, HL blocks, LH blocks, and HH blocks through the above filtering process. For simplicity, these blocks will be referred to as subbands LL, HL, LH, and HH as needed below. Here, L represents low frequency and H represents high frequency, with the first letter corresponding to vertical filtering and the second letter corresponding to horizontal filtering. For example, "HH" indicates a subband with high frequency in both the vertical and horizontal directions. If the input plane 400 is considered an image, subband LL in subband 413 is an image with its resolution reduced by half both vertically and horizontally. Furthermore, the areas of subbands HH, HL, and LH represent high-frequency component data.

[0039] The frequency transform unit 102 in this embodiment receives 16 pieces of plane data, GL, GH, GH01, GH23, RL, RH, BL, BH, and W0 to W7, as input plane 400 in FIG. 6. The frequency transform unit 102 then performs wavelet transform on each piece of plane data to generate subbands 413. Note that wavelet transforms generally allow the subband LL obtained in the immediately preceding transform to be recursively used as a transform target. Therefore, the frequency transform unit 102 in this embodiment may perform multi-stage wavelet transforms.

[0040] The code amount control unit 103 determines the target code amount to be allocated to each picture and each plane according to the compression rate set by the user, and transmits the determined target code amount to the quantization unit 106. At this time, the code amount is allocated equally to each of the R, G, B, and W color components. However, when allocating the code amount for each component data in the same color plane, the code amount is allocated using the following relationship: G component: GL>GH>GH01=GH23 R component: RL>RH B component: BL>BH W component: W0=W1=W2=W3=W4=W5=W6=W7

[0041] As described above, component GL corresponds to the low-frequency component of the second-stage frequency components, component GH corresponds to the high-frequency component of the second-stage GH, and components GH01 and GH23 correspond to the high-frequency components of the first-stage frequency transform. Therefore, more codes are allocated to the low-frequency components that have a large impact on image quality, and fewer codes are allocated to the high-frequency components that have a smaller impact, thereby enabling efficient encoding that maintains image quality. Similarly, components RL and BL correspond to low-frequency components, and components RH and BH correspond to high-frequency components, so more codes are allocated to components BL and RL, which have a large impact on image quality, and fewer codes are allocated to components RH and BH. The number of codes allocated depends on the quantization parameter set in the quantization unit 104, as described below.

[0042] The quantization unit 104 quantizes the transformation coefficients sent from the frequency conversion unit 102 using a quantization parameter determined based on the target code amount set by the code amount control unit 103, and sends the quantized transformation coefficients to the entropy coding unit 105.

[0043] The entropy coding unit 105 performs entropy coding such as EBCOT (Embedded Block Coding with Optimized Truncation) on the wavelet coefficients and quantization parameters quantized by the quantization unit 104 for each subband, and outputs the coded data. The output destination is generally a recording medium, but may also be on a network, and the type of destination is not particularly important.

[0044] Next, the decoding of the coded image data generated by the above procedure will be explained. Fig. 2 is a block diagram showing the configuration of an image decoding device according to this embodiment.

[0045] As shown in the figure, the image decoding device in this embodiment includes an entropy decoding unit 200, an inverse quantization unit 201, an inverse frequency transform unit 202, and a Bayer transform unit 203.

[0046] The entropy decoding unit 200 entropy decodes coded image data by EBCOT (Embedded Block Coding with Optimized Truncation) or the like, decodes wavelet coefficients and quantization parameters in the subbands of each plane, and outputs the dequantized wavelet coefficients and quantization parameters to the inverse quantization unit 202. 201 Transfer to.

[0047] The inverse quantization unit 201 inverse quantizes the restored wavelet transform coefficients sent from the entropy decoding unit 200 using the quantization parameter, and transfers the data obtained by the inverse quantization to the frequency inverse transformation unit 202 .

[0048] The frequency inverse transform unit 202 performs a frequency inverse transform on the frequency transform coefficients restored by the inverse quantization unit 201, reconstructs 16 plane data of GL, GH, GH01, GH23, RL, RH, BL, BH, and W0 to W7, and transfers them to the Bayer transform unit 203.

[0049] Bayer Conversion Unit 203 The frequency inverse transform unit 202 performs inverse plane transform on GL, GH, GH01, GH23, RL, RH, BL, BH, and W0 to W7, which are independently reconstructed by the frequency inverse transform unit 202. 203 The Bayer conversion unit restores R0, R1, G0 to G3, B0, B1, and W0 to W7 based on the data obtained by the inverse plane conversion. 203 rearranges R0, R1, G0 to G3, B0, B1, and W0 to W7 according to the Dual Bayer+HDR array, recombines the 4×4 pixel area of the original RAW image data, and outputs it.

[0050] Here, the restoration of the G components G0 to G4 can be calculated according to the following conversion formula. GL01 = (GL+GH) / 2 GL23 = (GL-GH) / 2 G0 = (GL01+GH01) / 2 G1 = (GL01-GH01) / 2 G2 = (GL23+GH23) / 2 G3 = (GL23-GH23) / 2 In addition, the restoration of R0 and R1 of the R component can be calculated according to the following conversion formula. R0 = (RL+RH) / 2 R1 = (RL-RH) / 2 Furthermore, the restoration of B0 and B1 of the B component can be calculated according to the following conversion formula. B0 = (BL+BH) / 2 B1 = (BL-BH) / 2 In addition, since the W component does not undergo plane transformation, the data obtained by inverse plane transformation can be used as is.

[0051] In the above-described embodiment, when the RAW image array is a Dual Bayer+HDR array with a repeating pattern of 4x4 pixels as shown in FIG. 4, the RAW image is converted into 16 planes of data with uniform pixel spacing and then encoded, thereby achieving highly efficient encoding. In this embodiment, the vertical and horizontal resolutions of each plane are the same, so the sizes of line buffers and other components required for frequency conversion can be standardized, and the circuit configuration can be designed assuming the same processing time for each plane. Furthermore, since each plane is spaced three pixels apart horizontally and vertically from the original RAW image, quantization is possible assuming the same separation characteristics during frequency conversion. Note that GL and GH may be generated all at once using the following formula instead of being generated stepwise when performing plane conversion of the G component: GL = G0 + G1 + G2 + G3 GH = (G0 + G1) - (G2 + G3)

[0052] In the above example, the image sensor in the image capturing unit 100 has been described as having a filter with a dual Bayer+HDR array as shown in Fig. 4. Here, the processing of the plane conversion unit 101 when N × N pixels (in the embodiment, N = 4) including a white pixel dedicated to luminance in addition to the three primary colors of RGB are used as a repeating pattern will be described with reference to the flowchart in Fig. 8.

[0053] In S1, the plane conversion unit 101 receives data of an N×N pixel area, which is a unit of a repeating pattern in the RAW image data.

[0054] In S2, the plane conversion unit 101 calculates the low-frequency and high-frequency component values for each of the three primary colors R, G, and B in the input N×N pixel area. No calculation is performed for the W pixel.

[0055] In S3, the plane conversion unit 101 stores the low-frequency component values, high-frequency component values calculated from each of the three primary colors, as well as all W pixel values within the N×N pixel region, in the corresponding plane buffers (16 plane buffers in this embodiment).

[0056] In S4, the plane conversion unit 101 determines whether conversion of the entire RAW image has been completed. If not, the plane conversion unit 101 returns the process to S1 to perform conversion on the next N×N pixel area. If it is determined that conversion of the entire RAW image has been completed, the plane conversion unit 101 proceeds to S5.

[0057] In S5, the plane conversion unit 101 outputs the plane data stored in the plane buffer to the RAW coding unit 110 (the frequency conversion unit 102 thereof) in accordance with a preset order.

[0058] The RAW encoding unit 110 simply performs encoding processing for each plane in accordance with a given target code amount, and therefore a description thereof will be omitted.

[0059] As a result of the above, the low-frequency component planes and high-frequency component planes of the three primary colors generated by the plane conversion unit 101 are planes configured with component values of the same cycle in the RAW image. Furthermore, for example, the W0 plane, W1 plane, ... of the W component can also be planes configured with pixel values of the same cycle, enabling efficient encoding.

[0060] [Second embodiment] Next, an image encoding device according to a second embodiment will be described with reference to Figures 1 and 7. The configuration of the second embodiment is similar to that of the first embodiment, but the conversion process performed by the plane conversion unit 101 is different. As a result, the allocation of the code amount performed by the code amount control unit 103 is also different. Other operations are similar to those of the first embodiment, and therefore will not be described again.

[0061] The plane conversion unit 101 in the second embodiment performs the following plane conversion for each color component, after dividing the 4x4 pixel array into R components {R0, R1}, G components {G0, G1, G2, G3}, B components {B0, B1}, and W components {W0, W1} as shown in Fig. 7. Note that the conversion of the G, R, and B components is the same as in the first embodiment, except that conversion of the W component is performed.

[0062] First, the plane conversion of the G component will be described. In the plane conversion of the G component within a 4×4 pixel area, the plane conversion unit 101 performs the following conversion using G0, G1, G2, and G3 present within that area. GL01 = G0 + G1 GH01 = G0 - G1 GL23 = G2 + G3 GH23 = G2 - G3 Next, the plane conversion unit 101 executes the following conversion using GL01 and GL23 obtained by the above calculation. GL = GL01 + GL23 GH = GL01 - GL23 Then, the plane conversion unit 101 transmits GL, GH, GH01, and GH23 obtained by the above calculation to the frequency conversion unit 102.

[0063] Next, the plane conversion of the R component will be described. In plane conversion of the R component within a 4×4 pixel region, the plane conversion unit 101 performs the following conversion using R0 and R1 within that region. RL = R0 + R1 RH = R0 - R1 The plane conversion unit 101 transmits the RL and RH obtained by the above conversion to the frequency conversion unit 102.

[0064] Next, the plane conversion of the B component will be described. In the plane conversion of the B component within a 4×4 pixel area, the plane conversion unit 101 performs the following conversion using B0 and B1 within that area. BL = B0 + B1 BH = B0 - B1 The plane conversion unit 101 transmits the BL and BH obtained by the above conversion to the frequency conversion unit 102 .

[0065] Finally, we will explain the plane transformation of the W component. As shown in Figure 7, there are four W0s and four W1s in a 4x4 pixel region. In this embodiment, the 4x4 pixel region is divided into four 2x2 pixel regions, and each 2x2 pixel region is divided into a group consisting of W0s located on the first line and a group consisting of W1s located on the second line, and the following transformation is performed. WL = ΣW0 + ΣW1 WH = ΣW0 - ΣW1 Here, ΣW0 represents the sum of four W0s, and ΣW1 represents the sum of four W1s. Therefore, WL is the sum of four W0s and four W1s, which is equivalent to taking the average of W0 and W1, and therefore corresponds to constructing the low-frequency component of frequency conversion and generating an intermediate pixel between W0 and W1. On the other hand, WH corresponds to the high-frequency component of frequency conversion.

[0066] Furthermore, the pixels in the WL plane are arranged in a uniform horizontal and vertical interval, which is an appropriate arrangement for separating high-frequency components from low-frequency components in the subsequent frequency conversion unit 102. This makes it possible to drop high-frequency components while appropriately retaining low-frequency components that have a large impact on image quality when quantizing in the quantization unit 104.

[0067] The frequency transform unit 102 performs a wavelet transform on each of the ten plane data, GL, GH, GH01, GH23, RL, RH, BL, BH, WL, and WH, input from the plane transform unit 101, and then sends the transform coefficients generated for each subband to the quantization unit 104.

[0068] The code amount control unit 103 determines the target code amount to be allocated to each picture and each plane according to the compression rate set by the user, and transmits the determined target code amount to the quantization unit 106. At this time, the code amount is allocated equally to each RGBW color component, and when allocating the code amount between planes of the same color, the code amount is allocated using the following relationship: G component: GL > GH > GH01 = GH23 R component: RL > RH B component: BL > BH W component: WL > WH

[0069] As mentioned above, GL corresponds to the low-frequency component of the two-stage frequency conversion, GH corresponds to the high-frequency component of the two-stage GH conversion, and GH01 and GH23 correspond to the high-frequency components of the first stage of frequency conversion. Therefore, by allocating more codes to the low-frequency components that have a large impact on image quality and reducing the amount of code allocated to the high-frequency components that have a smaller impact, efficient coding can be performed while maintaining image quality. Similarly, RL, BL, and WL correspond to low-frequency components, and RH, BH, and WH correspond to high-frequency components, so more codes are allocated to BL, RL, and WL, which have a large impact on image quality, and fewer codes are allocated to RH, BH, and WH.

[0070] In the second embodiment described above, RAW image data in a Dual Bayer+HDR array is converted into data of 10 planes with equal pixel spacing, and then encoded, thereby achieving highly efficient encoding.

[0071] The above describes an example in which RAW image data in the Dual Bayer+HDR array shown in Figure 7 is converted into 10 planes and encoded. As shown in Figure 7, there are four W0s and W1s in a 4x4 pixel area. However, the W0s and W1s obtained when the encoded data obtained by the above plane conversion are decoded are the average values of these four original W0s and W1s, so the four W0s in the 4x4 pixel area can only be reproduced as having the same value (average value), and the W1s also having the same value (average value).

[0072] Therefore, the following describes the conversion process of the W component by the plane conversion unit 101, which enables the reproduction of the four W0 and W1 pixels in the original 4×4 pixel area. Note that the plane conversion of the other components is assumed to be the same as above.

[0073] First, the plane conversion unit 101 subdivides the 4x4 pixel region in the RAW image data of the Dual Bayer+HDR array into four 2x2 pixel subregions (0) to (3), namely, a 2x2 pixel subregion (0) including G0 and G1, a 2x2 pixel subregion (1) including R0 and R1, a 2x2 pixel subregion (2) including B0 and B1, and a 2x2 pixel subregion (3) including G2 and G3.

[0074] Then, for one sub-region (i) (i=0, 1, 2, or 3), the plane conversion unit 101 performs the following conversion. WL(i) = W0 + W1 WH(i) = W0 - W1 The plane conversion unit 101 performs the above conversion for sub-regions (0) to (3) and transmits WL(i) and WH(i) obtained for each sub-region to the frequency conversion unit 102. Ultimately, the plane conversion unit 101 converts the input RAW image data into 16 plane data: GL, GH, GH01, GH23, RL, RH, BL, BH, WL(0), WH(0), WL(1), WH(1), WL(2), WH(2), WL(3), and WH(3), and transmits the data to the frequency conversion unit 102.

[0075] By performing processing in this manner, as in the first embodiment, the area of each plane becomes the same as 16 planes, and it is possible to standardize the size of the required line buffers, etc., and to realize a circuit configuration that considers the same processing time for each plane. Furthermore, in the WL and WH planes, pixels are arranged horizontally and vertically at intervals of one pixel compared to the original RAW image, so that aliasing during frequency conversion due to pixel skipping can be reduced compared to the other planes, allowing for more efficient encoding.

[0076] In the above embodiment, an example of R, G, and B as the three primary colors has been described, but Y (yellow), M (magenta), and C (cyan) may also be used, and therefore the present invention is not limited to RGB. Also, in the above embodiment, an example has been described in which white (W color) is arranged in a checkerboard pattern, but any color may be used as long as the luminance can be detected, and for example, when R, G, and B filters are used as the three primary colors, a yellow filter may be used as a luminance filter.

[0077] (Other Examples) The present invention can also be realized by supplying a program that realizes one or more functions of the above-described embodiments to a system or device via a network or a storage medium, and having one or more processors in the computer of the system or device read and execute the program.The present invention can also be realized by a circuit (e.g., ASIC) that realizes one or more functions.

[0078] The invention is not limited to the above-described embodiments, and various changes and modifications can be made without departing from the spirit and scope of the invention. Accordingly, the following claims are appended to apprise the public of the scope of the invention. [Explanation of symbols]

[0079] 100...imaging unit, 101...plane conversion unit, 102...frequency conversion unit, 103...code amount control unit, 104...quantization unit, 105...entropy coding unit, 110...RAW coding unit, 200...entropy decoding unit, 201...inverse quantization unit, 202...frequency inverse conversion unit, 203...Bayer conversion unit

Claims

1. 1. An image encoding device that encodes RAW image data obtained from an image sensor having a color filter array in which a plurality of filters for each of three primary colors and a plurality of filters of a specific color for luminance are arranged within an N×N pixel area, and the filters of the N×N pixel area are repeated, A conversion means for converting the RAW image data into a plurality of planes, each of which is composed of a single color component; encoding means for encoding each plane obtained by said converting means; The conversion means For each component representing the three primary colors, a plane composed of low-frequency component data and a plane composed of high-frequency component data are generated by referencing pixel values of the same component within the N×N pixel region; For each component representing the three primary colors, a plane composed of low-frequency component data and a plane composed of high-frequency component data are generated by referencing the pixel values of the same component diagonally adjacent within the N×N pixel region. An image encoding device comprising:

2. 2. The image encoding device according to claim 1, wherein the specific color is white.

3. 3. The image encoding device according to claim 1, wherein the filters of the specific color are arranged in a checkerboard pattern within the pixel area, and color filters are arranged diagonally adjacent to each other for each component representing the three primary colors.

4. 4. The image encoding device according to claim 1, wherein the conversion means generates, for the specific color component, a plurality of planes consisting of pixel values at the same relative positions in the N×N pixel area.

5. The image coding device according to any one of claims 1 to 3, characterized in that the conversion means divides the specific color components into two groups based on their relative positions within the NxN pixel region, calculates the average value of each group, and generates a low-frequency plane and a high-frequency plane based on the average value.

6. 4. The image coding device according to claim 1, wherein the conversion means divides the specific color components into a plurality of groups based on their relative positions within the N×N pixel region, and generates a low-frequency plane and a high-frequency plane for each of the plurality of groups.

7. The three primary colors are R, G, and B, the number of the G color filters arranged in the N×N pixel region is twice the number of the R color and B color filters; The conversion means For the R color and the B color, a first-order low frequency value and a first-order high frequency value are calculated by referring to pixel values of the same color component within the N×N pixel region, and a plane composed of the first-order low frequency values and a plane composed of the first-order high frequency values are generated; For the G color, a first-order low-frequency value and a first-order high-frequency value are calculated based on pixel values of the same color component within the N×N pixel region, and a second-order low-frequency value and a second-order high-frequency value are calculated based on the first-order low-frequency value, and a plane based on the first-order high-frequency value, a plane based on the second-order high-frequency value, and a plane based on the second-order low-frequency component are generated.

7. The image encoding device according to claim 1, wherein the first and second inputs are input to the image encoding unit.

8. the N×N pixel region is a 4×4 pixel region, Within the pixel region, eight white filters are arranged in a checkerboard pattern, four G filters are arranged diagonally adjacent to each other, and two R and B filters are arranged diagonally adjacent to each other, The conversion means For the R and B colors, two pixel values of the same color component that are diagonally adjacent within the 4×4 pixel region are referenced to calculate a first-order low frequency value and a first-order high frequency value, and two planes are generated: a plane composed of the first-order low frequency values and a plane composed of the first-order high frequency values; For the G color, two first-order low-frequency values and two first-order high-frequency values are calculated based on the values of four diagonally adjacent pixels in the 4×4 pixel region, and further, a second-order low-frequency value and a second-order high-frequency value are calculated based on the two first-order low-frequency values, and two planes corresponding respectively to the two first-order high-frequency values, a plane of the second-order high-frequency values, and a plane of the second-order low-frequency components are generated.

8. The image encoding device according to claim 7.

9. 9. The image coding device according to claim 8, wherein the conversion means divides four diagonally adjacent pixel values in the 4×4 pixel region into two groups of two adjacent pixels, and calculates a low-frequency value and a high-frequency value for each group, thereby calculating the two first-order low-frequency values and the two first-order high-frequency values.

10. the N×N pixel region is a 4×4 pixel region, In the pixel region, eight filters of the specific color, four filters of G color, and two filters each of R and B are arranged, 5. The image encoding device according to claim 4, wherein said conversion means generates, for the specific color, eight planes made up of pixel values at the same relative positions in said 4x4 pixel region.

11. the N×N pixel region is a 4×4 pixel region, In the pixel region, eight filters of the specific color, four filters of G color, and two filters each of R and B are arranged, the conversion means divides the eight pixel values of the 4×4 pixel area for the specific color into two groups having the same relative positional relationship, and generates two planes, a low frequency plane and a high frequency plane, based on the average value of each group; 6. The image encoding device according to claim 5.

12. the N×N pixel region is a 4×4 pixel region, In the pixel region, eight filters of the specific color, four filters of G color, and two filters each of R and B are arranged, 7. The image coding device according to claim 6, wherein the conversion means generates a plane composed of low-frequency component data and a plane composed of high-frequency components from each of the four 2x2 pixel regions contained in the 4x4 pixel region, thereby generating a total of eight planes.

13. The encoding means a transform means for performing a wavelet transform on a plane to be coded; quantization means for quantizing the subbands obtained by the conversion means in accordance with a quantization parameter corresponding to the type including the color of the plane to be coded; and entropy coding means for entropy-coding the quantized data obtained by the quantization means.

13. The image encoding device according to claim 1, wherein the first and second inputs are input to the image encoding unit.

14. 1. A control method for an image encoding device that encodes RAW image data obtained from an image sensor having a color filter array in which a plurality of filters for each of three primary colors and a plurality of filters of a specific color for luminance are arranged within an N×N pixel area, and the filters of the N×N pixel area are repeated, comprising: a conversion step of converting the RAW image data into a plurality of planes, each of which is composed of a single color component; an encoding step of encoding each plane obtained in the conversion step, The converting step comprises: For each component representing the three primary colors, a plane composed of low-frequency component data and a plane composed of high-frequency component data are generated by referencing pixel values of the same component within the N×N pixel region; For each component representing the three primary colors, a plane composed of low-frequency component data and a plane composed of high-frequency component data are generated by referencing the pixel values of the same component diagonally adjacent within the N×N pixel region.

2. A control method for an image encoding device comprising:

15. A program that, when read and executed by a computer, causes the computer to execute each step of the method according to claim 14.

Citation Information

Patent Citations

  • Image processor, electronic camera, and image processing program

    JP2003125209A

  • System and method for encoding mosaiced image data by using color transformation of invertible

    JP2006121669A

  • Solid-state imaging device, method of processing signal of the same, and image capturing apparatus

    JP2010136226A

  • Image processing apparatus, imaging apparatus, image processing method, and program

    JP2017085556A