Image processing device and method, and imaging device
The image processing device enhances subject area extraction by using size and distance criteria to refine the process, ensuring accurate separation of subjects and backgrounds for tailored image processing.
Patent Information
- Application Number
- JP2024143347
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Filing Date
- 2024-08-23
- Publication Date
- 2025-12-01
- Estimated Expiration
- 2040-03-19
AI Technical Summary
Existing image processing techniques struggle to accurately extract subject areas with precision, especially when different processing is required for subjects and backgrounds, leading to inconsistent image effects.
An image processing device that includes detection, selection, and extraction means to identify a first subject based on size and distance criteria, using distance information to refine the subject area extraction process.
Improves the accuracy of extracting subject areas from images, allowing for precise application of different processing effects to subjects and backgrounds.
Smart Images

Figure 0007778196000001 
Figure 0007778196000002 
Figure 0007778196000003
Abstract
Description
[Technical Field]
[0001] The present invention relates to an image processing device and method, and an imaging device, and more particularly to an image processing technique for images captured by a digital camera or the like. [Background technology]
[0002] Conventionally, image processing for correcting image clarity or brightness is applied to a portion of an image. In such processing, there is a technique for determining the subject region based on, for example, the size of a face, by referring to information on subject detection in order to extract the portion of the image to which the processing is applied.
[0003] Patent Document 1 discloses a technique for selecting a reference face from detected human faces and determining a human region that encompasses all faces based on the size of the reference face. [Prior art documents] [Patent documents]
[0004] [Patent Document 1] Japanese Patent Application Laid-Open No. 2005-208732 Summary of the Invention [Problem to be solved by the invention]
[0005] The configuration disclosed in Patent Document 1 is suitable for determining a rough area, such as a rectangle, that is likely to contain a person, and it is difficult to extract the subject area with the precision required to apply image processing to each area.
[0006] Furthermore, for example, when it is desired to perform different image processing on the subject and the background, it may not be possible to extract the subject area with sufficient accuracy depending on the scene, and the desired image processing effect may not be obtained.
[0007] The present invention has been made in consideration of the above problems, and has as its object to improve the accuracy of extracting a subject area from an image. [Means for solving the problem]
[0008] In order to achieve the above object, an image processing device of the present invention includes a detection means for detecting a predetermined subject from an image, a selection means for selecting a first subject from the detected subjects that satisfies a predetermined first condition, an acquisition means for acquiring distance information indicating subject distance for each of a plurality of regions that constitute the image, and an extraction means for extracting a first region that includes the first subject and a second region that does not include the first subject, based on the first subject, wherein the first condition is that the subject is the smallest of the subjects that have a size equal to or larger than a predetermined first size, and the extraction means extracts the first subject from the image based on the distance information when the image satisfies a predetermined second condition. and having a distance within a first range from the first object. The method is characterized in that a region is extracted as the first region. [Effects of the Invention]
[0009] According to the present invention, it is possible to improve the accuracy of extracting a subject area from an image. [Brief explanation of the drawings]
[0010] [Figure 1] 1 is a block diagram showing an example of the configuration of an imaging apparatus according to an embodiment of the present invention. [Figure 2] FIG. 2 is a diagram showing an example of the arrangement of an imaging unit in the embodiment. [Figure 3] FIG. 2 is a diagram showing an example of an input image according to the embodiment. [Figure 4] 1 is a flowchart showing processing in an embodiment. [Figure 5] 5A and 5B are diagrams showing examples of subject detection results in the embodiment. [Figure 6] FIG. 4 is a diagram showing an example of a defocus map in the embodiment. [Figure 7]FIG. 4 is a diagram showing an example of a region map according to the embodiment. [Figure 8] FIG. 4 is a diagram showing an example of a region map according to the embodiment. [Figure 9] 1A and 1B are diagrams showing an example of a human silhouette generation method and an area map according to an embodiment. DETAILED DESCRIPTION OF THE INVENTION
[0011] Hereinafter, embodiments will be described in detail with reference to the accompanying drawings. Note that the following embodiments do not limit the scope of the invention claimed. Although multiple features are described in the embodiments, not all of these multiple features are necessarily essential to the invention, and multiple features may be combined arbitrarily. Furthermore, in the accompanying drawings, the same reference numerals are used to designate the same or similar components, and redundant explanations will be omitted.
[0012] In this embodiment, a use case will be described in which contrast is corrected for each region of an image captured using an imaging device such as a digital camera, using various types of image information and distance information. In the following description, an imaging device will be described as an example of an image processing device, but the image processing device of this embodiment is not limited to an imaging device and may be, for example, a personal computer (PC) or the like.
[0013] Fig. 1 is a block diagram showing an example of the configuration when an image processing device of the present invention is applied to an imaging device 100. In Fig. 1, a control unit 101 is, for example, a CPU, which reads out an operation program for each block of the imaging device 100 from a ROM 102, expands it into a RAM 103, and executes it to control the operation of each block of the imaging device 100. The ROM 102 is a rewritable non-volatile memory, and stores parameters necessary for the operation of each block in addition to the operation program for each block of the imaging device 100. The RAM 103 is a rewritable volatile memory, and is used as a temporary storage area for data output in the operation of each block of the imaging device 100.
[0014] The optical system 104 forms an image of a subject on the imaging unit 105. The imaging unit 105 is an imaging element such as a CCD or CMOS sensor, and photoelectrically converts the optical image formed on the imaging unit 105 by the optical system 104, and outputs the resulting analog image signal to the A / D conversion unit 106. The A / D conversion unit 106 applies A / D conversion processing to the input analog image signal, and outputs the resulting digital image data to the RAM 103 for storage.
[0015] The image processing unit 107 applies various image processing to the image data stored in the RAM 103, including white balance adjustment, reduction / enlargement, normal signal processing such as noise reduction processing and development processing, and stores the processed image back in the RAM 103.
[0016] The recording medium 108 is a removable memory card or the like, and images processed by the image processing unit 107 stored in the RAM 103, images A / D converted by the A / D conversion unit 106, etc. are recorded as recorded images.
[0017] The display unit 109 is a display device such as an LCD, and displays various information, such as a through display of the subject image captured by the imaging unit 105 . The operation unit 110 is a plurality of operation members including a release switch, and is used to set the shooting mode, focal length, aperture value, exposure time, etc., as well as to instruct operations such as autofocus and release. The evaluation value acquisition unit 111 calculates evaluation values such as subject distance and defocus value from the obtained image data.
[0018] The subject detection unit 112 detects whether a specific subject exists in the image from the obtained image data, and if so, outputs the position (coordinates in the image) and size of the subject. Specific subjects can be various things such as a person's face or entire body, various animals, vehicles, etc.
[0019] FIG. 2 is a diagram showing the pixel arrangement configuration of the imaging unit 105 of FIG. 1. As shown in FIG. 2, a plurality of pixels 202 are regularly arranged two-dimensionally in the imaging unit 105. Each pixel 202 includes one microlens 201 and a pair of photoelectric conversion units 203, 204. Note that the planar shape of the photoelectric conversion units 203, 204 is not limited to this, and other planar shapes and arrangements may also be used. Furthermore, the number of photodiodes provided in each pixel 202 is not limited to two, and may be three or more (for example, four).
[0020] In this embodiment, signals are read out from the photoelectric conversion units 203 and 204, which are regularly arranged two-dimensionally. The signals read out from the photoelectric conversion units 203 of multiple pixels 202 (hereinafter referred to as "signals A") are connected to form image A, and the signals read out from the photoelectric conversion units 204 of multiple pixels 202 (hereinafter referred to as "signals B") are connected to form image B, thereby outputting images A and B, which are parallax images. Furthermore, by adding the signals A and B output from the photoelectric conversion units 203 and 204 of the same pixel 202, respectively, an A+B signal that can be used as a signal for a still image can be generated, and the generated A+B signal is recorded on the recording medium 108. Note that the method of reading out signals from the photoelectric conversion units 203 and 204 is not limited to this. For example, the A signal and the A+B signal may be read out from each pixel 202, and the A signal may be subtracted from the A+B signal to obtain the B signal.
[0021] By configuring the imaging unit 105 as shown in FIG. 2, a pair of light beams passing through different regions of the pupil of the optical system 104 can be formed as a pair of optical images, which can be output as image A and image B. Note that the method of acquiring image A and image B is not limited to the above, and various methods can be used. For example, images A and B may be images having parallax with respect to each other acquired by imaging devices such as multiple cameras installed at spatial intervals. Furthermore, parallax images acquired by an imaging device such as a single camera having multiple optical systems and imaging units may be respectively acquired as image A and image B.
[0022] 3 shows examples of images captured by the imaging device 100, and as an example, shows a case where multiple people are photographed as main subjects in different positions. In image 300 shown in FIG. 3(a), people 301, 302, and 303 are photographed side by side near the imaging device 100. In image 310 shown in FIG. 3(b), people 313, 311, and 312 are photographed from the foreground as seen from the imaging device 100 at different positions in the depth direction. In image 320 shown in FIG. 3(c), people 321, 322, 323, and 324 are photographed at positions far from the imaging device 100.
[0023] Next, the process of extracting a person area and a foreground area in this embodiment will be described with reference to the flowchart in Fig. 4. As a specific example of the process, images 300, 310, and 320 shown in Fig. 3 will be used as examples, with the image size being 6000 pixels wide x 4000 pixels high.
[0024] In S401, human faces are detected as subjects included in the input image. Specific detection methods include detecting faces using a known method such as that disclosed in Japanese Patent Application Laid-Open No. 10-162118, and outputting the coordinates and size of each detected face in the image.
[0025] 5 shows an example of performing the process of S401 on images 300, 310, and 320 shown in FIG. 3, and depicts a frame 500 indicating each face based on the position (center coordinates of the rectangle) and size (length of one side of the rectangle) of each face. Here, the detected face sizes are assumed to be 670 pixels for person 301 in image 300, 700 pixels for person 302, and 680 pixels for person 303. Similarly, the detected face sizes are assumed to be 660 pixels for person 311, 610 pixels for person 312, and 800 pixels for person 313 in image 310. Furthermore, the detected face sizes are assumed to be 320 pixels for person 321, 280 pixels for person 322, 350 pixels for person 323, and 120 pixels for person 324 in image 320.
[0026] In S402, a pair of parallax images, image A and image B, are acquired from the imaging unit 105, and data representing the spatial distribution of defocus values in the imaging range is output based on the acquired parallax images. In the following description, the data representing the spatial distribution of defocus values is referred to as a defocus map. The defocus value is a type of distance information, since it is the amount of focus deviation from the distance at which the optical system 104 is focused. Note that, as a method for acquiring the defocus value, a method for calculating the phase difference between parallax images, such as the method disclosed in Japanese Patent Application Laid-Open No. 2008-15754, can be used.
[0027] Defocus maps for images 300, 310, and 320 are shown in Fig. 6. The defocus map in Fig. 6 is expressed as a continuous grayscale, with the smaller the defocus value, the whiter the image (the higher the pixel value), and the area of person 302 in image 300 is expressed in gray, indicating an in-focus area (defocus value of zero). In image 310, the areas of people 312 and 313 have relatively large defocus values compared to the area of person 311. In image 320, the person in focus is far away and the depth of field is deep, so it can be seen that the defocus value of the background areas other than the people is smaller than in images 300 and 310.
[0028] In S403, one face to be used as a reference (hereinafter referred to as "reference face") is selected from the faces detected in each image. The selection method involves first selecting all detected faces that have a predetermined size Th1 (for example, Th1 = 10% of the long side of the image) or more, and then selecting the smallest face from among them as the reference face. In the example shown in this embodiment, the image size is 6000 pixels wide, so the predetermined size Th1 is set to 600 pixels, and faces of 600 pixels or more are selected.
[0029] Person 301 is selected as the reference face in image 300, and person 312 is selected as the reference face in image 310. In image 320, all of the people are smaller than Th1, so no reference face is selected here.
[0030] In S404, it is determined whether or not to refer to the defocus value to extract the person region. Specifically, if all faces detected in S401 are smaller than a predetermined size Th1, the defocus value is not referenced. However, if even one face with a size equal to or larger than Th1 is detected, the defocus value is referenced. This is because if all faces are small, i.e., if people are far from the image capture device 100, the image will be pan-focused, which tends to reduce the defocus value overall, making it difficult to separate the regions based on the defocus value. In this embodiment, in the case of images 300 and 310, it is determined that the defocus value should be referenced, and the process proceeds to S405. In the case of image 320, since all faces are smaller than Th1, it is determined that the defocus value should not be referenced, and the process proceeds to S408.
[0031] However, the method of determination is not limited to this, and various image information such as shooting conditions may be used. For example, when the F-number is narrowed to a predetermined value or more, the depth of field is similarly deep and the defocus value is likely to be small, and when the ISO sensitivity is high, it becomes difficult to separate the parallax between images A and B from noise, and the accuracy of the defocus value is likely to be insufficient. Therefore, when the F-number or ISO sensitivity is above a predetermined value, the defocus value may not be referenced.
[0032] In S405, the defocus value of the reference face is calculated. Specifically, in the defocus map generated in S402, the average value Def_base of the defocus values in the face detection frame of the reference face selected in S404 is calculated.
[0033] In S406, the range of defocus values for extraction as a person region is determined. To determine the range, upper and lower limit values of the defocus value that will be treated as the same distance as the reference face are calculated based on the defocus value Def_base of the reference face calculated in S405. The lower limit value, i.e., the threshold value on the far side of the reference face, is set to Def_base-5%, and the upper limit value, i.e., the threshold value on the near side of the reference face, is set to Def_base+5%. Therefore, a region with a defocus value of Def_base±5% is determined to be a person region.
[0034] The method for determining the range of defocus values indicating a person region is not limited to the method of comparing with a threshold calculated based on the average defocus value of the face region as described above, but may be calculated based on other evaluation values. For example, various methods are possible, such as a statistical method in which a defocus histogram is generated, the range is expanded while adding counts from the average defocus value of the reference face, and the values at the time when the integrated count value reaches a predetermined percentage (for example, 5% of the total count number) are set as the upper and lower limit values.
[0035] In S407, a region map showing the person region and the foreground region is generated based on the threshold calculated in S406. First, in the defocus map generated in S402, the region where the defocus value is within the range calculated in S406 is extracted as the person region, and the region on the foreground side where the defocus value is greater than the threshold is extracted as the foreground region, and a binarized version is generated.
[0036] In S406, even if a face's size is less than Th1, it is extracted as a person area as long as the defocus value is within the range calculated in S406. This is because a face located at approximately the same distance as the reference face (in this embodiment, a defocus value within Def_base ±5%) is presumed to be the face of a person related to the reference face, rather than a stranger unrelated to the reference face. This concept is important when the reference face is approximately equal to Th1. For example, an image may contain face A, whose size is slightly larger than Th1, and face B, whose size is slightly smaller than Th1, side by side. In such a case, face A will not be corrected as a person area, and face B will be corrected as a background area, resulting in an unnatural image.
[0037] Therefore, in this embodiment, a face that exists within a certain distance range from the reference face is determined to be a person area. In the example given above, face B is determined to be a person area and is not corrected. On the other hand, if face B exists at the same size but face A does not exist at the same size (if face B is not within a predetermined distance range from the reference image), face B is not determined to be a person area. This is because face B's size does not exceed Th1 by a little.
[0038] The results of extraction in S407 for image 300 are shown in Fig. 7. In image 300, people 301, 302, and 303 are located at approximately the same distance and have similar defocus values, so all three people are extracted as person regions 701, as shown in Fig. 7.
[0039] FIG. 8(a) shows a person region extracted from image 310. Person 312 selected as the reference face and person 311 located at a close distance are extracted as person regions. However, person 313, the foreground person, is not extracted as a person region because it is located farther away from person 312. FIG. 8(b) shows the result of extracting the foreground region of image 310. A value larger than the range calculated in S406, i.e., person 313 located in the foreground, is extracted. A region map is generated by integrating region 710 extracted as the person region and region 711 extracted as the foreground region. FIG. 8(c) shows the result, showing that all three people are extracted. To better fit the region map to the contours of the subject, shaping processing may be performed by referencing the pixel values of images 300 and 310, as described in, for example, JP 2017-11652 A.
[0040] Meanwhile, in S408, it is determined whether or not all detected faces are the faces of the main subject. As a method of determination, first, from the detected faces, those having a predetermined size Th2 or more (for example, Th2 = 5% of the long side of the image) are selected, and the smallest face among them is set as the reference face. Therefore, in the case of image 320, person 321, which is the smallest face among faces having a size of 6000 x 0.05 = 300 pixels or more, is selected as the reference face. Furthermore, the size of the face selected as the reference face is multiplied by a predetermined ratio (for example, 0.8) to calculate a size threshold of 300 x 0.8 = 240 pixels, and only faces having a size equal to or greater than the calculated threshold are determined to be main faces.
[0041] Person 321, the reference face, and person 323, who has a larger face, are determined to be the main face, and person 322, who has a smaller face than person 321, is also determined to be the main face because the size of his face is equal to or greater than the threshold of 240 pixels. On the other hand, person 324 is not determined to be the main face because his face is small enough to be below the threshold. This is to avoid the problem of the determination results being inconsistent when there are multiple people with slightly different face sizes, or the problem of people who are not actually main being determined to be main if Th2 is set too small.
[0042] In S409, a region map is generated based on the determination result of S408. Specifically, a silhouette of a simple human model is first assigned to each person determined to be a main face. FIG. 9 shows an example of the silhouette assignment in S409. FIG. 9(a) is a diagram showing an input image 800 and a frame 801 indicating a detected face, and FIG. 9(b) is a diagram showing an example in which a silhouette image 810 consisting of a face portion 902 and a torso portion 903 is generated for an area where a person is assumed to exist based on the position and size of the detected face. Note that the information to be referenced is not limited to the face detection result, and may also be the detection result of an area including the entire body of a person.
[0043] Figure 9(c) shows the results of performing the above process on image 320, generating silhouettes of the three main faces detected in S408. By performing shaping processing on image 320 with reference to pixel values, such as that described in JP 2017-11652 A, the region map can be further adapted to fit the contours of the subject, resulting in the final region map. An example of the resulting region map is shown in Figure 9(d).
[0044] In S410, an image in which the contrast of only a portion of the image is corrected is generated using the region map generated in S407 or S409. Here, the portion of the image 300 refers to the background region 702, which does not include a person, in the generated region map shown in FIG. 7. Specifically, the correction method involves first generating an image in which the contrast of the entire image 300 is corrected. Existing techniques, such as a method for correcting local contrast without changing the brightness of the low-frequency region, as described in JP 2019-28537 A, may be used for contrast correction. This image is then combined with an image developed without contrast correction by selecting an image without contrast correction for the person region 701 and an image with contrast correction for the background region based on the region map shown in FIG. 7. In this case, the two images may be weighted and added together. The weight of the weighted addition may also vary depending on the position within the image. For example, the weights of the uncorrected image and the corrected image may be equal at the boundary of the region, and the weights of the uncorrected image for the person region 701 and the corrected image for the background region may increase with increasing distance from the boundary of the region. This makes it possible to generate an image in which the contrast of only the background region 702 in the region map is corrected. Images 310 and 320 are also processed in the same manner based on the region map.
[0045] In this way, the subject detection information and distance information are used to extract a person area from the captured image, an area map is generated, and gradation processing is performed for each area, after which the process ends.
[0046] As described above, according to this embodiment, when there are multiple subjects, by extracting an area containing only the main subject using the smallest face among faces having a predetermined size as a reference, it is possible to obtain the desired effect in image processing for each area.
[0047] In this embodiment, as described with reference to Fig. 2, a configuration has been described in which distance information is generated based on the phase difference between multiple object images generated by light beams arriving from different regions of the pupil of the imaging optical system, but other configurations or means may be used instead or in combination. For example, by using a configuration in which distance can be measured using a TOF (Time Of Flight) camera or ultrasound, it is possible to improve distance measurement performance for objects with little change in pattern.
[0048] In addition, in this embodiment, the case where the main subject is a person and the contrast of the background is corrected has been described, but the area to be corrected and the correction method are not limited to this and other combinations may be used. For example, the main subject may be an animal, a building, a vehicle, etc., and the present invention can be applied by using a technology to detect the subject.
[0049] In addition, in this embodiment, a case has been described in which the background region is corrected and the person region is not corrected, but the present invention is not limited to this and can be applied to various processes that utilize the divided (extracted) region information. For example, after region division as in this embodiment, it can be used in a process in which the person region is corrected and the background region is not corrected (for example, a process in which the contrast of the skin region is reduced to impart a skin-beautifying effect).
[0050] Furthermore, image processing is not limited to contrast correction, and may also correct brightness, hue, saturation, sharpness, etc. For example, by correcting brightness only in the person area, it is possible to adjust the face of a person that has become dark due to backlighting or the like after shooting.
[0051] <Other embodiments> The present invention may be applied to a system made up of a plurality of devices, or to an apparatus made up of a single device.
[0052] The present invention can also be realized by supplying a program that realizes one or more functions of the above-described embodiments to a system or device via a network or a storage medium, and having one or more processors in the computer of the system or device read and execute the program.The present invention can also be realized by a circuit (e.g., ASIC) that realizes one or more functions.
[0053] The invention is not limited to the above-described embodiments, and various changes and modifications can be made without departing from the spirit and scope of the invention. Accordingly, the following claims are appended to apprise the public of the scope of the invention. [Explanation of symbols]
[0054] 100: imaging device, 101: control unit, 102: ROM, 103: RAM, 104: optical system, 105: imaging unit, 106: A / D conversion unit, 107: image processing unit, 108: recording medium, 109: display unit, 110: operation unit, 111: evaluation value acquisition unit, 112: subject detection unit, 201: microlens, 202: pixel, 203, 204: photoelectric conversion unit
Claims
1. a detection means for detecting a predetermined subject from an image; a selection means for selecting a first subject that satisfies a predetermined first condition from the detected subjects; an acquisition means for acquiring distance information indicating a subject distance for each of a plurality of regions constituting the image; an extraction unit that extracts a first region including the first subject and a second region not including the first subject based on the first subject; the first condition is that the subject is the smallest among the subjects having a size equal to or larger than a predetermined first size; The image processing device is characterized in that, when the image satisfies a predetermined second condition, the extraction means extracts from the image, based on the distance information, an area that includes the first subject and is at a distance within a first range from the first subject, as the first area.
2. 2. The image processing device according to claim 1, wherein the extraction means extracts, as the second region, an area farther from the first subject than the first range of distances when the image satisfies a predetermined second condition.
3. 3. The image processing device according to claim 1, wherein the second condition includes at least one of the following: at least one of the detected subjects is larger than a predetermined size; the F-number when the image was captured is not narrower than a predetermined value; and the ISO sensitivity when the image was captured is lower than a predetermined sensitivity.
4. The image processing device according to any one of claims 1 to 3, characterized in that, when the image does not satisfy the second condition, the extraction means also extracts, as the first region, a region of the detected subjects other than the first subject that has a size equal to or greater than a threshold based on the size of the first subject.
5. 5. The image processing device according to claim 1, wherein the subject is any one of a face, a person, an animal, a building, and a vehicle.
6. 6. The image processing device according to claim 1, further comprising an image processing unit for performing image processing on the first area or the second area.
7. 7. The image processing device according to claim 1, further comprising image processing means for performing different image processing on the first area and the second area.
8. 8. The image processing device according to claim 6, wherein the image processing includes at least one of brightness, contrast, hue, saturation, and sharpness.
9. an imaging means for capturing an image; The image processing device according to any one of claims 1 to 8, The image pickup device is characterized in that the image processing device processes the image captured by the image pickup means.
10. a detection step in which a detection means detects a predetermined subject from an image; a selection step in which a selection means selects a first subject that satisfies a predetermined first condition from the detected subjects; an acquisition step in which an acquisition means acquires distance information indicating a subject distance for each of a plurality of regions constituting the image; The extraction means extracts a first region including the first subject based on the first subject; and an extraction step of extracting a second region that does not include the first subject, the first condition is that the subject is the smallest among the subjects having a size equal to or larger than a predetermined first size; An image processing method characterized in that, in the extraction step, when the image satisfies a predetermined second condition, an area from the image that includes the first subject and is at a distance from the first subject within a first range is extracted as the first area based on the distance information.
11. A program for causing a computer to function as each of the means of the image processing apparatus according to any one of claims 1 to 8.
Citation Information
Patent Citations
Image processing apparatus and method
JP2005208732A
Camera
JP2007300221A
Imaging apparatus, its control method, program, and storage medium
JP2008011264A
Imaging apparatus and imaging method, and imaging control program
JP2009253925A
Electronic camera and image processing apparatus
JP2010104061A