Image processing apparatus, image processing method, and computer program

The image processing device enhances detection of personal belongings and region extraction by combining segmentation and distance information maps, addressing the limitations of existing methods in complex scenes.

JP2026069798APending Publication Date: 2026-04-27CANON KK
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
CANON KK
Filing Date
2024-10-15
Publication Date
2026-04-27

AI Technical Summary

Technical Problem

Existing image processing methods struggle to accurately detect personal belongings attached to a person and extract regions where the foreground and background blend together, such as in complex scenes with human hair.

Method used

An image processing device that combines segmentation maps from image matting and distance information to create a composite map, using gain waveforms and confidence maps to weight and blend these maps based on detection accuracy.

Benefits of technology

Accurately detects items attached to a person and performs precise region extraction even in areas where the foreground and background blend together, improving detection accuracy in complex scenes.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2026069798000001_ABST
    Figure 2026069798000001_ABST
Patent Text Reader

Abstract

For example, the present invention provides an image processing device that can accurately detect objects attached to a person while also performing accurate region extraction even in areas where the foreground and background blend together. [Solution] The image processing apparatus includes: a segmentation means that outputs a segmentation map based on matting results; a distance information output means that outputs distance information; a main subject detection means that detects a main subject; a distance map output means that outputs a distance map based on the distance information of the main subject, based on the detection results by the main subject detection means and the distance information; and a composite map generation means that performs a composite of the segmentation map and the distance map.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to an image processing apparatus, an image processing method, a computer program, and the like.

Background Art

[0002] Conventionally, a method for detecting a region of a specific subject in an image has been known. For example, as a method for extracting a region of a subject, a technique called image matting that can perform high-precision object extraction, as in Non-Patent Document 1, is known.

[0003] Also, as a technique for acquiring two-dimensional information indicating the defocus amount distribution of the image of each pixel in an image, a technique for estimating distance information from a single image using machine learning, as in Non-Patent Document 2, is known.

Prior Art Documents

Non-Patent Documents

[0004]

Non-Patent Document 1

Non-Patent Document 2

Summary of the Invention

Problems to be Solved by the Invention

[0005] However, when using image matting for person detection, there was a problem in that it could not accurately detect, for example, the belongings attached to the person. On the other hand, when using distance information, there was a problem in that it could not accurately extract areas of complex scenes where the foreground and background blend together, such as a person's hair.

[0006] This invention has been made in view of the above-mentioned problems, and one of its objectives is to provide an image processing device that can accurately detect, for example, personal belongings attached to a person, while also performing accurate region extraction even in areas where the foreground and background blend together. [Means for solving the problem]

[0007] In an image processing device, A segmentation means that outputs a segmentation map based on mating results, Distance information output means for outputting distance information, A main subject detection means for detecting the main subject, A distance map output means outputs a distance map based on the distance information of the main subject, based on the detection result by the main subject detection means and the distance information. A composite map generation means for combining the segmentation map and the distance map, It is characterized by having the following features. [Effects of the Invention]

[0008] According to the present invention, it is possible to realize an image processing device that can accurately detect, for example, items attached to a person, while also performing accurate region extraction even in areas where the foreground and background blend together. [Brief explanation of the drawing]

[0009] [Figure 1] This is a functional block diagram showing an example configuration of the image processing device 100 of Embodiment 1 of the present invention. [Figure 2] This is a functional block diagram showing an example configuration of the region extraction unit 111 of the image processing device 100 of Embodiment 1. [Figure 3] This is a flowchart illustrating an example of the operation of the distance map output unit 203 in Embodiment 1. [Figure 4] (A) to (C) are diagrams illustrating the distance information values ​​of the face region according to Embodiment 1. [Figure 5] This is a flowchart illustrating an example of the operation of the composite map generation unit 204 in Embodiment 1. [Figure 6] (A) is a diagram showing an example of a distance map output by the distance map output unit 203, (B) is a diagram showing an example of a segmentation result (segmentation map) output by the segmentation output unit 201, and (C) is a diagram showing an example of a composite map output by the composite map generation unit. [Figure 7] This is a functional block diagram showing an example configuration of the region extraction unit 111 of the image processing apparatus of Embodiment 2. [Figure 8] (A) is a figure showing an example of the gain waveform of the gain table of the segmentation processing unit 703, (B) is a figure showing the histogram of the segmentation output in the hair-thin region of region 611 in Figure 6(B), and (C) is a figure showing the histogram when the gain waveform 801 is applied to the output of the segmentation processing unit in the hair-thin region mentioned above. [Figure 9] This is a functional block diagram showing an example configuration of the region extraction unit 111 of the image processing apparatus of Embodiment 3. [Figure 10] This is a flowchart illustrating an example of an image processing method in the composite map generation unit 904 of Embodiment 3. [Figure 11] This is a functional block diagram showing an example configuration of the region extraction unit 111 of the image processing apparatus of Embodiment 4. [Figure 12] This is a flowchart illustrating an example of the operation of the first distance map output unit 1103 and the second distance map output unit 1104 of Embodiment 4. [Figure 13] Figures (A) to (E) illustrate the distance information of the face region in Embodiment 4. [Figure 14] This is a flowchart illustrating an example of the operation of the composite map generation unit 1105 in Embodiment 4. [Figure 15] (A) to (E), (A’), and (D’) are diagrams for explaining the operations of steps S1401 to S1407 in Embodiment 4.

Embodiments for Carrying Out the Invention

[0010] Hereinafter, embodiments of the present invention will be described with reference to the drawings. However, the present invention is not limited to the following embodiments. In each figure, the same members or elements are denoted by the same reference numerals, and duplicate explanations are omitted or simplified.

[0011] <Embodiment 1> Hereinafter, the image processing apparatus in Embodiment 1 of the present invention will be described. FIG. 1 is a functional block diagram showing a configuration example of the image processing apparatus 100 according to Embodiment 1 of the present invention.

[0012] Note that some of the functional blocks shown in FIG. 1 are realized by causing a CPU or the like as a computer included in the image processing apparatus (not shown) to execute a computer program stored in a memory as a storage medium (not shown).

[0013] However, some or all of them may be realized by hardware. As the hardware, a dedicated circuit (ASIC), a processor (reconfigurable processor, DSP), or the like can be used.

[0014] Also, each of the functional blocks shown in FIG. 1 does not have to be built in the same housing, and may be constituted by separate devices connected via signal paths. Note that the above description regarding FIG. 1 also applies to FIGS. 2, 7, 9, and 11 in the same manner.

[0015] Note that the image processing apparatus includes, for example, electronic devices having an imaging function such as a digital still camera, a digital movie camera, a smartphone with a camera, a tablet computer with a camera, an in-vehicle camera, a drone camera, and a camera mounted on a robot.

[0016] In Figure 1, the control unit 101 has a built-in CPU as a computer and reads computer programs for the operation of each block of the image processing device 100 from the ROM 102 (described later), expands them into the RAM 103 (described later), and executes them. In this way, it controls the operation of each block of the image processing device 100.

[0017] ROM 102 is an electrically erasable and recordable non-volatile memory that stores the operation programs for each block of the image processing device 100, as well as parameters necessary for the operation of each block. RAM 103 is a rewritable volatile memory that is used as a temporary storage area for data output by the operation of each block of the image processing device 100.

[0018] The optical system 104 consists of a lens group including a zoom lens and a focus lens, and forms an image of the subject on the imaging unit 105, which will be described later. The imaging unit 105 is an image sensor such as a CCD or CMOS sensor, and it performs photoelectric conversion of the optical image formed on the imaging unit 105 by the optical system 104, and outputs the resulting analog image signal to the A / D conversion unit 106.

[0019] In this embodiment, for example, a Bayer array color filter of RGB is placed on the light-receiving surface of the image sensor, and one of the R, G, or B signals is obtained from each pixel. The A / D conversion unit 106 converts the input analog image signal into a digital image signal and outputs the obtained digital image data to the RAM 103.

[0020] The image processing unit 107 applies various image processing functions, such as white balance adjustment, noise reduction, color interpolation (debayering), and gamma processing, to the image data stored in the RAM 103. The image processing unit 107 may also simultaneously generate a low-resolution image (e.g., VGA size) for thumbnail display (hereinafter referred to as a thumbnail image) along with the high-resolution image (also called the main image).

[0021] The recording unit 108 is a removable memory card or the like, and the main image and thumbnail image processed by the image processing unit 107 are associated and recorded as a recorded image via the RAM 103.

[0022] The display unit 109 is a display device such as an LCD, and displays the main image, thumbnail images, and user interface (GUI) for receiving user instructions stored in the RAM 103 and recording unit 108.

[0023] The region extraction unit 111, as will be described in detail later, outputs the region of a person as a likelihood map using machine learning or other methods for the main image or thumbnail image processed by the image processing unit 107.

[0024] The face detection unit 112 detects the face region of a person in the main image or thumbnail image processed by the image processing unit 107 using known methods such as machine learning.

[0025] Figure 2 is a functional block diagram showing an example configuration of the region extraction unit 111 of the image processing apparatus 100 of Embodiment 1.

[0026] The segmentation output unit 201 outputs a segmentation result (segmentation map) that can be matted for high-precision person detection to the main image or thumbnail image processed by the image processing unit 107. In this case, a method called image matting, as described in Non-Patent Document 1, etc., may be used.

[0027] Furthermore, the segmentation output unit 201 functions as a segmentation means that executes a segmentation step that outputs a segmentation map based on the mating results.

[0028] This makes it possible to extract complex scene regions where foreground and background elements, such as human hair or animal fur, blend together. The output of this segmentation output unit 201 is input to the composite map generation unit 204.

[0029] Furthermore, the composite map generation unit 204 functions as a composite map generation means that combines a segmentation map and a distance map. The detailed processing content of the composite map generation unit 204 will be described later.

[0030] The distance information output unit 202 outputs image distance information from the main image or thumbnail image processed by the image processing unit 107, using, for example, a method for estimating image distance information as described in Non-Patent Document 2.

[0031] The distance information of the image output from the distance information output unit 202 is input to the distance map output unit 203. Here, the distance information output unit 202 functions as a distance information output means that executes a distance information output step that outputs distance information.

[0032] Next, Figure 3 is a flowchart illustrating an example of the operation of the distance map output unit 203 in Embodiment 1. The CPU and other components within the control unit 101 execute computer programs stored in memory, sequentially performing each step of the flowchart in Figure 3. The processing flow in Figure 3 is executed periodically.

[0033] In step S301 of Figure 3, the face detection unit 112 performs face detection. Next, the process proceeds to step S302. If a face is detected, the process proceeds to step S303. If no face is detected, the process flow shown in Figure 3 ends.

[0034] Next, in step S303, it is determined whether or not the face is that of the main subject. That is, if there is one face detection frame or two or more face detection frames, it is determined whether or not the face in question is that of the main subject, and if it is the main subject, the process proceeds to step S304. Here, step S303 functions as a main subject detection step (main subject detection means) for detecting the main subject.

[0035] If there are two or more face detection frames in step S303 and the target face is not the main subject, repeat step S303 until the main subject's face is found. If the answer in step S303 is Yes, proceed to step S304 and obtain the distance information value of the face region.

[0036] Figures 4(A) to 4(C) are diagrams illustrating the distance information values ​​of the face region according to Embodiment 1. Figure 4(B) shows an example of distance information values, where the brightness is displayed higher the closer the distance from the camera. In step S304, the average value of the distance information of the face region 410 is obtained as the distance information value of the face region.

[0037] Next, the process proceeds to step S305, where a gain waveform table (gain table) is generated based on the distance information of the face region. That is, a gain table is generated based on the average value of the distance information of the face region 410 acquired in step S304. An example of the gain waveform of the gain table is shown in Figure 4(A). In Figure 4(A), the horizontal axis represents the distance information value and the vertical axis represents the gain value, with 401 showing the gain waveform.

[0038] The position at distance t3 indicates the peak of the gain value, which is the average value of the distance information for the face region 410. The gain waveform 401 decreases monotonically towards distance t1 on the near side, and the gain becomes 0 at distance t1. On the far side, it decreases monotonically towards distance t4, and the gain becomes 0 at distance t4. t2 shows the distance information for region 411 in Figure 4(B).

[0039] After generating the gain table in step S305 of Figure 3, the process proceeds to step S306, where a distance map is generated by referring to the gain table. Specifically, the distance map is generated by applying the gain waveform shown in Figure 4(A), which was generated in step S305, to the distance information of the image output from the distance information output unit 202.

[0040] Here, steps S304 to S306 function as a distance map output step in the distance map output unit 203, which outputs a distance map based on the distance information of the main subject, based on the detection result from the main subject detection step and the distance information. The distance map output unit 203 also functions as a distance map output means. The processing flow shown in Figure 3 ends after step S306.

[0041] Figure 4(C) shows the distance information (hereinafter referred to as the distance map) after applying the gain waveform 401 to the distance information in Figure 4(B). In Figure 4(B), face region 410 is displayed with a lower brightness because it is farther away from region 411, but in Figure 4(C), face region 420 is displayed with a higher brightness than region 421.

[0042] Thus, in this embodiment, the weight of the face region, which is a more important region, is increased. In other words, the distance map output means performs weighting processing on the output of the distance information output means so that the weight of the main subject becomes larger.

[0043] In this embodiment, since the face region is considered important, a distance map is created based on the average of the distance information of the face region 410, but this is not limited to this. For example, the gain value may be set to be greater than or equal to a predetermined value between the maximum and minimum values ​​of the distance information of the face region.

[0044] Alternatively, instead of detecting the face region, the entire human body region may be detected, and distance information for the human body region may be referenced. Furthermore, if animals or other specified objects are set as the main subject and can be detected, distance information for the region of the animals or specified objects may be referenced.

[0045] Let's return to the explanation of Figure 2. The distance map output by the distance map output unit 203 and the segmentation result (segmentation map) output by the segmentation output unit 201 are input to the composite map generation unit 204.

[0046] Figure 5 is a flowchart illustrating an example of the operation of the composite map generation unit 204 in Embodiment 1. The CPU and other components within the control unit 101 execute computer programs stored in memory, sequentially performing each step of the flowchart in Figure 5. The processing flow in Figure 5 is executed periodically.

[0047] In step S501, the distance map value (pixel value of the distance map) output by the distance map output unit 203 is obtained. Next, the process proceeds to step S502, and if the distance map value (pixel value of the distance map) is smaller than the threshold, the process proceeds to step S503, where the segmentation result (segmentation map) output by the segmentation output unit 201 is selected, and the process proceeds to step S505.

[0048] If the pixel value of the distance map is greater than the threshold in step S502, it is determined to be No, and the process proceeds to step S504, where the distance map result output by the distance map output unit 203 is selected, and the process proceeds to step S505.

[0049] Here, we will use Figure 6 to explain the operation in steps S502 to S504. Figure 6(A) shows an example of a distance map output by the distance map output unit 203, and Figure 6(B) shows an example of a segmentation result (segmentation map) output by the segmentation output unit 201. Figure 6(C) shows an example of a composite map output by the composite map generation unit.

[0050] Note that 601 in Figure 6(A) represents the hair portion of the subject. Generally, distance maps cannot detect areas in complex scenes where foreground and background elements, such as human hair or animal fur, blend together. Therefore, while it would be preferable to obtain a value close to the face area, accurate detection is not possible.

[0051] On the other hand, 611 in Figure 6(B) is also a portion of the subject's hair, but because it is a segmentation output using image matting, it can correctly detect areas such as a person's hair.

[0052] In Figure 6(A), 602 represents an object attached to a person (e.g., a bag). Because it is at the same distance as the person, it can be correctly detected in the distance map. On the other hand, 612 in Figure 6(B) is also an object attached to a person, but because it is a segmentation output that detects the human body region, it cannot be correctly detected.

[0053] As shown in Figure 6(C), a segmentation result (segmentation map) is used for the hair area of ​​person 621, and a distance map is used for the area of ​​the belongings associated with person 622.

[0054] By performing the processing in steps S502 to S504 in this way, it becomes possible to accurately detect objects attached to a person while also performing region extraction even in areas where the foreground and background blend together, such as hair.

[0055] Let's return to the explanation of the flowchart in Figure 5. In step S505, it is determined whether or not it is the last pixel. If it is not the last pixel, the process returns to step S501 and repeats the operations from steps S501 to S504. If it is determined in step S505 that it is the last pixel, the processing flow in Figure 5 is terminated.

[0056] Thus, in this embodiment, the composite map generation means combines the segmentation map and the distance map based on the values ​​of the distance map. Here, steps S501 to S505 function as composite map generation steps that combine the segmentation map and the distance map.

[0057] In this embodiment, the segmentation result (segmentation map) and the distance map result were switched according to the distance map value, but this is not limited to this. For example, any other method that can extract regions where foreground and background blend, such as hair, using segmentation, and regions where belongings attached to a person are extracted using a distance map, would also be acceptable.

[0058] In other words, the system uses image recognition to distinguish between areas where the foreground and background blend, such as hair, and objects associated with a person. If the former is identified, the segmentation result (segmentation map) should be selected. If the latter is identified, the distance map should be selected. The system is not limited by its configuration or method.

[0059] In steps S503 and S504, one of the segmentation map and the distance map are selected and combined, respectively. However, when combining them, it is also possible to increase the blending ratio (weighting) of each, similar to alpha blending.

[0060] Furthermore, the composite map generation means performs synthesis by increasing the synthesis ratio of the distance map when the distance map value is less than or equal to a predetermined value, and by increasing the synthesis ratio of the segmentation map when the distance map value is greater than the predetermined value, but is not limited to this.

[0061] As described above, according to Embodiment 1, it is possible to accurately detect, for example, items attached to a person, while also performing accurate region extraction even in areas where the foreground and background blend together, such as a person's hair.

[0062] <Embodiment 2> Embodiment 2 of the present invention generates a composite map using the processed output of the segmentation result (segmentation map). This improves the accuracy of segmentation in areas where the foreground and background blend together, such as hair. Furthermore, it becomes possible to freely control the segmentation accuracy according to the application.

[0063] Furthermore, the block diagram showing an example configuration of the image processing device 100 in Figure 1, the flowchart of the region extraction unit 111 in Figure 3, and the operation of the composite map generation unit 204 in Figure 5 are the same as those described in Embodiment 1, so a detailed explanation will be omitted.

[0064] Figure 7 is a functional block diagram showing an example configuration of the region extraction unit 111 of the image processing apparatus of Embodiment 2. Note that Embodiment 2 differs from Embodiment 1 in some aspects of the configuration of the region extraction unit 111 shown in Figure 2.

[0065] The segmentation output unit 701 is equivalent to the segmentation output unit 201 described in Figure 2 of Embodiment 1, so a detailed explanation is omitted. The output of this segmentation output unit 701 is supplied to the segmentation processing unit 703. Here, the segmentation processing unit 703 functions as a segmentation processing means for processing the segmentation map.

[0066] Figure 8(A) shows an example of the gain waveform of the gain table of the segmentation processing unit 703, where the horizontal axis represents the input to the segmentation processing unit 703, the vertical axis represents the output to the segmentation processing unit 703, and 801 represents the gain waveform.

[0067] i3 represents the maximum input value, and o3 represents the maximum output value. Gain waveform 801 outputs the maximum value o3 if the input value is greater than i2, and the minimum value 0 if the input value is less than i1. If the input value is between i1 and i2, it outputs a value between the minimum value 0 and the maximum value o3.

[0068] Figure 8(B) shows the histogram of the segmentation output for region 611 in Figure 6(B) in the hair region. The horizontal axis represents the output value, with 0 at the left and the maximum input value i3 at the right. The vertical axis represents the amount of the histogram, with larger values ​​indicating a greater number of pixels corresponding to that output value.

[0069] Figure 8(C) shows a histogram of the output of the segmentation processing unit in the hair region described above, when the gain waveform 801 is applied. The horizontal axis represents the output value, with 0 at the left end and the maximum input value o3 at the right end.

[0070] As shown in Figure 8(C), by applying the gain waveform 801, the amount of intermediate values ​​is reduced compared to Figure 8(B), resulting in a region extraction result with sharper edges. Furthermore, the gain waveform 801 may be made changeable by the user, for example.

[0071] Let's return to the explanation of Figure 7(B). The segmented output after processing, output by the segmentation processing unit 703, is input to the composite map generation unit 705. Note that the operation of the distance information output unit 702, distance map output unit 704, and composite map generation unit 705 is the same as that of the distance information output unit 202, distance map output unit 203, and composite map generation unit 204 described in Figure 2 of Embodiment 1, so a detailed explanation will be omitted.

[0072] In this embodiment 2, the segmentation output is processed by applying the gain waveform 801, and the composite map generation means performs a composite of the processed map output by the segmentation processing means and the distance map.

[0073] This configuration allows for clearer region extraction results. Therefore, it is possible to improve the accuracy of segmentation in areas where the foreground and background blend together, such as hair. Furthermore, by changing the gain waveform 801, the segmentation accuracy can be arbitrarily controlled according to the application.

[0074] <Embodiment 3> In Embodiment 3 of the present invention, a confidence map is output from the distance information output, and segmentation is used in areas where the confidence level of the distance information is low.

[0075] The following describes the image processing apparatus in Embodiment 3 of the present invention. The block diagram showing an example configuration of the image processing apparatus 100 in Figure 1 and the flowchart of the region extraction unit 111 in Figure 3 are the same as those described in Embodiment 1, so a detailed explanation will be omitted.

[0076] Furthermore, since Embodiment 3 differs from Embodiment 1 in the configuration and operation of the region extraction unit 111 shown in Figure 2, it will be explained using Figure 9. Figure 9 is a functional block diagram showing an example of the configuration of the region extraction unit 111 of the image processing apparatus in Embodiment 3.

[0077] Furthermore, since the segmentation output unit 901 and the distance map output unit 903 are equivalent to those described in Figure 2 of Embodiment 1, the segmentation output unit 201 and the distance map output unit 203, respectively, a detailed explanation will be omitted.

[0078] In the distance information output unit 902 of Embodiment 3, distance information of an image is output, and a distance information confidence map relating to the reliability of the outputted distance information is also output. That is, the distance information output unit 902, as a means of outputting distance information, outputs a distance information confidence map relating to the reliability of the distance information.

[0079] The distance information of the image output from the distance information output unit 902 is input to the distance map output unit 903, and the distance information confidence map output from the distance information output unit 902 is input to the composite map generation unit 904.

[0080] Figure 10 is a flowchart illustrating an example of the image processing method in the composite map generation unit 904 of Embodiment 3. The operation of the composite map generation unit 904 will be explained using the flowchart in Figure 10.

[0081] Furthermore, the CPU and other components within the control unit 101 execute the computer program stored in memory, sequentially performing the operations at each step of the flowchart in Figure 10. Note that the processing flow in Figure 10 is executed periodically.

[0082] In step S1001, the value of the distance information confidence map output from the distance information output unit 902 is obtained. Next, the process proceeds to step S1002, where it is determined whether the confidence map value is below a threshold. If it is determined that the value of the distance information confidence map is less than the threshold, the process proceeds to step S1005, where the segmentation result is selected.

[0083] If it is determined in step S1002 that the value of the distance information confidence map is greater than the threshold, the process proceeds to step S1003. The operations of steps S1003 to S1007 are equivalent to the operations of steps S501 to S505 in Figure 5 of Embodiment 1, so a detailed explanation is omitted.

[0084] In the flowchart of Figure 10, if the value of the distance information confidence map is greater than a predetermined value and the value of the distance map is greater than a predetermined threshold, the synthesis ratio of the distance map is increased; otherwise, the synthesis ratio of the segmentation map is increased.

[0085] As described above, a confidence map is output from the distance information output. For areas with low confidence in the distance information, the segmentation result (segmentation map) is selected, and for areas with high confidence in the distance information, the distance map result is selected.

[0086] Furthermore, as in Figure 5, instead of selecting either the segmentation map or the distance map in step S1005, the blending ratio (weighting) of each can be increased during the blending process, similar to alpha blending.

[0087] In this embodiment, the composite map generation unit 904, acting as a composite map generation means, synthesizes a segmentation map and a distance map based on the values ​​of the distance information confidence map and the values ​​of the distance map. Therefore, the accuracy of region extraction can be improved.

[0088] In this embodiment, a confidence map of distance information was output, but this is not the only option. For example, a segmentation confidence map, which is a confidence map of segmentation, may be output and input to the composite map generation unit.

[0089] In other words, the segmentation means may have a segmentation confidence map output means that outputs a segmentation confidence map. The composite map generation means may perform the synthesis based on at least one of the values ​​of the segmentation confidence map and the values ​​of the distance information confidence map.

[0090] For example, the synthesis may be performed such that the synthesis ratio of the distance map is increased in areas where the segmentation confidence map value is low, and the synthesis ratio of the segmentation map is increased in areas where the segmentation confidence is high.

[0091] Alternatively, both a confidence map for distance information and a confidence map for segmentation can be generated, and either the segmentation result (segmentation map) or the distance map result can be selected based on the one with the higher confidence level.

[0092] <Embodiment 4> Embodiment 4 of the present invention compares the segmentation output and the distance map output and generates a composite map based on the comparison result. That is, in Embodiment 4, the values ​​of the segmentation map and the values ​​of the distance map are compared and composite is performed based on the comparison result.

[0093] The following describes the image processing apparatus in Embodiment 4 of the present invention. The block diagram showing an example configuration of the image processing apparatus 100 in Figure 1 is the same as that described in Embodiment 1, so a detailed explanation is omitted. Compared to Embodiment 1, the operation of the region extraction unit 111 in Figure 2, the flowchart of the region extraction unit 111 in Figure 3, and the composite map generation unit 204 in Figure 5 are different.

[0094] Figure 11 is a functional block diagram showing an example configuration of the region extraction unit 111 of the image processing apparatus of Embodiment 4. The segmentation output unit 1101 and the distance information output unit 1102 are equivalent to the segmentation output unit 201 and the distance information output unit 202 in Figure 2 of Embodiment 1, respectively, so a detailed explanation is omitted.

[0095] The output of the segmentation output unit 1101 is input to the composite map generation unit 1105. The distance information of the image output from the distance information output unit 1102 is input to the first distance map output unit 1103 and the second distance map output unit 1104.

[0096] Figure 12 is a flowchart illustrating an example of the operation of the first distance map output unit 1103 and the second distance map output unit 1104 of Embodiment 4. The CPU and other components within the control unit 101 execute computer programs stored in memory, sequentially performing each step of the flowchart in Figure 12. The processing flow in Figure 12 is executed periodically.

[0097] In step S1201, the face detection unit 112 performs face detection. Next, the process proceeds to step S1202. If a face is detected, the process proceeds to step S1203. If no face is detected, the process flow shown in Figure 12 ends.

[0098] In step S1203, it is determined whether or not the face of the main subject is present. That is, if there is one face detection frame or two or more face detection frames, it is determined whether the target face is the face of the main subject. If yes, the process proceeds to step S1204.

[0099] If in step S1203 there are two or more face detection frames and the target face is not the main subject, step S1203 is repeated until the main subject's face is found. Next, in step S1204, the distance information value of the face region is obtained.

[0100] Figures 13(A) to 13(E) are diagrams illustrating the distance information of the face region in Embodiment 4. Figure 13(A) is a diagram showing an example of the distance information value of the face region, where the brightness increases as the distance from the camera decreases. The distance information value of the face region is obtained by taking the average value of the distance information of the face region 1300.

[0101] Next, the process proceeds to steps S1205 and S1207. In step S1205, a first gain table is generated based on the distance information of the face region. That is, a first gain table is generated based on the average value of the distance information of the face region 1300 obtained in step S1204.

[0102] Figure 13(B) shows an example of the gain waveform of the first gain table, with the horizontal axis representing distance information values ​​and the vertical axis representing gain values, and 1302 showing the gain waveform. The position at distance t2 is the average value of the distance information of the face region 1300, and the position at distance t3 is the average value of the distance information of the torso region 1301, with distances t2 and t3 being the peaks of the gain values.

[0103] The gain waveform 1302 decreases monotonically towards distance t1 on the near side, and the gain becomes 0 below distance t1. On the far side, the gain decreases monotonically towards distance t4, and the gain becomes 0 above distance t4. The gain waveform 1302 is used to perform a first weighting process on the distance information.

[0104] Figure 13(C) shows the distance information after applying the gain waveform 1302 to the distance information in Figure 13(A) (hereinafter referred to as the first distance map). In Figure 13(A), the face region 1300 is displayed with lower brightness because it is farther away than the torso region 1301, but in Figure 13(C), the distance value of the face region 1303 is equivalent to that of region 1304. This makes it possible to increase the weight of the face region, which is an important region.

[0105] Let's return to the explanation of the flowchart in Figure 12. After generating the first gain table in step S1205, we proceed to step S1206, where we generate the first distance map by referring to the first gain table.

[0106] Specifically, the first distance map output unit 1103 generates a first distance map by applying the first gain waveform generated in step S1205 to the distance information of the image output from the distance information output unit 1102.

[0107] Specifically, in step S1206, the first distance map output unit 1103, acting as a distance map output means, outputs a first distance map that has undergone a first weighting process to increase the weight of the main subject, relative to the output of the distance information output unit 1102, acting as a distance information output means. The processing flow shown in Figure 12 is then terminated.

[0108] On the other hand, in step S1207, a second gain table is generated based on the distance information of the face region. That is, a second gain table is generated based on the average value of the distance information of the face region 1300 obtained in step S1204. The gain waveform of the second gain table is different from the gain waveform of the first gain table.

[0109] Figure 13(D) shows an example of the gain waveform of the second gain table, with the horizontal axis representing distance information values ​​and the vertical axis representing gain values, and 1305 showing the gain waveform. In Figure 13(D), the position of distance t2 is the average value of the distance information of the face region 1300, and the position of distance t3 is the average value of the distance information of the torso region 1301, with distances t2 and t3 being the peaks of the gain values.

[0110] The gain waveform 1305 decreases monotonically towards distance t1' on the near side, and the gain becomes 0 below distance t1'. The slope of the gain decrease from distance t2 to t1' is gentler compared to the gain waveform of the first gain table. The gain waveform 1305 is used to perform a second weighting process on the distance information.

[0111] On the far-distance side, the gain decreases monotonically towards distance t4, and becomes 0 at distances above t4. Figure 13(E) shows the distance information (hereinafter referred to as the distance map) after applying the gain waveform 1305 to the distance information in Figure 13(A).

[0112] In Figure 13(A), the face region 1300 is displayed with lower brightness because it is farther away from the torso region 1301, but in Figure 13(E), the brightness of the face region 1306 is displayed as being equivalent to that of region 1307. This makes it possible to increase the weight of the face region, which is an important region.

[0113] Let's return to the explanation of the flowchart in Figure 12. After generating the second gain table in step S1207, we then proceed to step S1208, where we generate the second distance map by referring to the second gain table.

[0114] Specifically, the second distance map is generated in the second distance map output unit 1104 by applying the gain waveform generated in step S1207 to the distance information of the image output from the distance information output unit 1102.

[0115] Specifically, in step S1208, the second distance map output unit 1104, acting as a distance map output means, outputs a second distance map obtained by applying a second weighting process based on the main subject detection criterion to the output of the distance information output unit 1102, acting as a distance information output means. Note that the second weighting process is different from the first weighting process. The processing flow shown in Figure 12 is then terminated.

[0116] Let's return to the explanation of Figure 11. The first distance map output by the first distance map output unit 1103, the second distance map output by the second distance map output unit 1104, and the segmentation results output by the segmentation output unit 1101 are input to the composite map generation unit 1105.

[0117] Figure 14 is a flowchart illustrating an example of the operation of the composite map generation unit 1105 of Embodiment 4. The CPU and other components within the control unit 101 execute computer programs stored in memory, sequentially performing each step of the flowchart in Figure 14. The processing flow in Figure 14 is executed periodically.

[0118] Next, an example of the operation of the composite map generation unit 1105 will be explained using the flowchart in Figure 14. In step S1401, the pixel values ​​of the segmentation result (segmentation map) output by the segmentation output unit 1101 are acquired.

[0119] Next, the process proceeds to step S1402, where the pixel values ​​of the first distance map output by the first distance map output unit 1103 are obtained. Then, the process proceeds to step S1403, where the difference between the segmentation result (segmentation map) and the distance map is calculated. That is, the difference between the pixel values ​​of the segmentation result (segmentation map) from step S1401 and the pixel values ​​of the distance map from step S1402 is calculated.

[0120] In step S1404, it is determined whether the difference value calculated in step S1403 is greater than or equal to the threshold. If it is greater, the process proceeds to step S1405 and the count is incremented. Conversely, if it is less than the threshold, the process proceeds to step S1406.

[0121] In step S1406, it is determined whether or not it is the last pixel. If it is not the last pixel, the process returns to step S1401 and repeats steps S1401 to S1405. If it is determined in step S1406 that it is the last pixel, the process proceeds to step S1407, where the composite ratio is calculated from the count value incremented in S1405.

[0122] Figure 15 (A) to (E), (A'), and (D') are diagrams illustrating the operation of steps S1401 to S1407 of Embodiment 4. Figure 15(A) shows an example of the segmentation result (segmentation map) output by the segmentation output unit 1101.

[0123] Figure 15(B) shows an example of a first distance map output by the first distance map output unit 1103, and Figure 15(C) shows an example of a second distance map output by the second distance map output unit 1104. Figure 15(D) shows an example of a difference map, which is the result of calculating the difference between the segmentation result (segmentation map) of step S1402 and the pixel values ​​of the first distance map.

[0124] Figure 15(A) shows 1501, which represents the hair portion of the subject. Because this is a segmentation output using image matting, it can correctly detect areas such as a person's hair.

[0125] On the other hand, the distance maps in Figures 15(B) and (C) fail to detect areas in complex scenes where the foreground and background blend together, such as with people's hair or animal fur. Ideally, the maps should produce values ​​close to the face area, but they are unable to detect it correctly.

[0126] As shown in Figure 15(A), if the segmentation output is normal, it is preferable to use the segmentation output using image matting, which accurately extracts the hair region, rather than using the distance map in Figures 15(B) and (C).

[0127] Therefore, in this embodiment, the characteristics of the composite ratio in step S1407 are as shown in Figure 15(E). That is, in Figure 15(E), the horizontal axis shows the count value counted up in step S1405, and the vertical axis shows the proportion of the composite ratio of the segmentation result (segmentation map).

[0128] The composite ratio characteristic 1502 lowers the composite ratio of the segmentation result (segmentation map) as the count value incremented in step S1405 is larger (i.e., the proportion of pixels with large differences in the entire image is larger). Conversely, the lower the count value incremented in step S1405, the higher the composite ratio of the segmentation result (segmentation map).

[0129] Therefore, for example, in the difference map example in Figure 15(D), since there are few pixels with differences, the count value in step S1405 is also small, and the composite ratio of the segmentation result (segmentation map) in Figure 15(A) is controlled to be high.

[0130] On the other hand, if there are missing items, as shown in Figure 15(A'), the difference map will look like Figure 15(D'), and the count value in step S1405 will be larger. Therefore, the composite ratio of the segmentation result (segmentation map) in Figure 15(A') will be controlled to be low.

[0131] In step S1408, the combined map generation unit 1105 combines the segmentation result output from the segmentation output unit 1101 and the second distance map output from the second distance map output unit, based on the combined ratio calculated in step S1407. Then, the processing flow shown in Figure 14 is completed.

[0132] In this embodiment, the difference between the segmentation result (segmentation map) and the distance map value is taken, and the composite ratio of the segmentation result (segmentation map) and the distance map result is calculated based on the difference value and then composited. That is, the segmentation map and the second distance map are composited based on the comparison result between the segmentation map and the first distance map.

[0133] However, this is not the only method; any method that can extract regions where foreground and background elements blend, such as hair, using segmentation, and regions associated with a person's belongings using distance mapping, may also be used.

[0134] Although the present invention has been described in detail above based on its preferred embodiments, the present invention is not limited to the above embodiments, and various modifications and combinations of the above embodiments are possible in accordance with the spirit of the present invention, and these are not excluded from the scope of the present invention. Furthermore, some of the above embodiments may be combined as appropriate.

[0135] Furthermore, the present invention includes, for example, a system that implements the functions of the above embodiment using at least one processor such as a CPU, memory, and circuitry (e.g., an ASIC). Alternatively, multiple processors may be used for distributed processing.

[0136] Furthermore, in order to implement some or all of the control in the above embodiment, a computer program that implements the functions of the above embodiment may be supplied to an image processing device or the like via a network or various storage media.

[0137] The computer (or CPU, MPU, etc.) in the image processing device may read and execute the program. In that case, the program and the storage medium storing the program constitute the present invention. The present invention includes the following combinations.

[0138] (Configuration 1) An image processing apparatus characterized by comprising: a segmentation means that outputs a segmentation map based on mating results; a distance information output means that outputs distance information; a main subject detection means that detects a main subject; a distance map output means that outputs a distance map based on the distance information of the main subject, based on the detection results by the main subject detection means and the distance information; and a composite map generation means that performs a composite of the segmentation map and the distance map.

[0139] (Configuration 2) The image processing apparatus according to Configuration 1, characterized in that the composite map generation means performs a composite of the segmentation map and the distance map based on the values ​​of the distance map.

[0140] (Configuration 3) The image processing apparatus according to Configuration 2, characterized in that the composite map generation means increases the composite ratio of the distance map when the value of the distance map is less than or equal to a predetermined value, and increases the composite ratio of the segmentation map when the value of the distance map is greater than a predetermined value.

[0141] (Configuration 4) An image processing apparatus according to any one of Configurations 1 to 3, wherein the apparatus has a segmentation processing means for processing the segmentation map, and the composite map generation means performs a composite of the processed map output by the segmentation processing means and the distance map.

[0142] (Configuration 5) The image processing apparatus according to any one of Configurations 1 to 4, characterized in that the distance map output means performs weighting processing on the output of the distance information output means so that the weight of the main subject increases.

[0143] (Configuration 6) An image processing apparatus according to any one of Configurations 1 to 5, characterized in that the distance information output means outputs a distance information confidence map relating to the confidence level of the distance information, and the composite map generation means performs a composite of the segmentation map and the distance map based on the values ​​of the distance information confidence map and the values ​​of the distance map.

[0144] (Configuration 7) The image processing apparatus according to Configuration 6, characterized in that the composite map generation means increases the composite ratio of the distance map when the value of the distance information confidence map is greater than a predetermined value and the value of the distance map is greater than a predetermined threshold, and increases the composite ratio of the segmentation map otherwise.

[0145] (Configuration 8) The segmentation means has a segmentation confidence map output means that outputs a segmentation confidence map, The image processing apparatus according to configuration 6, characterized in that the composite map generation means performs the synthesis based on at least one of the values ​​of the segmentation confidence map and the values ​​of the distance information confidence map.

[0146] (Configuration 9) The image processing apparatus according to Configuration 8, characterized in that the composite map generation means increases the composite ratio of the distance map in areas where the value of the segmentation confidence map is low, and increases the composite ratio of the segmentation map in areas where the confidence of the segmentation is high.

[0147] (Configuration 10) The image processing apparatus according to any one of Configurations 1 to 9, characterized in that the composite map generation means compares the values ​​of the segmentation map with the values ​​of the distance map and performs the synthesis based on the comparison result.

[0148] (Configuration 11) The image processing apparatus according to any one of Configurations 1 to 10, characterized in that the distance map output means outputs a first distance map obtained by applying a first weighting process to the output of the distance information output means such that the weight of the main subject increases, and a second distance map obtained by applying a second weighting process different from the first weighting process, and the composite map generation means performs a composite between the segmentation map and the second distance map based on the comparison result of the segmentation map and the first distance map.

[0149] (Method) An image processing method characterized by comprising: a segmentation step of outputting a segmentation map based on mating results; a distance information output step of outputting distance information; a main subject detection step of detecting a main subject; a distance map output step of outputting a distance map based on the distance information of the main subject, based on the detection results from the main subject detection step and the distance information; and a composite map generation step of combining the segmentation map and the distance map.

[0150] (program) A computer program for controlling each means of the image processing apparatus described in any one of configurations 1 to 11 by a computer. [Explanation of symbols]

[0151] 100: Image processing device 101: Control Unit 102:ROM 103:RAM 104:Optical system 105: Imaging Unit 106: A / D conversion unit 107: Image Processing Unit 108: Records Department 109:Display section 111: Area extraction part 112: Face detection unit

Claims

1. A segmentation means that outputs a segmentation map based on mating results, Distance information output means for outputting distance information, A main subject detection means for detecting the main subject, A distance map output means outputs a distance map based on the distance information of the main subject, based on the detection result by the main subject detection means and the distance information. A composite map generation means for combining the segmentation map and the distance map, An image processing apparatus characterized by having

2. The image processing apparatus according to claim 1, characterized in that the composite map generation means performs a composite of the segmentation map and the distance map based on the values ​​of the distance map.

3. The image processing apparatus according to claim 2, characterized in that the composite map generation means performs the synthesis such that when the value of the distance map is less than or equal to a predetermined value, the synthesis ratio of the distance map is increased, and when the value of the distance map is greater than a predetermined value, the synthesis ratio of the segmentation map is increased.

4. The system includes a segmentation processing means for processing the aforementioned segmentation map, The image processing apparatus according to claim 1, characterized in that the composite map generation means performs a composite of the processed map output by the segmentation processing means and the distance map.

5. The image processing apparatus according to claim 1, characterized in that the distance map output means performs weighting processing on the output of the distance information output means so that the weight of the main subject increases.

6. The distance information output means outputs a distance information confidence map relating to the confidence level of the distance information, The image processing apparatus according to claim 1, characterized in that the composite map generation means performs a composite of the segmentation map and the distance map based on the values ​​of the distance information confidence map and the values ​​of the distance map.

7. The image processing apparatus according to claim 6, characterized in that the composite map generation means increases the composite ratio of the distance map when the value of the distance information confidence map is greater than a predetermined value and the value of the distance map is greater than a predetermined threshold, and increases the composite ratio of the segmentation map otherwise.

8. The segmentation confidence map output means outputs a segmentation confidence map in the segmentation means, The image processing apparatus according to claim 6, characterized in that the composite map generation means performs the synthesis based on at least one of the values ​​of the segmentation confidence map and the values ​​of the distance information confidence map.

9. The image processing apparatus according to claim 8, characterized in that the composite map generation means performs the synthesis such that it increases the synthesis ratio of the distance map in areas where the value of the segmentation confidence map is low, and increases the synthesis ratio of the segmentation map in areas where the confidence of the segmentation is high.

10. The image processing apparatus according to claim 1, characterized in that the composite map generation means compares the values ​​of the segmentation map with the values ​​of the distance map and performs the synthesis based on the comparison result.

11. The distance map output means outputs a first distance map obtained by applying a first weighting process to the output of the distance information output means such that the weight of the main subject increases, and a second distance map obtained by applying a second weighting process different from the first weighting process. The image processing apparatus according to claim 1, characterized in that the composite map generation means performs synthesis between the segmentation map and the second distance map based on the comparison result of the segmentation map and the first distance map.

12. A segmentation step that outputs a segmentation map based on the mating results, Distance information output step, which outputs distance information, A main subject detection step for detecting the main subject, A distance map output step outputs a distance map based on the distance information of the main subject, based on the detection result from the main subject detection step and the distance information. A composite map generation step which involves combining the segmentation map and the distance map, An image processing method characterized by having the following features.

13. A computer program for controlling each means of the image processing apparatus described in any one of claims 1 to 11 by a computer.