Image processing apparatus

The image processing device uses depth and saliency maps, along with skeletal analysis, to enhance the accuracy of subject overlap determination and scoring, addressing the challenges of blur and depth ambiguity in complex scenes.

JP2026012395APending Publication Date: 2026-01-23NIKON CORP
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
JP2025185358
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Filing Date
2025-11-04
Publication Date
2026-01-23

AI Technical Summary

Technical Problem

Existing image processing technologies struggle to accurately determine the overlap state and separation of subjects in complex scenes, particularly when there is significant blur or depth ambiguity, leading to inaccuracies in identifying foreground and background objects.

Method used

An image processing device and method that utilizes depth information, saliency maps, and skeletal analysis to extract main subjects and determine overlap states by calculating scores based on proximity and overlap with other subjects, employing techniques like time-of-flight methods and saliency map generation to enhance accuracy.

Benefits of technology

Improves the accuracy of identifying foreground subjects and calculating recommendation scores for images, effectively handling complex scenes with significant blur or depth ambiguity, thereby enhancing the quality of image processing outcomes.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2026012395000001_ABST
    Figure 2026012395000001_ABST
Patent Text Reader

Abstract

To determine an overlapping state between subjects.SOLUTION: The image processing device includes an extractor that extracts a first subject from image data including subjects, a detector that detects depth information at a plurality of positions in the image data, and a determiner that determines a state of the first subject based on the depth information at each of a predetermined position overlapping the first subject and a predetermined position around the first subject among the plurality of positions at which the depth information is detected, and a degree of closeness between the respective positions.SELECTED DRAWING: Figure 2
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] The present invention relates to an image processing device and an image processing method. [Background technology]

[0002] The imaging device disclosed in Patent Document 1 below is an imaging device that takes group photos. [Prior art documents] [Patent documents]

[0003] [Patent Document 1] Japanese Patent Application Laid-Open No. 2006-140695 Summary of the Invention

[0004] The image processing device of the first disclosed technique includes an extraction unit that extracts a first subject from image data including the subject, a detection unit that detects depth information at a plurality of positions in the image data, and a determination unit that determines a state of the first subject (for example, separation of the first subject from other subjects) based on the depth information at a predetermined position overlapping the first subject and predetermined positions around the first subject among the plurality of positions where the depth information has been detected, and the degree of proximity of the respective positions. The determination unit may also determine an overlap state between the first subject and other subjects.

[0005] The image processing device of the second disclosed technique includes an extraction unit that extracts a first subject and a second subject based on a saliency map and depth information related to image data including the subjects, and a determination unit that determines a state of the first subject (for example, separation between the first subject and the second subject) based on the depth information related to the first subject and the second subject and a degree of proximity between images of the first subject and the second subject. The determination unit may also determine an overlap state between the first subject and the second subject.

[0006] In an image processing method of a third disclosed technique, an image processing device executes an extraction process for extracting a first subject from image data including the subject, a detection process for detecting depth information at a plurality of positions in the image data, and a determination process for determining a state of the first subject (for example, separation of the first subject from other subjects) based on the depth information at a predetermined position overlapping the first subject and predetermined positions around the first subject among the plurality of positions where the depth information has been detected, and the degree of proximity of each of the positions. The determination process may also determine an overlap state between the first subject and other subjects.

[0007] In the image processing method of the fourth disclosed technique, an image processing device executes an extraction process of extracting a first subject and a second subject based on a saliency map and depth information related to image data including the subjects, and a determination process of determining a state of the first subject (for example, separation between the first subject and the second subject) based on the depth information related to the first subject and the second subject and a degree of proximity of images of the first subject and the second subject. The determination process may also determine an overlap state between the first subject and the second subject. [Brief explanation of the drawings]

[0008] [Figure 1] FIG. 1 is a block diagram illustrating an example of a hardware configuration of an imaging apparatus according to a first embodiment. [Figure 2] FIG. 2 is a block diagram of an example of a functional configuration of the image processing apparatus according to the first embodiment. [Figure 3] FIG. 3 is a flowchart of an example of an image processing procedure performed by the image processing apparatus according to the first embodiment. [Figure 4] FIG. 4 is an explanatory diagram showing an example of image data acquired by the imaging unit. [Figure 5] FIG. 5 is an explanatory diagram showing an example of a depth map obtained from the image data shown in FIG. [Figure 6] FIG. 6 is an explanatory diagram showing an example of a defocus value plot graph. [Figure 7] FIG. 7 is an explanatory diagram illustrating an example of a saliency map. [Figure 8] FIG. 8 is an explanatory diagram showing an example of determining the proximity using the AF range-finding point. [Figure 9] FIG. 9 is an explanatory diagram showing a first example of score calculation. [Figure 10] FIG. 10 is an explanatory diagram showing a second score calculation example. [Figure 11] FIG. 11 is an explanatory diagram showing an example of a score displayed by the output unit. [Figure 12] FIG. 12 is a block diagram of an example of a functional configuration of the image processing apparatus according to the second embodiment. [Figure 13] FIG. 13 is a flowchart illustrating an example of an image processing procedure by the image processing apparatus according to the second embodiment. [Figure 14] FIG. 14 is an explanatory diagram showing an example of dividing a subject. [Figure 15] FIG. 15 is an explanatory diagram showing an example of the occurrence of foreground overlap. [Figure 16] FIG. 16 is an explanatory diagram showing an example of improving the foreground overlap. [Figure 17] FIG. 17 is an explanatory diagram showing an example of changing the rapid-fire speed by the determining unit. DETAILED DESCRIPTION OF THE INVENTION [Example]

[0009] In the first embodiment, a subject that is adjacent to a main subject and is focused in front of the main subject (so-called front focus) is determined to be a foreground subject, and a score indicating the degree of recommendation of the main subject is calculated.

[0010] <Example of hardware configuration of imaging device> 1 is a block diagram illustrating an example of a hardware configuration of an imaging device according to a first embodiment. The imaging device 100 is a device capable of capturing still images or moving images, and specifically, is, for example, a digital camera, a digital video camera, a smartphone, a tablet, a personal computer, or a game console. In FIG. 1, a digital camera will be described as an example of the imaging device.

[0011] The imaging device 100 includes a processor 101, a storage device 102, a drive unit 103, an optical system 104, an imaging element 105, an AFE (Analog Front End) 106, an LSI (Large Scale Integration) 107, an operation device 108, a sensor 109, a display device 110, a communication IF (Interface) 111, and a bus 112. The processor 101, the storage device 102, the drive unit 103, the LSI 107, the operation device 108, the sensor 109, the display device 110, and the communication IF 111 are connected to the bus 112.

[0012] The processor 101 controls the imaging device 100. The storage device 102 serves as a working area for the processor 101. The storage device 102 is a non-transitory or temporary recording medium that stores various programs and data. Examples of the storage device 102 include a read-only memory (ROM), a random access memory (RAM), a hard disk drive (HDD), and a flash memory. A plurality of storage devices 102 may be implemented in the imaging device 100, and at least one of the storage devices 102 may be detachable from the imaging device 100.

[0013] The drive unit 103 drives and controls the optical system 104. The drive unit 103 has a drive circuit 103a and a drive source 103b. The drive circuit 103a controls the drive source 103b in accordance with instructions from the processor 101. The drive source 103b is, for example, a motor, and, under the control of the drive circuit 103a, moves the zoom lens 141b and the focusing lens 141c in the optical system 104 in the optical axis direction and controls the opening and closing of the diaphragm 142.

[0014] The optical system 104 includes a plurality of lenses (a lens 141a, a zooming lens 141b, and a focusing lens 141c) arranged in the optical axis direction, and an aperture 142. The optical system 104 collects subject light and outputs the collected light to the image sensor 105.

[0015] The image sensor 105 receives subject light from the optical system 104 and converts it into an electrical signal. The image sensor 105 may be, for example, an XY address type solid-state image sensor (e.g., a CMOS (Complementary Metal-Oxide Semiconductor)) or a progressive scan type solid-state image sensor (e.g., a CCD (Charge Coupled Device)).

[0016] A plurality of light receiving elements (pixels) are arranged in a matrix on the light receiving surface of the image sensor 105. A plurality of types of color filters, each of which transmits light of a different color component, are arranged in a predetermined color array (for example, a Bayer array) on the pixels of the image sensor 105. Therefore, each pixel of the image sensor 105 outputs an analog electrical signal corresponding to each color component through color separation by the color filter.

[0017] The AFE 106 is an analog front-end circuit that performs signal processing on an analog electrical signal from the image sensor 105. The AFE 106 sequentially performs gain adjustment of the electrical signal, analog signal processing (correlated double sampling, black level correction, etc.), A / D conversion processing, and digital signal processing (defective pixel correction, etc.) to generate RAW image data and output it to the LSI. The above-mentioned drive unit 103, optical system 104, image sensor 105, and AFE 106 constitute the image sensor 120.

[0018] The LSI 107 is an integrated circuit that executes specific processes such as image processing such as color interpolation, white balance adjustment, edge enhancement, gamma correction, and gradation conversion, as well as encoding, decoding, and compression / expansion processes, on the RAW image data from the AFE 106. Specifically, the LSI 107 may be realized by a PLD (Programmable Logic Device) such as an ASIC (Application Specific Integrated Circuit) or an FPGA (Field-Programmable Gate Array).

[0019] The operation device 108 inputs commands and data. Examples of the operation device 108 include various buttons including a release button, switches, dials, and a touch panel. The sensor 109 is a device that detects information, and includes, for example, an AF (Automatic Focus) sensor, an AE (Automatic Exposure) sensor, a gyro sensor, an acceleration sensor, a temperature sensor, and a depth sensor. The AF sensor is, for example, an AF sensor using a re-imaging phase difference method. The AF sensor can also be an AF sensor using an image plane phase difference detection method, which enables focus detection through processing by a processor or the like (not shown). The display device 110 displays image data and setting screens. The display device 110 includes a rear monitor on the rear of the imaging apparatus 100 and an electronic viewfinder. The communication IF 111 connects to a network and transmits and receives data.

[0020] The processor 101, the storage device 102, the LSI 107, the operation device 108, the sensor 109, the display device 110, the communication IF 111, and the bus 112 are collectively referred to as an image processing unit 130. Furthermore, a device consisting of the image processing unit 130, excluding the image capturing unit 120 from the imaging device 100, is referred to as an image processing device.

[0021] <Example of functional configuration of image processing device> Fig. 2 is a block diagram illustrating an example of a functional configuration of the image processing apparatus according to the first embodiment. Fig. 3 is a flowchart illustrating an example of an image processing procedure by the image processing apparatus according to the first embodiment. The image processing apparatus 200 is configured, for example, by the image processing unit 130 illustrated in Fig. 1 and includes an acquiring unit 201, an extracting unit 202, a detecting unit 203, a determining unit 204, a calculating unit 205, and an outputting unit 206. Specifically, the acquiring unit 201, the extracting unit 202, the detecting unit 203, the determining unit 204, the calculating unit 205, and the outputting unit 206 are realized, for example, by causing the processor 101 to execute a program stored in the storage device 102 illustrated in Fig. 1 or by the LSI 107.

[0022] In FIG. 3, the image processing device 200 performs the following operations: acquisition of image data by the acquisition unit 201 (step S301); extraction of a subject by the extraction unit 202 (step S302); detection of depth information by the detection unit 203 (step S303); determination of an overlap state by the determination unit 204 (step S304); calculation of a score by the calculation unit 205 (step S305); and output of the result by the output unit 206 (step S306).

[0023] Returning to FIG. 2, the acquisition unit 201 acquires image data including a subject. Specifically, for example, the acquisition unit 201 acquires image data including a subject by capturing a still image using the imaging unit 120, by capturing a live view image before capturing an image, or by acquiring the image data from the storage device 102. Note that the image data may include one or more subjects. The acquisition unit 201 also acquires a depth map of the image data, which is a live view image, using a depth sensor, which is an example of the sensor 109. The acquired depth map is stored in the storage device 102 in association with the image data.

[0024] The acquisition unit 201 may also perform face detection on the acquired image data and acquire face detection information from the image data, including the number of detected face images, their positions within the image data, and facial expressions. The acquired face detection information is stored in the storage device 102 in association with the image data.

[0025] Fig. 4 is an explanatory diagram showing an example of image data acquired by the imaging unit 120. Fig. 4 shows image data 400 displayed on the display device 110. Note that 105 rectangles in the center of the screen of the display device 110 are AF distance measuring points 410. The image data 400 includes subjects 401 to 403 representing people. When the image data 400 is captured, the acquisition unit 201 stores the AF distance measuring points 410 in the storage device 102 as focus information.

[0026] 5 is an explanatory diagram showing an example of a defocus value plot graph. The defocus value plot graph 500 is a graph in which the defocus value for each AF ranging point 410 is plotted, and is an example of a defocus map. The horizontal axis represents the AF ranging point number that uniquely identifies the AF ranging point 410, and the vertical axis represents the defocus value. Plot 501 represents the defocus value of subject 401, and plots 502 and 503 represent the defocus values ​​of subjects 402 and 403, respectively. Since plots 502 and 503 have negative defocus values, it can be seen that subjects 402 and 403 are located closer to the subject 401.

[0027] Fig. 6 is an explanatory diagram showing an example of a depth map obtained from the image data 400 shown in Fig. 4. For convenience of explanation, in Fig. 5, the depth map 600 is shown with shading of the subjects 401-403 and background in a pixel region group, which is a collection of multiple pixel regions 610 shown as rectangles. However, the depth map 600 is composed of the depth values ​​of each pixel region 610 (the darker the value, the shallower the value, and the lighter the value, the deeper the value). The pixel region 610 is a collection of one or more pixels. The depth value is, for example, the distance from the objective lens or sensor surface in front of the depth sensor to the person.

[0028] 6 has been described as the depth map 600, the acquisition unit 201 may also acquire a defocus map from the image data 400. In the case of a defocus map, the map indicates the defocus value for each AF ranging point 410. The acquisition unit 201 may also acquire a defocus value plot graph from the defocus map. Alternatively, the acquisition unit 201 may create a depth map using the defocus map.

[0029] Returning to FIG. 2, the acquisition unit 201 may acquire a saliency map indicating the distribution of saliency of the image data 400 from the image data 400. The saliency map is a method for calculating a general-purpose gaze point that does not rely on human knowledge or experience. Specifically, for example, the acquisition unit 201 decomposes the color image data 400 into three features, namely, color, brightness, and edge direction, generates Gaussian pyramids for each, performs a center-surround operation using these three Gaussian pyramids to generate a feature map, and then normalizes and integrates all the feature maps to generate the saliency map. In the saliency map, areas with properties different from surrounding areas are extracted as "highly salient (attention-grabbing) areas."

[0030] 7 is an explanatory diagram showing an example of a saliency map. The saliency map 700 of the first embodiment detects contours for each of the subjects 401 to 403. This makes it possible to identify, for example, the boundary between the subject 401 and the subject 402 with a large amount of blur, and separate the object representing the subject 401 from the object representing the subject 402.

[0031] Returning to FIG. 2, the extraction unit 202 extracts a main subject from image data 400 including a plurality of subjects 401 to 403 acquired by the acquisition unit 201. Here, the subject 401 is the main subject. Specifically, the extraction unit 202 extracts the main subject 401 based on the orientation of the object in the image data 400. For example, the extraction unit 202 extracts the object with reference to the depth map 600.

[0032] For example, when extracting a person as an object, extraction unit 202 extracts an area using depth values, identifies an object with a shape that can be recognized as a person, and extracts the person object (subjects 401 to 403). Furthermore, extraction unit 202 extracts skeleton information for each object extracted as a person, which is a combination of multiple (e.g., 18) nodes that serve as skeleton points and links that connect the nodes.

[0033] The extraction unit 202, for example, identifies the real-space parts of the person reflected in the object area (left and right hands, head, neck, left and right shoulders, left and right elbows, left and right knees, left and right feet, etc.) based on the depth values ​​of each part in the object area and the features of the object's shape, and sets a skeleton point at the center position of each part.

[0034] Furthermore, the extraction unit 202 may use a feature dictionary stored in the storage device 102 to compare features determined from the object area with features for each body part registered in the feature dictionary, thereby identifying each body part of the person in the object area and setting a skeleton point at the center of each body part. Examples of skeleton points include the left and right hands, head, neck, left and right shoulders, left and right elbows, left and right knees, and left and right feet. The number of skeleton points increases or decreases depending on the posture and range of the object.

[0035] Furthermore, when extracting a person object, the extraction unit 202 may use the face detection information acquired by the acquisition unit 201 to extract an object including an image in which a face has been detected as a person object.

[0036] Furthermore, the extraction unit 202 may extract the main subject 401 as a person object based on the saliency map 700 obtained from the image data 400.

[0037] In addition, the extraction unit 202 may extract, as the main subject 401, an object that overlaps with an AF ranging point 410 selected manually (by user operation of the operation device 108) from among a plurality of AF ranging points 410 set for focus detection of the subjects 401 to 403 in the image data 400.

[0038] Similarly, the extraction unit 202 may extract, as the main subject, an object that overlaps with an AF ranging point 410 selected automatically (by the autofocus function) from among a plurality of AF ranging points 410 set for focus detection of subjects 401 to 403 in the image data 400. In this case, the extraction unit 202 may extract, as the main subject, the subject that includes the largest number of AF ranging points 410 and has a small amount of defocus. In the example of Fig. 4, subject 401 is extracted as the main subject.

[0039] The detection unit 203 detects depth information at multiple positions within the image data 400. The depth information is a depth value for each pixel region 610 obtained from the depth map 600, or a defocus value for each AF ranging point 410 at which focus has been detected among the AF ranging points 410 in the defocus map (or the defocus value plot graph 500). Note that the detection unit 203 may measure distances at multiple positions within the image data 400 using, for example, a time-of-flight method to obtain the depth information.

[0040] The determination unit 204 determines the overlap state between the main subject 401 and other subjects based on the detected depth information at the predetermined position overlapping the main subject 401 and predetermined positions around the main subject 401, and the degree of proximity of each position. Determining the overlap state means, for example, determining whether the main subject overlaps with other subjects. Furthermore, if the main subject overlaps with other subjects, the determination unit 204 may detect the degree of overlap.

[0041] The predetermined position is the pixel region 610 when the depth map 600 is used, and is the AF ranging point 410 when the defocus map is used. Therefore, the depth information of the predetermined position is the depth value of the pixel region 610 when the depth map 600 is used, and is the defocus value of the AF ranging point 410 when the defocus map is used. The depth information identifies the position of the subject in the depth direction.

[0042] The predetermined position overlapping with the main subject 401 is, among the multiple pixel areas 610 or AF ranging points 410, a pixel area 610 or AF ranging point 410 that is included in the main subject 401. Furthermore, the predetermined position around the main subject 401 is, for example, a pixel area 610 or AF ranging point 410 outside the main subject 401 that is adjacent to a pixel area 610 or AF ranging point 410 that overlaps with the outline of the main subject 401.

[0043] The degree of proximity is determined by a predetermined position overlapping the main subject 401 and a predetermined position around the main subject 401. Specifically, for example, the determination unit 204 detects the shortest distance between each predetermined position overlapping the outline of the main subject 401 and a predetermined position overlapping the outline of an object present around the main subject 401. The distance is, for example, the Manhattan distance (the number of pixel regions 610 or AF measuring points 410). The case of the AF measuring points 410 will be taken as an example.

[0044] Fig. 8 is an explanatory diagram showing an example of proximity determination using AF ranging points 410. Fig. 8 shows an example of proximity determination between a main subject 401 and a subject 402. The AF ranging points 410 that overlap with the outline of the main subject 401 are designated as dotted AF ranging points 411, and the AF ranging points 410 that overlap with the outline of the subject 402 are designated as hatched AF ranging points 412. Furthermore, the AF ranging points 410 that overlap with the outline of the subject 403 are designated as hatched AF ranging points 413. Note that the AF ranging points 410 that include both the main subject 401 and the subject 402 are considered to belong to the subject with the larger area.

[0045] The shortest distance between each of the AF measuring points 411 and each of the AF measuring points 412 is 1. The determination unit calculates a statistical value of the shortest distance (for example, an average value, a maximum value, a minimum value, a mode value, or a median value) as a value indicating the overlap state between the main subject 401 and the subject 402. If the value indicating the overlap state is equal to or less than a threshold value, the determination unit determines that the subject 402, which is located in front of the main subject 401, is an overlapping subject in front of the main subject 401. If the threshold value is, for example, "2," then the subject 402 is an overlapping subject in front of the main subject 401.

[0046] On the other hand, the shortest distance between AF measuring point 411 and AF measuring point 413 is 6. Therefore, the determination unit 204 determines that subject 403, which is positioned closer to the main subject 401 than the main subject 401, is not an overlapping subject in front of the main subject 401. The same applies even if the predetermined position is the pixel area 610.

[0047] The determination unit 204 may detect the AF ranging point 411 that is closest to and located in front of the predetermined position from among the AF ranging points 411 that are located around the predetermined position (outside the main subject) that overlaps with the outline of the main subject 401, and determine the overlap state of the main subject by calculating the Manhattan distance from the predetermined position.

[0048] Furthermore, the determination unit 204 may acquire depth information of AF ranging point 411 that is within a predetermined distance (for example, within "2") outside the main subject from a predetermined position that overlaps with the outline of the main subject 401, and determine the overlapping state of the main subject. For example, when the threshold value indicating the overlapping state is "3," if the determination unit 204 does not detect an AF ranging point 411 that is located in front of the predetermined position within a predetermined distance of "2," the determination unit 204 can determine that the main subject 401 is not a foreground overlapping subject.

[0049] In this way, the smaller the value indicating the overlap state, the higher the possibility that the subject 402 around the main subject 401 will be determined as a foreground subject, and the larger the value indicating the overlap state, the higher the possibility that the subject 403 around the main subject 401 will not be determined as a foreground subject. Note that in the example of Example 1, it is not necessarily necessary to recognize and detect the subject 402 and the subject 403 as other subjects.

[0050] In other words, even without detecting subjects 402 and 403, the determination unit 204 can determine the overlap state between the main subject 401 and other subjects based on the depth information at each of the predetermined positions overlapping with the main subject 401 and predetermined positions around the main subject 401, and the degree of proximity at each position.

[0051] Therefore, the overlap state determination may be a determination of whether the main subject overlaps with some kind of object. Furthermore, if it is determined that the main subject overlaps with another subject (object), an object (for example, a character string or a circle) may be displayed on the display unit of the image data to indicate that the main subject overlaps with the other subject. In this case, the object may be displayed superimposed on the main subject that overlaps with the other subject.

[0052] The determination unit 204 may determine the degree of proximity using skeletal information of the main subject 401. Specifically, for example, if a skeleton point that should be present on the main subject 401 is missing and the position where that skeleton point should be located overlaps with the subject 402, the determination unit 204 may determine that the subject 402 is an obscured subject in front of the main subject 401.

[0053] 4, the skeleton point for the right hand of main subject 401 is missing, and the position where that skeleton point should be located is the head of subject 402. Determination unit 204 estimates the position where the skeleton point for the right hand of main subject 401 should be located based on the node and link indicating the skeleton point for the right arm of main subject 401.

[0054] If the position of the estimation result overlaps with the subject 402, the determination unit 204 determines that the subject 402 is an obscured subject in front of the main subject 401. Furthermore, if an AF ranging point 412 that is in front of the main subject is detected at the position of the estimation result, the determination unit 204 may determine that the subject 402 is an obscured subject in front of the main subject 401.

[0055] The calculation unit 205 calculates a score indicating the degree of recommendation for each subject extracted by the extraction unit 202. Specifically, the calculation unit 205 calculates a foreground overlap score according to a value indicating the presence or absence of a foreground overlapping subject and the overlap state determined by the determination unit, and a score indicating the level of recommendation.

[0056] For example, when the determination unit 204 determines that no foreground object exists, the calculation unit 205 assigns a high score to the object because there is no area of ​​the object that is obscured by the foreground object. On the other hand, when the determination unit 204 determines that a foreground object exists, the calculation unit 205 assigns a low score to the object because there is an area of ​​the object that is obscured by the foreground object.

[0057] 4 and 8, main subject 401 has subject 402 as a foreground subject overlapping main subject 401, and therefore the score is lower than when subject 402 is not present. Calculation unit 205 may also calculate the score according to the skeletal part where foreground overlap occurs, using skeletal information of main subject 401. For example, calculation unit 205 may assign a lower score when foreground overlap occurs near the head than when foreground overlap occurs near the legs.

[0058] Furthermore, the calculation unit 205 may further calculate the score based on at least one of the size, posture, position, and defocus amount of the subject in the image data 400, and the size of the subject compared to other subjects.

[0059] 9 is an explanatory diagram showing score calculation example 1. The score related to the size of the subject (hereinafter referred to as the size score) is the ratio "V1 / V0" obtained by dividing the vertical width V1 of the human subject 901 identified by the face detection information and skeletal information by the vertical width V0 of the background of the image data 900. Size scores are also calculated for the other human subjects 902 to 904.

[0060] The score related to the posture of the subject (hereinafter referred to as the pose score) is calculated for each of the subjects 901-904 based on the skeletal information of the subjects 901-904, who are people identified by the face detection information and skeletal information. Specifically, for example, the pose score increases the higher the position of the hands of the subjects 901-904 in the vertical direction, and, if both hands are shown, the further apart the hands are. For example, the pose score is highest when the subject is making a V-sign.

[0061] The score related to the subject arrangement (position score) is calculated for each of the subjects 901-904 based on the arrangement of the subjects 901-904 in the imaging range of the person identified by the face detection information and skeletal information. Specifically, for example, the position score is high when the subjects 901-904 are arranged near the center of the imaging range, and low when the subjects 901-904 are arranged near the four corners of the imaging range. Furthermore, when multiple subjects are included in one image data, a score can be assigned to the image data as a whole, focusing on the overall balance of the subject arrangement, degree of dispersion, etc.

[0062] The score related to the defocus amount of the subject (hereinafter referred to as the focus score) is calculated for each of the subjects 901-904 based on the face detection information of the human subjects 901-904 identified by the face detection information and skeletal information, the depth information obtained from the depth map 600, and the focus information. Taking the image data 400 as an example, the focus information is information related to the focus state based on the defocus amount at the position of the AF ranging point 410. Specifically, for example, the focus score increases as the area around the eyes of the subject's face becomes more in focus.

[0063] 10 is an explanatory diagram showing score calculation example 2. The score relating to the size of a subject compared to other subjects (hereinafter referred to as the prominence score) is a score indicating the relative sizes of the subjects 901 to 904 based on the vertical widths V1 to V4 of the subjects 901 to 904. Specifically, for example, for image data 900, the prominence score value cs is calculated using the following formula (1):

[0064] cs=V# / (V1+V2+V3+V4) (1) However, # is a value between 1 and 4.

[0065] In this way, the calculation unit 205 calculates the foreground overlap score, size score, pose score, position score, focus score, and prominence score for each subject, and calculates a total score of the foreground overlap score, size score, pose score, position score, focus score, and prominence score as a score indicating the recommendation level of the subject.

[0066] The overall score can be expressed, for example, by any combination of these elements and using a regression equation of a simple sum or a weighted linear sum. A normalization technique can also be used to adjust the reference values ​​of each element. In this case, each normalized element can be weighted and expressed as a regression equation of a simple sum or a weighted linear sum.

[0067] The calculation method for each score related to foreground overlap, subject size, pose, position, focus, and prominence may be changed depending on the shooting scene. For example, if the shooting scene is a finish line scene of a marathon, a high pose score may be assigned to image data including a pose in which the subject's arms are stretched out to the sides. Furthermore, for human subjects whose faces are not detected due to foreground overlap, the foreground overlap score may be lowered accordingly.

[0068] The output unit 206 outputs the determination result by the determination unit 204 and the calculation result by the calculation unit 205. The determination result is the presence or absence of a foreground subject. The calculation results are the foreground subject score, size score, pose score, focus score, prominence score, and overall score.

[0069] Fig. 11 is an explanatory diagram showing an example of displaying scores by the output unit 206. In Fig. 11, scores SC1 to SC3 are displayed in association with each of the subjects 401 to 403.

[0070] As described above, according to the first embodiment, it is possible to determine the overlap state of the main subject and to identify whether the main subject is a foreground subject covered by other subjects. Also, it is possible to improve the accuracy of the score indicating the recommendation level of an image including a subject based on the overlap state and the presence or absence of a foreground subject. [Example]

[0071] Next, a second embodiment will be described. In the second embodiment, an example of foreground overlap determination that is effective when the amount of blur of a subject that is adjacent to the main subject and is focused in front of the main subject (so-called foreground focus) is too large and the depth map 600 or defocus map cannot determine the foreground overlap. In the second embodiment, differences from the first embodiment will be mainly described. Also, the same parts as in the first embodiment will be assigned the same reference numerals, and their description will be omitted.

[0072] <Example of functional configuration of image processing device> Fig. 12 is a block diagram illustrating an example of a functional configuration of an image processing apparatus according to a second embodiment. Fig. 13 is a flowchart illustrating an example of an image processing procedure by the image processing apparatus according to the second embodiment. The image processing apparatus 1200 includes an acquisition unit 201, an extraction unit 1202, a division unit 1203, a determination unit 1204, a calculation unit 205, and an output unit 206. Specifically, the extraction unit 1202, the division unit 1203, the determination unit 1204, and the calculation unit 205 are realized by, for example, causing the processor 101 to execute a program stored in the storage device 102 illustrated in Fig. 1 or by the LSI 107.

[0073] In FIG. 13, the image processing device 200 performs the following steps: acquisition of image data 400 by the acquisition unit 201 (step S1301), object extraction by the extraction unit 1202 (step S1302), object segmentation by the division unit 1203 (step S1303), determination of the overlap state by the determination unit 1204 (step S1304), calculation of a score by the calculation unit 205 (step S1305), and output of the result by the output unit 206 (step S1303).

[0074] The extraction unit 1202 extracts the main subject 401 and the subject 402 based on the saliency map 700 and depth information related to the image data 400 including the multiple subjects 401 to 403. The depth information is either defocus information or depth information representing the depth information of the image data 400, as described in the first embodiment.

[0075] If the difference in depth information between adjacent subjects is outside the allowable range, they are extracted as separate subjects, such as main subject 401 and subject 402. If the difference in depth information between adjacent subjects is within the allowable range, they are extracted as a single subject. However, by using the saliency map 700, the contours of each of adjacent subjects are detected, so they are extracted as separate subjects even if the difference in depth information is within the allowable range. Therefore, the extraction unit 1202 can extract subjects that are too blurry to be extracted using the defocus map or depth map 600.

[0076] Furthermore, the extraction unit 1202 extracts the main subject 401 based on the orientation of the object in the image data 400. For example, the extraction unit 1202 extracts the object by referring to the depth map 600, as in the first embodiment.

[0077] The dividing unit 1203 divides the subject extracted by the extracting unit 1202. Specifically, for example, when the difference in depth information between multiple subjects is within an allowable range, the subjects may be extracted as one subject even when the saliency map 700 is used.

[0078] 14 is an explanatory diagram showing an example of dividing a subject. Assume that the extraction unit 1202 extracts a subject 1430 from image data 1400. In this case, the division unit 1203 uses skeletal information to identify that the subject 1430 includes multiple human subjects 402 and 1403, and sets a boundary between the subjects 402 and 1403. This boundary may be the midpoint between the skeleton points of the adjacent subject 402 and the skeleton points of the subject 1403. The division unit 1203 sets a line passing through this midpoint as the boundary, and divides the subject 1430 into the subjects 402 and 1403.

[0079] The dividing unit 1203 may further use face detection information to divide the subject 1430 into the subject 402 and the subject 1403. By using the face detection information, it is possible to identify the number and positions of people in the subject 1430, thereby improving the division accuracy.

[0080] The determination unit 1204 determines the overlap state between the main subject 401 and the subject 402 based on the depth information regarding the main subject 401 and the subject 402 and the degree of proximity between the main subject 401 and the subject 402. Specifically, for example, similar to the first embodiment, the determination unit 1204 detects the shortest distance between each predetermined position overlapping with the outline of the main subject 401 and a predetermined position overlapping with the outline of the subject 402. The determination unit 1204 calculates a statistical value of the shortest distance as a value indicating the overlap state between the main subject 401 and the subject 402.

[0081] Furthermore, when subject division is performed by the division unit 1203, the determination unit 1204 determines the overlap state between the main subject 401 and the subject 402 based on the depth information regarding the main subject 401 and one of the subjects 402 and 1403, for example, the subject 402 that is closer to the main subject 401, and the degree of proximity between the main subject 401 and the subject 402.

[0082] Regarding which of subject 402 and subject 1403 is closest to main subject 401, for example, division unit 1203 may compare the shortest distance between the outline of main subject 401 and the outline of subject 402 with the shortest distance between the outline of main subject 401 and the outline of subject 1403, and determine the shorter subject as the specific subject closest to main subject 401. Furthermore, division unit 1203 may compare the shortest distance between the skeleton points of main subject 401 and the skeleton points of subject 402 with the shortest distance between the skeleton points of main subject 401 and the skeleton points of subject 1403, and determine the shorter subject as the specific subject closest to main subject 401.

[0083] In this way, the smaller the value indicating the overlap state, the more likely it is that the subject 402 around the main subject 401 will be determined to be a foreground overlapping subject, and the larger the value indicating the overlap state, the more likely it is that the subject 403 around the main subject 401 will not be determined to be a foreground overlapping subject.

[0084] According to the second embodiment, the saliency map 700 is used to extract the object, and therefore it is possible to extract an object that has too much blur to be extracted using the defocus map or the depth map 600. Therefore, it is possible to suppress the loss of overlapping determination target candidates, and to improve the accuracy of the score calculated by the calculation unit. [Example]

[0085] Next, a third embodiment will be described. The third embodiment is an example of determining the overlap state of moving subjects in the second embodiment. The third embodiment will be described mainly with respect to differences from the first and second embodiments. The same parts as those in the first and second embodiments are denoted by the same reference numerals, and the description thereof will be omitted.

[0086] Fig. 15 is an explanatory diagram showing an example of foreground overlap. Fig. 15 shows, as an example, an example in which foreground overlap occurs in image data 1502, the frame following image data 1501, when imaging device 100 is shooting continuously at 10 fps. Image data 1501 includes, as images, a first object 1510, which is a moving object (a runner running in the direction of arrow A), and a second object 1520, which is a stationary object (for example, a signboard or postbox). Image data 1502 shows a state in which, due to the movement of first object 1510, a portion of first object 1510 is hidden by second object 1520, resulting in foreground overlap caused by second object 1520.

[0087] Fig. 16 is an explanatory diagram showing an example of improving foreground overlap. Fig. 16 shows an example in which, while the imaging device 100 is continuously shooting at 10 [fps], after acquiring image data 1501, the continuous shooting speed (frame rate) is slowed down from 10 [fps] to 5 [fps] to capture image data 1602, so as to prevent foreground overlap as shown in Fig. 15. In this way, the imaging device 100 changes the continuous shooting speed in advance to prevent foreground overlap from occurring.

[0088] FIG. 17 is an explanatory diagram showing an example of changing the continuous shooting speed by the determination unit 1204. The lower left vertex of image data 1501 is the origin, and the x-coordinate value of the lower right vertex is x3. First, the information in image data 1501 will be described. A first rectangle 1710 is rectangular data in which the first subject 1510 is inscribed. The length of the vertical side of first rectangle 1710 in real space is H [m], and the length of the horizontal side in real space is W [m]. Furthermore, the coordinate values ​​of the upper left vertex of first rectangle 1710 are (x0, y0). Furthermore, the moving speed of first subject 1510 in the direction of arrow A is a [m / s].

[0089] The second rectangle 1720 is rectangular data inscribed with the second subject 1520. With the origin O as the reference, the coordinate values ​​of the upper left vertex of the second rectangle 1720 are (x1, y1) and the coordinate values ​​of the lower right vertex are (x2, y2). The frame rate before the change is b [fps]. It is assumed that this information in the image data 1501 is obtained by known CV calculations up to the frames prior to the image data 1501.

[0090] First, determination unit 1204 determines whether second subject 1520 in image data 1501 will be in front of first subject 1510 in the next frame. The condition for second subject 1520 to be in front of first subject 1510 is satisfied when one of the following conditions is met: (i) in the next frame, the x-coordinate value x10 of the upper right vertex of first rectangle 1710 is within the width (x1, x2) of second rectangle 1720; and (ii) in the next frame, when the x-coordinate value x10 of the upper right vertex of first rectangle 1710 is on the positive side of the x-coordinate value x2 of the lower right vertex of second rectangle 1720, the x-coordinate value x20 of the upper left vertex of first rectangle 1710 is smaller than the x-coordinate value x2 of the lower right vertex of second rectangle 1720.

[0091] The following explanation shows an example of how to prevent the second subject 1520 from being obscured in front of the first subject 1510 by increasing or decreasing the frame rate b [fps] under the above-mentioned condition (i). Specifically, for example, the determination unit 1204 calculates the distance a / b [m] that the first subject 1510 will move by the time the next frame is displayed. Also, the determination unit 1204 calculates the x-coordinate value x10 of the upper right vertex of the first rectangle 1710 in the next frame using the following formula (2).

[0092] x10=x0+W+a / b (2)

[0093] The determining unit 1204 determines whether the x coordinate value x10 calculated by the above formula (2) is included in the width (x1, x2) of the second rectangle 1720 using the following formula (3).

[0094] x1 <x10<x2···(3)

[0095] When the above formula (3) is satisfied, second subject 1520 overlaps first subject 1510 in front of it. Therefore, imaging section 120 adjusts frame rate b [fps] so as to satisfy the following formula (4).

[0096] b>a / (x1-x0-W) (4)

[0097] In the above example, the frame rate b [fps] is increased, but the frame rate b [fps] may be adjusted to decrease. Specifically, the image capturing unit 120 adjusts the frame rate b [fps] so as to satisfy the following formula (5).

[0098] b

[0099] As a result, even by lowering the frame rate b [fps], it is possible to prevent the second subject 1520 from overlapping the first subject 1510 in front of the first subject 1510. Note that when the frame rate b [fps] is lowered, adjustment may be made so that the first subject 1510 fits within the frame of the image data 1501 in the next frame, that is, so that the x-coordinate value x10 of the upper right vertex of the first rectangle 1710 does not exceed the x-coordinate value x3 of the image data 1501. Furthermore, although the third embodiment illustrates an example in which the first subject 1510 moves in the +X direction, the present invention is not limited to this. For example, the first subject 1510 may move in the -X direction or the +Y direction.

[0100] In this way, when at least one of the first subject 1510 and the second subject 1520 is moving, the determination unit 1204 determines the overlap state between the first subject 1510 and the second subject 1520 after a predetermined time has elapsed, based on the speed and direction of movement of the moving first subject 1510 and the size of the first subject 1510 and the second subject 1520.

[0101] ​If it is determined that first subject 1510 and second subject 1520 overlap, imaging section 120 controls the imaging timing to be earlier or later so that first subject 1510 and second subject 1520 do not overlap. Therefore, it is possible to prevent second subject 1520 from overlapping first subject 1510 in front of it.

[0102] If it is determined that the first subject 1510 and the second subject 1520 will overlap after a predetermined time has elapsed, the image capturing timing may be left unchanged and the image data captured after the predetermined time may not be saved in the storage device. In this way, by not saving image data in which overlapping occurs, the storage capacity of the storage device can be reduced.

[0103] In the above example, the first subject 1510 moves in the direction of arrow A from left to right, but the present invention can also be applied to cases where the first subject 1510 moves from right to left. In this case, the inequality signs in the above formulas (1) to (6) are reversed. The present invention can also be applied to cases where the subject moves not only left and right, but also up and down or diagonally.

[0104] According to the third embodiment, the occurrence of foreground overlap can be reduced by predicting the foreground overlap in advance and adjusting the frame rate, and this also increases the number of image data with high scores.

[0105] The present invention is not limited to the above, and may be implemented by any combination thereof. Furthermore, other aspects conceivable within the scope of the technical concept of the present invention are also included within the scope of the present invention. For example, in this embodiment, the determination unit determines the overlap state between the main subject and other subjects, but instead of the determination unit, a detection unit may detect the overlap state between the main subject and other subjects. [Explanation of symbols]

[0106] 100 imaging device, 120 imaging unit, 130 image processing unit, 200 image processing device, 201 acquisition unit, 202 extraction unit, 203 detection unit, 204 determination unit, 205 calculation unit, 206 output unit, 400 image data, 401 to 403 object, 410 to 413 AF measuring point, 500 depth map, 600 defocus value plot graph, 700 saliency map, 900 image data, 901 to 904 object, 1200 image processing device, 1202 extraction unit, 1203 division unit, 1204 determination unit, 1400 image data, 1403 object, 1430 object, 1501 image data, 1502 image data, 1510 object, 1520 object, 1602 image data

Claims

[Claim 1] an extracting unit that extracts a first object from image data including the object; a detection unit that detects depth information at a plurality of positions within the image data; a determination unit that determines a state of the first subject based on the depth information at a predetermined position overlapping the first subject and a predetermined position around the first subject among the plurality of positions at which the depth information is detected, and based on the degree of proximity of the respective positions; An image processing device comprising:

Citation Information

Patent Citations

  • Imaging apparatus

    JP2006140695A