Image processing system, imaging apparatus, method for producing learning model, image processing method and program
Patent Information
- Application Number
- JP2022210202
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2022-12-27
- Publication Date
- 2026-01-08
AI Technical Summary
Existing methods for automatic focusing in cameras, particularly when photographing people or animals, require detection of multiple facial organ points, leading to increased processing time and memory usage, especially when using heat maps for estimating coordinates, which is inefficient and costly.
An image processing device that detects a specific type of region, such as eyes, by estimating the distance between these regions and a secondary point like the nose, allowing for quicker detection and focus adjustment by selecting the region closest to the camera based on learned neural networks.
This approach reduces detection time and memory requirements while improving accuracy and robustness in focusing on the eyes, enabling more impressive facial photographs.
Smart Images

Figure 00000000_0000_ABST
Abstract
Description
[Technical field]
[0001] The present invention relates to an image processing device, an imaging device, a method for producing a learning model, an image processing method, and a program, and in particular to an automatic focus adjustment function. [Background technology]
[0002] Current cameras have an automatic focus adjustment (AF) function. With the AF function, the photographic lens is adjusted so that the subject is automatically brought into focus. In particular, when photographing people or animals, there is a need to focus on the eyes or the pupils of the eyes (hereinafter sometimes simply referred to as the eyes). This configuration is advantageous for taking impressive photographs of faces.
[0003] Patent Document 1 discloses an eye detection AF mode. In this mode, eyes are detected from a captured image. Then, focus adjustment is performed so that the focus is on the eyes. Patent Document 1 aims to achieve good focusing on the eyes. To this end, Patent Document 1 discloses detecting the direction of the face and focusing on the eye that is easier to detect based on the direction of the face. Specifically, the focus is adjusted to the eye that is closer to the photographer (camera).
[0004] In recent years, many methods have been proposed for detecting objects in captured images. Among them, a method using a multi-layered neural network called a deep net (also called a deep neural net or deep learning) has been actively researched. In this method, features of objects in an image are learned. Then, the learning results are used to recognize the position or type of the object. For example, Non-Patent Document 1 discloses a method for detecting facial organs from an image using a deep net. In addition, Non-Patent Document 2 lists a direct regression method and a heat map method as methods for estimating facial organ points. [Prior art documents] [Patent documents]
[0005] [Patent Document 1] JP 2012-123301 A [Non-patent literature]
[0006] [Non-Patent Document 1] J. Deng et al. "RetinaFace: Single-shot Multi-level Face Localization in the Wild", CVPR 2020. [Non-Patent Document 2] K. Khabarlak et al. "Fast Facial Landmark Detection and Applications: A Survey", CVPR 2021. Summary of the Invention [Problem to be solved by the invention]
[0007] In the method described in Patent Document 1, a wireframe based on many facial organ points (such as eyes, mouth, nose, chin, forehead, eyebrows, and eyebrow gap) is used to estimate the face direction. This requires the detection of many organ points. Therefore, there is a problem that the detection process time increases in proportion to the number of organ points detected, and memory usage increases. In particular, when a heat map method is used to detect face organ points, one heat map for estimating x and y coordinates is required to estimate one organ point. This tends to increase the detection process time. Furthermore, in order to perform such an estimation, it is costly to prepare correct answer data corresponding to each organ point.
[0008] An object of the present invention is to quickly detect, among specific types of parts in an image, parts that are closer to an imaging device. [Means for solving the problem]
[0009] According to an embodiment of the present invention, an image processing device includes: A first position estimation means for detecting a first type of part in an image; a distance estimation means for estimating a distance between the first type of site and a second type of site in the image; a selection means for selecting one of the first type sites from the detected plurality of first type sites based on the distance estimated for each of the first type sites, and outputting information indicating the selected site; Equipped with. Effect of the Invention
[0010] Of the specific types of parts in an image, those closer to the imaging device can be detected quickly. [Brief description of the drawings]
[0011] [Figure 1] FIG. 2 is a diagram showing an example of the hardware configuration of a camera according to an embodiment. [Diagram 2] FIG. 1 is a block diagram showing an example of the functional arrangement of an image processing apparatus according to an embodiment. [Diagram 3] 1 is a flowchart showing a processing flow of an image processing method according to an embodiment. [Figure 4] FIG. 2 is a diagram showing an example of the structure of a neural network used in an embodiment. [Diagram 5] FIG. 2 is a diagram showing an example of the functional configuration of a learning device according to an embodiment. [Figure 6] Schematic diagram showing the relationship between face orientation and the distance between the eyes and nose. [Figure 7] FIG. 4 is a schematic diagram showing the relationship between learning images and correct answer information. [Figure 8] 1 is a flowchart showing a process flow of a learning method according to an embodiment. [Figure 9] FIG. 2 is a diagram showing an example of the structure of a neural network used in an embodiment. [Figure 10] 11 is a flowchart showing the flow of a process for creating teacher data. [Figure 11] 1 is a flowchart showing a processing flow of an image processing method according to an embodiment. [Figure 12] FIG. 2 is a diagram showing an example of the structure of a neural network used in an embodiment. [Figure 13] Schematic diagram showing a correct map of eye center positions close to the camera. [Figure 14] FIG. 4 is a schematic diagram showing the relationship between learning images and correct answer information. DETAILED DESCRIPTION OF THE PREFERRED EMBODIMENTS
[0012] Hereinafter, the embodiments will be described in detail with reference to the attached drawings. Note that the following embodiments do not limit the invention according to the claims. Although the embodiments describe a number of features, not all of these features are essential to the invention, and the features may be combined in any manner. Furthermore, in the attached drawings, the same reference numbers are used for the same or similar configurations, and duplicated descriptions are omitted.
[0013] In the following, a camera, which is an example of an imaging device according to the present invention, will be described. However, the embodiment described below is applicable to any electronic device that detects an object region in an image or video. Such electronic devices include imaging devices such as digital cameras or digital video cameras, as well as personal computers, mobile phones, drive recorders, robots, drones, and the like that have a camera function. However, electronic devices are not limited to these. Such electronic devices can have an image processing unit 104, which will be described later.
[0014] <Hardware configuration> 1 shows an example of the hardware configuration of a camera, which is an image capturing device according to one embodiment. The camera according to this embodiment includes an image capturing unit 101, a RAM 102, a ROM 103, an image processing unit 104, an input / output unit 105, and a control unit 106. These units are configured to be able to communicate with each other, and are connected to each other via a bus or the like.
[0015] The imaging unit 101 can capture an image. The imaging unit 101 can have a photographing lens, an imaging element, an A / D converter, an aperture control unit, and a focus control unit. The photographing lens can include a fixed lens, a zoom lens, a focus lens, an aperture, and an aperture motor. The imaging element converts an optical image of a subject into an electric signal. The imaging element can include a CCD or a CMOS. The A / D converter converts an analog signal into a digital signal. The aperture control unit changes the aperture opening diameter by controlling the operation of the aperture motor. In this way, the aperture control unit can control the aperture of the photographing lens. The focus control unit controls the focus state of the photographing lens by driving the focus lens. The focus control unit can control the operation of the focus motor based on the phase difference between a pair of focus detection signals obtained from the imaging element. Here, the focus control unit can perform control so that the focus is set on a designated AF area.
[0016] In the imaging unit 101, the imaging element converts the subject image formed on the imaging surface of the imaging element by the photographing lens into an electrical signal. The A / D converter generates image data by applying A / D conversion processing to the obtained electrical signal. The image data obtained in this manner is supplied to the RAM 102.
[0017] The RAM 102 stores image data obtained by the imaging unit 101. The RAM 102 can also store image data to be displayed on the input / output unit 105. The RAM 102 has a storage capacity sufficient to store a predetermined number of still images or a moving image for a predetermined period of time. The RAM 102 can also serve as a memory for displaying images (video memory). In this case, the RAM 102 can supply image data for display to the input / output unit 105. The RAM 102 is, for example, a volatile memory.
[0018] The ROM 103 is a non-volatile memory. The ROM 103 is a storage device such as a magnetic storage device or a semiconductor memory. The ROM 103 can store programs used for the operation of the image processing unit 104 and the control unit 106. The ROM 103 can also store data that is stored for a long period of time.
[0019] The image processing unit 104 performs image processing to detect an object region from an image. The image processing unit 104 can detect object candidate regions from an image, select an object region from the object candidate regions, and output the results. The image processing unit 104 can detect a portion (object region) in an image for a specific type of object, that is, a specific type of portion in an image. The object to be detected can be a specific type of object such as a person, an animal, or a vehicle, or a specific type of local portion included in these objects such as a head, a face, or an eye. In this embodiment, the image processing unit 104 outputs the position and size of the candidate region for the specific type of object, as well as a likelihood representing the object-likeness, as the detection result. Based on these pieces of information, the object region in the image is detected. The configuration and operation of the image processing unit 104 will be described in detail later.
[0020] The input / output unit 105 includes an input device for a user to input instructions to the camera 100, and a display that can display characters or images. The input device can include one or more of a switch, a button, a key, and a touch panel. The display can be an LCD or an organic EL display. The input through the input device can be detected by the control unit 106 through a bus. At this time, the control unit 106 controls each unit to realize an operation according to the input. The touch sensing surface of the touch panel can be used as the display surface of the display. The type of the touch panel is not limited, and can be, for example, a resistive film type, a capacitive type, or an optical sensor type. The input / output unit 105 may display a live view image. That is, the input / output unit 105 can sequentially display image data obtained by the imaging unit 101.
[0021] The control unit 106 is a processor. The control unit 106 may be a central processing unit (CPU). The control unit 106 can realize the functions of the camera 100 by executing a program stored in the ROM 103. The control unit 106 can also perform aperture control, focus control, and exposure control by controlling the imaging unit 101. For example, the control unit 106 executes an automatic exposure (AE) process that automatically determines exposure conditions (shutter speed or accumulation time, aperture value, and sensitivity). Such AE processing can be performed based on subject brightness information of image data obtained by the imaging unit 101. The control unit 106 can also automatically set an AF area by using the detection result of the object area by the image processing unit 104. In this way, the control unit 106 can realize tracking AF processing for an arbitrary subject area. Note that the AE processing can be performed based on brightness information of the AF area. Furthermore, the control unit 106 can also perform image processing (such as gamma correction processing or auto white balance (AWB) adjustment processing). These image processing can be performed based on pixel values of the AF area.
[0022] The control unit 106 can also perform display control by controlling the input / output unit 105. For example, the control unit 106 can superimpose an index (e.g., a rectangular frame surrounding the area) indicating the position of the object area or the AF area on the display image. The display of such an index can be performed based on the detection result of the object area or the AF area in the image processing unit 104.
[0023] <Functional configuration> FIG. 2(A) shows a functional configuration of an image processing unit 104 which is an image processing device according to an embodiment. The image processing unit 104 includes an input unit 210, an extraction unit 220, a first part estimation unit 230, a distance estimation unit 240, and a selection unit 250. The image processing unit 104 may be a processing unit independent of the control unit 106. For example, the image processing unit 104 may be realized by dedicated hardware such as an ASIC. The image processing unit 104 may also be a processor having a memory or connected to the memory. In this case, the processor executes a program stored in the memory to realize the functions of each unit shown in FIG. 2(A) to (C), etc. On the other hand, the function of the image processing unit 104 may be realized by the control unit 106. In this way, some or all of the functions of the image processing unit 104 can be realized by a combination of a processor and a memory, or by dedicated hardware. The image processing device according to an embodiment may be configured by a plurality of information processing devices connected via a network, for example.
[0024] The input unit 210 acquires an image. This image may be an image captured by the imaging unit 101. For example, the input unit 210 can acquire frame images included in a time-series moving image acquired by the imaging unit 101. For example, the input unit 210 can acquire image data of 1600×1200 pixels in real time at 60 frames per second.
[0025] The extraction unit 220 extracts features from the image acquired by the input unit 210. In this embodiment, the features of the image are represented as a map. Meanwhile, the method of extracting the features is not particularly limited. The extraction unit 220 can extract the features by processing according to predetermined parameters. For example, the extraction unit 220 can extract the features by processing using a neural network. The parameters used in the processing performed by the extraction unit 220, such as the weight parameters of the neural network, can be determined by learning as described later. In addition, the extraction unit 220 may calculate a feature vector by aggregating the colors or textures of pixels.
[0026] First part estimation section 230 detects a first type of part in an image. For example, first part estimation section 230 can detect the position of the first type of part in an image. Specifically, first part estimation section 230 can determine the position of the first type of part based on the likelihood of the first type of part for each position in the image. First part estimation section 230 can detect the first type of part based on the feature amount of the image extracted by extraction section 220.
[0027] In the embodiment described below, the first type of part is a person's eyes. In this case, the extraction unit 220 can extract information for estimating the eye area from the image as a feature. This feature may be, for example, a center position map indicating the likelihood that a position on the image is the center of a frame indicating the eye area. Based on such a feature, the first part estimation unit 230 can estimate the center position of the eye.
[0028] The feature amount may also be a size map indicating an estimated width and height of a frame indicating the eye region for a position on the image. Based on such feature amount, the first part estimation unit 230 can estimate the width and height of a frame corresponding to the estimated center position of the eye.
[0029] Through such processing, first part estimation section 230 can select a frame that has a high likelihood of representing the center of the eye frame from among multiple frames. First part estimation section 230 can select, for example, one or two frames.
[0030] The distance estimation unit 240 estimates the distance between the first type of part and the second type of part in the image. The distance estimation unit 240 can detect the first type of part based on the feature amount of the image extracted by the extraction unit 220. This feature amount may be a distance map. This distance map can indicate the distance to the second part for the first type of part in the image. Such a distance map may indicate the distance to the second part (e.g., Euclidean distance) for the position on the image. Such a distance map can be obtained by the extraction unit 220 using a learning model, as described later. In this way, the extraction unit 220 can extract information for estimating the distance between the eyes and the nose from the image as a feature amount. Then, the distance estimation unit 240 can estimate the distance from the center position of each of the two frames selected by the first part estimation unit 230 to the nose. In this embodiment, the first type of part and the second part are parts that belong to the same object (e.g., a person).
[0031] The selection unit 250 selects one first type part from the detected multiple first type parts based on the distance estimated for each first type part. Then, the selection unit 250 outputs information indicating the selected part. The selection unit 250 can select one from the multiple first type parts based on the priority order. For example, the selection unit 250 can select one first type part so that the distance estimated for the selected one first type part is longer than the distance estimated for the other parts of the multiple first type parts. Specifically, the selection unit 250 can select the first type part so that the distance estimated by the distance estimation unit 240 is the smallest. In this embodiment, the selection unit 250 selects the frame that is closer to the nose from the two eye frames. The eye indicated by the frame selected in this way is located closer to the camera 100 than the other eye. The selection unit 250 can output information indicating the position of the selected first type part to the control unit 106.
[0032] The control unit 106 controls the imaging unit 101 so as to focus on one of the first type parts selected by the selection unit 250. For example, the control unit 106 can perform AF processing so as to focus on the eye frame selected by the selection unit 250.
[0033] An example will be described below in which the image processing unit 104 detects noteworthy eye frames from a moving image captured by the camera 100. In this example, AF processing is performed on the detected frames. FIG. 3 is a flowchart showing the flow of processing in an image processing method according to an embodiment. In the following description, each process (step) is preceded by an S, and the notation of the process (step) is omitted. However, the image processing unit 104 does not need to perform all of the processes shown in this flowchart.
[0034] In S300, the input unit 210 acquires one frame image from a time-series video captured by the imaging unit 101. The input unit 210 inputs the acquired image to the extraction unit 220. The input unit 210 may acquire multiple frame images sequentially. In this case, the process shown in Fig. 3 is performed on each frame image. The image acquired in S300 is, for example, bitmap data expressed in RGB 8 bits.
[0035] In S301, the extraction unit 220 extracts image features by processing the image acquired from the input unit 210. Fig. 6(A) and (B) are schematic diagrams showing the relationship between the face direction and the distance between the eyes and the nose. Fig. 6(A) shows an input image 600 in which a person's face facing rightward is captured. The input image 600 includes an eye closer to the camera (right eye 601), an eye farther from the camera (left eye 602), and a nose 603. Fig. 6(B) shows an input image 604 in which a person's face facing forward is captured. The input image 604 includes a right eye 605, a left eye 606, and a nose 607.
[0036] As shown in FIG. 6B, when a person faces the camera, a right eye 605 and a left eye 606 are located at approximately the same distance from the camera. Also, the distance (42) between the right eye and the nose and the distance (41) between the left eye and the nose are approximately the same. On the other hand, as shown in FIG. 6A, when a person faces the right direction, the distance (49) between the right eye, which is closer to the camera, and the nose is relatively large, and the distance (35) between the left eye, which is farther from the camera, and the nose is relatively small. Thus, the distance between the eye, which is closer to the camera, and the nose is shorter than the distance between the eye, which is farther from the camera, and the nose. In this embodiment, this relationship is utilized to select the eye, which is closer to the camera.
[0037] In S302, first part estimation section 230 detects eyes based on the feature amount extracted by extraction section 220. First part estimation section 230 can detect one or two eyes. First part estimation section 230 can select a frame having a likelihood greater than a preset threshold. Furthermore, when a plurality of such frames are detected, first part estimation section 230 can select one or two eye frames in descending order of likelihood.
[0038] In the example described below, the extraction unit 220 is realized by a learning model having learned parameters. Such a learning model can be realized, for example, by using a neural network. The structure of the neural network used is not particularly limited. FIG. 4 shows an example of a network structure realized by using a neural network and a schematic diagram of an output result. FIG. 4 shows an input image 400, a person's eyes 401 and 402, and a person's nose 403. FIG. 4 also shows a center position map 404 indicating the magnitude of the likelihood representing the central position of the eyes. The center position map 404 shows the magnitudes of likelihood 405 and 406 for the parts of the image 400 corresponding to the person's eyes 401 and 402.
[0039] FIG. 4 also shows a size map 407 representing the width of the eyes. The size map 407 indicates the widths 408 and 409 of the eyes 401 and 402 of the person in the image 400. FIG. 4 further shows a size map 410 representing the height of the eyes. The size map 410 indicates the heights 411 and 412 of the eyes 401 and 402 of the person in the image 400. FIG. 4 also shows a distance map 413 representing the distance between the eyes and the nose. The distance map 413 indicates the distances 414 and 415 from the eyes to the nose of the person. A display result 416 shows a result of superimposing an eye frame 417 detected by the first part estimation unit 230 on the image 400. As described later, the frame 417 is selected by the selection unit 250 from the frame corresponding to the eye 401 and the frame corresponding to the eye 402.
[0040] The neural network shown in FIG. 4 has a network structure for face organ detection described in Non-Patent Document 1. In this network, an image 400 is input to a network 420 called a backbone. Then, the network 420 outputs intermediate features. The intermediate features output from the network 420 are input to networks 421 to 423 for each task. FIG. 4 shows a network 421 for estimating the position of the eyes, a network 422 for estimating the size of the eyes, and a network 423 for estimating the distance between the eyes and the nose. The network 421 outputs a center position map 404. The network 422 outputs two size maps 407 and 410. The network 423 outputs a distance map 413. In this embodiment, each map 404, 407, 410, and 413 is a two-dimensional array and is represented by a grid.
[0041] In FIG. 4, the center position map 404 indicates that the closer to the center of the circle, the higher the likelihood, and that the location of the eye is high. The size maps 407 and 410 indicate the inference results of the width and height of the eye when each position is the center position of the person's eyes. In FIG. 4, the length of the arrow indicates the magnitude of the value. As shown in FIG. 4, the inferred values of the width and height of the eye are shown at the positions corresponding to the center positions of the eyes in the size maps 407 and 410. The distance map 413 indicates the inference results of the distance from each position to the nose when each position is the center position of the person's eyes. In FIG. 4, the length of the arrow indicates the magnitude of the value. As shown in FIG. 4, the inferred values of the distance from each position to the nose are shown at the positions corresponding to the center positions of the eyes in the distance map 413. In the example of FIG. 4, four maps are used to estimate the eye close to the camera 100.
[0042] The frame indicating the eye region can be defined by the center coordinates, width, and height of a rectangle surrounding the eye. As described above, the center position map 404 indicates the likelihood of the center position of the eye. The first part estimation unit 230 selects an element having a value exceeding a preset threshold in the center position map 404 as a center position candidate of the eye frame. When the elements selected as the center position candidates are adjacent to each other, the first part estimation unit 230 can select an element having a higher likelihood than the others among the adjacent elements as the center position of the eye. Note that the resolution of the center position map 404 may be lower than the resolution of the original image 400. In this case, the center position of the eye on the image 400 is obtained by converting the center position of the eye on the center position map to match the size of the image 400. In addition, the first part estimation unit 230 obtains the width and height of the eye frame from the elements of the size maps 407 and 410 corresponding to the detected center position of the eye. In this way, the first part estimation unit 230 can determine the eye frame.
[0043] In S303, selection unit 250 determines whether or not first part estimation unit 230 has detected one or more eyes. If it is determined that one or more eyes have not been detected, the process in Fig. 3 ends. If it is determined that one or more eyes have been detected, the process proceeds to S304.
[0044] In S304, selection unit 250 determines whether or not first part estimation unit 230 has detected two eyes. If it is determined that two eyes have not been detected, the process proceeds to S309. In S309, selection unit 250 selects the frame of one eye detected in S302 as the AF area. If it is determined that two eyes have been detected in S304, the process proceeds to S305.
[0045] In S305, the selection unit 250 estimates the distance between the eye and the nose for each of the two eyes detected in S302. Specifically, the selection unit 250 obtains the distance between the eye and the nose shown in the distance map obtained in S301, which corresponds to the center position of the two eye frames. For example, the selection unit 250 can obtain the distance from each eye to the nose based on the element of the distance map, which corresponds to the x coordinate and the y coordinate of the center position of the two eye frames obtained in S302.
[0046] The selection unit 250 can select a method of selecting one of the two parts based on the difference in distance between each of the two eyes and the nose. In S306, the selection unit 250 calculates the absolute value of the difference in the distance value for each of the two eyes obtained in S305. Then, the selection unit 250 judges whether the calculated absolute value is equal to or greater than a set threshold value. If it is judged that the absolute value of the difference in the distance between the two eyes is equal to or greater than the threshold value, the process proceeds to S307. If it is judged that the absolute value is less than the threshold value, the process proceeds to S308.
[0047] The process of S307 is performed when the distance difference between each of the two eyes and the nose is large. In S307, one of the two eyes is selected according to the first method. Specifically, the selection unit 250 selects one of the two eyes based on a comparison of the distance between each of the two eyes and the nose. For example, the selection unit 250 selects one of the two eyes such that the distance between the eye to be selected and the nose is greater than the distance between the eye not selected and the nose. Then, the selection unit 250 selects the frame of the selected eye as the AF area.
[0048] The process of S308 is performed when the difference in distance between each of the two eyes and the nose is small, for example, when the distances are almost the same. In S308, one of the two eyes is selected according to a second method different from the first method. In this case, the distances from the camera to the two eyes are approximately equal, so there is no significant difference whether the focus is on either eye. Therefore, the selection unit 250 can select one of the two eyes by any method. For example, the selection unit 250 may select an eye detected from a position close to the center of the image. Then, the selection unit 250 selects the frame of the selected eye as the AF area.
[0049] In S310, the control unit 106 performs AF processing so as to focus on one eye frame, which is the AF area selected by the selection unit 250. As the AF processing method, for example, a phase difference detection method can be used.
[0050] <Learning Method> The learning of the neural network as shown in Fig. 4 can be performed as follows. Fig. 5 shows the functional configuration of a learning device according to an embodiment. The learning device 500 learns parameters of a learning model used by the image processing unit 104 to perform processing. For example, the learning device 500 can generate parameters of a learning model used in particular by the extraction unit 220 and supply them to the camera 100. The learning device 500 includes a data storage unit 510, a data acquisition unit 520, an image acquisition unit 530, an object estimation unit 540, a data creation unit 550, an error calculation unit 560, and a learning unit 570.
[0051] The data storage unit 510 stores the learning data used for learning. The learning data used in this embodiment includes a set of a learning image and correct answer information about the eyes of a person in the learning image. The correct answer information in this embodiment includes the coordinates of the center position of the eye, the size (width and height) of the eye, and the distance between the eye and the nose. The data storage unit 510 stores learning data having a sufficient amount and variation for the learning device 500 to perform learning. The correct answer information may be information input by a person who has viewed the learning image. The correct answer information may also be an image obtained by the image processing device performing a detection process on the learning image. The correct answer information does not need to be generated in real time. For this reason, such an image processing device may generate the correct answer information using an algorithm that requires a lot of time and a large amount of calculation.
[0052] The data acquisition unit 520 acquires the learning data stored in the data storage unit 510. The image acquisition unit 530 acquires the learning image from the data acquisition unit 520. The image acquired by the image acquisition unit 530 is input to the target estimation unit 540. The image acquisition unit 530 may perform data expansion. For example, the image acquisition unit 530 may rotate, enlarge, reduce, add noise to the image, or change the brightness or color of the image. Such data expansion is expected to improve the robustness of the processing by the image processing unit 104. When performing data expansion involving a geometric transformation such as rotation, enlargement, or reduction of an image, the image acquisition unit 530 may perform a conversion process according to the geometric transformation on the correct answer information of the learning data. The image acquisition unit 530 may input the image obtained by such data expansion to the target estimation unit 540.
[0053] The object estimation unit 540 performs processing equivalent to that of the extraction unit 220 of the image processing unit 104. That is, the object estimation unit 540 can acquire a map indicating the distance to the second part for the first type of part in the training image obtained by inputting the training image to the training model. In this embodiment, the object estimation unit 540 outputs a distance map indicating the center position map of the eyes of the person in the training image, the size map of the eyes, and the distance between the eyes and the nose, based on the training image input by the image acquisition unit 530. In this embodiment, the object estimation unit 540 can generate these maps using the training model. For example, the object estimation unit 540 can generate the map using the neural network shown in FIG. 4.
[0054] The data creating unit 550 generates teacher data based on the correct answer information acquired from the data acquiring unit 520. This teacher data is used as a target value for the output of the object estimating unit 540. For example, the data creating unit 550 can generate, as teacher data, a correct answer map of the center positions of the eyes, a correct answer map of the sizes of the eyes, and a correct answer map of the distance between the eyes and the nose.
[0055] 7(A)-(E) are schematic diagrams of a training image, correct answer information, and eye center position correct map. FIG. 7(A) shows a training image 700 showing a person's face. FIG. 7(B) shows correct answer information 710 for the training image 700. FIG. 7(C) shows correct answer information 720 expressed in coordinates of the correct answer map. FIG. 7(A) further shows an enlarged view 730 of the eye center position correct answer map near the person's eyes. FIG. 7(D) shows an enlarged view 740 of the eye size correct answer map near the person's eyes. FIG. 7(E) shows an enlarged view 750 of the eye-nose distance correct answer map near the person's eyes.
[0056] In the examples of Fig. 7(A) to (E), a learning image 700 in which a person appears and the correct answer information shown in Fig. 7(B) are given as learning data. The correct answer information includes the center coordinates (X, Y) = (900, 300) of the person's eyes, the size of the eyes (20), and the distance between the eyes and the nose (85). The center coordinates of the eyes, the size of the eyes, and the distance between the eyes and the nose shown in Fig. 7(B) are indicated by the coordinates, size, and distance on the learning image 700. In this example, the size of the learning image 700 (and the image acquired by the input unit 210) is 1600 x 1200 pixels. The distance is the Euclidean distance between the eyes and the nose. However, the distance may be another type of distance such as the Manhattan distance or the Chebyshev distance. The correct answer information may include position and size information for two or more eyes. In this case, the positions or areas corresponding to each eye in each map can be labeled.
[0057] In the examples of Figs. 7(A) to (E), the teacher data indicates the absolute distance between the eyes and the nose. However, the teacher data may indicate the relative distance. The relative distance between one eye and the nose represents the relative value of the distance between one eye and the nose with respect to the distance between the other eye and the nose. For example, when two eyes of one person are shown in the learning image 700, the average value of the distance between the eyes and the nose for each eye can be calculated. Then, the relative distance for each eye can be calculated by subtracting the average value from the distance for each eye. In this case, when the distance is relatively large, the relative distance is a positive value. When the distance is relatively small, the relative distance is a negative value. When the distances are approximately equal, the relative distance is close to 0. The distance indicated by the teacher data may be a normalized value. For example, the absolute distance can be normalized by dividing it by the size of the face. When learning is performed using such teacher data, the image processing unit 104 can generate a distance map indicating the relative distance or the normalized absolute distance. The distance estimation unit 240 can estimate the distance using a learning model trained using such a distance map. That is, the distance between a first type part and a second part estimated by the distance estimation unit 240 may be a relative distance with respect to a distance between another first type part and a second part. Also, the distance between a first type part and a second part estimated by the distance estimation unit 240 may be a normalized distance.
[0058] In addition, when only one eye of one person is shown in the learning image 700, the distance of this eye indicated by the teacher data may be a positive value obtained by multiplying the face size by a constant. For example, when the size of the face on the image is 240, the teacher data may indicate a value 72 obtained by multiplying 240 by 0.3. The distance indicated by the teacher data may be a positive constant independent of the face size. Furthermore, learning may be controlled so that learning using the learning image 700 in which only one eye of one person is shown is not performed uniformly.
[0059] The eye center position correct map is matrix data of the same size as the eye center position map output by the target estimation unit 540. In this embodiment, the size of the center position correct map is 320×240. Thus, compared with the learning image, the center position correct map has a size of 1 / 5 in both length and width. In this example, the size correct map and distance correct map also have the same size as the center position correct map. Therefore, in order to obtain the correct answer information 720 represented by the coordinates of the correct answer map shown in FIG. 7(C), the values indicating the center coordinates of the eye and the size of the eye on the learning image are reduced to 1 / 5. In the example of FIG. 7(C), the center coordinates of the eye are (X,Y)=(180,60), and the size of the eye is 4.
[0060] The eye center position correct map 730 is a map obtained by labeling correct cases at the eye center positions. As shown in FIG. 7(A), the center position correct map 730 is labeled with a heat map of a circular region centered at coordinates (X,Y)=(180,60) and with a diameter equal to the eye size (4). In this example, the correct map has values in the range of 0 to 1. Therefore, the center of the circular region has a maximum value of 1. Also, the value gradually decreases toward the edge of the circular region.
[0061] The eye size correct answer map 740 is a map obtained by labeling the correct cases in the eye region. As shown in FIG. 7(D), in the size correct answer map 740, labels are added within a frame whose center is the eye center coordinate (X, Y) = (180, 60) and whose side length is the same as the eye size. The black thick frame shown in the size correct answer map 740 indicates the center position of the eye. The value of the label can be determined according to the size of the eye. For example, as shown in FIG. 7(D), the value (0.1) obtained by dividing the eye size (20) by the maximum size (200) can be used as the label value. In addition, an empty value can be assigned to elements outside the frame so as not to contribute to learning.
[0062] The distance correct answer map 750 is a map obtained by labeling the eye region with the correct case. The position of the distance correct answer map 750 corresponding to the first type of part (eye in this example) in the learning image is given information indicating the distance between the first type of part and the second type of part (nose in this example). As shown in FIG. 7(E), the distance correct answer map 750 is given a label within a frame whose center is the center coordinate (X, Y) = (180, 60) of the eye and whose side length is the same as the size of the eye. The black thick frame shown in the distance correct answer map 750 indicates the center position of the eye. The value of the label can be determined according to the distance between the eye and the nose. For example, as shown in FIG. 7(E), the value (0.2) obtained by dividing the distance (85) by the maximum size (425) can be used as the value of the label. In addition, an empty value can be assigned to the elements outside the frame so as not to contribute to learning.
[0063] The error calculation unit 560 calculates a center position error, which is the error between the eye center position map output by the object estimation unit 540 and the eye center position correct map created by the data creation unit 550. The error calculation unit 560 also calculates a size error, which is the error between the eye size map output by the object estimation unit 540 and the eye size correct map created by the data creation unit 550. The error calculation unit 560 also calculates a distance error, which is the error between the distance map output by the object estimation unit 540 and the distance correct map created by the data creation unit 550.
[0064] The learning unit 570 updates the parameters of the learning model used by the object estimation unit 540 based on the error calculated by the error calculation unit 560. For example, the learning unit 570 can update the parameters of the learning model so that the center position error, the size error, and the distance error are reduced. The parameter update method is not particularly limited. The learning unit 570 can update the parameters of the learning model used by the object estimation unit 540 using, for example, an error backpropagation method. In this way, the learning unit 570 can learn the parameters of the learning model based on the map indicating the distance to the second part for the first type of part in the learning image obtained by the object estimation unit 540 and the distance answer map.
[0065] The parameters used by the object estimation unit 540, which have been updated by such learning, are supplied to the camera 100. Then, the image processing unit 104 can estimate the center position, eye size, and distance by processing according to the supplied parameters. For example, the extraction unit 220 can generate a center position map, a size map, and a distance map by processing using a neural network according to the supplied parameters.
[0066] 8 is a flow chart illustrating a training method according to an embodiment of the present invention, which produces a training model having trained parameters for use in estimating distances between a first type of feature and a second type of feature in an image.
[0067] In S801, the data acquiring unit 520 acquires learning data stored in the data storage unit 510. In S802, the image acquiring unit 530 acquires learning images from the data acquiring unit 520.
[0068] In S803, the object estimation unit 540 performs an inference process on the learning image acquired from the image acquisition unit 530. The object estimation unit 540 can output an eye center position map, an eye size map, and an eye-nose distance map. In S804, the data creation unit 550 creates an eye center position correct map, an eye size correct map, and an eye-nose distance correct map from the correct answer information acquired in S801 according to the above-mentioned method.
[0069] In S805, the error calculation unit 560 calculates a center position error based on the eye center position correct map created in S804 and the eye center position map obtained in S803. The error calculation unit 560 also calculates a size error based on the eye size correct map created in S804 and the eye size map obtained in S803. The error calculation unit 560 also calculates a distance error based on the distance correct map created in S804 and the distance map obtained in S803.
[0070] In S806, the learning unit 570 learns the parameters used by the object estimation unit 540 so that the center position error, size error, and distance error calculated in S805 are reduced. In S807, the learning unit 570 determines whether or not to continue learning. If the learning is to be continued, the processes from S801 onwards are repeated. If the learning is not to be continued, the process shown in FIG. 8 ends. The method of determining whether or not to continue learning is not particularly limited. For example, it is possible to determine whether or not to continue learning based on whether the number of learning times has reached a predetermined number of times or whether the learning time has reached a predetermined time. The process shown in FIG. 8 can be repeatedly performed using various learning data.
[0071] In the above-described embodiment, the eye closer to the camera 100 among the multiple eyes is detected based on the distance between the eye and the nose. However, the detection target (first type of part) is not limited to the eye. For example, the detection target may be a part that exists in multiple on the face, such as an ear. When the detection target is an ear, the selection unit 250 may select one of the two ears detected by the first part estimation unit 230 based on the distance between the ear and the nose. For example, the selection unit 250 may select the ear farther from the nose among the two ears on the image. The ear selected in this manner is located closer to the camera 100 than the other ear.
[0072] Also, in the above-described embodiment, one eye is selected based on the distance between the eye and the nose. However, the distance used by the selection unit 250 is not limited to the distance between the eye and the nose. For example, instead of the nose, another second part can be used. For example, the second part may be a part that is equidistant from both eyes when facing forward. In this specification, being equidistant from both eyes includes being approximately equidistant from both eyes. For example, the second part may be a part on a center line extending from the top to the bottom of the face. Specific examples of the second part include the mouth, the chin, the top of the head, and the space between the eyebrows. The selection unit 250 can select an eye that is farther away from these parts based on the distance between these parts and the eye. Also, the subject is not limited to a person. If the subject is a bird, the eye that is closer to the camera 100 can be selected similarly using the distance between the eye and the tip of the beak.
[0073] Furthermore, in the above embodiment, if the difference in distance between the two eyes is determined to be less than the threshold in S306, the eye closer to the center of the image is selected in S308. On the other hand, a tracking process for tracking a subject in a video may be used. In this case, in S308, the selection unit 250 can select an eye closer to the coordinates of the eye tracked in the previous frame. For example, one of the two eyes may be selected based on the eye tracking result in the past frame. The selection unit 250 may have selected one first type part from a plurality of first type parts (eyes in this example) detected from a past image captured at a time before the image. In this case, in S308, the selection unit 250 can select one of the two first type parts based on the position of one first type part selected from a plurality of first type parts detected from the past image. For example, the selection unit 250 can select the eye closer to the one eye selected from the plurality of eyes detected from the past image. Also, in S308, the selection unit 250 may select the larger eye.
[0074] Also, in the above embodiment, if one or more eyes are detected according to the determination in S303, AF processing is performed to focus on the eyes. On the other hand, if one or more eyes are not detected, AF processing may be performed to focus on a part other than the eyes. For example, the extraction unit 220 may detect a face frame, a head frame, or a whole body frame from the image. In this case, the extraction unit 220 may perform AF processing to preferentially focus on a smaller frame.
[0075] In the above example, it is assumed that one person is included in the image. Therefore, in S304, it is determined whether two eyes are detected. If multiple people are included, three or more eyes may be detected. However, in this case, the eye closer to the camera can be selected using a similar method. For example, the first part estimation unit 230 may detect a third part in the image. This third part may be, for example, a face, a head, or a person. In addition, the first part estimation unit 230 may determine a frame corresponding to the third part, such as a face frame, a head frame, or a person area. Such a determination can be made using the feature amount of the image extracted by the extraction unit 220.
[0076] In this case, the first part estimation unit 230 can detect the first type of part from inside the detected frame. For example, the first part estimation unit 230 can detect only the eyes included in the face frame, head frame, or person area. The first part estimation unit 230 may also detect only two eyes included in one specific face frame, head frame, or person area. This specific one face frame, head frame, or person area may be the largest face frame, head frame, or person area in the image. According to this method, even when eyes of multiple people are detected, the selection target can be limited to two eyes of one person.
[0077] In this embodiment, among the multiple eyes, the eye that is closer to the camera is preferentially selected. However, the priority order is not limited to this. For example, the eye that is farther from the camera may be preferentially selected. Also, the eye that is included in the face frame may be preferentially selected.
[0078] As described above, in this embodiment, a first type of part closer to the imaging device is selected based on the distance between the first type of part (e.g., eyes) and a second type of part (e.g., nose). According to the method of this embodiment, the number of parts required for detection can be reduced compared to the case where the posture of the subject (e.g., face direction) is estimated. Therefore, a first type of part closer to the imaging device can be detected quickly. In addition, the preparation cost of correct answer data used for learning to detect each part can be reduced.
[0079] In particular, in this embodiment, a learning model is used to estimate the distance between the first type of part and the second type of part. In such a configuration, even if the nose is hidden, the distance between the eyes and the nose can be estimated based on information about other parts of the face (e.g., parts around the eyes). This improves the accuracy and robustness of the process of detecting the eyes close to the imaging device.
[0080] In addition, according to the method of detecting the eyes close to the imaging device based on the distance between the eyes and the nose (or other second part) as in the present embodiment, the detection accuracy is improved compared to the method of selecting the eyes using only information about the eyes (e.g., the size of the eyes). In particular, when the eyes are partially hidden, the detection accuracy is improved according to the method of the present embodiment.
[0081] Also, as in this embodiment, by detecting the eye closest to the imaging device and focusing on the detected eye, it is possible to take an impressive facial photograph. That is, according to the configuration of this embodiment, it is possible to take a more impressive facial photograph compared to the case where the largest eye in the image (e.g., the eye that is more open) is focused on in order to improve the focus accuracy. Also, according to the configuration of this embodiment, it is possible to take a more impressive facial photograph compared to the case where the part in the image that has the highest likelihood of being an eye is focused on.
[0082] <Variation 1> The method of estimating the distance between a first type of part (e.g., eyes) and a second type of part (e.g., nose) in an image is not limited to the above method. For example, instead of using the distance map as described above, the second type of part in the image may be detected. According to such a detection result, the distance between the first type of part and the second type of part can be calculated. In the following modified example, in addition to the first type of part, the center coordinates of the second type of part are estimated. Then, based on the estimated center coordinates of the first type of part and the second type of part, the distance between the first type of part and the second type of part is calculated. Then, based on the calculated distance, the first type of part to be prioritized is selected. In the following description, the first type of part is a person's eyes, and the second type of part is a nose.
[0083] Fig. 2(B) shows the functional configuration of image processing unit 104, which is an image processing device according to this modification. In this modification, image processing unit 104 has extraction unit 271 instead of extraction unit 220. Also, image processing unit 104 has distance calculation unit 273 instead of distance estimation unit 240. Furthermore, image processing unit 104 has second part estimation unit 272. The other configurations are similar to the configurations shown in Fig. 2(A), and detailed descriptions of these configurations will be omitted.
[0084] The extraction unit 271 extracts features from the image acquired by the input unit 210, similarly to the extraction unit 220. FIG. 9 shows an example of the structure of a neural network used in this modification. The neural network shown in FIG. 9 outputs a center position map 918 indicating the center position of the nose in addition to a center position map 904 indicating the center position of the eye and two size maps 907 and 910 indicating the width and height of the frame surrounding the eye. In FIG. 9, 900 to 912, 916 to 917, and 920 to 922 are the same as 401 to 412, 416 to 417, and 420 to 422 in FIG. 4. The features output from the network 920 are input to the network 923. Then, the network 923 outputs the center position map 918. The center position map 918 shows the magnitude 919 of the likelihood of the nose of the person. The nose center position map 918 indicates that the closer to the center of the circle, the higher the likelihood, and that the location of the nose is high, similar to the eye center position map 904. The extraction unit 271 can extract the feature amount by processing using a neural network as shown in FIG.
[0085] Second part estimation section 272 detects a second part in the image. For example, second part estimation section 272 can determine the position of the second part based on the likelihood of the second part for each position in the image. Second part estimation section 272 can detect the second part based on the feature amount of the image extracted by extraction section 271. For example, second part estimation section 272 can select one position with a high likelihood shown in the nose center position map output by extraction section 271 as the center position of the nose. This process can be performed in the same manner as first part estimation section 230. Note that, as described above, when first part estimation section 230 detects a first type of part from within a predetermined frame (for example, a face frame), second part estimation section 272 may detect a second part from within the same frame.
[0086] Distance calculation section 273 determines the distance between a first type of part detected from an image by first part estimation section 230 and a second part detected from an image by second part estimation section 272. For example, distance calculation section 273 can calculate the distance between each of the center positions of two eye frames detected by first part estimation section 230 and the position of the nose detected by second part estimation section 272. This distance may be a Euclidean distance or another type of distance.
[0087] The image processing method according to this modification can be performed according to the flowchart shown in FIG. 3. In this modification, the following process is performed in S305. First, the second part estimation unit 272 estimates the center position of the nose as described above. Furthermore, the distance calculation unit 273 calculates the distance between the eye and the nose for each of the two eyes detected in S302 as described above. For example, the distance calculation unit 273 can calculate the Euclidean distance between the coordinates (x coordinate and y coordinate) of the center position of the frame obtained in S302 and the coordinates (x coordinate and y coordinate) of the estimated center position of the nose.
[0088] In this modified example, a case will be described where the distance between one eye and the nose is 49 and the distance between the other eye and the nose is 33. In this case, the difference between the two distances is 16. If the threshold value used in S306 is 15, this difference is greater than or equal to the threshold, and the process proceeds to S307. In this case, in S307, the eye that is farther away from the nose is selected, in this example, the eye with the distance to the nose of 49.
[0089] Next, a learning method of a neural network as shown in FIG. 9 will be described. The learning process can be performed using the learning device 500 shown in FIG. 5. Detailed description of the configuration already described will be omitted. In the case of this modified example, the object estimation unit 540 performs processing equivalent to that of the extraction unit 271. That is, the object estimation unit 540 can output an eye center position map, an eye size map, and a nose center position map of a person in a learning image.
[0090] The data creation unit 550 generates an eye center position correct map, an eye size correct map, and a nose center position correct map as teacher data. As an example, a case will be described in which a learning image 700 in which a person is shown and the correct answer information shown in FIG. 7(B) are used as learning data. The correct answer information indicates the center coordinates (X, Y)=(960, 360) of the nose, which are coordinates on the learning image 700. The nose center position correct answer map (not shown), which is the teacher data for the nose center position, is matrix data of the same size as the nose center position map output by the target estimation unit 540. In this example, the size of the nose center position correct answer map is 320×240, which is 1 / 5 the size of the learning image in both length and width. Therefore, in order to obtain the correct answer information 720 shown in FIG. 7(C) expressed by the coordinates of the correct answer map, the values indicating the eye center coordinates and eye size on the learning image are reduced to 1 / 5. 7C, the coordinates of the center of the nose are (X, Y) = (192, 72). The correct nose center position map can be created in the same way as the correct eye center position map 730, according to the coordinates of the center of the nose.
[0091] The error calculation unit 560 calculates the eye center position error and the eye size error as described above. Furthermore, the error calculation unit 560 calculates the nose center position error, which is the error between the nose center feature map output by the object estimation unit 540 and the nose center position correct map created by the data creation unit 550. The learning unit 570 can update the parameters used by the object estimation unit 540 based on the error calculated by the error calculation unit 560 in this manner.
[0092] The processing of the learning method according to this modified example can be performed according to the flowchart shown in Fig. 8. In this modified example, in S803, the object estimation unit 540 outputs an eye center position map, an eye size map, and a nose center position map by inference processing on the learning image acquired from the image acquisition unit 530. In S804, the data creation unit 550 creates an eye center position correct answer map, an eye size correct answer map, and a nose center position correct answer map from the correct answer information acquired in S801 according to the above-mentioned method. In S805, the error calculation unit 560 calculates the eye center position error, eye size error, and nose center position error as described above. Other processing can be performed as already described.
[0093] As described above, in this modification, the distance between a first type of part (e.g., eyes) and a second type of part (e.g., nose) is directly calculated. With this configuration, when the second type of part is clearly visible, it is expected that the accuracy of detecting the first type of part closer to the imaging device can be improved.
[0094] <Variation 2> In this modification, a first type of part closer to the camera is directly detected. Fig. 2(C) shows the functional configuration of image processing unit 104, which is an image processing device according to this modification. In this modification, image processing unit 104 has extraction unit 281 and first part estimation unit 282 in addition to input unit 210. The function of input unit 210 has already been described.
[0095] The extraction unit 281 extracts a feature amount from the image acquired by the input unit 210, similarly to the extraction unit 220. The first part estimation unit 282 detects a proximate part, which is a part of a first type in the image (for example, an eye) that is closer to an imaging device that captured the image than other parts of the first type in the image. The first part estimation unit 282 can detect a proximate part based on the feature amount of the image extracted by the extraction unit 281. For example, the extraction unit 281 can generate a center position map indicating the center position of the eye that is closer to the camera. The first part estimation unit 282 can select one frame according to the likelihood indicated by the center position map. The extraction unit 281 can also generate two size maps indicating the width and height of the frame surrounding the eye, as already described. Then, the first part estimation unit 282 outputs information indicating the detection result.
[0096] First part estimation unit 282 can detect such nearby parts using a learning model trained to estimate nearby parts in an image. In this modification, extraction unit 281 is also realized by a learning model having learned parameters. In one embodiment, when generating teacher data used for learning this parameter, the center position of the second part (e.g., nose) or the distance between the first type of part and the second part is used. This configuration allows first part estimation unit 282 to estimate the eye closer to the camera with higher accuracy.
[0097] 11 is a flowchart showing the flow of processing of the image processing method in this modified example. S1100 and S1110 are performed in the same manner as S300 and S310. In S1101, the extraction unit 281 extracts features from the image acquired in S1100. In this example, the extraction unit 281 outputs an eye center position map indicating the center positions of the eyes closest to the camera, and a size map of the eyes closest to the camera as features. The eye center position map in this modified example indicates the likelihood that a nearby part exists in the image.
[0098] FIG. 12 shows an example of the structure of a neural network used in this modification. The neural network shown in FIG. 12 outputs a center position map showing the center position of the eye closest to the camera, and two size maps showing the width and height of the frame surrounding the eye. In FIG. 12, 1207 to 1211, 1216, 1217, 1220, and 1222 are the same as 407 to 411, 416, 417, 420, and 422 in FIG. 4. FIG. 12 shows an input image 1200. The input image 1200 shows the eye 1201 of a person closer to the camera, the eye 1202 of a person farther from the camera, and the nose 1203 of the person. The feature amount output from the network 1220 is input to the network 1221. Then, as shown in FIG. 12, the network 1221 outputs a center position map 1204 showing the center position of the eye of the person closer to the camera. FIG. 12 also shows the magnitude 1205 of the likelihood of the person's eye being closer to the camera. Similar to the center position map 904, the center position map 1204 indicates that the closer to the center of the circle the higher the likelihood, and that the location of the eyes of a person close to the camera is high. The extraction unit 281 can extract the feature amount by processing using a neural network as shown in FIG.
[0099] In S1102, first part estimation section 282 selects one eye frame based on the likelihood indicated by the eye center position map. For example, first part estimation section 282 can select an eye frame that has the highest likelihood and that exceeds a preset threshold. This process can be performed in the same way as first part estimation section 230, except that only one eye frame is selected.
[0100] In S1103, control unit 106 determines whether or not one or more eyes have been detected. If it is determined that one or more eyes have not been detected, the process in Fig. 11 ends. If it is determined that one or more eyes have been detected, control unit 106 performs AF processing on one eye frame selected by first part estimation unit 282 in S1110.
[0101] <Learning Method> Next, a learning method of a neural network as shown in FIG. 12 will be described. The learning process can be performed using the learning device 500 shown in FIG. 5. Detailed description of the configuration already described will be omitted. In the case of this modified example, the object estimation unit 540 performs processing equivalent to that of the extraction unit 281. That is, the object estimation unit 540 can obtain a size map of eyes closer to the camera, which indicates the likelihood that a part closer to the imaging device that captured the image exists than other parts of the first type in the learning image, obtained by inputting the learning image into the learning model. In this embodiment, the object estimation unit 540 can output a center position map indicating the center position of the eye closer to the camera in the learning image, and a size map of the eye closer to the camera of the person.
[0102] 14(A) to (C) are schematic diagrams of a learning image and correct answer information. FIG. 14(A) shows a learning image 1400 showing a person's face. FIG. 14(B) shows correct answer information 1410 for the learning image 1400. The information indicated by the correct answer information 1410 is based on the coordinates on the learning image 1400. FIG. 14(C) shows correct answer information 1420 expressed by the coordinates of the correct answer map. The learning data used for learning in this example includes the learning image 1400 and the correct answer information 1410. The correct answer information 1410 indicates the center coordinates (X, Y)=(900, 380) of the person's first eye (eye 1), the size of the first eye (25), the center coordinates (X, Y)=(960, 340) of the person's second eye (eye 2), and the size of the second eye (20). The correct answer information 1410 further indicates the center coordinates (X, Y) of the person's face (910, 440), the size of the face (240), and the center coordinates of the nose (985, 440). Here, the center coordinates of the first eye, the center coordinates of the second eye, the center coordinates of the face, and the center coordinates of the nose are indicated by coordinates on the learning image 1400. The size of the first eye, the size of the second eye, and the size of the face indicate the sizes on the learning image 1400.
[0103] The data creation unit 550 generates teacher data based on the correct answer information acquired from the data acquisition unit 520. This teacher data is used as a target value for the output of the object estimation unit 540. In this example, the data creation unit 550 can generate a correct answer map of the center position of the eye closest to the camera and a correct answer map of the size of the eye closest to the camera as teacher data. A specific generation method will be described later with reference to FIG. 10.
[0104] The processing of the learning method according to this modification can be performed according to the flowchart shown in Fig. 8. This method can produce a learning model having parameters trained to be used to estimate the position of a first type of part in an image that is closer to an imaging device that captured the image than other first type parts in the image.
[0105] In this modified example, in S803, the object estimation unit 540 outputs a center position map of the eye closest to the camera and a size map of the eye closest to the camera by inference processing on the training image acquired from the image acquisition unit 530. In S804, the data creation unit 550 creates a center position correct map of the eye closest to the camera and a size correct map of the eye closest to the camera from the correct answer information acquired in S801. This correct answer information indicates the distance between each of a first type of multiple parts (eyes in this example) in the training image and a second part (nose in this example). The data creation unit 550 can generate a center position correct answer map based on such distances.
[0106] The process of S804 in this modification can be performed according to the flowchart shown in FIG. 10. FIG. 10 shows the flow of the teacher data creation process. In S1501, the data creation unit 550 calculates a first distance between the center coordinates of the first eye and the center coordinates of the nose. The data creation unit 550 also calculates a second distance between the center coordinates of the second eye and the center coordinates of the nose. In the example shown in FIG. 14(B), the first distance is (86) and the second distance is (65). In this example, Euclidean distance is used as the first distance and the second distance. However, other types of distances may be used. For example, the first distance and the second distance may be normalized distances normalized by dividing by a value such as an image size.
[0107] In S1502, the data creation unit 550 calculates the difference between the first distance and the second distance. In the example shown in Fig. 14(B), the first distance is greater than the second distance, and the absolute value of the difference is (21).
[0108] In S1503, the data creation unit 550 judges whether the absolute value of the difference in distance calculated in S1502 is equal to or greater than a threshold. In this example, (10) is used as the threshold. If it is judged that the absolute value of the difference in distance is equal to or greater than the threshold, the process proceeds to S1504. If it is judged that the absolute value of the difference in distance is less than the threshold, the process proceeds to S1505.
[0109] In S1504, the data creation unit 550 selects one of a plurality of parts of the first type (eyes in this example) in the learning image. The distance between the selected part and a second part (nose in this example) is longer than the distance between the other parts of the first type and the second part. For example, the data creation unit 550 can select the eye closer to the nose among the two eyes. In the example shown in FIG. 14(B), the first distance is greater than the second distance, so the first eye is selected. Then, the data creation unit 550 generates a center position correct answer map of the eye closer to the camera based on the position of the selected eye. The data creation unit 550 can obtain a correct answer map in which a label is given to a position corresponding to the selected part. For example, the data creation unit 550 can assign a label to a position corresponding to the first eye in the center position correct answer map.
[0110] FIG. 13 shows a schematic diagram of a center position correct map of the eye closest to the camera. FIG. 13(A) shows a learning image 1600 in which a person's face facing rightward is shown. The learning image 1600 shows an eye 1601 closer to the camera and an eye 1602 farther from the camera. FIG. 13(C) also shows a center position correct map 1606 of the eye closest to the camera corresponding to the learning image 1600, and a likelihood magnitude 1607 of the eye closest to the camera. If the learning image shows a person's face facing rightward, the process of S1504 is performed. In this case, as shown in FIG. 13(B), a label is added only to the position corresponding to the selected one eye in the center position correct map 1606 corresponding to the learning image 1600.
[0111] In S1505, the data creating unit 550 generates a center position correct map of the eye close to the camera based on the positions of each of the two eyes. In this example, for example, the data creating unit 550 can assign labels to the positions of the center position correct map corresponding to the first eye and the second eye.
[0112] FIG. 13B shows a learning image 1603 in which a person's face facing forward is captured. The learning image 1600 captures eyes 1604 and 1605 that are at approximately the same distance from the camera. FIG. 13D also shows a center position correct answer map 1608 of the eye closest to the camera corresponding to the learning image 1603, and likelihood magnitudes 1609 and 1610 of the eye closest to the camera. If the learning image captures a person's face facing forward, the process of S1505 is performed. In this case, as shown in FIG. 13D, labels are added to positions corresponding to the two eyes in the center position correct answer map 1608 corresponding to the learning image 1603.
[0113] The method of labeling in S1504 and S1505 is similar to the method of labeling the eye center position correct map 730 shown in FIG. 7(A), and therefore a description thereof will be omitted.
[0114] In addition, in S804, the data creation unit 550 also creates a size correct answer map of the eye closest to the camera. The method of creating the size correct answer map has already been described. For example, in S1504, the data creation unit 550 can assign a label to an area in the size correct answer map corresponding to one selected eye. In addition, in S1505, the data creation unit 550 can assign labels to areas in the size correct answer map corresponding to each of the two detected eyes.
[0115] In S805, the error calculation unit 560 calculates a center position error based on the center position correct map of the eye closest to the camera created in S804 and the center position map of the eye closest to the camera obtained in S803. In addition, the error calculation unit 560 calculates a size error based on the size correct map of the eye closest to the camera created in S804 and the size map of the eye closest to the camera obtained in S803. In S806, the learning unit 570 learns the parameters used by the target estimation unit 540 based on the center position error and the eye size error. The specific learning method has already been described. For example, the learning unit 570 may perform a learning process according to the error backpropagation method. Other processes can be performed as already described. In this way, the learning unit 570 can learn the parameters of the learning model based on the map obtained by inputting a learning image into the learning model trained to estimate the adjacent part in the image and the size correct map of the eye closest to the camera.
[0116] As described above, in this modification, the first type of part (e.g., eyes) closer to the imaging device is directly estimated, thereby detecting the first type of part closer to the imaging device. With this configuration, even if the second part is hidden, it is expected that the accuracy of detecting the first type of part closer to the imaging device can be improved.
[0117] (Other Examples) In the above embodiment, the first type of part is detected and the size of the first type of part is detected. However, in order to quickly detect parts closer to the imaging device, it is not essential to detect the size of the first type of part. Therefore, it is not essential that the neural networks shown in Figs. 4, 9, and 12 have networks 422, 822, and 1222 for detecting the size of the eyes. It is also not essential to perform learning using a size answer map.
[0118] The present invention can also be realized by a process in which a program for implementing one or more of the functions of the above-described embodiments is supplied to a system or device via a network or a storage medium, and one or more processors in a computer of the system or device read and execute the program. The present invention can also be realized by a circuit (e.g., ASIC) that implements one or more of the functions.
[0119] The disclosure of this specification includes the following image processing device, imaging device, method, and program. (Item 1) A first position estimation means for detecting a first type of part in an image; a distance estimation means for estimating a distance between the first type of site and a second type of site in the image; a selection means for selecting one of the first type sites from the detected plurality of first type sites based on the distance estimated for each of the first type sites, and outputting information indicating the selected site; An image processing device comprising: (Item 2) 2. The image processing device according to claim 1, wherein the distance estimation means estimates the distance using a map indicating the distance to the second part for the first type of part in the image, the map being obtained using a learning model. (Item 3) Further comprising a second position estimation means for detecting a second part in the image, 2. The image processing device according to claim 1, wherein the distance estimation means determines a distance between the first type of part detected from the image and the second type of part detected from the image. (Item 4) 4. The image processing device according to any one of items 1 to 3, characterized in that the distance estimated for the selected one first type part is longer than the distance estimated for other parts of the plurality of first type parts. (Item 5) The first position estimation means detects two parts of the first type; 5. The image processing device according to any one of items 1 to 4, wherein the selection means selects a method for selecting one of the two parts based on a difference in distance between each of the two parts and the second part. (Item 6) The selection means is if the difference is equal to or greater than a threshold, selecting a first method of selecting the one site based on a comparison of distances between each of the two sites and the second site; If the difference is less than a threshold, selecting a second method different from the first method. 6. The image processing device according to item 5, (Item 7) the selection means selects one of the first type of regions from a plurality of the first type of regions detected from a previous image captured at a time prior to the image; 7. The image processing device according to claim 6, wherein the selection means, in the second method, selects one of the two parts based on the position of one of the first type parts selected from the plurality of first type parts detected from the past image. (Item 8) 8. The image processing device according to any one of items 1 to 7, wherein the second region is a region equidistant from a plurality of regions of the first type. (Item 9) 9. The image processing device according to any one of items 1 to 8, wherein the second part is a nose, a mouth, a chin, a region between the eyebrows, or a top of the head. (Item 10) An acquisition means for acquiring an image; a first position estimation means for detecting a proximate portion, which is a first type of portion in the image and is a portion closer to an imaging device that captured the image than other portions of the first type in the image, by using a learning model having parameters trained to estimate the proximate portion in the image, and outputting information indicating the detection result; An image processing device comprising: (Item 11) Item 11. The image processing device according to item 10, wherein the learning model outputs a map indicating the likelihood that the adjacent part exists in the image. (Item 12) Item 10 or 11, characterized in that the parameters are learned based on a map obtained by inputting a learning image including a plurality of parts of the first type into the learning model, and a correct answer map in which a label is assigned to a position corresponding to a part selected from the plurality of parts of the first type in the learning image such that a distance between the part and a second part is longer than a distance between other parts of the first type and the second part. (Item 13) 13. The image processing device according to any one of items 1 to 12, wherein the first type of part is an eye or an ear. (Item 14) 14. The image processing device according to any one of items 1 to 13, characterized in that the first position estimation means detects a frame corresponding to a third part in the image and detects the first type of part from inside the frame. (Item 15) Item 15. The image processing device according to item 14, wherein the third part is a face, a head, or a person. (Item 16) An imaging means for capturing an image; a first position estimation means for detecting a first type of part in the image; a distance estimation means for estimating a distance between the first type of site and a second type of site in the image; a selection means for selecting one of the first type sites from the detected plurality of first type sites based on the distance estimated for each of the first type sites; A focus control means for controlling the imaging means so as to focus on a selected one of the first type parts; An imaging device comprising: (Item 17) 1. A method for producing a learned model having trained parameters for use in estimating a distance between a first type of feature and a second type of feature in an image, comprising: acquiring a learning image and a correct answer map in which information indicating a distance between the first type of site and the second type of site is provided at a position corresponding to the first type of site in the learning image; A step of learning parameters of the learning model based on a map indicating a distance to the second part for the first type of part in the learning image, the map being obtained by inputting the learning image into the learning model, and the correct answer map; The method according to claim 1, further comprising: (Item 18) 1. A method for producing a learned model having parameters trained for use in estimating the location of a first type of feature in an image that is closer to an imaging device that captured the image than other features of the first type in the image, the method comprising: acquiring a learning image including a plurality of parts of the first type, and a ground truth map in which a label is assigned to a position corresponding to a part selected from the plurality of parts of the first type in the learning image such that the distance between the part and a second part is longer than the distance between the other parts of the first type and the second part; A step of learning parameters of the learning model based on a map indicating a likelihood that a part of the training image is closer to an imaging device that captured the image than other parts of a first type in the training image, the map being obtained by inputting the training image into the learning model, and the correct answer map; The method according to claim 1, further comprising: (Item 19) An image processing method performed by an image processing device, comprising: detecting a first type of site in the image; estimating a distance between the first type of feature and a second type of feature in the image; selecting one of the first type sites from the detected plurality of first type sites based on the distance estimated for each of the first type sites, and outputting information indicating the selected site; 13. An image processing method comprising: (Item 20) 16. A program for causing a computer to function as the image processing device according to any one of items 1 to 15.
[0120] The invention is not limited to the above-described embodiments, and various modifications and variations are possible without departing from the spirit and scope of the invention. Accordingly, the following claims are appended to apprise the public of the scope of the invention. [Explanation of symbols]
[0121] 104: image processing unit, 106: control unit, 210: input unit, 220: extraction unit, 230: first part estimation unit, 240: distance estimation unit, 250: selection unit, 271: extraction unit, 272: second part estimation unit, 273: distance calculation unit, 281: extraction unit, 282: first part estimation unit, 500: learning device, 510: data storage unit, 520: data acquisition unit, 530: image acquisition unit, 540: object estimation unit, 550: data creation unit, 560: error calculation unit, 570: learning unit
Claims
1. A first position estimation means for detecting a first type of part in an image; a distance estimation means for estimating a distance between the first type of part and a second type of part in the image; a selection means for selecting one of the first type regions from the detected plurality of first type regions based on the distance estimated for each of the first type regions, and outputting information indicating the selected region; An image processing device comprising:
2. The image processing device according to claim 1 , wherein the distance estimation means estimates the distance using a map indicating a distance to the second part for the first type of part in the image, the map being obtained using a learning model.
3. a second position estimation means for detecting a second part in the image; The image processing device according to claim 1 , wherein the distance estimation means determines a distance between the first type of part detected from the image and the second type of part detected from the image.
4. The image processing device according to claim 1 , wherein the distance estimated for the selected one of the first type parts is longer than the distance estimated for other parts of the plurality of first type parts.
5. The first position estimation means detects two parts of the first type; 2. The image processing device according to claim 1, wherein the selection means selects a method for selecting one of the two regions based on a difference in distance between each of the two regions and the second region.
6. The selection means is if the difference is equal to or greater than a threshold, selecting a first method of selecting the one site based on a comparison of distances between each of the two sites and the second site; If the difference is less than a threshold, selecting a second method different from the first method.
6. The image processing device according to claim 5,
7. the selection means selects one of the first type of regions from a plurality of the first type of regions detected from a past image captured at a time prior to the image; 7. The image processing device according to claim 6, wherein the selection means, in the second method, selects the one of the two parts based on a position of one of the first type parts selected from a plurality of the first type parts detected from the past image.
8. The image processing device according to claim 1 , wherein the second region is a region equidistant from the plurality of regions of the first type.
9. The image processing device according to claim 1 , wherein the second part is a nose, a mouth, a chin, a region between the eyebrows, or a top of the head.
10. An acquisition means for acquiring an image; a first position estimation means for detecting a proximate portion, which is a first type of portion in the image and is closer to an imaging device that captured the image than other first type portions in the image, by using a learning model having parameters trained to estimate the proximate portion in the image, and outputting information indicating the detection result; An image processing device comprising:
11. The image processing device according to claim 10 , wherein the learning model outputs a map indicating the likelihood that the adjacent portion exists in the image.
12. The image processing device of claim 11, characterized in that the parameters are learned based on a map obtained by inputting a training image including a plurality of parts of the first type into the learning model, and a correct answer map in which labels are assigned to positions corresponding to parts selected from among the plurality of parts of the first type in the training image such that a distance between the part and a second part is longer than a distance between other parts of the first type and the second part.
13. The image processing device according to claim 1 , wherein the first type of part is an eye or an ear.
14. 2 . The image processing device according to claim 1 , wherein the first position estimation means detects a frame corresponding to a third part in the image, and detects the first type of part from within the frame.
15. The image processing device according to claim 14 , wherein the third part is a face, a head, or a person.
16. An imaging means for capturing an image; a first position estimation means for detecting a first type of part in the image; a distance estimation means for estimating a distance between the first type of part and a second type of part in the image; a selection means for selecting one of the first type regions from the detected plurality of first type regions based on the distance estimated for each of the first type regions; a focus control means for controlling the imaging means so as to focus on a selected one of the first type parts; An imaging device comprising:
17. 1. A method for producing a learned model having trained parameters for use in estimating a distance between a first type of feature and a second type of feature in an image, comprising: acquiring a learning image and a correct answer map in which information indicating a distance between the first type of site and the second type of site is provided at a position corresponding to the first type of site in the learning image; A step of learning parameters of the learning model based on a map indicating a distance to the second region for the first type of region in the learning image, the map being obtained by inputting the learning image into the learning model, and the correct answer map; The method according to claim 1, further comprising:
18. 1. A method for producing a learned model having parameters trained to be used to estimate the location of a first type of feature in an image that is closer to an imaging device that captured the image than other features of the first type in the image, the method comprising: acquiring a learning image including a plurality of parts of the first type, and a ground truth map in which a label is assigned to a position corresponding to a part selected from the plurality of parts of the first type in the learning image such that a distance between the part and a second part is longer than a distance between the other parts of the first type and the second part; a step of learning parameters of the learning model based on a map indicating a likelihood that a part of the training image is closer to an imaging device that captured the image than other parts of a first type in the training image, the map being obtained by inputting the training image into the learning model, and the correct answer map; The method according to claim 1, further comprising:
19. An image processing method performed by an image processing device, comprising: detecting a first type of site in the image; estimating a distance between the first type of feature and a second type of feature in the image; selecting one of the first type sites from the detected plurality of first type sites based on the distance estimated for each of the first type sites, and outputting information indicating the selected site; 13. An image processing method comprising:
20. A program for causing a computer to function as the image processing device according to any one of claims 1 to 15.