Imaging device, control method for imaging device, and program
The information processing device addresses the inefficiency in manual selection of subject images for machine learning by automating the process, enabling efficient data collection and supervised learning through focus area identification and adjustable frames.
Patent Information
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- FUJIFILM CORP
- Filing Date
- 2024-05-23
- Publication Date
- 2026-04-24
AI Technical Summary
Existing imaging devices face challenges in efficiently collecting data for machine learning, particularly in selecting specific subject images manually from captured images, which is time-consuming and inefficient.
An information processing device that automatically identifies and outputs specific subject data for machine learning by setting a focus target area, allowing for supervised learning and displaying the focus area distinguishably, with adjustable frames and coordinates for image extraction.
Facilitates easier and more efficient collection of data for machine learning by automating the selection and labeling of specific subject images, enhancing the training process for image classification models.
Smart Images

Figure 0007851356000001 
Figure 0007851356000002 
Figure 0007851356000003
Abstract
Description
[Technical Field]
[0001] The technology disclosed herein relates to an information processing device, a learning device, an imaging device, a control method for an information processing device, and a program. [Background technology]
[0002] International Publication No. 2008 / 133237 discloses an imaging device for capturing images of an object space. This imaging device is characterized by comprising a subject feature point learning means, a subject feature point learning information storage unit, a shooting candidate image information acquisition means, an image search processing means, and a shooting condition adjustment means. The subject feature point learning means detects an image of a predetermined subject from image information obtained by capturing an object space, and extracts subject feature point learning information that indicates the feature points of the image of the subject. The subject feature point learning information storage unit stores the subject feature point learning information. The shooting candidate image information acquisition means acquires a shooting candidate image, which is an image that is a candidate for capture. The image search processing means determines from the acquired shooting candidate image whether the shooting candidate image feature point information that indicates the feature points of at least one subject image contained in the shooting candidate image contains feature points that match the feature points indicated by the subject feature point learning information that was previously stored in the subject feature point learning information storage unit. If the shooting condition adjustment means determines, as a result of the determination, that the candidate image feature point information contains feature points that match the feature points indicated by the subject feature point learning information, it instructs the shooting condition optimization means to optimize the shooting conditions for the subject in the candidate image that corresponds to the candidate image feature point information.
[0003] Japanese Patent Publication No. 2013-80428 discloses a program that causes a computer to perform an acquisition step of acquiring first learning data adapted by a first device through learning, and a data conversion step of converting the acquired first learning data into learning data in a data format that conforms to the data format of the second learning data, based on the data format of the second learning data adapted by a second device through learning. [Overview of the project]
[0004] One embodiment of the technology disclosed herein provides an information processing device that can collect data for machine learning more easily than when specific subject images for machine learning are manually selected from captured images obtained by an image sensor. [Means for solving the problem]
[0005] A first aspect of the technology of this disclosure is an information processing device comprising a processor and a memory connected to or built into the processor, wherein when imaging is performed by an image sensor with a focus operation that sets a specific subject as the focus target area, the information processing device outputs specific subject data relating to a specific subject image that shows the specific subject in the captured image obtained by imaging as data to be used for machine learning.
[0006] A second aspect of the technology of this disclosure is an information processing device according to the first aspect, wherein the machine learning is supervised machine learning, and the processor assigns labels, which are information relating to specific subject images, to specific subject data, and outputs the specific subject data as training data used for supervised machine learning.
[0007] A third aspect of the technology of this disclosure is an information processing device according to the first or second aspect, wherein the processor displays a displayable moving image based on a signal output from an image sensor on a monitor in a manner that allows the focus target area to be distinguished from other image areas, and the specific subject image is an image corresponding to the position of the focus target area in the captured image.
[0008] A fourth aspect of the technology of this disclosure is an information processing device relating to a third aspect, wherein the processor displays a frame surrounding the focus target area in a displayable moving image, thereby displaying the focus target area in a manner that makes it distinguishable from other image areas.
[0009] A fifth aspect of the technology of this disclosure is an information processing device according to the fourth aspect, wherein the position of the frame can be changed according to a given position change instruction.
[0010] A sixth aspect of the technology of this disclosure is an information processing device according to the fourth or fifth aspect, wherein the size of the frame is changeable according to a given resizing instruction.
[0011] A seventh aspect of the technology of this disclosure is an information processing device relating to any one of the first to sixth aspects, wherein the processor outputs an captured image and the coordinates of the focus target area as data to be used for machine learning.
[0012] An eighth aspect of the technology of this disclosure is an information processing device according to the first or second aspect, wherein the processor displays a display video on a monitor based on a signal output from an image sensor, accepts the designation of a focus target area in the display video, and extracts a specific subject image from a predetermined area including the focus target area, based on an area where the similarity evaluation value indicating the degree of similarity to the focus target area is within a first predetermined range.
[0013] A ninth aspect of the technology of this disclosure is an information processing device relating to the eighth aspect, wherein the processor displays the focus target area in a manner that makes it distinguishable from other image areas.
[0014] A tenth aspect of the technology of this disclosure is an information processing device according to the eighth or ninth aspect, wherein at least one of the focus target area and the specific subject image is defined in units of divided areas obtained by dividing a predetermined area.
[0015] An eleventh aspect of the technology of this disclosure is an information processing device relating to any one of the eighth to tenth aspects, wherein the similar evaluation value is a value based on the focus evaluation value used in the focus operation.
[0016] The twelfth aspect according to the technology of the present disclosure is an information processing apparatus according to any one of the eighth to eleventh aspects, where the similarity evaluation value is a color evaluation value based on the color information of a predetermined region.
[0017] The thirteenth aspect according to the technology of the present disclosure is an information processing apparatus according to any one of the eighth to twelfth aspects, where when the difference between a display specific subject image indicating a specific subject in a display moving image and the specific subject image exceeds a second predetermined range, the processor performs an abnormality detection process, and the display specific subject image is determined based on the similarity evaluation value.
[0018] The fourteenth aspect according to the technology of the present disclosure is an information processing apparatus according to any one of the first to thirteenth aspects, where the specific subject data includes the coordinates of the specific subject image, and the processor outputs the captured image and the coordinates of the specific subject image as data for use in machine learning. <000009%>
[0019] The fifteenth aspect according to the technology of the present disclosure is an information processing apparatus according to any one of the first to fourteenth aspects, where the specific subject data is a specific subject image cut out from the captured image, and the processor outputs the cut-out specific subject image as data for use in machine learning.
[0020] The sixteenth aspect according to the technology of the present disclosure is an information processing apparatus according to any one of the first to fifteenth aspects, where the processor stores data in a memory and performs machine learning using the data stored in the memory.
[0021] The seventeenth aspect according to the technology of the present disclosure is a learning device including a reception device that receives data output from an information processing apparatus according to any one of the first to fifteenth aspects, and an arithmetic device that performs machine learning using the data received by the reception device. <0000!05>
[0022] An eighteenth aspect of the technology of this disclosure is an imaging device comprising an information processing device according to any one of the first to sixteenth aspects, and an image sensor.
[0023] A 19th aspect of the technology of this disclosure is an imaging device according to an 18th aspect, wherein the image sensor takes images at multiple focus positions, and the processor outputs the coordinates of a specific subject image obtained from a focus image that is in focus on a specific subject as the coordinates of a specific subject image in an out-of-focus image that is not in focus on the specific subject, for the multiple images obtained by the imaging.
[0024] A 20th aspect of the technology of this disclosure is a control method for an information processing device, which includes outputting specific subject data relating to a specific subject image showing a specific subject in an image obtained by imaging with a specific subject as the focus target area when imaging is performed by an image sensor with a focus operation on a specific subject as the focus target area, as data to be used for machine learning.
[0025] A 21st aspect of the technology of this disclosure is a program that causes a computer to perform a process that includes outputting specific subject data relating to a specific subject image showing a specific subject in an image obtained by imaging with a specific subject as the focus area, as data to be used for machine learning, when imaging is performed by an image sensor with a focus operation that sets a specific subject as the focus area. [Brief explanation of the drawing]
[0026] [Figure 1] This is a schematic diagram showing an example of a training data generation system. [Figure 2] This is a perspective view showing an example of the front view of an imaging device. [Figure 3] This is a rear view showing an example of the external appearance of the rear side of the imaging device. [Figure 4] This is a block diagram of the imaging device. [Figure 5]This is a rear view of an imaging device showing an example of how the label selection screen is displayed on the monitor when the training data acquisition mode is selected. [Figure 6] This is a rear view of an imaging device showing an example of a configuration in which the AF frame is superimposed on the live view image displayed on the monitor. [Figure 7] This is a rear view of an imaging device showing an example of how the position of the AF frame is changed according to the position of the subject's face. [Figure 8] This is a rear view of an imaging device showing an example of how to change the size of the AF frame to match the position of the subject's face. [Figure 9] This is an explanatory diagram showing an example of the position coordinates of the AF frame. [Figure 10] This is an explanatory diagram showing an example of how training data output from the information processing device according to the first embodiment is stored in a database. [Figure 11] This flowchart shows an example of the flow of the training data generation process performed by the information processing device according to the first embodiment. [Figure 12] This is a rear view of an imaging device showing an example of how the position and size of the AF frame are changed to match the position of the subject's left eye. [Figure 13] This is an explanatory diagram illustrating an example of how the information processing device according to the second embodiment extracts a specific subject image from the exposure image according to the distance between the focus positions of each divided region. [Figure 14] This is a schematic diagram showing an example of the arrangement of pixels included in the photoelectric conversion element of an imaging device having an information processing device according to the second embodiment. [Figure 15] Figure 14 is a conceptual diagram showing an example of the incident light characteristics of the first and second phase difference pixels included in the photoelectric conversion element. [Figure 16] This flowchart shows an example of the flow of the training data generation process performed by the information processing device according to the second embodiment. [Figure 17]This is an explanatory diagram showing an example of how the information processing device according to the third embodiment extracts a specific subject image from the exposed image according to the color difference of each divided region. [Figure 18] This flowchart shows an example of the flow of the training data generation process performed by the information processing device according to the third embodiment. [Figure 19] This flowchart shows an example of the flow of the training data generation process performed by the information processing device according to the fourth embodiment. [Figure 20] This is an explanatory diagram illustrating an example of how the information processing device according to the fifth embodiment outputs warning information to a learning device when the difference in size between the live view image and the actual exposure image exceeds a predetermined size range. [Figure 21] This is an explanatory diagram illustrating an example of how the information processing device according to the fifth embodiment outputs warning information to a learning device when the degree of difference in the center position of a specific subject image between the live view image and the actual exposure image exceeds a predetermined position range. [Figure 22A] This flowchart shows an example of the flow of the training data generation process performed by the information processing device according to the fifth embodiment. [Figure 22B] This is a continuation of the flowchart shown in Figure 22A. [Figure 23] This is an explanatory diagram showing an example of how the information processing device according to the sixth embodiment determines the position coordinates of a specific subject image. [Figure 24] This is an explanatory diagram showing an example of how training data output from the information processing device according to the sixth embodiment is stored in a database. [Figure 25] This flowchart shows an example of the flow of the training data generation process performed by the information processing device according to the sixth embodiment. [Figure 26] This is an explanatory diagram showing an example of training data for extracting and outputting a specific subject image from the main exposure image. [Figure 27]This block diagram shows an example of how a training data generation program is installed on a controller in an imaging device from a storage medium on which the training data generation program is stored. [Modes for carrying out the invention]
[0027] Hereinafter, an example of an embodiment of the imaging device and the method of operating the imaging device relating to the technology of this disclosure will be described with reference to the attached drawings.
[0028] First, let's explain the terminology used in the following explanation.
[0029] CPU stands for "Central Processing Unit". RAM stands for "Random Access Memory". NVM stands for "Non-Volatile Memory". IC stands for "Integrated Circuit". ASIC stands for "Application Specific Integrated Circuit". PLD stands for "Programmable Logic Device". FPGA stands for "Field-Programmable Gate Array". SoC stands for "System-on-a-chip". SSD stands for "Solid State Drive". USB stands for "Universal Serial Bus". HDD stands for "Hard Disk Drive". EEPROM stands for "Electrically Erasable and Programmable Read Only Memory". EL stands for "Electro-Luminescence". I / F stands for "Interface". UI stands for "User Interface." TOF stands for "Time of Flight." fps stands for "frames per second." MF stands for "Manual Focus." AF stands for "Auto Focus." In the following, for the sake of explanation, a CPU is used as an example of a "processor" related to the technology disclosed herein, but the "processor" related to the technology disclosed herein may be a combination of multiple processing units, such as a CPU and a GPU. When a combination of a CPU and a GPU is applied as an example of a "processor" related to the technology disclosed herein, the GPU operates under the control of the CPU and is responsible for executing image processing.
[0030] In this specification, “perpendicular” means not only perfect perpendicularity but also perpendicularity that includes errors generally accepted in the art to which the art of this disclosure pertains.
[0031] In the following explanation, when the term "image" is used instead of "image data" (other than "image displayed on the monitor"), "image" also includes the meaning of "data representing an image (image data)." In this specification, "subject in image" means a subject that is included as an image within the image.
[0032] [First Embodiment] As an example, as shown in Figure 1, the training data generation system 10 includes an imaging device 12, a learning device 14, and a database 16 connected to the learning device 14.
[0033] The imaging device 12 is, for example, a digital camera. The imaging device 12 is connected to the learning device 14 via a communication network such as the Internet. The imaging device 12 has two operating modes for its imaging system: a normal imaging mode and a training data imaging mode. In normal imaging mode, the imaging device 12 operates the mechanical shutter 48 (see Figure 4) to store the image formed on the light-receiving surface 24A (see Figure 4) of the image sensor 24 (hereinafter referred to as the "main exposure image") in memory. In training data imaging mode, the imaging device 12 outputs data related to images showing specific subjects within the main exposure image (hereinafter referred to as "specific subject images") to the learning device 14 as data to be used for machine learning. Hereinafter, data related to specific subject images will also be referred to as "specific subject data". Machine learning includes, for example, deep learning and convolutional neural networks.
[0034] The learning device 14 is, for example, a computer. The database 16 has storage such as an HDD or EEPROM and stores the data received by the learning device 14.
[0035] The data used in machine learning is, for example, training data used to build a model in machine learning. In this embodiment, the training data is labeled image data that includes specific subject data and labels that are information about the specific subject images. The learning device 14 constructs a classification model that classifies the class of subjects in an image by performing supervised machine learning using the training data.
[0036] In the example shown in Figure 1, the user 11 of the imaging device 12 (hereinafter simply referred to as "user 11") sets the imaging device 12 to training data imaging mode and sequentially images specific subjects A, B, and C. Before imaging specific subject A, user 11 selects the label LA indicating "face" via the reception unit 60 (see Figure 4) in the imaging device 12. The imaging device 12 associates the specific subject data related to the specific subject image SA in the main exposure image PA obtained by imaging specific subject A with the label LA and outputs it to the learning device 14 as training data 17A. The learning device 14 receives the training data 17A and stores the specific subject data related to the specific subject image SA with the label LA in the database 16.
[0037] Similarly, before imaging the specific subject B, user 11 selects a label LB indicating "car" via the reception unit 60 (see Figure 4) in the imaging device 12. The imaging device 12 associates the specific subject data related to the specific subject image SB in the main exposure image PB obtained by imaging the specific subject B with the label LB and outputs it to the learning device 14 as training data 17B. The learning device 14 receives the training data 17B and stores the specific subject data related to the specific subject image SB with the label LB in the database 16.
[0038] Furthermore, before imaging the specific subject C, user 11 selects a label LC indicating "flower" via the reception unit 60 (see Figure 4) in the imaging device 12. The imaging device 12 associates the specific subject data related to the specific subject image SC in the main exposure image PC obtained by imaging the specific subject C with the label LC and outputs it to the learning device 14 as training data 17C. The learning device 14 receives the training data 17C and stores the specific subject data related to the specific subject image SC with the label LC in the database 16.
[0039] Here, the exposure images PA, PB, and PC are examples of “imported images” relating to the technology of this disclosure. Specific subjects A, B, and C are examples of “specific subjects” relating to the technology of this disclosure. Specific subject images SA, SB, and SC are examples of “specific subject images” relating to the technology of this disclosure. Specific subject data is an example of “specific subject data” relating to the technology of this disclosure. In the following description, when it is not necessary to distinguish between the exposure images PA, PB, and PC, they will be collectively referred to as “exposure image P.” Also, in the following description, when it is not necessary to distinguish between specific subjects A, B, and C, they will be referred to as “specific subjects” without any designation. Also, in the following description, when it is not necessary to distinguish between the specific subject images SA, SB, and SC, they will be collectively referred to as “specific subject image S.”
[0040] Labels LA, LB, and LC are examples of “labels” relating to the technology of this disclosure. Training data 17A, 17B, and 17C are examples of “data” and “training data” relating to the technology of this disclosure. In the following description, when it is not necessary to distinguish between labels LA, LB, and LC, they will be collectively referred to as “label L.” Similarly, in the following description, when it is not necessary to distinguish between training data 17A, 17B, and 17C, they will be collectively referred to as “training data 17.”
[0041] As an example, as shown in Figure 2, the imaging device 12 is a digital camera with interchangeable lenses and an omitted reflex mirror. The imaging device 12 comprises an imaging device body 20 and an interchangeable lens 22 that is interchangeably attached to the imaging device body 20. Here, an example of the imaging device 12 is given as a digital camera with interchangeable lenses and an omitted reflex mirror, but the technology of this disclosure is not limited to this, and it may also be a digital camera with a fixed lens, a digital camera that does not omit a reflex mirror, or a digital camera built into various electronic devices such as smart devices, wearable terminals, cell observation devices, ophthalmic observation devices, or surgical microscopes.
[0042] The imaging device body 20 is equipped with an image sensor 24. The image sensor 24 includes a photoelectric conversion element 80 (see Figure 14). The image sensor 24 has a light-receiving surface 24A (see Figure 14). The image sensor 24 is positioned within the imaging device body 20 such that the center of the light-receiving surface 24A coincides with the optical axis OA.
[0043] The image sensor 24 is a CMOS image sensor. When the interchangeable lens 22 is attached to the imaging device body 20, subject light representing the subject passes through the interchangeable lens 22 and is imaged onto the image sensor 24, and image data representing the image of the subject is generated by the image sensor 24. Here, the image sensor 24 is an example of an "image sensor" related to the technology of this disclosure.
[0044] In this embodiment, a CMOS image sensor is used as an example for the image sensor 24. However, the technology of this disclosure is not limited to this, and the technology of this disclosure can be applied even if the image sensor 24 is another type of image sensor, such as a CCD image sensor.
[0045] The top surface of the imaging device body 20 is provided with a release button 26 and a dial 28. The dial 28 is operated when setting the operating mode of the imaging device 12. The operating modes of the imaging device 12 include an imaging system operating mode that includes a normal imaging mode and a training data imaging mode, and a playback system operating mode that includes a playback mode.
[0046] The release button 26 functions as both an imaging preparation instruction unit and an imaging instruction unit, and can detect two stages of pressing: an imaging preparation instruction state and an imaging instruction state. The imaging preparation instruction state refers to the state in which the button is pressed from the standby position to an intermediate position (half-press position), for example, while the imaging instruction state refers to the state in which the button is pressed beyond the intermediate position to the final pressed position (full-press position). Hereafter, the state in which the button is pressed from the standby position to the half-press position will be referred to as the "half-press state," and the state in which the button is pressed from the standby position to the full-press position will be referred to as the "full-press state." Furthermore, hereafter, the operation in which the release button 26 is pressed to the final pressed position (full-press position) will also be referred to as the "main exposure operation." The "main exposure operation" may be performed by another method, such as touching the touch panel monitor 3 described later.
[0047] As an example, as shown in Figure 3, a touch panel monitor 30 and instruction keys 32 are provided on the back of the imaging device body 20.
[0048] The touch panel monitor 30 comprises a monitor 34 and a touch panel 36 (see also Figure 4). An example of the monitor 34 is an organic EL display. The monitor 34 may be other types of displays, such as an inorganic EL display or a liquid crystal display, instead of an organic EL display. Note that the monitor 34 is an example of a "monitor" related to the technology of this disclosure.
[0049] Monitor 34 displays images and / or text information. Monitor 34 is used to display live view images obtained by continuous imaging when the imaging device 12 is in the imaging system's operating mode. Live view image imaging (hereinafter also referred to as "live view image imaging") is performed according to a frame rate of, for example, 60fps. However, the frame rate for live view image imaging is not limited to 60fps; it may be higher or lower than 60fps.
[0050] Here, "live view image" refers to a displayable moving image based on image data obtained by capturing data with the image sensor 24. Here, the live view image is an example of a "displayable moving image" related to the technology of this disclosure. The live view image is also generally referred to as a through image. The monitor 34 is also used to display the exposed image P. Furthermore, the monitor 34 is also used to display the playback image when the imaging device 12 is in playback mode, and to display menu screens, etc.
[0051] The touch panel 36 is a transparent touch panel and is superimposed on the surface of the display area of the monitor 34. The touch panel 36 receives instructions from the user 11 by detecting contact with an object such as a finger or stylus pen.
[0052] In this embodiment, an out-cell type touch panel display in which the touch panel 36 is superimposed on the surface of the display area of the monitor 34 is given as an example of the touch panel monitor 30, but this is merely one example. For example, an on-cell type or in-cell type touch panel display can also be used as the touch panel monitor 30.
[0053] The instruction key 32 accepts various instructions. Here, "various instructions" refers to, for example, instructions to display a menu screen where various menus can be selected, instructions to select one or more menus, instructions to confirm the selection, instructions to clear the selection, instructions to zoom in, zoom out, and various other instructions such as frame-by-frame playback. These instructions may also be given via the touch panel 36.
[0054] As an example, as shown in Figure 4, the imaging device 12 is equipped with mounts 37 and 38. Mount 37 is provided on the imaging device body 20. Mount 38 is provided on the interchangeable lens 22 at a position opposite to that of mount 37. The interchangeable lens 22 is interchangeably mounted on the imaging device body 20 by coupling mount 38 to mount 37.
[0055] As an example, as shown in Figure 4, the imaging lens 40 comprises an objective lens 40A, a focusing lens 40B, and an aperture 40C. The objective lens 40A, focusing lens 40B, and aperture 40C are arranged in that order along the optical axis OA from the subject side (object side) to the imaging device body 20 side (image side).
[0056] The interchangeable lens 22 also includes a slide mechanism 42, a motor 44, and a motor 46. The focus lens 40B is mounted on the slide mechanism 42 so as to be slidable along the optical axis OA. The motor 44 is connected to the slide mechanism 42, and the slide mechanism 42 operates by receiving power from the motor 44, moving the focus lens 40B along the optical axis OA.
[0057] Aperture 40C is a variable aperture. A motor 46 is connected to aperture 40C, and aperture 40C adjusts the exposure by operating under the power of motor 46. The components and / or operating method of interchangeable lens 22 can be changed as needed.
[0058] Motors 44 and 46 are connected to the imaging device body 20 via the mount 38, and their drive is controlled according to commands from the imaging device body 20. In this embodiment, stepping motors are used as an example of motors 44 and 46. Therefore, motors 44 and 46 operate in synchronization with pulse signals according to commands from the imaging device body 20. In the example shown in Figure 4, motors 44 and 46 are provided on the interchangeable lens 22, but this is not limited to this, and either one of motors 44 or 46 may be provided on the imaging device body 20, or both motors 44 and 46 may be provided on the imaging device body 20.
[0059] In normal imaging mode, the imaging device 12 is selectively set to either MF mode or AF mode according to instructions given to the imaging device body 20. MF mode is an operation mode in which the focus is adjusted manually. In MF mode, for example, the user 11 operates the focus ring (not shown) of the interchangeable lens 22, causing the focus lens 40B to move along the optical axis OA by an amount corresponding to the amount the focus ring is operated, thereby adjusting the focus.
[0060] In AF mode, when the release button 26 is half-pressed, the imaging device body 20 calculates the focus position according to the subject distance and adjusts the focus by moving the focus lens 40B toward the calculated focus position. Subsequently, when the release button 26 is fully pressed, the imaging device body 20 performs the main exposure operation (described later). Here, the focus position refers to the position of the focus lens 40B on the optical axis OA when in focus.
[0061] In training data imaging mode, the imaging device 12 is set to AF mode. For the sake of explanation, the control that adjusts the focus lens 40B to the focus position will also be referred to as "AF control." Also, for the sake of explanation, the calculation of the focus position will also be referred to as "AF calculation."
[0062] The imaging device body 20 is equipped with a mechanical shutter 48. The mechanical shutter 48 is a focal-plane shutter and is positioned between the aperture 40C and the light-receiving surface 24A. The mechanical shutter 48 operates by receiving power from a drive source such as a motor (not shown). The mechanical shutter 48 has a light-shielding mechanism (not shown) that blocks subject light that passes through the imaging lens 40 and is imaged on the light-receiving surface 24A of the image sensor 24. The imaging device 12 performs the main exposure operation in accordance with the timing of when the mechanical shutter 48 opens and closes the light-shielding mechanism. The main exposure operation refers to the operation of acquiring image data of the image (main exposure image P) imaged on the light-receiving surface 24A and storing it in memory. Note that the main exposure operation is an example of "imaging" according to the technology of this disclosure.
[0063] The imaging device body 20 includes a controller 50 and a UI device 52. The controller 50 controls the entire imaging device 12. The UI device 52 is a device that presents information to the user 11 and receives instructions from the user 11. The UI device 52 is connected to the controller 50 via a bus line 58, and the controller 50 acquires various information from the UI device 52 and controls the UI device 52. Note that the controller 50 is an example of an "information processing device" related to the technology of this disclosure.
[0064] The controller 50 includes a CPU 50A, an NVM 50B, a RAM 50C, a control I / F 50D, and an input I / F 50E. The CPU 50A, NVM 50B, RAM 50C, control I / F 50D, and input I / F 50E are interconnected via a bus line 58.
[0065] CPU 50A is an example of a "processor" related to the technology of this disclosure. CPU 50A controls the entire imaging device 12. NVM 50B is an example of a "memory" related to the technology of this disclosure. An example of NVM 50B is EEPROM. However, EEPROM is merely an example, and for example, ferroelectric memory may be used instead of EEPROM, or any non-volatile memory that can be mounted on the imaging device 12 may be used. RAM 50C is a volatile memory used as a work area during the execution of various programs.
[0066] Various programs 51 are stored in the NVM 50B. The CPU 50A reads the necessary programs 51 from the NVM 50B and executes the read programs 51 on the RAM 50C, thereby comprehensively controlling the imaging device 12.
[0067] The control interface 50D is a device with an FPGA and is connected to the image sensor 24. The CPU 50A controls the image sensor 24 via the control interface 50D. The control interface 50D is also connected to motors 44 and 46 via mounts 37 and 38, and the CPU 50A controls motors 44 and 46 via the control interface 50D.
[0068] The input interface 50E is connected to the image sensor 24. The input interface 50E receives image data output from the image sensor 24. The controller 50 generates exposure image data representing the exposure image P by applying known signal processing to the image data, such as white balance adjustment, sharpness adjustment, gamma correction, color space conversion processing, and color difference correction.
[0069] An external interface (I / F) 54 is connected to bus line 58. The external interface 54 is a device with an FPGA. An external device (not shown), such as a USB memory or memory card, is connected to the external interface 54. The external interface 54 is responsible for the exchange of various information between the CPU 50A and the external device. The CPU 50A stores the exposure image data in the external device via the external interface 54.
[0070] Furthermore, a communication interface 56 is connected to the bus line 58. The communication interface 56 is connected to the learning device 14 via a communication network such as the Internet. In training data imaging mode, the CPU 50A outputs training data 17 to the learning device 14 via the communication interface 56.
[0071] The UI device 52 includes a touch panel monitor 30 and a reception unit 60. The monitor 34 and the touch panel 36 are connected to a bus line 58. Therefore, the CPU 50A displays various information on the monitor 34 and operates according to various instructions received by the touch panel 36.
[0072] The reception unit 60 includes a touch panel 36 and a hard key unit 62. The hard key unit 62 consists of multiple hard keys, including a release button 26, a dial 28, and an indicator key 32. The hard key unit 62 is connected to a bus line 58, and the CPU 50A operates according to the various instructions received by the hard key unit 62.
[0073] In the example shown in Figure 4, for illustrative purposes, a single bus is shown as bus line 58, but multiple buses are also possible. Bus line 58 may be a serial bus, or a parallel bus including a data bus, address bus, and control bus, etc.
[0074] The various programs 51 stored in the NVM 50B include a training data generation program 51A. When the imaging device 12 is set to training data imaging mode, the CPU 50A reads the training data generation program 51A from the NVM 50B and executes the read training data generation program 51A on the RAM 50C, thereby operating as a training data generation unit 53. The training data generation unit 53 performs training data generation processing. The training data generation processing performed by the training data generation unit 53 is described in detail below.
[0075] As an example, as shown in Figure 5, in training data acquisition mode, the training data generation unit 53 displays a label selection screen 64 on the touch panel monitor 30. The label selection screen 64 displays a message 64A that says "Please select a label to assign to the subject" and a table 64B listing multiple label candidates.
[0076] The first column of Table 64B displays label candidates that represent relatively broad attributes (hereinafter also referred to as "major label candidates"). Major label candidates include, for example, "person," "vehicle," and "building." The other columns of Table 64B display label candidates that represent attributes that are subdivisions of the major label candidates in the first column (hereinafter also referred to as "minor label candidates"). For example, if the major label candidate is "person," minor label candidates include "face," "male," "female," and "child." User 11 selects any label candidate from Table 64B by touching the touch panel 36 with an indicator.
[0077] When imaging a specific subject A shown in Figure 1, as an example, as shown in Figure 5, the user 11 selects the label "face" from the label candidates listed in Table 64B via the touch panel monitor 30. Note that the label candidates listed in Figure 5 are just examples and are not limited to these. Also, the method of displaying the label candidates is not limited to these. In the example shown in Figure 5, one small label candidate is selected, but a large label candidate may be selected, or multiple small label candidates may be selected.
[0078] The training data generation unit 53 receives the selected label L. The training data generation unit 53 stores the received label L in the RAM 50C.
[0079] As an example, as shown in Figure 6, after receiving label L, the training data generation unit 53 displays a live view image 66 on the monitor 34 based on the imaging signal output from the image sensor 24. In training data imaging mode, the training data generation unit 53 also superimposes an AF frame 68 on the center of the monitor 34 where the live view image 66 is displayed. The AF frame 68 is a frame used in AF mode to display the area to be focused on (hereinafter referred to as the "focus target area") on the live view image 66 in a way that distinguishes it from other image areas. The AF frame 68 is an example of a "frame" related to the technology of this disclosure. The focus target area is also an example of a "focus target area" related to the technology of this disclosure. The imaging signal is also an example of a "signal" related to the technology of this disclosure.
[0080] The AF frame 68 includes a rectangular frame line 68A and four triangular arrows 68B-U, 68B-D, 68B-R, and 68B-L positioned on all four sides of the frame line 68A. Hereafter, unless it is necessary to distinguish between the triangular arrows 68B-U, 68B-D, 68B-R, and 68B-L, they will be collectively referred to as "triangular arrow 68B".
[0081] User 11 can give a position change instruction to the training data generation unit 53 by touching the triangular arrows 68B on the touch panel 36 with an indicator, thereby moving the position of the AF frame 68 in the direction indicated by each triangular arrow 68B. The training data generation unit 53 changes the position of the AF frame 68 on the monitor 34 according to the given position change instruction. Here, the position change instruction is an example of a "position change instruction" related to the technology of this disclosure. Note that the triangular arrows 68B displayed on the touch panel 36 are merely one example of a means for receiving position change instructions from User 11, and the means are not limited as long as it is possible to receive position change instructions from User 11 via the reception unit 60.
[0082] For example, in Figure 6, user 11 touches the triangular arrows 68B-U and 68B-L on the touch panel 36 with an indicator, thereby giving a position change instruction to the training data generation unit 53 to move the AF frame 68 so that the frame line 68A surrounds the area indicating the face of a specific subject A. As a result, the AF frame 68 moves to the position shown in Figure 7, for example.
[0083] Furthermore, the user 11 can give a resize instruction to the training data generation unit 53 to change the size of the frame line 68A by performing a pinch-in or pinch-out operation on the frame line 68A displayed on the touch panel monitor 30. As an example, as shown in Figure 8, if the zoom magnification of the imaging lens 40 is lower than in the example shown in Figure 7, the user 11 gives a resize instruction to the training data generation unit 53 to reduce the size of the frame line 68A so that it surrounds the area showing the face of a specific subject A. The training data generation unit 53 changes the size of the frame line 68A on the monitor 34 according to the given resize instruction. Note that the resize instruction is just one example of a "resize instruction" related to the technology of this disclosure. Note that the pinch-in and pinch-out operation is merely one example of a means for receiving a resize instruction from the user 11, and the means are not limited as long as a position change instruction from the user 11 can be received via the reception unit 60.
[0084] User 11 changes the position and size of the AF frame 68 and then performs an AF operation by pressing the release button 26 to the half-press position. Here, the AF operation is an example of a "focus operation" related to the technology of this disclosure. When an AF operation is performed, the training data generation unit 53 designates the area surrounded by the frame line 68A in the live view image 66 as the focus target area F.
[0085] The training data generation unit 53 acquires position coordinates indicating the position of the focus target area F. As an example, as shown in Figure 9, the position coordinates of the focus target area F are the lower right corner Q of the frame line 68A, with the lower left corner of the live view image 66 as the origin O(0,0). 1A Coordinates (X 1A ,Y 1A ) and the upper left corner Q of frame line 68A 2A Coordinates (X 2A ,Y 2A ) is expressed as follows. The training data generation unit 53 stores the acquired position coordinates of the focus target area F in the RAM 50C. Note that the position coordinates are an example of "coordinates" related to the technology of this disclosure.
[0086] When user 11 performs an AF operation and then presses the release button 26 to its fully depressed position, the imaging device 12 performs the main exposure operation, and the training data generation unit 53 extracts an image indicating the focus target area F as a specific subject image SA from the main exposure image PA. As an example, as shown in Figure 10, the training data generation unit 53 associates specific subject data and label LA related to the specific subject image SA and outputs it to the learning device 14 as training data 17A. The specific subject data related to the specific subject image SA includes the main exposure image PA and the position coordinates indicating the position of the specific subject image SA within the main exposure image PA, i.e., the position coordinates of the focus target area F.
[0087] Similarly, in training data imaging mode, when user 11 moves the AF frame 68 to surround a specific subject B and then instructs the imaging device 12 to perform AF operation and main exposure operation, the training data generation unit 53 extracts an image indicating the focus target area F as the specific subject image SB from the main exposure image PB. The training data generation unit 53 associates the specific subject data and label LB related to the specific subject image SB and outputs it to the learning device 14 as training data 17B. The specific subject data related to the specific subject image SB includes the main exposure image PB and position coordinates indicating the position of the specific subject image SB within the main exposure image PB.
[0088] Similarly, in training data imaging mode, when user 11 moves the AF frame 68 to surround a specific subject C and then instructs the imaging device 12 to perform AF operation and main exposure operation, the training data generation unit 53 extracts an image indicating the focus target area F as the specific subject image SC from the main exposure image PC. The training data generation unit 53 associates the specific subject data and label LC related to the specific subject image SC and outputs it to the learning device 14 as training data 17C. The specific subject data related to the specific subject image SC includes the main exposure image PC and position coordinates indicating the position of the specific subject image SC within the main exposure image PC.
[0089] The learning device 14 includes a computer 15 and an input / output interface 14D. The input / output interface 14D is connected to the communication interface 56 of the imaging device 12 in a communicative manner. The input / output interface 14D receives training data 17 from the imaging device 12. The computer 15 stores the training data 17 received by the input / output interface 14D in the database 16. The computer 15 also reads the training data 17 from the database 16 and performs machine learning using the read training data 17.
[0090] Computer 15 is equipped with a CPU 14A, an NVM 14B, and RAM 14C. The CPU 14A controls the entire learning device 14. An example of NVM 14B is an EEPROM. However, EEPROM is merely an example; for example, a ferroelectric memory could be used instead of EEPROM, or any non-volatile memory that can be installed in the learning device 14. RAM 14C is a volatile memory used as a work area during the execution of various programs.
[0091] The NVM14B stores the learning execution program 72. The CPU14A reads the learning execution program 72 from the NVM14B and executes the read learning execution program 72 on the RAM14C, thereby operating as a learning execution unit 76. The learning execution unit 76 constructs a supervised learning model by training the neural network 74 using the training data 17 according to the learning execution program 72.
[0092] Next, the operation of the imaging device 12 according to this first embodiment will be explained with reference to Figure 11. Figure 11 shows an example of the flow of the training data generation process performed by the training data generation unit 53. The training data generation process is realized when the CPU 50A executes the training data generation program 51A. The training data generation process is started when the imaging device 12 is set to training data imaging mode.
[0093] In the training data generation process shown in Figure 11, first, in step ST101, the training data generation unit 53 displays a label selection screen 64 on the touch panel monitor 30, for example, as shown in Figure 5. After this, the training data generation process proceeds to step ST102.
[0094] In step ST102, the training data generation unit 53 determines whether label L is selected on the touch panel monitor 30. If label L is selected in step ST102, the determination is affirmed, and the training data generation process proceeds to step ST103. If label L is not selected in step ST102, the determination is denied, and the training data generation process proceeds to step ST101.
[0095] In step ST103, the training data generation unit 53 displays the live view image 66 on the touch panel monitor 30. After this, the training data generation process proceeds to step ST104.
[0096] In step ST104, the training data generation unit 53 superimposes the AF frame 68 onto the live view image 66 displayed on the touch panel monitor 30. After this, the training data generation process proceeds to step ST105.
[0097] In step ST105, the training data generation unit 53 changes the position and size of the AF frame 68 according to the position change and size change instructions from the user 11. The user 11 gives position change and size change instructions via the reception unit 60 so that the area indicating a specific subject in the live view image 66 is surrounded by the frame line 68A of the AF frame 68. After this, the training data generation process proceeds to step ST106.
[0098] In step ST106, the training data generation unit 53 determines whether or not an AF operation has been performed. If an AF operation has been performed in step ST106, the determination is affirmed, and the training data generation process proceeds to step ST107. If an AF operation has not been performed in step ST106, the determination is denied, and the training data generation process proceeds to step ST105.
[0099] In step ST107, the training data generation unit 53 obtains the position coordinates of the focus target area F indicated by the AF frame 68. After this, the training data generation process proceeds to step ST108.
[0100] In step ST108, the training data generation unit 53 determines whether or not the main exposure has been performed. If the main exposure has been performed in step ST108, the determination is affirmed, and the training data generation process proceeds to step ST109. If the main exposure has not been performed in step ST108, the determination is denied, and the training data generation process proceeds to step ST106.
[0101] In step ST109, the training data generation unit 53 acquires the main exposure image P. After this, the training data generation process proceeds to step ST110.
[0102] In step ST110, the training data generation unit 53 extracts an image from the exposure image P that indicates the focus target area F as a specific subject image S. After this, the training data generation process proceeds to step ST111.
[0103] In step ST111, the training data generation unit 53 outputs the specific subject data and label L to the learning device 14. The specific subject data includes the main exposure image P and the position coordinates of the specific subject image S, i.e., the position coordinates of the focus target area F. The learning device 14 stores the received specific subject data and label L as training data 17 in the database 16. This completes the training data generation process.
[0104] As described above, in this first embodiment, when the exposure operation, which involves focusing on a specific subject as the focus target area, is performed by the image sensor 24, the training data generation unit 53 outputs specific subject data relating to the specific subject image S in the exposed image P obtained by the exposure operation as training data 17 to be used for machine learning. Therefore, with this configuration, the training data 17 to be used for machine learning can be collected more easily than when the specific subject image S is manually extracted from the exposed image P obtained by imaging with the image sensor 24.
[0105] Furthermore, in this first embodiment, the machine learning is supervised machine learning. The training data generation unit 53 assigns a label L, which is information about a specific subject image S, to the specific subject data, and outputs the specific subject data as training data 17 to be used for supervised machine learning. Therefore, with this configuration, the training data 17 necessary for supervised machine learning can be collected.
[0106] Furthermore, in this first embodiment, the training data generation unit 53 displays a live view image 66 on the monitor 34 based on the imaging signal output from the image sensor 24. The training data generation unit 53 displays the focus target area F in the live view image 66 in a manner that can be distinguished from other image areas using the AF frame 68. The specific subject image S is an image corresponding to the position of the focus target area F in the exposure image P. Therefore, with this configuration, the specific subject image S can be extracted more easily than when the specific subject image S is unrelated to the position of the focus target area F.
[0107] Furthermore, in this first embodiment, the training data generation unit 53 displays an AF frame 68 surrounding the focus target area F on the live view image 66, thereby displaying the focus target area F in a manner that makes it distinguishable from other image areas. Therefore, with this configuration, the user 11 can more easily recognize the specific subject image S compared to when the AF frame 68 is not displayed.
[0108] Furthermore, in this first embodiment, the position of the AF frame 68 can be changed according to a given position change instruction. Therefore, with this configuration, the user 11 can move the focus target area F more freely than when the position of the AF frame 68 is fixed.
[0109] Furthermore, in this first embodiment, the size of the AF frame 68 can be changed according to a given resizing instruction. Therefore, with this configuration, compared to the case where the size of the AF frame 68 is fixed, the user 11 can freely change the size of the focus target area F.
[0110] Furthermore, in this first embodiment, the specific subject data includes the position coordinates of the specific subject image S. The training data generation unit 53 outputs the exposure image P and the position coordinates of the focus target area F, i.e., the position coordinates of the specific subject image S, as training data 17 to be used for machine learning. Therefore, this configuration has the advantage of requiring fewer processing steps compared to the case where the specific subject image S is extracted and output.
[0111] Furthermore, in this first embodiment, the learning device 14 includes an input / output interface 14D that receives specific subject data output from the controller 50 of the imaging device 12, and a computer 15 that performs machine learning using the specific subject data received by the input / output interface 14D. The imaging device 12 includes a controller 50 and an image sensor 24. Accordingly, with this configuration, the learning device 14 can easily collect training data 17 for learning compared to the case where the specific subject image S to be used for learning is manually selected from the exposure image P obtained by imaging with the image sensor 24.
[0112] In the first embodiment described above, as an example, as shown in Figure 1, one user 11 acquires training data 17A, 17B, and 17C by imaging multiple specific subjects A, B, and C using the same imaging device 12. However, the technology of this disclosure is not limited to this. Multiple users may each image different subjects using different imaging devices 12, and the training data 17 may be output from the multiple imaging devices 12 to the same learning device 14. In this case, since the training data 17 acquired by multiple users is output to the same learning device 14, the learning device 14 can efficiently collect the training data 17.
[0113] Furthermore, in the first embodiment described above, the training data generation unit 53 uses the lower right corner Q of the frame line 68A as the position coordinate of the specific subject image S. 1A and upper left corner Q 2A The coordinates of the frame line 68A are output, but the technology of this disclosure is not limited thereto. The training data generation unit 53 may output the coordinates of the upper right corner and the lower left corner of the frame line 68A. Alternatively, the training data generation unit 53 may output the coordinates of one corner of the frame line 68A and the lengths of the vertical and horizontal sides that constitute the frame line 68A. Alternatively, the training data generation unit 53 may output the coordinates of the center of the frame line 68A and the lengths from the center to the vertical and horizontal sides. Furthermore, although the position coordinates of the specific subject image S are expressed as coordinates with the lower left corner of the live view image 66 as the origin, the technology of this disclosure is not limited thereto, and other corners of the live view image 66 may be used as the origin, or the center of the live view image 66 may be used as the origin.
[0114] [Second Embodiment] This second embodiment differs from the first embodiment in that the focus target area F, which is designated by being surrounded by the AF frame 68, is not extracted as a specific subject image S. The differences from the first embodiment will be described in detail below. In the following description, the same reference numerals are used for components and functions similar to those in the first embodiment, and their descriptions are omitted.
[0115] As an example, as shown in Figure 12, in this second embodiment, the touch panel monitor 30 displays a live view image 66 based on the imaging signal output from the image sensor 24, and an AF frame 68 is superimposed on the live view image 66. In the example shown in Figure 12, the training data generation unit 53 receives position change instructions and size change instructions from the user 11 via the reception unit 60 in the live view image 66, and places the AF frame 68 on the image showing the left eye of the specific subject A. After this, the AF operation is performed, and the area of the left eye of the specific subject A surrounded by the frame line 68A is designated as the focus target area F. The training data generation unit 53 receives the designation of the focus target area F in the live view image 66. After this, the imaging device 12 performs the main exposure operation, and the training data generation unit 53 acquires the main exposure image P that is in focus on the focus target area F.
[0116] As an example, as shown in Figure 13, in the exposure image P obtained by imaging, the training data generation unit 53 sets a candidate region 78 that includes the focus target region F. The candidate region 78 is a region that is a candidate for extracting a specific subject image S. Note that the candidate region 78 is an example of a "predetermined region" related to the technology of this disclosure.
[0117] Candidate region 78 is divided, for example, into a 9x9 matrix. In the following, for the sake of clarity, each divided region is assigned a reference numeral according to its position, as shown in Figure 13, in order to distinguish them. For example, the divided region located in the first row and first column of candidate region 78 is assigned the reference numeral D11, and the divided region located in the second row and first column of candidate region 78 is assigned the reference numeral D21. Furthermore, when it is not necessary to distinguish between divided regions, they are collectively referred to as "divided region D". Note that divided region D is an example of a "divided region" relating to the technology of this disclosure.
[0118] The divided region D55, located at the center of candidate region 78, coincides with the focus target region F. In other words, the position and size of the focus target region F are specified in units of divided region D.
[0119] As an example, as shown in Figure 14, the image sensor 24 includes a photoelectric conversion element 80. The photoelectric conversion element 80 has a plurality of photosensitive pixels arranged in a matrix, and the light-receiving surface 24A is formed by these photosensitive pixels. The photosensitive pixels are pixels that have a photodiode PD, which converts the received light into electrical signals and outputs an electrical signal corresponding to the amount of light received. Image data for each divided region D is generated based on the electrical signals output from the plurality of photodiodes PD.
[0120] A color filter is placed inside the photodiode PD. The color filter includes a G filter corresponding to the G (green) wavelength range, which contributes most to obtaining the luminance signal, an R filter corresponding to the R (red) wavelength range, and a B filter corresponding to the B (blue) wavelength range.
[0121] The photoelectric conversion element 80 has two types of photosensitive pixels: phase difference pixels 84 and non-phase difference pixels 86, which are different from the phase difference pixels 84. Generally, the non-phase difference pixels 86 are also called normal pixels. The photoelectric conversion element 80 has three types of photosensitive pixels as non-phase difference pixels 86: R pixels, G pixels, and B pixels. The R pixels, G pixels, B pixels, and phase difference pixels 84 are regularly arranged with a predetermined periodicity in both the row direction (for example, the horizontal direction when the bottom surface of the imaging device body 20 is in contact with a horizontal surface) and the column direction (for example, the vertical direction which is perpendicular to the horizontal direction). The R pixels are pixels corresponding to photodiode PDs on which R filters are placed, the G pixels and phase difference pixels 84 are pixels corresponding to photodiode PDs on which G filters are placed, and the B pixels are pixels corresponding to photodiode PDs on which B filters are placed.
[0122] Multiple phase-difference pixel lines 82A and multiple non-phase-difference pixel lines 82B are arranged on the light-receiving surface 24A. Phase-difference pixel lines 82A are horizontal lines that include phase-difference pixels 84. Specifically, phase-difference pixel lines 82A are horizontal lines in which phase-difference pixels 84 and non-phase-difference pixels 86 are mixed. Non-phase-difference pixel lines 82B are horizontal lines that include only multiple non-phase-difference pixels 86.
[0123] On the light-receiving surface 24A, phase-difference pixel lines 82A and a predetermined number of non-phase-difference pixel lines 82B are arranged alternately along the column direction. The "determined number of lines" here refers to, for example, 2 lines. Although 2 lines are used as an example of the predetermined number of lines here, the technology of this disclosure is not limited to this, and the predetermined number of lines may be 3 or more lines, or it may be tens of lines, tens of lines, or hundreds of lines, etc.
[0124] The phase difference pixel line 82A is arranged in the column direction with two rows skipped from the first row to the last row. Some of the pixels in the phase difference pixel line 82A are phase difference pixels 84. Specifically, the phase difference pixel line 82A is a horizontal line in which phase difference pixels 84 and non-phase difference pixels 86 are arranged periodically.
[0125] The phase difference pixels 84 are broadly divided into first phase difference pixels 84-L and second phase difference pixels 84-R. In the phase difference pixel line 82A, the first phase difference pixels 84-L and the second phase difference pixels 84-R are arranged alternately as G pixels at intervals of several pixels in the line direction.
[0126] The first phase difference pixels 84-L and the second phase difference pixels 84-R are arranged to appear alternately in the column direction. In the example shown in Figure 14, in the fourth column, the first phase difference pixels 84-L, the second phase difference pixels 84-R, the first phase difference pixels 84-L, and the second phase difference pixels 84-R are arranged in that order from the first row along the column direction. That is, the first phase difference pixels 84-L and the second phase difference pixels 84-R are arranged alternately from the first row along the column direction. Also in the example shown in Figure 14, in the tenth column, the second phase difference pixels 84-R, the first phase difference pixels 84-L, the second phase difference pixels 84-R, and the first phase difference pixels 84-L are arranged in that order from the first row along the column direction. That is, the second phase difference pixels 84-R and the first phase difference pixels 84-L are arranged alternately from the first row along the column direction.
[0127] As an example, as shown in Figure 15, the first phase difference pixel 84-L comprises a light-shielding member 88-L, a microlens 90, and a photodiode PD. In the first phase difference pixel 84-L, the light-shielding member 88-L is positioned between the microlens 90 and the light-receiving surface of the photodiode PD. The left half in the row direction of the light-receiving surface of the photodiode PD (the left side when viewing the subject from the light-receiving surface, or in other words, the right side when viewing the light-receiving surface from the subject) is shielded by the light-shielding member 88-L.
[0128] The second phase-difference pixel 84-R comprises a light-shielding member 88-R, a microlens 90, and a photodiode PD. In the second phase-difference pixel 84-R, the light-shielding member 88-R is positioned between the microlens 90 and the light-receiving surface of the photodiode PD. The right half of the photodiode PD's light-receiving surface in the row direction (the right side when viewing the subject from the light-receiving surface, or in other words, the left side when viewing the light-receiving surface from the subject) is shielded by the light-shielding member 88-R. For the sake of explanation, in the following, when it is not necessary to distinguish between the light-shielding members 88-L and 88-R, they will be referred to as "light-shielding member 88".
[0129] The light beam passing through the exit pupil of the imaging lens 40 is broadly divided into left-region passing light 92L and right-region passing light 92R. Left-region passing light 92L refers to the left half of the light beam passing through the exit pupil of the imaging lens 40 when viewed from the phase difference pixel 84 side toward the subject side, and right-region passing light 92R refers to the right half of the light beam passing through the exit pupil of the imaging lens 40 when viewed from the phase difference pixel 84 side toward the subject side. The light beam passing through the exit pupil of the imaging lens 40 is divided into left and right by the microlens 90, light-shielding member 88-L, and light-shielding member 88-R, which function as pupil division parts, and the first phase difference pixel 84-L receives the left-region passing light 92L as subject light, and the second phase difference pixel 84-R receives the right-region passing light 92R as subject light. As a result, the photoelectric conversion element 80 generates a first phase difference image data corresponding to the subject image corresponding to the light passing through the left region 92L, and a second phase difference image data corresponding to the subject image corresponding to the light passing through the right region 92R.
[0130] The training data generation unit 53 acquires first phase difference image data for one line from the first phase difference pixels 84-L located on the same phase difference pixel line 82A, and second phase difference image data for one line from the second phase difference pixels 84-R located on the same phase difference pixel line 82A, among the phase difference pixels 84 that image the focus target area F. The training data generation unit 53 measures the distance to the focus target area F based on the amount of shift α between the first phase difference image data for one line and the second phase difference image data for one line. Note that the method for deriving the distance to the focus target area F from the amount of shift α is a known technique, so a detailed explanation is omitted here.
[0131] The training data generation unit 53 derives the focus position of the focus lens 40B by performing AF calculations based on the measured distance to the focus target area F. Hereinafter, the focus position of the focus lens 40B derived based on the distance to the focus target area F will also be referred to as the "focus target area focus position". The training data generation unit 53 performs a focusing operation to align the focus lens 40B with the focus target area focus position.
[0132] Furthermore, for each divided region D, the training data generation unit 53 acquires first phase difference image data for one line from the first phase difference pixels 84-L located on the same phase difference pixel line 82A, and second phase difference image data for one line from the second phase difference pixels 84-R located on the same phase difference pixel line 82A. The training data generation unit 53 measures the distance to each divided region D based on the difference amount α between the first phase difference image data for one line and the second phase difference image data for one line.
[0133] The training data generation unit 53 derives the focus position of the focus lens 40B in each divided region D by performing AF calculations based on the measured distance to each divided region D. Hereinafter, the focus position of the focus lens 40B derived based on the distance to each divided region D will also be referred to as the "divided region focus position".
[0134] The training data generation unit 53 determines whether the distance from the focus point of the target area to the focus point of the divided area (hereinafter referred to as the "focus point distance") for each divided area D is less than a predetermined distance threshold. The training data generation unit 53 identifies divided areas D in which the focus point distance is less than the distance threshold as areas with a high degree of similarity to the target area F. Here, the distance threshold is a value derived in advance as a threshold for extracting a specific subject image S, for example, through actual equipment testing and / or computer simulation. The distance threshold may be a fixed value or a variable value that is changed according to given instructions and / or conditions (for example, imaging conditions).
[0135] The distance between focus positions is an example of a "similar evaluation value" related to the technology of this disclosure. The focus target area focus position is an example of a "focus evaluation value" related to the technology of this disclosure. The distance threshold is an example of a "first predetermined range" related to the technology of this disclosure.
[0136] In the example shown in Figure 13, the training data generation unit 53 calculates the distance between focus points for 80 of the 81 divided regions D included in the candidate region 78, excluding the focus target region F (divided region D55). The training data generation unit 53 determines whether the calculated distance between focus points is less than the distance threshold. In Figure 13, the divided regions D shown by hatching are the divided regions for which the distance between focus points is determined to be less than the distance threshold, i.e., the divided regions identified as having a high degree of similarity to the focus target region F.
[0137] The training data generation unit 53 extracts specific subject images S from the exposure image P based on the identified divided regions D. In the example shown in Figure 13, the training data generation unit 53 extracts rectangular specific subject images S in units of divided regions D so as to completely surround the identified divided regions D.
[0138] Next, the operation of the imaging device 12 according to this second embodiment will be explained with reference to Figure 16. Figure 16 shows an example of the flow of the training data generation process according to the second embodiment.
[0139] In Figure 16, steps ST201 to ST209 are the same as steps ST101 to ST109 in Figure 11, so their explanation is omitted.
[0140] In step ST210, the training data generation unit 53 sets candidate regions 78 and divided regions D in the exposure image P. After this, the training data generation process proceeds to step ST211.
[0141] In step ST211, the training data generation unit 53 calculates the distance between the focus positions of each divided region D. After this, the training data generation process proceeds to step ST212.
[0142] In step ST212, the training data generation unit 53 identifies a divided region D in which the distance between focus positions is less than a distance threshold. After this, the training data generation process proceeds to step ST213.
[0143] In step ST213, the training data generation unit 53 extracts a specific subject image S from the exposure image P based on the identified segmented region D. The training data generation unit 53 also obtains the position coordinates of the extracted specific subject image S. After this, the training data generation process proceeds to step ST214.
[0144] In step ST214, the training data generation unit 53 outputs the specific subject data and label L to the learning device 14. The specific subject data includes the position coordinates of the main exposure image P and the specific subject image S. The learning device 14 stores the received specific subject data and label L as training data 17 in the database 16. This completes the training data generation process.
[0145] As described above, in this second embodiment, the training data generation unit 53 displays a live view image 66 based on the imaging signal output from the image sensor 24 on the touch panel monitor 30. In the live view image 66, the training data generation unit 53 receives a designation of the focus target area F from the user 11 via the reception unit 60. The training data generation unit 53 extracts a specific subject image S from the exposure image P based on the divided area D, which is among the candidate areas 78 including the focus target area F, and whose distance between focus positions, indicating the similarity to the focus target area F, is less than the distance threshold. Therefore, with this configuration, by the user 11 imaging a part of the specific subject A as the focus target area F, a specific subject image S showing the entire specific subject A is extracted from the exposure image P. This allows for the collection of training data 17 for learning with simpler operation compared to the case where the entire specific subject A must be designated as the focus target area F.
[0146] Furthermore, in this second embodiment, the training data generation unit 53 displays the focus target area F in a manner that makes it distinguishable from other image areas by displaying an AF frame 68 surrounding the focus target area F. Therefore, with this configuration, the user 11 can more easily recognize the specific subject image S compared to when the AF frame 68 is not displayed.
[0147] Furthermore, in this second embodiment, at least one of the focus target region F and the specific subject image S is defined in units of divided region D obtained by dividing the candidate region 78. Therefore, with this configuration, the processing required to extract the specific subject image S from the exposure image P becomes easier compared to the case where the candidate region 78 is not divided.
[0148] Furthermore, in this second embodiment, the distance from the focus target area focus position used in the focusing operation to the focus position of each divided area (distance between focus positions) is used as a similarity evaluation value indicating the degree of similarity to the focus target area F. Therefore, with this configuration, the training data generation unit 53 can extract a specific subject image S from the exposure image P more easily than when the focus target area focus position used in the focusing operation is not used.
[0149] In this second embodiment, as an example, the focus target region F includes one divided region D55, as shown in Figure 13. However, the focus target region F may be specified to include two or more divided regions D. Furthermore, the position and size of the candidate region 78 are not limited to the example shown in Figure 13; the candidate region 78 can be set to any position and size as long as it includes the focus target region F. Also, the number, position, and size of the divided regions D are not limited to the example shown in Figure 13 and can be changed as desired.
[0150] In the above-described second embodiment, as an example, a rectangular specific subject image S is illustrated as shown in FIG. 13. However, the technology of the present disclosure is not limited to this. The training data generation unit 53 may extract, as the specific subject image S, only the divided region D in the present exposure image P where the distance between the in-focus positions with respect to the focus target region F is less than the distance threshold, that is, the divided region D indicated by hatching in FIG. 13.
[0151] [Third Embodiment] This third embodiment is different from the second embodiment in that, as a similarity evaluation value, a color evaluation value based on the color information of the candidate region 78 is used instead of the distance between the in-focus positions. Hereinafter, the differences from the second embodiment will be described. In the following description, the same reference numerals are given to the same configurations and operations as those in the first and second embodiments, and the description thereof will be omitted.
[0152] As shown in FIG. 17 as an example, in the present exposure image P, a focus target region F, a candidate region 78, and a plurality of divided regions D are set in the same manner as in the second embodiment. The training data generation unit 53 calculates the RGB integrated value of each divided region D. The RGB integrated value is a value obtained by integrating the electrical signals for each of RGB of each divided region D. Further, the training data generation unit 53 calculates the RGB value indicating the color of each divided region D based on the RGB integrated value.
[0153] The training data generation unit 53 calculates the color difference between the focus target region F and each divided region D (hereinafter simply referred to as "color difference") with the color of the divided region D55 corresponding to the focus target region F as a reference. Note that when the RGB value of the focus target region F is (R F , G F , B F ) and the RGB value of the divided region D is (R D , G D , B D ), the color difference between the focus target region F and the divided region D is calculated using the following formula.
[0154] Color difference ={(R D - R F ) 2 +(GD -G F ) 2 +(B D -B F ) 2} 1 / 2
[0155] The training data generation unit 53 determines whether the calculated color difference for each divided region D is less than a predetermined color difference threshold. The training data generation unit 53 identifies divided regions D in which the color difference is less than the color difference threshold as regions with a high degree of similarity to the focus target region F. Here, the color difference threshold is a value derived in advance as a threshold for extracting a specific subject image S, for example, through actual equipment testing and / or computer simulation. The color difference threshold may be a fixed value or a variable value that is changed according to given instructions and / or conditions (for example, imaging conditions). Note that the RGB values are an example of "color information" relating to the technology of this disclosure. Also, the color difference is an example of "similarity evaluation value" and "color evaluation value" relating to the technology of this disclosure. Also, the color difference threshold is an example of a "first predetermined range" relating to the technology of this disclosure.
[0156] In the example shown in Figure 17, the training data generation unit 53 calculates the color difference for 80 of the 81 divided regions D included in the candidate region 78, excluding the focus target region F (divided region D55). The training data generation unit 53 determines whether the calculated color difference is less than the color difference threshold. In Figure 17, the divided regions D shown by hatching are the divided regions for which the color difference was determined to be less than the color difference threshold, i.e., the divided regions identified as having a high degree of similarity to the focus target region F.
[0157] The training data generation unit 53 extracts rectangular images of specific subjects S from the exposure image P in units of divided regions D, so as to completely surround the identified divided regions D.
[0158] Next, the operation of the imaging device 12 according to this third embodiment will be explained with reference to Figure 18. Figure 18 shows an example of the flow of the training data generation process according to the third embodiment.
[0159] In Figure 18, steps ST301 to ST309 are the same as steps ST101 to ST109 in Figure 11, so their explanation is omitted. Also, in Figure 18, step ST310 is the same as step ST210 in Figure 16, so its explanation is omitted.
[0160] In step ST311, the training data generation unit 53 calculates the color difference for each divided region D. After this, the training data generation process proceeds to step ST312.
[0161] In step ST312, the training data generation unit 53 identifies the divided region D in which the color difference is less than the color difference threshold. After this, the training data generation process proceeds to step ST313.
[0162] In step ST313, the training data generation unit 53 extracts a specific subject image S from the exposure image P based on the identified segmented region D. The training data generation unit 53 also obtains the position coordinates of the extracted specific subject image S. After this, the training data generation process proceeds to step ST314.
[0163] In step ST314, the training data generation unit 53 outputs the specific subject data and label L to the learning device 14. The specific subject data includes the position coordinates of the main exposure image P and the specific subject image S. The learning device 14 stores the received specific subject data and label L as training data 17 in the database 16. This completes the training data generation process.
[0164] As described above, in this third embodiment, the color difference between the focus target area F and each divided area D is used as the similarity evaluation value. Therefore, with this configuration, the training data generation unit 53 can extract a specific subject image S from the exposure image P more easily than when the color difference between the focus target area F and each divided area D is not used.
[0165] In this third embodiment, the training data generation unit 53 used the color difference between the focus target area F and each divided area D as the similarity evaluation value, but the technology of this disclosure is not limited to this. In addition to the color difference between the focus target area F and each divided area D, or instead of the color difference, the training data generation unit 53 may use the difference in saturation between the focus target area F and each divided area D as the similarity evaluation value.
[0166] [Fourth Embodiment] In this fourth embodiment, the training data generation unit 53 extracts a specific subject image S from the exposure image P using both the distance between focus positions and the color difference. The configuration of the imaging device 12 according to this fourth embodiment is the same as that of the first embodiment, so its description is omitted. Also, the method for calculating the distance between focus positions and the color difference according to this fourth embodiment is the same as that of the second and third embodiments, so its description is omitted.
[0167] The operation of the imaging device 12 according to this fourth embodiment will be explained with reference to Figure 19. Figure 19 shows an example of the flow of the training data generation process according to the fourth embodiment.
[0168] In Figure 19, steps ST401 to ST409 are the same as steps ST101 to ST109 in Figure 11, so their explanation is omitted. Also, in Figure 19, step ST410 is the same as step ST210 in Figure 16, so its explanation is omitted.
[0169] In step ST411, the training data generation unit 53 calculates the distance between the focus positions of each divided region D. After this, the training data generation process proceeds to step ST412.
[0170] In step ST412, the training data generation unit 53 calculates the color difference for each divided region D. After this, the training data generation process proceeds to step ST413.
[0171] In step ST413, the training data generation unit 53 identifies a divided region D in which the distance between focus positions is less than the distance threshold and the color difference is less than the color difference threshold. After this, the training data generation process proceeds to step ST414.
[0172] In step ST414, the training data generation unit 53 extracts a specific subject image S from the exposure image P based on the identified segmented region D. The training data generation unit 53 also obtains the position coordinates of the extracted specific subject image S. After this, the training data generation process proceeds to step ST415.
[0173] In step ST415, the training data generation unit 53 associates specific subject data with label L and outputs it to the learning device 14. The learning device 14 stores the received specific subject data and label L as training data 17 in the database 16. This completes the training data generation process.
[0174] As described above, in this fourth embodiment, both the distance between focus positions and the color difference are used as similar evaluation values. Therefore, with this configuration, the training data generation unit 53 can extract the specific subject image S from the exposure image P with greater accuracy compared to when neither the distance between focus positions nor the color difference is used.
[0175] [Fifth Embodiment] This fifth embodiment is effective, for example, when the specific subject is a moving object. In this fifth embodiment, if the specific subject moves between the AF operation and the main exposure operation, and the reliability of the specific subject image S extracted from the main exposure image P is determined to be low, warning information indicating low reliability is added to the specific subject data. The fifth embodiment will be described below with reference to Figures 20 to 22. Note that the configuration of the imaging device 12 according to this fifth embodiment is the same as that of the first embodiment described above, so the description will be omitted.
[0176] As an example, as shown in Figure 20, when user 11 performs an AF operation, the training data generation unit 53 acquires one frame from the live view images 66 that are continuously captured at a frame rate of, for example, 60 fps. The training data generation unit 53 extracts an image indicating a specific subject (hereinafter referred to as "live view specific subject image LS") from the one frame of the live view image 66 based on the focus position distance described in the second embodiment and / or the color difference described in the third embodiment. The live view specific subject image LS is an example of a "specific subject image for display" related to the technology of this disclosure.
[0177] The training data generation unit 53 generates the lower right corner Q of the extracted live view specific subject image LS. 1L and the upper left corner Q 2L The coordinates of the live view specific subject image LS are determined as the position coordinates. The training data generation unit 53 also determines the size of the live view specific subject image LS and the center point Q of the live view specific subject image LS based on the position coordinates of the live view specific subject image LS. CL Coordinates (X CL ,Y CL The coordinates of the center of the live view specific subject image LS (hereinafter referred to as "the center coordinates of the live view specific subject image LS") are determined.
[0178] Subsequently, when user 11 performs the actual exposure operation, the training data generation unit 53 acquires the actual exposure image P. The training data generation unit 53 extracts the specific subject image S from the actual exposure image P in the same manner as when the live view specific subject image LS was extracted.
[0179] The training data generation unit 53 generates the lower right corner Q of the extracted specific subject image S. 1E and the upper left corner Q 2E The coordinates of are determined as the position coordinates of the specific subject image S. The training data generation unit 53 also determines the size of the specific subject image S and the center point Q of the specific subject image S based on the position coordinates of the specific subject image S. CE Coordinates (X CE ,Y CE We determine the coordinates of the center of the image S of the specific subject (hereinafter referred to as "center coordinates of the image S of the specific subject").
[0180] The training data generation unit 53 calculates the size difference between the live view specific subject image LS and the specific subject image S by comparing their sizes. As an example, as shown in Figure 20, if the calculated size difference exceeds a predetermined size range, the training data generation unit 53 outputs warning information to the learning device 14, along with the specific subject data and label L, warning that the confidence level of the extracted specific subject image S is low. The size difference is an example of the "difference" related to the technology of this disclosure. The predetermined size range is an example of the "second predetermined range" related to the technology of this disclosure. The process of outputting warning information is an example of the "anomaly detection process" related to the technology of this disclosure.
[0181] Furthermore, the training data generation unit 53 calculates the degree of difference in the center positions of the live view specific subject image LS and the specific subject image S by comparing the center coordinates of the live view specific subject image LS and the specific subject image S. As an example, as shown in Figure 21, if the calculated degree of difference in center positions exceeds a predetermined position range, the training data generation unit 53 outputs warning information to the learning device 14, along with the specific subject data and label L, warning that the confidence level of the extracted specific subject image S is low. Note that the degree of difference in center positions is an example of the "degree of difference" related to the technology of this disclosure. Also, the predetermined position range is an example of the "second predetermined range" related to the technology of this disclosure.
[0182] The operation of the imaging device 12 according to this fifth embodiment will be explained with reference to Figures 22A and 22B. Figures 22A and 22B show an example of the flow of the training data generation process according to the fifth embodiment.
[0183] In Figure 22A, steps ST501 to ST507 are the same as steps ST101 to ST107 in Figure 11, so their explanation is omitted.
[0184] In step ST508, the training data generation unit 53 acquires one frame from the live view image 66. After this, the training data generation process proceeds to step ST509.
[0185] In step ST509, the training data generation unit 53 sets candidate regions 78 and divided regions D in the acquired 1-frame live view image 66. After this, the training data generation process proceeds to step ST510.
[0186] In step ST510, the training data generation unit 53 calculates the distance between the focus positions and / or the color difference of each divided region D. After this, the training data generation process proceeds to step ST511.
[0187] In step ST511, the training data generation unit 53 identifies a divided region D that satisfies "distance between focus positions < distance threshold" and / or "color difference < color difference threshold". After this, the training data generation process proceeds to step ST512.
[0188] In step ST512, the training data generation unit 53 extracts a live view specific subject image LS from a single frame of live view image 66 based on the identified segmented region D. After this, the training data generation process proceeds to step ST513.
[0189] In step ST513, the training data generation unit 53 calculates the position coordinates, size, and center coordinates of the live view specific subject image LS. After this, the training data generation process proceeds to step ST514.
[0190] In step ST514, the training data generation unit 53 determines whether or not the main exposure has been performed. If the main exposure has been performed in step ST514, the determination is affirmed, and the training data generation process proceeds to step ST515. If the main exposure has not been performed in step ST514, the determination is denied, and the training data generation process proceeds to step ST506.
[0191] In step ST515, the training data generation unit 53 acquires the exposure image P. After this, the training data generation process proceeds to step ST516.
[0192] In step ST516, the training data generation unit 53 sets candidate regions 78 and divided regions D in the exposure image P. After this, the training data generation process proceeds to step ST517.
[0193] In step ST517, the training data generation unit 53 calculates the distance between the focus positions and / or the color difference of each divided region D. After this, the training data generation process proceeds to step ST518.
[0194] In step ST518, the training data generation unit 53 identifies a divided region D that satisfies "distance between focus positions < distance threshold" and / or "color difference < color difference threshold". After this, the training data generation process proceeds to step ST519.
[0195] In step ST519, the training data generation unit 53 extracts a specific subject image S from the exposure image P based on the identified segmented region D. After this, the training data generation process proceeds to step ST520.
[0196] In step ST520, the training data generation unit 53 calculates the position coordinates, size, and center coordinates of the specific subject image S. After this, the training data generation process proceeds to step ST521.
[0197] In step ST521, the training data generation unit 53 calculates the size difference between the live view specific subject image LS and the specific subject image S by comparing their sizes. After this, the training data generation process proceeds to step ST522.
[0198] In step ST522, the training data generation unit 53 determines whether the calculated size difference is within the predetermined size range. If the size difference is within the predetermined size range in step ST522, the determination is affirmed, and the training data generation process proceeds to step ST523. If the size difference exceeds the predetermined size range in step ST522, the determination is denied, and the training data generation process proceeds to step ST526.
[0199] In step ST523, the training data generation unit 53 calculates the degree of difference between the center positions of the live view specific subject image LS and the specific subject image S by comparing the center position of the live view specific subject image LS with the center position of the specific subject image S. After this, the training data generation process proceeds to step ST524.
[0200] In step ST524, the training data generation unit 53 determines whether the calculated difference in center positions is within the predetermined position range. If the difference in center positions is within the predetermined position range in step ST524, the determination is affirmed, and the training data generation process proceeds to step ST525. If the difference in center positions exceeds the predetermined position range in step ST524, the determination is denied, and the training data generation process proceeds to step ST526.
[0201] In step ST525, the training data generation unit 53 outputs the specific subject data and label L to the learning device 14. The specific subject data includes the exposed image P and the position coordinates of the specific subject image S. Meanwhile, in step ST526, the training data generation unit 53 outputs warning information to the learning device 14 in addition to the specific subject data and label L. This completes the training data generation process.
[0202] As described above, according to this fifth embodiment, the training data generation unit 53 outputs warning information to the learning device 14 if the size difference between the live view specific subject image LS extracted from the live view image 66 and the specific subject image S extracted from the exposure image P exceeds a predetermined size range, or if the difference in the center position between the live view specific subject image LS and the specific subject image S exceeds a predetermined position range. Therefore, specific subject data relating to the specific subject image S that is judged to have low reliability is output to the learning device 14 with warning information attached, thus improving the quality of the training data 17 compared to the case where no warning information is attached.
[0203] In the fifth embodiment described above, the training data generation unit 53 adds warning information to the specific subject data relating to the specific subject image S that it determines to have a low reliability and outputs it to the learning device 14. However, the technology of this disclosure is not limited to this. The training data generation unit 53 does not have to output the specific subject data relating to the specific subject image S that it determines to have a low reliability to the learning device 14. Alternatively, the training data generation unit 53 may add a confidence score indicating the reliability of the specific subject image S to the specific subject data and output it to the learning device 14. In this case, the learning device 14 may refer to the confidence score and not accept the specific subject data with a low confidence score.
[0204] [Sixth Embodiment] In this sixth embodiment, the training data generation unit 53 causes the image sensor 24 to perform the main exposure operation at multiple focus positions, thereby acquiring not only the main exposure image P (hereinafter also referred to as the "focused image") in focus on the focus target area F, but also the main exposure image P (hereinafter also referred to as the "out-of-focus image") that is out of focus on the focus target area F. The training data generation unit 53 not only outputs specific subject data relating to the specific subject image S captured in the focused image as training data 17, but also outputs specific subject data relating to the specific subject image S captured in the out-of-focus image as training data 17. The sixth embodiment will now be described with reference to Figures 23 to 25. Note that the configuration of the imaging device 12 according to this sixth embodiment is the same as that of the first embodiment described above, so its description will be omitted.
[0205] As an example, as shown in Figure 23, the training data generation unit 53 causes the image sensor 24 to perform the main exposure operation at multiple focus positions, including a focus position derived by performing AF calculations based on the distance to the focus target area F. For example, when imaging is performed with the position of the left eye of a specific subject A as the focus target area F (see Figure 12), the training data generation unit 53 causes the image sensor 24 to perform the main exposure operation at five focus positions, including a focus position derived based on the distance to the focus target area F. Note that the five focus positions are an example of "multiple focus positions" related to the technology of this disclosure.
[0206] As a result, the image sensor 24 outputs not only the in-focus image P3, which is in focus on the specific subject A, but also out-of-focus images P1, P2, P4, and P5, which are not in focus on the specific subject A. Out-of-focus images P1 and P2 are front-focused images that are in focus on a subject closer to the imaging device 12 than the specific subject A. Out-of-focus images P4 and P5 are back-focused images that are in focus on a subject further away from the imaging device 12 than the specific subject A. Note that the in-focus image P3 is an example of an "in-focus image" according to the technology of this disclosure. Out-of-focus images P1, P2, P4, and P5 are examples of "out-of-focus images" according to the technology of this disclosure.
[0207] The training data generation unit 53 extracts a specific subject image S from the focused image P3 based on the distance between focus positions described in the second embodiment and / or the color difference described in the third embodiment. The training data generation unit 53 also determines the position coordinates of the extracted specific subject image S.
[0208] As an example, as shown in Figure 24, the training data generation unit 53 associates the focused image P3 with the position coordinates of the specific subject image S and the label L, and outputs it to the learning device 14 as training data 17-3.
[0209] Furthermore, the training data generation unit 53 associates each out-of-focus image P1, P2, P4, or P5 with the position coordinates of the specific subject image S extracted from the in-focus image P3 and the label L, and outputs them to the learning device 14 as training data 17-1, 17-2, 17-4, or 17-5. That is, the training data generation unit 53 outputs the position coordinates of the specific subject image S extracted from the in-focus image P3 as the position coordinates of the specific subject image S in the out-of-focus images P1, P2, P4, or P5. The learning device 14 receives the training data 17-1 to 17-5 and stores them in the database 16.
[0210] The operation of the imaging device 12 according to this sixth embodiment will be explained with reference to Figure 25. Figure 25 shows an example of the flow of the training data generation process according to this sixth embodiment.
[0211] In Figure 25, steps ST601 to ST607 are the same as steps ST101 to ST107 in Figure 11, so their explanation is omitted.
[0212] In step ST608, the training data generation unit 53 determines whether or not the exposure operation has been performed. If the exposure operation has been performed in step ST608, the determination is affirmed, and the exposure operation is performed at multiple focus positions, including a focus position based on the distance to the focus target area F, and the training data generation process proceeds to step ST609. If the exposure operation has not been performed in step ST608, the determination is denied, and the training data generation process proceeds to step ST606.
[0213] In step ST609, the training data generation unit 53 acquires multiple main exposure images P1 to P5. Of the multiple main exposure images P1 to P5, main exposure image P3 is in focus, while main exposure images P1, P2, P4, and P5 are out of focus. After this, the training data generation process proceeds to step ST610.
[0214] In step ST610, the training data generation unit 53 sets candidate regions 78 and divided regions D in the focused image P3. After this, the training data generation process proceeds to step ST611.
[0215] In step ST611, the training data generation unit 53 calculates the distance between focus positions and / or color difference for each divided region D. After this, the training data generation process proceeds to step ST612.
[0216] In step ST612, the training data generation unit 53 identifies a divided region D in which the distance between focus positions is less than the distance threshold and / or the color difference is less than the color difference threshold. After this, the training data generation process proceeds to step ST613.
[0217] In step ST613, the training data generation unit 53 extracts a specific subject image S from the main exposure image (focused image) P3 based on the identified segmented region D. After this, the training data generation process proceeds to step ST614.
[0218] In step ST614, the training data generation unit 53 obtains the position coordinates of the specific subject image S. After this, the training data generation process proceeds to step ST615.
[0219] In step ST615, the specific subject data and label L are associated and output to the learning device 14. The specific subject data includes each of the main exposure images P1 to P5 and the position coordinates of the specific subject image S extracted from the main exposure image P3. Therefore, in this sixth embodiment, five types of specific subject data are output by executing the training data generation process once. The learning device 14 associates the specific subject data with label L and stores it in the database 16. This completes the training data generation process.
[0220] As described above, in this sixth embodiment, the image sensor 24 performs the exposure operation at multiple focus positions. For each of the multiple exposure images P1 to P5 obtained by the exposure operation, the training data generation unit 53 outputs the position coordinates of the specific subject image S obtained from the focused image P3 as the position coordinates of the specific subject image S in each of the out-of-focus images P1, P2, P4, and P5. Therefore, with this configuration, compared to the case where the specific subject image S is extracted manually, the training data generation unit 53 can easily acquire specific subject data relating to the specific subject image S included in the focused image P3 and specific subject data relating to the specific subject image S included in each of the out-of-focus images P1, P2, P4, and P5.
[0221] Furthermore, with this configuration, the training data generation unit 53 can individually label multiple main exposure images P1 to P5 with a single selection of label L. This eliminates the need to manually label multiple main exposure images P1 to P5. Alternatively, the training data generation unit 53 may also label the main exposure images P1 to P5 after capture. In this case as well, it is desirable that the label L be applied to multiple continuously captured main exposure images P1 to P5 with a single selection of label L. If labels L are applied individually after capture, the out-of-focus images may become unrecognizable depending on their blurring. However, this problem can be resolved by applying the same label L to multiple continuously captured main exposure images P1 to P5 with a single selection of label L. In this case, it is desirable that the training data generation unit 53 applies the label L selected for the in-focus image P3 to the out-of-focus images P1, P2, P4, and P5, respectively.
[0222] In the sixth embodiment described above, the training data generation unit 53 outputs five types of specific subject data obtained by imaging at five focus positions during a single exposure operation, but the technology of this disclosure is not limited thereto. The number of focus positions for which the image sensor 24 performs imaging may be more or less than five. The training data generation unit 53 outputs a number of specific subject data corresponding to the number of focus positions.
[0223] Furthermore, in the sixth embodiment described above, the training data generation unit 53 may assign an AF evaluation value indicating the degree of out-of-focus to the specific subject data, including the out-of-focus images P1, P2, P4, and P5. The training data generation unit 53 may also assign a label indicating "in focus" or "out-of-focus" to the specific subject data based on the AF evaluation value. This improves the quality of the training data 17 compared to the case where no AF evaluation value is assigned.
[0224] In the first to sixth embodiments described above, the specific subject data includes the exposed image P and the position coordinates of the specific subject image S, but the technology of this disclosure is not limited thereto. As an example, as shown in Figure 26, the specific subject data may be a specific subject image S extracted from the exposed image P. The training data generation unit 53 outputs the specific subject image S extracted from the exposed image P as training data 17 to be used for machine learning, associating it with a label L. With this configuration, the size of the specific subject data to be output is smaller compared to the case where the exposed image P is output without being extracted. Specifically, "the training data generation unit 53 outputs the specific subject data as data to be used for machine learning" includes, for example, a storage process in which the training data generation unit 53 stores the position coordinates of the exposed image P and the specific subject image S, or an extraction process in which the specific subject image S is extracted from the exposed image P.
[0225] Furthermore, although the frame line 68A is rectangular in the first to sixth embodiments described above, the technology of this disclosure is not limited thereto, and the shape of the frame line 68A can be arbitrarily changed.
[0226] Furthermore, in the first to sixth embodiments described above, the area surrounded by the AF frame 68 is designated as the focus target area F, and the focus target area F is displayed in a manner that makes it distinguishable from other image areas, but the technology of this disclosure is not limited thereto. The training data generation unit 53 may, for example, display an arrow on the live view image 66 and designate the area indicated by the arrow as the focus target area F. Alternatively, the training data generation unit 53 may, for example, receive a designation of the focus target area F by sensing contact of an indicator to the touch panel 36, and display the designated focus target area F in a color that makes it distinguishable from other image areas.
[0227] Furthermore, in the first to sixth embodiments described above, the learning device 14 stores the training data 17 output from the imaging device 12 in a database 16 and performs machine learning using the training data 17 stored in the database 16, but the technology of this disclosure is not limited thereto. For example, the CPU 50A of the imaging device 12 may store the training data 17 it has acquired in an NVM 50B and perform machine learning using the training data 17 stored in the NVM 50B. With this configuration, the imaging device 12 can perform both the acquisition of training data 17 and learning, so fewer devices are required compared to when the acquisition of training data 17 and learning are performed by separate devices.
[0228] Furthermore, in the first to sixth embodiments described above, when the imaging device 12 is set to training data acquisition mode, the training data generation unit 53 displays the label selection screen 64 on the touch panel monitor 30 before the AF operation and the main exposure operation, allowing the user 11 to select label L. However, the technology of this disclosure is not limited thereto. The training data generation unit 53 may also display the label selection screen 64 on the touch panel monitor 30 after the image sensor 24 has acquired the main exposure image P, and then accept the selection of label L from the user 11.
[0229] Furthermore, in the first to sixth embodiments described above, the training data generation unit 53 associates specific subject data with labels L and outputs it to the learning device 14 as training data 17 for supervised machine learning, but the technology of this disclosure is not limited thereto. The training data generation unit 53 may output only specific subject data to the learning device 14. In this case, the user 11 may label the specific subject data on the learning device 14. Alternatively, labeling of the specific subject data may not be performed. In this case, the specific subject data may be used as training data for unsupervised machine learning or for conventional pattern recognition techniques.
[0230] Furthermore, although the first to sixth embodiments described above have described an example of a configuration in which the non-phase difference pixel group 86G and the phase difference pixel group 84G are used in combination, the technology of this disclosure is not limited thereto. For example, instead of the non-phase difference pixel group 86G and the phase difference pixel group 84G, an area sensor may be used in which phase difference image data and non-phase difference image data are selectively generated and read out. In this case, the area sensor has a plurality of photosensitive pixels arranged in two dimensions. For example, a pair of independent photodiodes without light-shielding members are used as photosensitive pixels in the area sensor. When non-phase difference image data is generated and read out, photoelectric conversion is performed by the entire area of the photosensitive pixels (the pair of photodiodes), and when phase difference image data is generated and read out (for example, when performing a passive distance measurement), photoelectric conversion is performed by one of the photodiodes of the pair. Here, one of the pair of photodiodes is a photodiode corresponding to the first phase difference pixel 84-L described in the above embodiment, and the other of the pair of photodiodes is a photodiode corresponding to the second phase difference pixel 84-R described in the above embodiment. While it is possible to selectively generate and read out phase difference image data and non-phase difference image data using all the photosensitive pixels included in the area sensor, the system is not limited to this, and may also be configured so that phase difference image data and non-phase difference image data are selectively generated and read out using some of the photosensitive pixels included in the area sensor.
[0231] Furthermore, in the first to sixth embodiments described above, a method for deriving the distance to the focus target region F was explained using the phase difference method as an example. However, the technology of this disclosure is not limited thereto, and a TOF method or a contrast method may also be used.
[0232] Furthermore, although the first to sixth embodiments described above described an example in which the training data generation program 51A is stored in the NVM 50B, the technology of this disclosure is not limited thereto. For example, as shown in Figure 27, the training data generation program 51A may be stored in a storage medium 100. The storage medium 100 is a non-temporary storage medium. An example of the storage medium 100 is any portable storage medium such as an SSD or a USB memory.
[0233] The training data generation program 51A stored in the storage medium 100 is installed in the controller 50. The CPU 50A executes the training data generation process according to the training data generation program 51A.
[0234] Alternatively, the training data generation program 51A may be stored in the memory of another computer or server device connected to the controller 50 via a communication network (not shown), and the training data generation program 51A may be downloaded and installed in the controller 50 in response to a request from the imaging device 12.
[0235] It is not necessary to store the entire training data generation program 51A in the storage unit of another computer or server device connected to the controller 50, or in the storage medium 100; it is acceptable to store only a portion of the training data generation program 51A.
[0236] In the example shown in Figure 4, the controller 50 is built into the imaging device 12. However, the technology of this disclosure is not limited to this, and for example, the controller 50 may be provided outside the imaging device 12.
[0237] In the example shown in Figure 4, CPU50A is a single CPU, but it could be multiple CPUs. Alternatively, a GPU could be used instead of CPU50A.
[0238] In the example shown in Figure 4, a controller 50 is illustrated, but the technology of this disclosure is not limited thereto, and devices including ASICs, FPGAs, and / or PLDs may be used instead of the controller 50. Alternatively, a combination of hardware and software configurations may be used instead of the controller 50.
[0239] The hardware resources used to perform the training data generation process described in the above embodiment include the following types of processors. Examples of processors include a CPU, which is a general-purpose processor that functions as a hardware resource for performing the training data generation process by executing software, i.e., a program. Other examples of processors include dedicated electrical circuits, which are processors with circuit configurations specifically designed to perform particular processing, such as FPGAs, PLDs, or ASICs. Each processor has built-in or connected memory, and each processor uses this memory to perform the training data generation process.
[0240] The hardware resources that perform the training data generation process may consist of one of these various processors, or a combination of two or more processors of the same or different types (for example, a combination of multiple FPGAs, or a combination of a CPU and an FPGA). Alternatively, the hardware resources that perform the training data generation process may consist of a single processor.
[0241] Examples of configurations using a single processor include, firstly, a configuration in which one or more CPUs and software are combined to form a single processor, and this processor functions as a hardware resource that performs the training data generation process. Secondly, there is a configuration using a processor that realizes the functions of the entire system, including multiple hardware resources that perform the training data generation process, on a single IC chip, such as SoCs. In this way, the training data generation process is realized using one or more of the above types of processors as hardware resources.
[0242] Furthermore, the hardware structure of these various processors can more specifically utilize electrical circuits that combine circuit elements such as semiconductor devices. Also, the training data generation process described above is merely an example. Therefore, it goes without saying that unnecessary steps can be deleted, new steps added, or the processing order rearranged, as long as it does not deviate from the main purpose.
[0243] The descriptions and illustrations presented above are detailed explanations of the technical aspects of this disclosure and are merely examples of the technical aspects. For example, the above descriptions of the structure, function, operation, and effect are examples of the structure, function, operation, and effect of the technical aspects of this disclosure. Therefore, it goes without saying that you may delete unnecessary parts, add new elements, or replace elements in the descriptions and illustrations presented above, as long as you do not deviate from the essence of the technical aspects of this disclosure. Furthermore, in order to avoid confusion and facilitate understanding of the technical aspects of this disclosure, explanations of common technical knowledge and the like that do not require special explanation to enable the implementation of the technical aspects of this disclosure have been omitted from the descriptions and illustrations presented above.
[0244] In this specification, "A and / or B" is synonymous with "at least one of A and B." That is, "A and / or B" means that it may be A alone, or B alone, or a combination of A and B. Furthermore, in this specification, the same concept as "A and / or B" applies when expressing three or more things linked by "and / or."
[0245] All documents, patent applications, and technical standards described herein are incorporated by reference to the same extent as if each individual document, patent application, and technical standard were specifically and individually noted as being incorporated by reference.
Claims
1. Processor and The memory connected to or built into the aforementioned processor, An imaging device comprising an image sensor connected to the processor, A display motion image based on the signal output from the aforementioned image sensor is displayed on the monitor. In response to receiving an instruction to designate a specific subject as the focus area in the aforementioned display video, the focus area is set in the aforementioned display video. In imaging by the image sensor, after the focusing operation on the focus target area is performed, the main exposure is performed, and when an image is obtained by the main exposure, the captured image is divided into a predetermined area including the focus target area into a plurality of divided areas. For each of the divided regions, a specific subject image representing the specific subject within the predetermined region is extracted based on the focus evaluation value used for the focusing operation and the evaluation value of the pixel information. The coordinates of the specific subject image within the captured image and the specific subject data including the captured image are output as data to be used for machine learning. Imaging device.
2. The aforementioned machine learning is supervised machine learning, The aforementioned processor, A label, which is information relating to the aforementioned specific subject image, is assigned to the aforementioned specific subject data. The aforementioned specific subject data is output as training data to be used in the supervised machine learning. The imaging apparatus according to claim 1.
3. The processor displays the focus target area in a manner that allows it to be distinguished from other image areas, while a displayable moving image based on the signal output from the image sensor is displayed on the monitor. The aforementioned specific subject image is an image corresponding to the position of the focus target region within the captured image. The imaging apparatus according to claim 1 or claim 2.
4. The processor displays the focus target area in the display video in a manner that allows it to be distinguished from other image areas by displaying a frame surrounding the focus target area. The imaging device according to claim 3.
5. The position of the frame can be changed according to a given position change instruction. The imaging apparatus according to claim 4.
6. The size of the aforementioned frame can be changed according to the given resizing instructions. The imaging apparatus according to claim 4 or claim 5.
7. The processor is For each of the divided regions, a first similarity evaluation value is calculated based on the focus evaluation value used in the focus operation, and a second similarity evaluation value is calculated based on the color information of each of the divided regions, which is the evaluation value of the pixel information. The aforementioned specific subject image is extracted based on a segmented region in which the first similarity evaluation value is within the first predetermined range and the second similarity evaluation value is within the second predetermined range. The imaging apparatus according to any one of claims 1 to 6.
8. The processor outputs warning information if the degree of difference between the display image of a specific subject showing the specific subject in the display video exceeds the second predetermined range. The aforementioned specific subject image for display is determined based on the first similarity evaluation value and the second similarity evaluation value. The imaging apparatus according to claim 7.
9. The image sensor captures images at multiple focus positions, The processor outputs, for each of the multiple captured images obtained by the imaging, the coordinates of the specific subject image obtained from the in-focus image that is in focus on the specific subject, as the coordinates of the specific subject image in the out-of-focus image that is not in focus on the specific subject. The imaging apparatus according to any one of claims 1 to 8.
10. A method for controlling an imaging device comprising a processor, a memory connected to or built into the processor, and an image sensor connected to the processor, The aforementioned processor, Display a moving image for display on a monitor based on the signal output from the aforementioned image sensor. In response to receiving an instruction to designate a specific subject as the focus area in the aforementioned display video, the focus area is set in the aforementioned display video. In imaging by the image sensor, after the focusing operation on the focus target area is performed, the main exposure is performed, and when an image is obtained by the main exposure, the captured image is divided into a predetermined area including the focus target area into a plurality of divided areas. For each of the divided regions, a specific subject image representing the specific subject within the predetermined region is extracted based on the focus evaluation value used for the focusing operation and the evaluation value of the pixel information. This includes outputting specific subject data, which includes the coordinates of the specific subject image within the captured image and the captured image, as data to be used for machine learning. A method for controlling an imaging device.
11. A program for causing a computer to perform processing, which is included in an imaging device comprising a computer and an image sensor connected to the computer, The aforementioned process is, Display a moving image for display on a monitor based on the signal output from the aforementioned image sensor. In response to receiving an instruction to designate a specific subject as the focus area in the aforementioned display video, the focus area is set in the aforementioned display video. In imaging by the image sensor, after the focusing operation on the focus target area is performed, the main exposure is performed, and when an image is obtained by the main exposure, the captured image is divided into a predetermined area including the focus target area into a plurality of divided areas. For each of the divided regions, a specific subject image representing the specific subject within the predetermined region is extracted based on the focus evaluation value used for the focusing operation and the evaluation value of the pixel information. This includes outputting specific subject data, which includes the coordinates of the specific subject image within the captured image and the captured image, as data to be used for machine learning. A program to execute a process.
Citation Information
Patent Citations
Automatic focusing device
JP1992346580A
Imaging apparatus, method, and program
JP2006333443A
Photographing apparatus, photographing method and program
JP2009055432A
Image processing device and image processing method, recognition device and recognition method, and program
JP2009069996A
Image processing apparatus, image processing method and program
JP2012133607A