Imaging support device, imaging device, imaging support method, and program
The imaging support device uses a processor and learned model to enhance focus control by identifying subject types and dividing areas based on probability, ensuring critical features are in focus, addressing the challenge of subject misfocus in existing technologies.
Patent Information
- Application Number
- JP2025066307
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2020-12-28
- Filing Date
- 2025-04-14
- Publication Date
- 2025-07-15
AI Technical Summary
Existing imaging technologies struggle to accurately control focus and distinguish subjects from other areas in an imaging range, leading to potential misfocus on important subject features.
An imaging support device that utilizes a processor to analyze images using a learned model to identify subject types and divide areas based on probability, adjusting focus and control methods accordingly.
Enhances focus control by accurately identifying and prioritizing subject areas, ensuring critical features are in focus, even in varying imaging conditions.
Smart Images

Figure 2025106533000001_ABST
Abstract
Description
Technical Field
[0001] The technology of the present disclosure relates to an imaging support device, an imaging device, an imaging support method, and a program.
Background Art
[0002] Japanese Patent Application Laid-Open No. 2015-046917 discloses an imaging unit that images a subject image formed through a photographing lens and outputs image data, a focusing determination unit that performs weighting determined according to at least one of the focus position of the subject and the characteristics of the subject area, and determines from which area in the photographed image to which area to perform focusing and the order of focusing according to the weighting result, a focus control unit that drives the photographing lens according to the order of focusing determined by the focusing determination unit, and a recording unit that records image data of a moving image and a still image based on the image data. The focusing determination unit determines a main subject / secondary subject or a main area / secondary area according to the weighting result, determines the order of focusing with one of the main subject / secondary subject or the main area / secondary area as the starting point and the other as the end point, and the recording unit records the image data of the moving image while the photographing lens is being driven by the focus control unit, and records the image data of the still image after the driving of the photographing lens is stopped by the focus control unit. A photographing device characterized by this is disclosed.
[0003] Japanese Patent Application Laid-Open No. 2012-128287 discloses a face detection means for detecting the position and size of a person's face from a captured image, a setting means for setting a first focus detection area where a person's face exists and a second focus detection area where the person's body is predicted to be located as seen from the position of the person's face as a focus detection area when detecting the in-focus state from an imaging optical system, and a focus adjustment means for moving the imaging optical system based on the signal output in the focus detection area to perform focus adjustment. The setting means is characterized in that when the size of the face detected by the face detection means is smaller than a predetermined size, the size of the second focus detection area is set to be larger than the size of the first focus detection area. A focus detection device characterized by this is disclosed.
Summary of the Invention
[0004] One embodiment of the technology disclosed herein provides an imaging support device, an imaging device, an imaging support method, and a program that can accurately control imaging by an image sensor for a subject.
Means for Solving the Problems
[0005] A first aspect of the technology disclosed herein includes a processor and a memory connected to or incorporated in the processor. The processor acquires the type of a subject based on an image obtained by imaging an imaging range including the subject with an image sensor, and outputs information indicating a division method of dividing a division region that can distinguish the subject from other regions in the imaging range according to the acquired type. This is an imaging support device.
[0006] A second aspect of the technology disclosed herein is the imaging support device according to the first aspect, in which the processor acquires the type based on an output result output from a learned model by providing an image to the learned model on which machine learning has been performed.
[0007] A third aspect of the technology disclosed herein is the imaging support device according to the second aspect, in which the learned model assigns an object within a bounding box applied to an image to a corresponding class, and the output result includes a value based on the probability that the object within the bounding box applied to the image belongs to a specific class.
[0008] A fourth aspect of the technology disclosed herein is the imaging support device according to the third aspect, in which the output result includes a value based on the probability that an object within the bounding box belongs to a specific class when a value based on the probability that an object exists within the bounding box is equal to or greater than a first threshold.
[0009] A fifth aspect of the technology disclosed herein is the imaging support device according to the third aspect or the fourth aspect, in which the output result includes a value equal to or greater than a second threshold among the values based on the probability that an object belongs to a specific class.
[0010] A sixth aspect of the technology according to the present disclosure is an imaging support device according to any one of the third to fifth aspects, in which the processor expands the bounding box when a value based on the probability that an object exists within the bounding box is less than a third threshold value.
[0011] A seventh aspect of the technology according to the present disclosure is an imaging support device according to any one of the third to sixth aspects, in which the processor changes the size of the divided area according to a value based on the probability that an object belongs to a specific class.
[0012] An eighth aspect of the technology according to the present disclosure is that a plurality of subjects are included in the imaging range, a learned model assigns each of a plurality of objects in a plurality of bounding boxes applied to an image to a corresponding class, and an output result includes object-by-object class information indicating each class to which the plurality of objects in the plurality of bounding boxes applied to the image belong, and the processor narrows down at least one subject surrounded by a divided area from the plurality of subjects based on the object-by-object class information. It is an imaging support device according to any one of the second to seventh aspects.
[0013] A ninth aspect of the technology according to the present disclosure is an imaging support device according to any one of the first to eighth aspects, in which the division method defines the number of divisions for dividing the divided area.
[0014] A tenth aspect of the technology according to the present disclosure is an imaging support device according to the ninth aspect, in which the divided area is defined by a first direction and a second direction intersecting the first direction, and the division method defines the number of divisions in the first direction and the number of divisions in the second direction.
[0015] An eleventh aspect of the technology according to the present disclosure is an imaging support device according to any one of the first to tenth aspects, in which the number of divisions in the first direction and the number of divisions in the second direction are defined based on the composition within the image of the subject image indicating the subject within the image.
[0016] In a twelfth aspect of the technology according to the present disclosure, when the focus of the focus lens can be adjusted by moving the focus lens that guides incident light to the image sensor along the optical axis, the processor, among a plurality of divided regions obtained by dividing the divided region by the number of divisions, outputs information that contributes to moving the focus lens to a focusing position corresponding to the divided region according to the acquired type, which is an imaging support device according to any one of the ninth to eleventh aspects.
[0017] In a thirteenth aspect of the technology according to the present disclosure, when the focus of the focus lens can be adjusted by moving the focus lens that guides incident light to the image sensor along the optical axis, the processor outputs information that contributes to moving the focus lens to a focusing position according to the acquired type, which is an imaging support device according to any one of the first to twelfth aspects.
[0018] A fourteenth aspect of the technology according to the present disclosure is an imaging support device according to any one of the first to thirteenth aspects, where the divided region is used for control related to imaging of a subject by an image sensor.
[0019] In a fifteenth aspect of the technology according to the present disclosure, the control related to imaging includes custom control, and the custom control recommends changing the control content of the control related to imaging according to the subject, and is a control that changes the control content according to a given instruction. The processor outputs information that contributes to changing the control content according to the acquired type, which is an imaging support device according to the fourteenth aspect.
[0020] A sixteenth aspect of the technology according to the present disclosure is an imaging support device according to the fifteenth aspect, where the processor further acquires the state of the subject based on the image and outputs information that contributes to changing the control content according to the acquired state and type.
[0021] A seventeenth aspect of the technology of the present disclosure is that when the focus of a focus lens that guides incident light to an image sensor is adjustable by moving the focus lens along the optical axis, custom control waits for a predetermined time when the position of the focus lens is deviated from the in-focus position where the subject is in focus, and then moves the focus lens toward the in-focus position, controls for setting an allowable range of the speed of the subject to be focused, and controls for setting which area among a plurality of areas within a divided region is to be given priority for focus adjustment. The imaging support device according to the fifteenth aspect or the sixteenth aspect includes at least one of the controls.
[0022] An eighteenth aspect of the technology of the present disclosure is an imaging support device according to any one of the first aspect to the seventeenth aspect, in which the divided region is a frame surrounding the subject.
[0023] A nineteenth aspect of the technology of the present disclosure is an imaging support device according to the eighteenth aspect, in which a processor causes a display to display a live view image based on an image and outputs information for causing a frame to be displayed within the live view image.
[0024] A twentieth aspect of the technology of the present disclosure is an imaging support device according to the eighteenth aspect or the nineteenth aspect, in which a processor adjusts the focus of a focus lens that guides incident light to an image sensor by moving the focus lens along the optical axis, and the frame is a focus frame that defines an area to be a focus candidate.
[0025] A twenty-first aspect of the technology of the present disclosure is an imaging device including a processor, a memory connected to or built in the processor, and an image sensor. The processor acquires the type of a subject based on an image obtained by imaging an imaging range including the subject with the image sensor, and outputs information indicating a division method for dividing a divided region that divides the subject from other regions in the imaging range according to the acquired type.
[0026] The 22nd aspect according to the technology of the present disclosure includes acquiring the type of a subject based on an image obtained by imaging an imaging range including the subject with an image sensor, and outputting information indicating a division method of dividing a division area that discriminates the subject from other areas of the imaging range according to the acquired type. This is an imaging support method.
[0027] The 23rd aspect according to the technology of the present disclosure is a program for causing a computer to execute a process including acquiring the type of a subject based on an image obtained by imaging an imaging range including the subject with an image sensor, and outputting information indicating a division method of dividing a division area that discriminates the subject from other areas of the imaging range according to the acquired type.
Brief Description of Drawings
[0028]
Figure 1
Figure 2
Figure 3
Figure 4
Figure 5
Figure 6
Figure 7
Figure 8
Figure 9
Figure 10
Figure 11
Figure 12
Figure 13
Figure 14
Figure 15
Figure 16
Figure 17
Figure 18
Figure 19
Figure 20
Figure 21
Figure 22
Figure 23
Figure 24
Figure 25
Figure 26
Figure 27
Figure 28
Figure 29
Figure 30
Figure 31A
Figure 31B
Figure 32
Figure 33
Figure 34
Figure 35
Figure 36
Figure 37
Figure 38
Figure 39
Figure 40
Embodiments for Carrying Out the Invention
[0029] Hereinafter, an example of an embodiment of an imaging support device, an imaging device, an imaging support method, and a program according to the technology of the present disclosure will be described with reference to the accompanying drawings.
[0030] First, the terms used in the following description will be explained.
[0031] The CPU refers to the abbreviation of "Central Processing Unit". The GPU refers to the abbreviation of "Graphics Processing Unit". The TPU refers to the abbreviation of "Tensor processing unit". The NVM refers to the abbreviation of "Non-volatile memory". The RAM refers to the abbreviation of "Random Access Memory". The IC refers to the abbreviation of "Integrated Circuit". The ASIC refers to the abbreviation of "Application Specific Integrated Circuit". The PLD refers to the abbreviation of "Programmable Logic Device". The FPGA refers to the abbreviation of "Field-Programmable Gate Array". The SoC refers to the abbreviation of "System-on-a-chip". The SSD refers to the abbreviation of "Solid State Drive". The USB refers to the abbreviation of "Universal Serial Bus". The HDD refers to the abbreviation of "Hard Disk Drive". The EEPROM refers to the abbreviation of "Electrically Erasable and Programmable Read Only Memory". The EL refers to the abbreviation of "Electro-Luminescence". The I / F refers to the abbreviation of "Interface". The UI refers to the abbreviation of "User Interface". The fps refers to the abbreviation of "frame per second". The MF refers to the abbreviation of "Manual Focus". The AF refers to the abbreviation of "Auto Focus". The CMOS refers to the abbreviation of "Complementary Metal Oxide Semiconductor". The LAN refers to the abbreviation of "Local Area Network". The WAN refers to the abbreviation of "Wide Area Network". The CNN refers to the abbreviation of "Convolutional Neural Network". The AI refers to the abbreviation of "Artificial Intelligence". The TOF refers to the abbreviation of "Time Of Flight".
[0032] In the description of this specification, "vertical" refers to vertical in the sense that, in addition to complete verticality, it includes an error generally acceptable in the technical field to which the technology of the present disclosure belongs and that does not go against the gist of the technology of the present disclosure. In the description of this specification, "orthogonal" refers to orthogonal in the sense that, in addition to complete orthogonality, it includes an error generally acceptable in the technical field to which the technology of the present disclosure belongs and that does not go against the gist of the technology of the present disclosure. In the description of this specification, "parallel" refers to parallel in the sense that, in addition to complete parallelism, it includes an error generally acceptable in the technical field to which the technology of the present disclosure belongs and that does not go against the gist of the technology of the present disclosure. In the description of this specification, "coincident" refers to coincident in the sense that, in addition to complete coincidence, it includes an error generally acceptable in the technical field to which the technology of the present disclosure belongs and that does not go against the gist of the technology of the present disclosure.
[0033] As an example, as shown in FIG. 1, the imaging system 10 includes an imaging device 12 and an imaging support device 14. The imaging device 12 is a device that images a subject. In the example shown in FIG. 1, as an example of the imaging device 12, an interchangeable-lens digital camera is shown. The imaging device 12 includes an imaging device main body 16 and an interchangeable lens 18. The interchangeable lens 18 is detachably attached to the imaging device main body 16. The interchangeable lens 18 is provided with a focus ring 18A. The focus ring 18A is operated by a user of the imaging device 12 (hereinafter simply referred to as "user") or the like when the user or the like manually adjusts the focus on the subject by the imaging device 12.
[0034] Note that in the present embodiment, an interchangeable-lens digital camera is exemplified as the imaging device 12, but this is merely an example, and it may be a fixed-lens digital camera, or a digital camera incorporated in various electronic devices such as a smart device, a wearable terminal, a cell observation device, an ophthalmic observation device, or a surgical microscope.
[0035] The imaging device body 16 is provided with an image sensor 20. The image sensor 20 is a CMOS image sensor. The image sensor 20 images an imaging range including at least one subject. When the interchangeable lens 18 is attached to the imaging device body 16, subject light indicating the subject passes through the interchangeable lens 18 and forms an image on the image sensor 20, and image data indicating an image of the subject is generated by the image sensor 20. The subject light is an example of the "incident light" according to the technology of the present disclosure.
[0036] Note that in the present embodiment, a CMOS image sensor is exemplified as the image sensor 20, but the technology of the present disclosure is not limited thereto, and other image sensors may be used.
[0037] On the upper surface of the imaging device body 16, a release button 22 and a dial 24 are provided. The dial 24 is operated when setting the operation mode of the imaging system, the operation mode of the playback system, etc. By operating the dial 24, in the imaging device 12, the imaging mode and the playback mode are selectively set as the operation modes.
[0038] The release button 22 functions as an imaging preparation instruction unit and an imaging instruction unit, and a two-stage pressing operation of an imaging preparation instruction state and an imaging instruction state can be detected. The imaging preparation instruction state refers to a state where, for example, the button is pressed from the standby position to the intermediate position (half-press position), and the imaging instruction state refers to a state where the button is pressed to the final pressing position (full-press position) beyond the intermediate position. Hereinafter, the state of "being pressed from the standby position to the half-press position" is referred to as the "half-press state", and the state of "being pressed from the standby position to the full-press position" is referred to as the "full-press state". Depending on the configuration of the imaging device 12, the imaging preparation instruction state may be a state where the user's finger touches the release button 22, and the imaging instruction state may be a state where the finger of the operating user moves away from the state of touching the release button 22.
[0039] On the back surface of the imaging device body 16, a touch panel display 32 and an instruction key 26 are provided.
[0040] The touch panel display 32 includes a display 28 and a touch panel 30 (see also FIG. 2). As an example of the display 28, an EL display (e.g., an organic EL display or an inorganic EL display) can be mentioned. The display 28 may be another type of display such as a liquid crystal display instead of an EL display.
[0041] The display 28 displays images and / or character information, etc. The display 28 is used for displaying a live view image obtained by continuously capturing an image for a live view image when the imaging device 12 is in the imaging mode (see FIG. 16 for the live view image 130). The imaging performed to obtain the live view image 130 (see FIG. 16) (hereinafter also referred to as "imaging for live view image") is performed, for example, according to a frame rate of 60 fps. 60 fps is merely an example, and a frame rate less than 60 fps or a frame rate exceeding 60 fps may also be used.
[0042] Here, the "live view image" refers to a moving image for display based on image data obtained by being captured by the image sensor 20. The live view image is generally also referred to as a through image.
[0043] The display 28 is also used for displaying a still image obtained by performing still image imaging when an instruction for still image imaging is given to the imaging device 12 via the release button 22. Further, the display 28 is also used for displaying a reproduced image and displaying a menu screen, etc. when the imaging device 12 is in the playback mode.
[0044] The touch panel 30 is a transmissive touch panel and is overlaid on the surface of the display area of the display 28. The touch panel 30 receives an instruction from the user by detecting contact with an indicator such as a finger or a stylus pen. Hereinafter, for convenience of explanation, the above-described "fully pressed state" also includes a state in which the user turns on the soft key for starting imaging via the touch panel 30.
[0045] Also, in this embodiment, as an example of the touch panel display 32, an out-cell type touch panel display in which the touch panel 30 is overlaid on the surface of the display area of the display 28 is cited, but this is merely an example. For example, an on-cell type or in-cell type touch panel display can also be applied as the touch panel display 32.
[0046] The instruction keys 26 receive various instructions. Here, the "various instructions" refer to, for example, an instruction to display a menu screen on which various menus can be selected, an instruction to select one or more menus, an instruction to confirm the selected content, an instruction to delete the selected content, and various instructions such as zoom in, zoom out, and frame advance. These instructions may also be given by the touch panel 30.
[0047] As will be described in detail later, the imaging device main body 16 is connected to the imaging support device 14 via the network 34. The network 34 is, for example, the Internet. The network 34 is not limited to the Internet, and may be a WAN and / or a LAN such as an intranet. Further, in the present embodiment, the imaging support device 14 is a server that provides a service according to a request from the imaging device 12 to the imaging device 12. Note that the server may be a mainframe used together with the imaging device 12 on-premises, or may be an external server realized by cloud computing. Further, the server may be an external server realized by network computing such as fog computing, edge computing, or grid computing. Here, as an example of the imaging support device 14, a server is mentioned, but this is merely an example, and at least one personal computer or the like may be used as the imaging support device 14 instead of the server.
[0048] As shown in FIG. 2 as an example, the image sensor 20 includes a photoelectric conversion element 72. The photoelectric conversion element 72 has a light receiving surface 72A. The photoelectric conversion element 72 is arranged in the imaging device main body 16 such that the center of the light receiving surface 72A coincides with the optical axis OA (see also FIG. 1). The photoelectric conversion element 72 has a plurality of photosensitive pixels arranged in a matrix, and the light receiving surface 72A is formed by the plurality of photosensitive pixels. The photosensitive pixel is a physical pixel having a photodiode (not shown), photoelectrically converts the received light, and outputs an electrical signal corresponding to the received light amount.
[0049] The interchangeable lens 18 includes an imaging lens 40. The imaging lens 40 has an objective lens 40A, a focus lens 40B, a zoom lens 40C, and a diaphragm 40D. The objective lens 40A, the focus lens 40B, the zoom lens 40C, and the diaphragm 40D are arranged in this order along the optical axis OA from the subject side (object side) to the imaging device main body 16 side (image side).
[0050] The interchangeable lens 18 also includes a control device 36, a first actuator 37, a second actuator 38, and a third actuator 39. The control device 36 controls the entire interchangeable lens 18 in accordance with an instruction from the imaging device main body 16. The control device 36 is, for example, a device having a computer including a CPU, NVM, and RAM. Here, a computer is exemplified, but this is merely an example, and a device including an ASIC, FPGA, and / or PLD may also be applied. Further, as the control device 36, for example, a device realized by a combination of a hardware configuration and a software configuration may be used.
[0051] The first actuator 37 includes a focus slide mechanism (not shown) and a focus motor (not shown). A focus lens 40B is attached to the focus slide mechanism so as to be slidable along the optical axis OA. Further, a focus motor is connected to the focus slide mechanism, and the focus slide mechanism moves the focus lens 40B along the optical axis OA by operating in response to the power of the focus motor.
[0052] The second actuator 38 includes a zoom slide mechanism (not shown) and a zoom motor (not shown). A zoom lens 40C is attached to the zoom slide mechanism so as to be slidable along the optical axis OA. Further, a zoom motor is connected to the zoom slide mechanism, and the zoom slide mechanism moves the zoom lens 40C along the optical axis OA by operating in response to the power of the zoom motor.
[0053] The third actuator 39 includes a power transmission mechanism (not shown) and a diaphragm motor (not shown). The diaphragm 40D has an aperture 40D1, and is a diaphragm whose aperture 40D1 size is variable. The aperture 40D1 is formed by a plurality of diaphragm blades 40D2. The plurality of diaphragm blades 40D2 are connected to the power transmission mechanism. Also, a diaphragm motor is connected to the power transmission mechanism, and the power transmission mechanism transmits the power of the diaphragm motor to the plurality of diaphragm blades 40D2. The plurality of diaphragm blades 40D2 operate by receiving the power transmitted from the power transmission mechanism to change the size of the aperture 40D1. The diaphragm 40D adjusts the exposure by changing the size of the aperture 40D1.
[0054] The focus motor, the zoom motor, and the diaphragm motor are connected to the control device 36, and the control device 36 controls the driving of the focus motor, the zoom motor, and the diaphragm motor. In this embodiment, as an example of the focus motor, the zoom motor, and the diaphragm motor, a stepping motor is adopted. Therefore, the focus motor, the zoom motor, and the diaphragm motor operate in synchronization with the pulse signal according to the command from the control device 36. Here, an example is shown in which the focus motor, the zoom motor, and the diaphragm motor are provided in the interchangeable lens 18, but this is merely an example, and at least one of the focus motor, the zoom motor, and the diaphragm motor may be provided in the imaging device main body 16. Note that the configuration and / or operation method of the interchangeable lens 18 can be changed as necessary.
[0055] In the imaging device 12, in the imaging mode, the MF mode and the AF mode are selectively set according to the instruction given to the imaging device main body 16. The MF mode is an operation mode in which the focus is manually adjusted. In the MF mode, for example, when the focus ring 18A or the like is operated by the user, the focus lens 40B moves along the optical axis OA by an amount of movement corresponding to the operation amount of the focus ring 18A or the like, and thereby the focus is adjusted.
[0056] In the AF mode, the imaging device main body 16 calculates the focusing position according to the subject distance, and adjusts the focus by moving the focus lens 40B toward the calculated focusing position. Here, the focusing position refers to the position on the optical axis OA of the focus lens 40B in the in-focus state. Note that hereinafter, for convenience of explanation, the control to adjust the focus lens 40B to the focusing position is also referred to as "AF control".
[0057] The imaging device main body 16 includes an image sensor 20, a controller 44, an image memory 46, a UI device 48, an external I / F 50, a communication I / F 52, a photoelectric conversion element driver 54, a mechanical shutter driver 56, a mechanical shutter actuator 58, a mechanical shutter 60, and an input / output interface 70. The image sensor 20 includes a photoelectric conversion element 72 and a signal processing circuit 74.
[0058] Connected to the input / output interface 70 are the controller 44, the image memory 46, the UI device 48, the external I / F 50, the photoelectric conversion element driver 54, the mechanical shutter driver 56, and the signal processing circuit 74. Also connected to the input / output interface 70 is the control device 36 of the interchangeable lens 18.
[0059] The controller 44 includes a CPU 62, an NVM 64, and a RAM 66. The CPU 62, the NVM 64, and the RAM 66 are connected via a bus 68, and the bus 68 is connected to the input / output interface 70.
[0060] In the example shown in FIG. 2, for the sake of illustration, one bus is shown as the bus 68, but it may be a plurality of buses. The bus 68 may be a serial bus or a parallel bus including a data bus, an address bus, a control bus, etc.
[0061] NVM64 is a non-volatile memory medium that stores various parameters and various programs. For example, NVM64 is an EEPROM. However, this is only an example, and instead of or together with the EEPROM, an HDD and / or an SSD, etc. may be applied as NVM64. Also, RAM66 temporarily stores various information and is used as a work memory.
[0062] CPU62 reads a necessary program from NVM64 and executes the read program on RAM66. CPU62 controls the entire imaging device 12 according to the program executed on RAM66. In the example shown in FIG. 2, the image memory 46, the UI-based device 48, the external I / F 50, the communication I / F 52, the photoelectric conversion element driver 54, the mechanical shutter driver 56, and the control device 36 are controlled by CPU62.
[0063] A photoelectric conversion element driver 54 is connected to the photoelectric conversion element 72. The photoelectric conversion element driver 54 supplies an imaging timing signal that defines the timing of imaging performed by the photoelectric conversion element 72 to the photoelectric conversion element 72 according to an instruction from CPU62. The photoelectric conversion element 72 performs reset, exposure, and output of an electrical signal according to the imaging timing signal supplied from the photoelectric conversion element driver 54. Examples of the imaging timing signal include a vertical synchronization signal and a horizontal synchronization signal.
[0064] When the interchangeable lens 18 is attached to the imaging device body 16, the subject light incident on the imaging lens 40 is imaged on the light receiving surface 72A by the imaging lens 40. The photoelectric conversion element 72 photoelectrically converts the subject light received by the light receiving surface 72A under the control of the photoelectric conversion element driver 54 and outputs an electrical signal corresponding to the amount of the subject light as analog image data indicating the subject light to the signal processing circuit 74. Specifically, the signal processing circuit 74 reads the analog image data from the photoelectric conversion element 72 in units of one frame and for each horizontal line in the sequential exposure reading method.
[0065] The signal processing circuit 74 generates digital image data by digitizing analog image data. In the following, for convenience of explanation, when there is no need to distinguish between the digital image data that is the target of internal processing in the imaging device main body 16 and the image indicated by the digital image data (that is, the image that is visualized based on the digital image data and displayed on the display 28 or the like), it is referred to as "captured image 108".
[0066] The mechanical shutter 60 is a focal plane shutter and is disposed between the aperture 40D and the light receiving surface 72A. The mechanical shutter 60 includes a front curtain (not shown) and a rear curtain (not shown). Each of the front curtain and the rear curtain includes a plurality of vanes. The front curtain is disposed closer to the subject side than the rear curtain.
[0067] The mechanical shutter actuator 58 is an actuator having a link mechanism (not shown), a solenoid for the front curtain (not shown), and a solenoid for the rear curtain (not shown). The solenoid for the front curtain is a driving source for the front curtain and is mechanically connected to the front curtain via the link mechanism. The solenoid for the rear curtain is a driving source for the rear curtain and is mechanically connected to the rear curtain via the link mechanism. The mechanical shutter driver 56 controls the mechanical shutter actuator 58 according to an instruction from the CPU 62.
[0068] The solenoid for the front curtain generates power under the control of the mechanical shutter driver 56, and selectively winds up and pulls down the front curtain by applying the generated power to the front curtain. The solenoid for the rear curtain generates power under the control of the mechanical shutter driver 56, and selectively winds up and pulls down the rear curtain by applying the generated power to the rear curtain. In the imaging device 12, the opening and closing of the front curtain and the opening and closing of the rear curtain are controlled by the CPU 62, so that the exposure amount to the photoelectric conversion element 72 is controlled.
[0069] In the imaging device 12, imaging for live view images and imaging for recording images for recording still images and / or moving images are performed in an exposure sequential readout method (rolling shutter method). The image sensor 20 has an electronic shutter function, and the imaging for live view images is realized by activating the electronic shutter function without operating the mechanical shutter 60 while keeping it fully open.
[0070] On the other hand, imaging with this exposure, that is, imaging for still images, is realized by activating the electronic shutter function and operating the mechanical shutter 60 so that the mechanical shutter 60 transitions from the front curtain closed state to the rear curtain closed state.
[0071] The imaging image 108 generated by the signal processing circuit 74 is stored in the image memory 46. That is, the signal processing circuit 74 stores the imaging image 108 in the image memory 46. The CPU 62 acquires the imaging image 108 from the image memory 46 and executes various processes using the acquired imaging image 108.
[0072] The UI device 48 includes a display 28, and the CPU 62 causes the display 28 to display various information. The UI device 48 also includes a reception device 76. The reception device 76 includes a touch panel 30 and a hard key unit 78. The hard key unit 78 is a plurality of hard keys including an instruction key 26 (see FIG. 1). The CPU 62 operates according to various instructions received by the touch panel 30. Here, the hard key unit 78 is included in the UI device 48, but the technology of the present disclosure is not limited to this. For example, the hard key unit 78 may be connected to the external I / F 50.
[0073] The external I / F 50 controls the exchange of various information with devices existing outside the imaging device 12 (hereinafter also referred to as "external devices"). As an example of the external I / F 50, a USB interface can be mentioned. External devices (not shown), such as smart devices, personal computers, servers, USB memories, memory cards, and / or printers, are directly or indirectly connected to the USB interface.
[0074] The communication I / F 52 controls the exchange of information between the CPU 62 and the imaging support device 14 (see FIG. 1) via the network 34 (see FIG. 1). For example, the communication I / F 52 transmits information corresponding to a request from the CPU 62 to the imaging support device 14 via the network 34. Further, the communication I / F 52 receives the information transmitted from the imaging support device 14 and outputs the received information to the CPU 62 via the input / output interface 70.
[0075] As shown in FIG. 3 as an example, a display control processing program 80 is stored in the NVM 64. The CPU 62 reads out the display control processing program 80 from the NVM 64 and executes the read display control processing program 80 on the RAM 66. The CPU 62 performs display control processing according to the display control processing program 80 executed on the RAM 66 (see FIG. 30).
[0076] By executing the display control processing program 80, the CPU 62 operates as an acquisition unit 62A, a generation unit 62B, a transmission unit 62C, a reception unit 62D, and a display control unit 62E. The specific processing contents by the acquisition unit 62A, the generation unit 62B, the transmission unit 62C, the reception unit 62D, and the display control unit 62E will be described later with reference to FIGS. 16, 17, 28, 29, and 30.
[0077] As an example, as shown in FIG. 4, the imaging support device 14 includes a computer 82 and a communication I / F 84. The computer 82 includes a CPU 86, a storage 88, and a memory 90. Here, the computer 82 is an example of the "computer" according to the technology of the present disclosure, the CPU 86 is an example of the "processor" according to the technology of the present disclosure, and the memory 90 is an example of the "memory" according to the technology of the present disclosure.
[0078] The CPU 86, the storage 88, the memory 90, and the communication I / F 84 are connected to a bus 92. In the example shown in FIG. 4, for the sake of illustration, one bus is shown as the bus 92, but a plurality of buses may be used. The bus 92 may be a serial bus or a parallel bus including a data bus, an address bus, a control bus, etc.
[0079] The CPU 86 controls the entire imaging support device 14. The storage 88 is a non-temporary storage medium and is a non-volatile storage device that stores various programs and various parameters, etc. Examples of the storage 88 include an EEPROM, an SSD, and / or an HDD, etc. The memory 90 is a memory in which information is temporarily stored and is used as a work memory by the CPU 86. Examples of the memory 90 include a RAM.
[0080] The communication I / F 84 is connected to the communication I / F 52 of the imaging device 12 via the network 34. The communication I / F 84 manages the exchange of information between the CPU 86 and the imaging device 12. For example, the communication I / F 84 receives the information transmitted from the imaging device 12 and outputs the received information to the CPU 86. Also, the information corresponding to the request from the CPU 62 is transmitted to the imaging device 12 via the network 34.
[0081] Incidentally, the learned model 93 is stored in the storage 88. The learned model 93 is a model generated by training a convolutional neural network, that is, a CNN 96 (see FIG. 5). The CPU 86 performs subject recognition using the learned model 93 based on the captured image 108.
[0082] Here, an example of how the learned model 93 is created (that is, an example of the learning stage) will be described with reference to FIGS. 5 and 6.
[0083] As shown in FIG. 5 as an example, the storage 88 stores a learning execution processing program 94 and a CNN 96. The CPU 86 reads out the learning execution processing program 94 from the storage 88 and operates as a learning stage calculation unit 87A, an error calculation unit 87B, and an adjustment value calculation unit 87C by executing the read learning execution processing program 94.
[0084] In the learning stage, a teacher data supply device 98 is used. The teacher data supply device 98 holds teacher data 100 and supplies the teacher data 100 to the CPU 86. The teacher data 100 includes a plurality of frames of learning images 100A and a plurality of correct answer data 100B. The correct answer data 100B is associated with each of the plurality of frames of learning images 100A one by one.
[0085] The learning stage calculation unit 87A acquires the learning image 100A. The error calculation unit 87B acquires the correct answer data 100B corresponding to the learning image 100A acquired by the learning stage calculation unit 87A.
[0086] CNN96 has an input layer 96A, a plurality of intermediate layers 96B, and an output layer 96C. The learning stage calculation unit 87A extracts feature data indicating the features of the subject specified from the learning image 100A by passing the learning image 100A through the input layer 96A, the plurality of intermediate layers 96B, and the output layer 96C. Here, the features of the subject refer to, for example, the outline, color tone, surface texture, part features, and overall comprehensive features. The output layer 96C determines to which cluster among the plurality of clusters the plurality of feature data extracted through the input layer 96A and the plurality of intermediate layers 96B belong, and outputs a CNN signal 102 indicating the determination result. Here, the cluster refers to collective feature data for each of the plurality of subjects. The determination result is, for example, information indicating the probabilistic correspondence relationship between the feature data and the class 110C (see FIG. 7) specified from the cluster (for example, information including information indicating the probability that the correspondence relationship between the feature data and the class 110C is correct). The class 110C refers to the type of the subject. Note that the class 110C is an example of the "type of subject" of the technology of the present disclosure.
[0087] The correct data 100B is data predetermined as data corresponding to the ideal CNN signal 102 output from the output layer 96C. The correct data 100B includes information in which the feature data and the class 110C are associated.
[0088] The error calculation unit 87B calculates an error 104 between the CNN signal 102 and the correct data 100B. The adjustment value calculation unit 87C calculates a plurality of adjustment values 106 that minimize the error 104 calculated by the error calculation unit 87B. The learning stage calculation unit 87A adjusts a plurality of optimization variables in the CNN96 using the plurality of adjustment values 106 so that the error 104 is minimized. Here, the plurality of optimization variables refer to, for example, a plurality of connection weights and a plurality of offset values included in the CNN96.
[0089] The learning stage calculation unit 87A optimizes the CNN 96 by adjusting a plurality of optimization variables in the CNN 96 using a plurality of adjustment values 106 calculated by the adjustment value calculation unit 87C so that the error 104 is minimized for each of the learning images 100A for a plurality of frames. Then, as shown in FIG. 6 as an example, the CNN 96 is optimized by adjusting a plurality of optimization variables, and thereby the learned model 93 is constructed.
[0090] As shown in FIG. 7 as an example, the learned model 93 has an input layer 93A, a plurality of intermediate layers 93B, and an output layer 93C. The input layer 93A is a layer obtained by optimizing the input layer 96A shown in FIG. 5, the plurality of intermediate layers 93B are layers obtained by optimizing the plurality of intermediate layers 96B shown in FIG. 5, and the output layer 93C is a layer obtained by optimizing the output layer 96C shown in FIG. 5.
[0091] In the imaging support device 14, the subject identification information 110 is output from the output layer 93C by giving the captured image 108 to the input layer 93A of the learned model 93. The subject identification information 110 is information including bounding box position information 110A, objectness score 110B, class 110C, and class score 110D. The bounding box position information 110A is position information capable of specifying the relative position of the bounding box 116 (see FIG. 9) in the captured image 108. Here, the captured image 108 is illustrated, but this is merely an example, and an image based on the captured image 108 (for example, a live view image 130) may be used.
[0092] The objectness score 110B is the probability that an object exists within the bounding box 116. Here, the probability that an object exists within the bounding box 116 is illustrated, but this is merely an example, and it may be a value obtained by finely adjusting the probability that an object exists within the bounding box 116, or any value based on the probability that an object exists within the bounding box 116.
[0093] Class 110C is the type of the subject. Class score 110D is the probability that the object existing within the bounding box 116 belongs to a specific class 110C. Here, the probability that the object existing within the bounding box 116 belongs to a specific class 110C is illustrated, but this is merely an example. It may be a value obtained by finely adjusting the probability that the object existing within the bounding box 116 belongs to a specific class 110C, or any value based on the probability that the object existing within the bounding box 116 belongs to a specific class 110C.
[0094] Also, as an example of a specific class 110C, a specific person, the face of a specific person, a specific automobile, a specific passenger aircraft, a specific bird, and a specific train, etc. may be mentioned. Note that there are a plurality of specific classes 110C, and a class score 110D is assigned to each class 110C.
[0095] When the captured image 108 is input to the learned model 93, as an example, as shown in FIG. 8, the anchor box 112 is applied to the captured image 108. The anchor box 112 is a set of a plurality of virtual frames each having a predetermined height and width. In the example shown in FIG. 8, as an example of the anchor box 112, a first virtual frame 112A, a second virtual frame 112B, and a third virtual frame 112C are shown. The first virtual frame 112A and the third virtual frame 112C have the same shape (rectangle in the example shown in FIG. 8) and the same size as each other. The second virtual frame 112B is a square. Within the captured image 108, the centers of the first virtual frame 112A, the second virtual frame 112B, and the third virtual frame 112C coincide with each other. Also, the outer frame of the captured image 108 is rectangular, and the first virtual frame 112A, the second virtual frame 112B, and the third virtual frame 112C are arranged within the captured image 108 in such a direction that the short side of the first virtual frame 112A, a specific side of the second virtual frame 112B, and the long side of the third virtual frame 112C are parallel to a specific side of the outer frame of the captured image 108.
[0096] In the example shown in FIG. 8, the captured image 108 includes a person image 109 showing a person, and a part of the person image 109 is included in any of the first virtual frame 112A, the second virtual frame 112B, and the third virtual frame 112C. Hereinafter, for convenience of explanation, when it is not necessary to distinguish between the first virtual frame 112A, the second virtual frame 112B, and the third virtual frame 112C, they are referred to as "virtual frames" without reference numerals.
[0097] The captured image 108 is divided into a plurality of virtual cells 114. When the captured image 108 is input to the learned model 93, the CPU 86 positions the center of the anchor box 112 with respect to the center of each cell 114 for each cell 114, and calculates an anchor box score and an anchor class score.
[0098] The anchor box score refers to the probability that the person image 109 is included in the anchor box 112. The anchor box score is a probability obtained based on the probability that the person image 109 is included in the first virtual frame 112A (hereinafter referred to as "the first virtual frame probability"), the probability that the person image 109 is included in the second virtual frame 112B (hereinafter referred to as "the second virtual frame probability"), and the probability that the person image 109 is included in the third virtual frame 112C (hereinafter referred to as "the third virtual frame probability"). For example, the anchor box score is the average value of the first virtual frame probability, the second virtual frame probability, and the third virtual frame probability.
[0099] The anchor class score refers to the probability that the subject (e.g., a person) indicated by the image (in the example shown in FIG. 8, the person image 109) contained in the anchor box 112 belongs to a specific class 110C (e.g., a specific person). The anchor class score is the probability that the subject indicated by the image contained in the first virtual frame 112A belongs to a specific class 110C (hereinafter referred to as the "first virtual frame class probability"), the probability that the subject indicated by the image contained in the second virtual frame 112B belongs to a specific class 110C (hereinafter referred to as the "second virtual frame class probability"), and the probability that the subject indicated by the image contained in the third virtual frame 112C belongs to a specific class 110C (hereinafter referred to as the "third virtual frame class probability"). For example, the anchor class score is the average value of the first virtual frame class probability, the second virtual frame class probability, and the third virtual frame class probability.
[0100] The CPU 86 calculates an anchor box value when the anchor box 112 is applied to the captured image 108. The anchor box value is a value obtained based on cell identification information, an anchor box score, an anchor class score, and an anchor box constant. Here, the cell identification information refers to information for specifying the width of the cell 114, the height of the cell 114, and the position of the cell 114 (e.g., two-dimensional coordinates capable of specifying the position within the captured image 108) in the captured image 108. The anchor box constant refers to a constant predetermined for the type of the anchor box 112.
[0101] The anchor box value is calculated according to the formula "(anchor box value) = (the number of all cells 114 existing in the captured image 108) × {(cell identification information), (anchor box score), (anchor class score)} × (anchor box constant)".
[0102] The CPU 86 calculates the anchor box value by applying the anchor box 112 for each cell 114. Then, the CPU 86 identifies the anchor box 112 having an anchor box value exceeding the anchor box threshold. The anchor box threshold may be a fixed value or a variable value that is changed according to an instruction given to the imaging support device 14 and / or various conditions.
[0103] The CPU 86 identifies the maximum virtual frame probability from the first virtual frame probability, the second virtual frame probability, and the third virtual frame probability regarding the anchor box 112 having an anchor box value exceeding the anchor box threshold. The maximum virtual frame probability refers to the one with the largest value among the first virtual frame probability, the second virtual frame probability, and the third virtual frame probability.
[0104] As shown in FIG. 9 as an example, the CPU 86 determines the virtual frame having the maximum virtual frame probability as the bounding box 116 and deletes the remaining virtual frames. In the example shown in FIG. 9, within the captured image 108, the first virtual frame 112A is determined as the bounding box 116, and the second virtual frame 112B and the third virtual frame 112C are deleted. The bounding box 116 is used as the AF area frame 134. The AF area frame 134 is an example of the "frame" and the "focus frame" according to the technology of the present disclosure. The AF area frame 134 is a frame surrounding the subject. Within the captured image 108, it is a frame surrounding the subject image (the person image 109 in the example shown in FIG. 8) indicating a specific subject (for example, a specific person). The AF area frame 134 refers to a focus frame that defines an area (hereinafter, also referred to as the "focus candidate area") that is a candidate for focusing the focus lens 40B in the AF mode. Here, the AF area frame 134 is illustrated, but this is merely an example and may be used as a focus frame that defines the focus candidate area in the MF mode.
[0105] The CPU 86 infers the class 110C and the class score 110D for each cell 114 within the bounding box 116, and associates the class 110C and the class score 110D with each cell 114. Then, the CPU 86 extracts, as the subject identification information 110, information including the inference results inferred for each cell 114 within the bounding box 116 from the learned model 93.
[0106] In the example shown in FIG. 9, as the subject identification information 110, there is shown information including the bounding box position information 110A regarding the bounding box 116 within the captured image 108, the objectness score 110B regarding the bounding box 116 within the captured image 108, a plurality of classes 110C obtained from each class 110C associated with each cell 114 within the bounding box 116, and a plurality of class scores 110D obtained from each class score 110D associated with each cell 114 within the bounding box 116.
[0107] The objectness score 110B included in the subject identification information 110 is the probability that a subject exists in the bounding box 116 specified from the bounding box position information 110A included in the subject identification information 110. For example, an objectness score 110B of 0% means that there is no subject image (e.g., the person image 109, etc.) indicating a subject in the bounding box 116, and an objectness score 110B of 100% means that a subject image indicating a subject surely exists in the bounding box 116.
[0108] In the subject identification information 110, the class score 110D exists for each class 110C. That is, each of the plurality of classes 110C included in the subject identification information 110 individually has one class score 110D. The class score 110D included in the subject identification information 110 is, for example, the average value of all the class scores 110D associated with all the cells 114 within the bounding box 116 for the corresponding class 110C.
[0109] Of the plurality of Class 110C, the CPU 86 determines that the Class 110C with the highest class score 110D is likely to be the class 110C to which the subject indicated by the subject image existing within the bounding box 116 belongs. The Class 110C determined to be likely is the Class 110C corresponding to the overall highest class score 110D among the respective class scores 110D associated with each cell 114 within the bounding box 116. The overall highest class score 110D refers to, for example, the Class 110C for which the average value of all the class scores 110D associated with all cells 114 is the highest among all the Class 110C associated with all cells 114 within the bounding box 116.
[0110] The method of recognizing a subject using the subject identification information 110 is a subject recognition method using AI (hereinafter, also referred to as the "AI subject recognition method"). As other subject recognition methods, there is a method of recognizing a subject by template matching (hereinafter, also referred to as the "template matching method"). In the template matching method, for example, as shown in FIG. 10, a subject recognition template 118 is used. In the example shown in FIG. 10, the subject recognition template 118, which is an image showing a reference subject (for example, a face created by machine learning or the like as a general face of a person), is used by the CPU 86. The CPU 86 detects the person image 109 included in the captured image 108 by performing a raster scan (in the example shown in FIG. 10, a scan along the straight arrow direction) of the captured image 108 with the subject recognition template 118 while changing the size of the subject recognition template 118. In the template matching method, the CPU 86 measures the difference between the entire subject recognition template 118 and a part of the captured image 108, and calculates the degree of coincidence (hereinafter, also simply referred to as the "degree of coincidence") between the subject recognition template 118 and a part of the captured image 108 based on the measured difference.
[0111] As an example, as shown in FIG. 11, the degree of coincidence is distributed radially from the pixel with the maximum degree of coincidence in the captured image 108. In the example shown in FIG. 11, it is distributed radially from the center of the outer frame 118A of the subject recognition template 118. In this case, the outer frame 118A arranged centering on the pixel having the maximum degree of coincidence in the captured image 108 is used as the AF area frame 134.
[0112] As described above, in the template matching method, the degree of coincidence with the subject recognition template 118 is calculated, whereas in the AI subject recognition method, the probability of the presence of a specific subject (for example, a specific class 110C) is inferred. Therefore, the indices contributing to the evaluation of the subject are completely different between the template matching method and the AI subject recognition method. In the template matching method, a single value similar to the class score 110D, that is, a single value in which the objectness score 110B and the class score 110D are mixed and the tendency of the class score 110D is stronger than the objectness score 110B is used as the degree of coincidence to contribute to the evaluation of the subject. On the other hand, in the AI subject recognition method, two values, the objectness score 110B and the class score 110D, contribute to the evaluation of the subject.
[0113] In the template matching method, even if the degree of matching changes when the conditions of the subject change, the degree of matching is distributed radially from the pixel having the maximum degree of matching within the captured image 108. Therefore, for example, if the degree of matching within the outer frame 118A is equal to or higher than a threshold corresponding to the class threshold used in the comparison with the class score 110D for measuring the reliability of the class score 110D, it may be recognized as a specific class 110C, or it may be recognized as a class 110C other than the specific class 110C. That is, as shown in FIG. 12 as an example, there are cases where the class 110C is correctly recognized and cases where it is misrecognized. If the captured image 108 obtained by capturing by the imaging device 12 is used in the template matching method under a constant environment (for example, an environment where conditions such as the subject and illumination are constant), the subject is considered to be almost correctly recognized. However, if the captured image 108 obtained by capturing by the imaging device 12 is used in the template matching method under a fluctuating environment (for example, an environment where conditions such as the subject and illumination fluctuate), the subject is more likely to be misrecognized than when imaging is performed under a constant environment.
[0114] On the other hand, in the AI subject recognition method, as shown in FIG. 12 as an example, an objectness threshold used in the comparison with the objectness score 110B for measuring the reliability of the objectness score 110B is used. In the AI subject recognition method, if the objectness score 110B is equal to or higher than the objectness threshold, it is determined that a subject image is included within the bounding box 116. Then, in the AI subject recognition method, on the premise that it is determined that a subject image is included within the bounding box 116, if the class score 110D is equal to or higher than the class threshold, it is determined that the subject belongs to a specific class 110C. Even if it is determined that a subject image is included within the bounding box 116, if the class score 110D is less than the class threshold, it is determined that the subject does not belong to a specific class 110C.
[0115] In addition, when subject recognition using the template matching method is performed on a zebra image showing a zebra using a horse image showing a horse as the subject recognition template 118, the zebra image is not recognized. On the other hand, when subject recognition using the AI subject recognition method is performed on the zebra image using the learned model 93 obtained by machine learning using a horse image as the learning image 100A, the objectness score 110B is equal to or higher than the objectness threshold, and the class score 110D is less than the class threshold, and it is determined that the zebra shown by the zebra image does not belong to a specific class 110C (here, as an example, a horse). That is, the zebra shown by the zebra image is recognized as a foreign object other than a specific subject.
[0116] The reason why the subject recognition results are different between the template matching method and the AI subject recognition method in this way is that in the template matching method, the difference in the patterns between a horse and a zebra affects the degree of coincidence, and the degree of coincidence is the only value contributing to subject recognition, whereas in the AI subject recognition method, it is probabilistically determined whether it is a specific subject from a set of partial feature data such as the overall shape of the horse, the face of the horse, four legs, and a tail.
[0117] In this embodiment, the AI subject recognition method is used for the control of the AF area frame 134. Examples of the control of the AF area frame 134 include adjustment of the size of the AF area frame 134 according to the subject, and selection of the division method for dividing the AF area frame 134 according to the subject.
[0118] In the AI subject recognition method, since the bounding box 116 is used as the AF area frame 134, when the objectness score 110B is high (for example, when the objectness score 110B is equal to or higher than the objectness threshold), the size of the AF area frame 134 can be said to be more reliable than the size of the AF area frame 134 when the objectness score 110B is low (for example, when the objectness score 110B is less than the objectness threshold). Conversely, when the objectness score 110B is low, the size of the AF area frame 134 is less reliable than the size of the AF area frame 134 when the objectness score 110B is high, and it can be said that there is a high possibility that the subject is outside the focus candidate area. In this case, for example, the AF area frame 134 is expanded or the AF area frame 134 is deformed.
[0119] By the way, in the conventionally known technique, when dividing the AF area frame 134, the number of divisions is changed according to the size of the subject image surrounded by the AF area frame 134. However, when the number of divisions is changed according to the size of the subject image, the dividing line 134A1 (see FIGS. 22 and 25 to 27), that is, the boundary line between the plurality of divided areas 134A (see FIGS. 22 and 25 to 27) obtained by division may overlap with a portion corresponding to an important position of the subject. Generally, it is expected that an important position of a subject has a higher tendency to be desired as a focusing target than a location that has nothing to do with the important position of the subject.
[0120] For example, as shown in FIG. 13, when the subject is a person 120, one of the important positions of the subject may be the pupil of the person 120. When the subject is a bird 122, one of the important positions of the subject may be the pupil of the bird 122. When the subject is an automobile 124, one of the important positions of the subject may be the position of the emblem on the front side of the automobile 124.
[0121] Also, when the imaging device 12 images a subject from the front side, the length of the subject in the depth direction, that is, the length from the reference position of the subject to the important position, varies depending on the subject. The reference position of the subject when the imaging device 12 images the subject from the front side is, for example, the tip of the nose of the person 120 when the subject is the person 120, the tip of the beak of the bird 122 when the subject is the bird 122, and between the roof and the windshield of the automobile 124 when the subject is the automobile 124. Thus, although the length from the reference position of the subject to the important position varies depending on the subject, when imaging is performed with the focus on the reference position of the subject, the image indicating the important position of the subject in the captured image 108 becomes blurred. In this case, if the portion corresponding to the important position of the subject is made to fit within the divided area 134A of the AF area frame 134 without missing, compared to the case where the portion corresponding to the important position of the subject does not fit within the divided area 134A of the AF area frame 134 and the portion corresponding to the reference position of the subject fits within the divided area 134A of the AF area frame 134, it becomes easier to focus on the important position of the subject in the captured image 108.
[0122] Therefore, in order to make the portion corresponding to the important position of the subject fit within the divided area of the AF area frame 134 without missing, as an example, as shown in FIG. 14, the storage 88 of the imaging support device 14 stores, in addition to the learned model 93, an imaging support processing program 126 and a division method derivation table 128. The imaging support processing program 126 is an example of the "program" according to the technology of the present disclosure.
[0123] As an example, as shown in FIG. 15, the CPU 86 reads the imaging support processing program 126 from the storage 88 and executes the read imaging support processing program 126 on the memory 90. The CPU 86 performs imaging support processing according to the imaging support processing program 126 executed on the memory 90 (see also FIGS. 31A and 31B).
[0124] The CPU 86 performs imaging support processing to obtain the class 110C based on the captured image 108, and outputs information indicating a division method for dividing a division area that can distinguish the subject from other areas in the imaging range according to the obtained class 110C. Further, the CPU 86 obtains the class 110C based on the subject identification information 110 output from the learned model 93 by providing the learned model 93 with the captured image 108. Here, the subject identification information 110 is an example of the "output result" according to the technology of the present disclosure.
[0125] The CPU 86 operates as a reception unit 86A, an execution unit 86B, a determination unit 86C, a generation unit 86D, an extension unit 86E, a division unit 86F, a derivation unit 86G, and a transmission unit 86H by executing the imaging support processing program 126. Specific processing contents by the reception unit 86A, the execution unit 86B, the determination unit 86C, the generation unit 86D, the extension unit 86E, the division unit 86F, the derivation unit 86G, and the transmission unit 86H will be described later with reference to FIGS. 17 to 28, FIGS. 31A, and 31B.
[0126] As an example, as shown in FIG. 16, in the imaging device 12, when the captured image 108 is stored in the image memory 46, the acquisition unit 62A acquires the captured image 108 from the image memory 46. The generation unit 62B generates a live view image 130 based on the captured image 108 acquired by the acquisition unit 62A. The live view image 130 is, for example, an image obtained by thinning out pixels from the captured image 108 according to a predetermined rule. In the example shown in FIG. 16, an image including an airliner image 132 showing an airliner is shown as the live view image 130. The transmission unit 62C transmits the live view image 130 to the imaging support device 14 via the communication I / F 52.
[0127] As an example, as shown in FIG. 17, in the imaging support device 14, the receiving unit 86A receives the live view image 130 transmitted from the transmitting unit 62C of the imaging device 12. The execution unit 86B executes subject recognition processing based on the live view image 130 received by the receiving unit 86A. Here, the subject recognition processing refers to processing that detects a subject by an AI subject recognition method (that is, processing that determines the presence or absence of a subject) and processing that identifies the type of the subject by the AI subject recognition method.
[0128] The execution unit 86B executes subject recognition processing using the learned model 93 in the storage 88. The execution unit 86B extracts subject identification information 110 from the learned model 93 by providing the live view image 130 to the learned model 93.
[0129] As an example, as shown in FIG. 18, the execution unit 86B stores the subject identification information 110 in the memory 90. The determination unit 86C acquires the objectness score 110B from the subject identification information 110 in the memory 90, and determines the presence or absence of a subject by referring to the acquired objectness score 110B. For example, if the objectness score 110B is a value greater than "0", it is determined that a subject exists, and if the objectness score 110B is "0", it is determined that no subject exists. When the determination unit 86C determines that a subject exists, the CPU 86 performs first determination processing. When the determination unit 86C determines that no subject exists, the CPU 86 performs first division processing.
[0130] As an example, as shown in FIG. 19, in the first determination process, the determination unit 86C acquires the objectness score 110B from the subject identification information 110 in the memory 90, and determines whether the acquired objectness score 110B is equal to or greater than the objectness threshold. Here, if the objectness score 110B is equal to or greater than the objectness threshold, the second determination process is performed by the CPU 86. If the objectness score 110B is less than the objectness threshold, the second division process is performed by the CPU 86. Note that the objectness score 110B equal to or greater than the objectness threshold is an example of the "output result" and the "value based on the probability that the object within the bounding box applied to the image belongs to a specific class" according to the technology of the present disclosure. In addition, the objectness threshold when the imaging support process shifts to the second division process is an example of the "first threshold" according to the technology of the present disclosure.
[0131] As an example, as shown in FIG. 20, in the second determination process, the determination unit 86C acquires the class score 110D from the subject identification information 110 in the memory 90, and determines whether the acquired class score 110D is equal to or greater than the class threshold. Here, if the class score 110D is equal to or greater than the class threshold, the third division process is performed by the CPU 86. If the class score 110D is less than the class threshold, the first division process is performed by the CPU 86. Note that the class score 110D equal to or greater than the class threshold is an example of the "output result" and the "value equal to or greater than the second threshold among the values based on the probability that the object belongs to a specific class" according to the technology of the present disclosure. In addition, the class threshold is an example of the "second threshold" according to the technology of the present disclosure.
[0132] As an example, as shown in FIG. 21, in the first division process, the generation unit 86D acquires the bounding box position information 110A from the subject identification information 110 in the memory 90, and generates the AF area frame 134 based on the acquired bounding box position information 110A. The division unit 86F calculates the area of the AF area frame 134 generated by the generation unit 86D, and divides the AF area frame 134 by a division method according to the calculated area. After the AF area frame 134 is divided by the division method according to the area, the CPU 86 performs the AF area frame-included image generation and transmission process.
[0133] The division method defines the number of divisions for dividing a divided area (hereinafter also simply referred to as a "divided area") that is distinguishable from other areas in the captured image 108 by being surrounded by the AF area frame 134, that is, the number of divisions for dividing the AF area frame 134. The divided area is defined by the vertical direction and the horizontal direction orthogonal to the vertical direction. The vertical direction corresponds to the row direction among the row direction and the column direction that define the captured image 108, and the horizontal direction corresponds to the column direction. The division method defines the number of divisions in the vertical direction and the number of divisions in the horizontal direction. Note that the row direction is an example of the "first direction" according to the technology of the present disclosure, and the column direction is an example of the "second direction" according to the technology of the present disclosure.
[0134] As an example, as shown in FIG. 22, in the first division process, the division unit 86F changes the number of divisions of the AF area frame 134, that is, the number of divided areas 134A of the AF area frame 134, within a predetermined number range according to the area of the AF area frame 134 (for example, within a range where the lower limit value is 2 and the upper limit value is 30). That is, the larger the area of the divided area 134A, the more the division unit 86F increases the number of divided areas 134A within the predetermined number range. Specifically, the larger the area of the divided area 134A, the more the division unit 86F increases the number of divisions in the vertical direction and the horizontal direction.
[0135] Further, the dividing unit 86F decreases the number of divided areas 134A within a predetermined number range as the number of divided areas 134A becomes smaller. Specifically, the dividing unit 86F decreases the number of divisions in the vertical and horizontal directions as the number of divided areas 134A becomes smaller.
[0136] As an example, as shown in FIG. 23, in the second division process, the generation unit 86D acquires the bounding box position information 110A from the subject identification information 110 in the memory 90, and generates the AF area frame 134 based on the acquired bounding box position information 110A. The expansion unit 86E expands the AF area frame 134 at a predetermined magnification (for example, 1.25 times). Note that the predetermined magnification may be a variable value that is changed according to an instruction given to the imaging support device 14 and / or various conditions, or may be a fixed value.
[0137] In the second division process, the dividing unit 86F calculates the area of the AF area frame 134 expanded by the expansion unit 86E, and divides the AF area frame 134 by a division method corresponding to the calculated area. After the AF area frame 134 is divided by the division method corresponding to the area, the CPU 86 performs an image generation and transmission process with the AF area frame.
[0138] As an example, as shown in FIG. 24, the division method derivation table 128 in the storage 88 is a table that associates the class 110C with the division method. There are a plurality of classes 110C, and the division method is uniquely determined for each of the plurality of classes 110C. In the example shown in FIG. 24, the division method A is associated with the class A, the division method B is associated with the class B, and the division method C is associated with the class C.
[0139] The derivation unit 86G acquires the class 110C corresponding to the class score 110D equal to or higher than the class threshold from the subject identification information 110 in the memory 90, and derives the division method corresponding to the acquired class 110C from the division method derivation table 128. The generation unit 86D acquires the bounding box position information 110A from the subject identification information 110 in the memory 90, and generates the AF area frame 134 based on the acquired bounding box position information 110A. The division unit 86F divides the AF area frame 134 generated by the generation unit 86D by the division method derived by the derivation unit 86G. The fact that the AF area frame 134 is divided by the division method derived by the derivation unit 86G means that the area surrounded by the AF area frame 134 as a division area that can be distinguished from other areas of the imaging range to identify the subject is divided by the division method derived by the derivation unit 86G.
[0140] Generally, for a subject, a three-dimensional subject image can be obtained when the subject is imaged from an oblique direction rather than from the front, compared to when it is imaged from the front. Thus, when imaging is performed in a composition that gives a sense of depth to a person viewing the live view image 130 (hereinafter also referred to as "viewer"), it is important to make the number of vertical divisions of the AF area frame 134 more than the number of horizontal divisions or the number of horizontal divisions more than the number of vertical divisions according to the class 110C.
[0141] Therefore, the number of vertical and horizontal divisions of the AF area frame 134 by the division method derived by the derivation unit 86G is defined based on the composition of the subject image surrounded by the AF area frame 134 in the live view image 130.
[0142] As an example, as shown in FIG. 25, in the case of a composition in which an airliner is imaged by the imaging device 12 from the lower diagonal side, the dividing unit 86F divides the AF area frame 134 so that the number of divisions in the horizontal direction of the AF area frame 134 is larger than the number of divisions in the vertical direction of the AF area frame 134 surrounding the airliner image 132 in the live view image 130. In the example shown in FIG. 25, the AF area frame 134 is divided into 24 in the horizontal direction, and the AF area frame 134 is divided into 10 in the vertical direction. Also, the AF area frame 134 is equally divided in each of the horizontal and vertical directions. In this way, the divided area surrounded by the AF area frame 134 is divided into 240 (= 24 × 10) divided areas 134A. In the example shown in FIG. 25, the 240 divided areas 134A are an example of the "plurality of divided areas" according to the technology of the present disclosure.
[0143] Also, as an example, as shown in FIG. 26, a face image 120A obtained by imaging the face of a person 120 (see FIG. 13) from the front side is surrounded by an AF area frame 134. When the face image 120A is located at a first predetermined position in the AF area frame 134, the dividing unit 86F divides the AF area frame 134 into an even number (in the example shown in FIG. 26, "8") of equal parts in the horizontal direction. The first predetermined position refers to a position where a portion corresponding to the nose of the person 120 is located at the center in the horizontal direction of the AF area frame 134. In this way, by dividing the AF area frame 134 into an even number of equal parts in the horizontal direction, compared with the case where it is divided into an odd number of equal parts in the horizontal direction, the portion 120A1 corresponding to the pupil in the face image 120A within the AF area frame 134 is less likely to be cut off and more likely to fit within one divided area 134A. In other words, this means that the portion 120A1 corresponding to the pupil in the face image 120A is less likely to cross the dividing line 134A1 compared with the case where the AF area frame 134 is divided into an odd number of equal parts in the horizontal direction. Note that the AF area frame 134 is also divided in the vertical direction. The number of divisions in the vertical direction is less than the number of divisions in the horizontal direction (in the example shown in FIG. 26, "4"). In the example shown in FIG. 26, the AF area frame 134 is divided in the vertical direction, but this is only an example, and it may not be divided in the vertical direction. The size of the divided area 134A is a size derived in advance by a simulator or the like as a size that allows the portion 120A1 corresponding to the pupil in the face image 120A to fit within one divided area 134A without being cut off.
[0144] Also, as an example, as shown in FIG. 27, an automobile image 124A obtained by imaging an automobile 124 (see FIG. 13) from the front side is surrounded by an AF area frame 134. When the automobile image 124A is located at a second predetermined position of the AF area frame 134, the dividing unit 86F equally divides the AF area frame 134 in the horizontal direction into an odd number (in the example shown in FIG. 27, "7"). The second predetermined position refers to a position where the center of the front view of the automobile (e.g., the emblem) is located at the center in the horizontal direction of the AF area frame 134. In this way, by equally dividing the AF area frame 134 into an odd number in the horizontal direction, compared with the case of equally dividing it into an even number in the horizontal direction, the portion 124A1 corresponding to the emblem in the automobile image 124A within the AF area frame 134 is less likely to be missing and is more likely to fit into one divided area 134A.
[0145] In other words, this means that, compared with the case where the AF area frame 134 is equally divided into an even number in the horizontal direction, the portion 124A1 corresponding to the emblem in the automobile image 124A is less likely to cross the dividing line 134A1.
[0146] Note that the AF area frame 134 is also equally divided in the vertical direction. The number of divisions in the vertical direction is less than the number of divisions in the horizontal direction (in the example shown in FIG. 27, "3"). In the example shown in FIG. 27, the AF area frame 134 is divided in the vertical direction, but this is only an example, and it may not be divided in the vertical direction. The size of the divided area 134A is a size derived in advance by a simulator or the like as a size that allows the portion 124A1 corresponding to the emblem in the automobile image 124A to fit into one divided area 134A without being missing.
[0147] As an example, as shown in FIG. 28, in the AF area frame-added image generation and transmission process, the generation unit 86D generates an AF area frame-added live view image 136 using the live view image 130 received by the reception unit 86A and the AF area frame 134 divided into a plurality of divided areas 134A by the division unit 86F. The AF area frame-added live view image 136 is an image in which the AF area frame 134 is superimposed on the live view image 130. The AF area frame 134 is arranged at a position surrounding the aircraft image 132 within the live view image 130. The transmission unit 86H transmits the AF area frame-added live view image 136 generated by the generation unit 86D to the imaging device 12 via the communication I / F 84 (see FIG. 4) as information for causing the display 28 of the imaging device 12 to display the live view image 130 based on the captured image 108 and to display the AF area frame 134 within the live view image 130. Since the division method derived by the derivation unit 86G can be specified from the AF area frame 134 divided into a plurality of divided areas 134A by the division unit 86F, the transmission of the AF area frame-added live view image 136 to the imaging device 12 means the output of information indicating the division method to the imaging device 12. Note that the AF area frame 134 divided into a plurality of divided areas 134A by the division unit 86F is an example of the "information indicating the division method" according to the technology of the present disclosure.
[0148] In the imaging device 12, the reception unit 62D receives the AF area frame-added live view image 136 transmitted from the transmission unit 86H via the communication I / F 52 (see FIG. 2).
[0149] As an example, as shown in FIG. 29, in the imaging device 12, the display control unit 62E causes the display 28 to display the live view image 136 with the AF area frame received by the receiving unit 62D. For example, the live view image 136 with the AF area frame has an AF area frame 134 displayed thereon, and the AF area frame 134 is divided into a plurality of divided areas 134A. For example, the plurality of divided areas 134A are selectively specified according to an instruction (e.g., a touch operation) received by the touch panel 30. The CPU 62 performs AF control so as to focus on the real space area corresponding to the specified divided area 134A. The CPU 62 performs AF control on the real space area corresponding to the specified divided area 134A. That is, the CPU 62 calculates the in-focus position for the real space area corresponding to the specified divided area 134A, and moves the focus lens 40B to the calculated in-focus position. Note that the AF control may be control of a so-called phase difference AF method or control of a contrast AF method. Also, an AF method based on a distance measurement result using the parallax of a pair of images obtained from a stereo camera, or an AF method using a distance measurement result of a TOF method using laser light or the like may be employed.
[0150] Next, the operation of the imaging system 10 will be described with reference to FIGS. 30 to 31B.
[0151] FIG. 30 shows an example of the flow of display control processing performed by the CPU 62 of the imaging device 12. In the display control processing shown in FIG. 30, first, in step ST100, the acquisition unit 62A determines whether the captured image 108 is stored in the image memory 46. In step ST100, if the captured image 108 is not stored in the image memory 46, the determination is negative, and the display control processing proceeds to step ST112. In step ST100, if the captured image 108 is stored in the image memory 46, the determination is affirmative, and the display control processing proceeds to step ST102.
[0152] In step ST102, acquisition unit 62A acquires captured image 108 from image memory 46. After the process of step ST102 is executed, the display control process proceeds to step ST104.
[0153] In step ST104, generation unit 62B generates live view image 130 based on captured image 108 acquired in step ST102. After the process of step ST104 is executed, the display control process proceeds to step ST106.
[0154] In step ST106, transmission unit 62C transmits live view image 130 generated in step ST104 to imaging support device 14 via communication I / F 52. After the process of step ST106 is executed, the display control process proceeds to step ST108.
[0155] In step ST108, reception unit 62D determines whether AF area frame-added live view image 136 transmitted from imaging support device 14 has been received by communication I / F 52 due to the execution of the process of step ST238 included in the imaging support process shown in FIG. 31B. In step ST108, if AF area frame-added live view image 136 has not been received by communication I / F 52, the determination is negative and the determination in step ST108 is made again. In step ST108, if AF area frame-added live view image 136 has been received by communication I / F 52, the determination is affirmative and the display control process proceeds to step ST110.
[0156] In step ST110, display control unit 62E causes display 28 to display AF area frame-added live view image 136 received in step ST108. After the process of step ST110 is executed, the display control process proceeds to step ST112.
[0157] In step ST112, the display control unit 62E determines whether or not a condition for ending the display control process (hereinafter also referred to as the "display control process end condition") is satisfied. Examples of the display control process end condition include a condition that the imaging mode set for the imaging device 12 has been canceled, or a condition that an instruction to end the display control process has been received by the reception device 76. In step ST112, if the display control process end condition is not satisfied, the determination is negative and the display control process proceeds to step ST100. In step ST112, if the display control process end condition is satisfied, the determination is positive and the display control process ends.
[0158] FIGS. 31A and 31B show an example of the flow of the imaging support process performed by the CPU 86 of the imaging support device 14. Note that the flow of the imaging support process shown in FIGS. 31A and 31B is an example of the "imaging support method" according to the technology of the present disclosure.
[0159] In the imaging support process shown in FIG. 31A, in step ST200, the reception unit 86A determines whether or not the live view image 130 transmitted by the execution of the process of step ST106 shown in FIG. 30 has been received by the communication I / F 84. In step ST200, if the live view image 130 has not been received by the communication I / F 84, the determination is negative and the imaging support process proceeds to step ST240 shown in FIG. 31B. In step ST200, if the live view image 130 has been received by the communication I / F 84, the determination is positive and the imaging support process proceeds to step ST202.
[0160] In step ST202, the execution unit 86B executes subject recognition processing based on the live view image 130 received in step ST200. The execution unit 86B executes subject recognition processing using the learned model 93. That is, the execution unit 86B gives the live view image 130 to the learned model 93, extracts subject specific information 110 from the learned model 93, and stores the extracted subject specific information 110 in the memory 90. After the process of step ST202 is executed, the imaging support process proceeds to step ST204.
[0161] In step ST204, the determination unit 86C determines whether a subject has been detected by the subject recognition processing. That is, the determination unit 86C acquires the objectness score 110B from the subject specific information 110 in the memory 90, and determines the presence or absence of the subject with reference to the acquired objectness score 110B. In step ST204, if the subject does not exist, the determination is negative and the imaging support process proceeds to step ST206. In step ST204, if the subject exists, the determination is positive and the imaging support process proceeds to step ST212.
[0162] In step ST206, the generation unit 86D acquires the bounding box position information 110A from the subject specific information 110 in the memory 90. After the process of step ST206 is executed, the imaging support process proceeds to step ST208.
[0163] In step ST208, the generation unit 86D generates the AF area frame 134 based on the bounding box position information 110A acquired in step ST206. After the process of step ST208 is executed, the imaging support process proceeds to step ST210.
[0164] In step ST210, the dividing unit 86F calculates the area of the AF area frame 134 generated in step ST208 or step ST220, and divides the AF area frame by a dividing method according to the calculated area (see FIG. 22). After the process of step ST210 is executed, the imaging support process proceeds to step ST236 shown in FIG. 31B.
[0165] In step ST212, the determination unit 86C acquires the objectness score 110B from the subject identification information 110 in the memory 90. After the process of step ST212 is executed, the imaging support process proceeds to step ST214.
[0166] In step ST214, the determination unit 86C determines whether or not the objectness score 110B acquired in step ST212 is equal to or greater than the objectness threshold. In step ST214, if the objectness score 110B is less than the objectness threshold, the determination is negative, and the imaging support process proceeds to step ST216. In step ST214, if the objectness score 110B is equal to or greater than the objectness threshold, the determination is positive, and the imaging support process proceeds to step ST222. Note that the objectness threshold when the imaging support process proceeds to step ST216 is an example of the "third threshold" according to the technology of the present disclosure.
[0167] In step ST216, the generation unit 86D acquires the bounding box position information 110A from the subject identification information 110 in the memory 90. After the process of step ST216 is executed, the imaging support process proceeds to step ST218.
[0168] In step ST218, the generation unit 86D generates the AF area frame 134 based on the bounding box position information 110A acquired in step ST216. After the process of step ST218 is executed, the imaging support process proceeds to step ST220.
[0169] In step ST220, the extension unit 86E expands the AF area frame 134 at a predetermined magnification (for example, 1.25 times). After the process of step ST220 is executed, the imaging support process proceeds to step ST210.
[0170] In step ST222, the determination unit 86C acquires the class score 110D from the subject identification information 110 in the memory 90. After the process of step ST222 is executed, the imaging support process proceeds to step ST224.
[0171] In step ST224, the determination unit 86C determines whether the class score 110D acquired in step ST222 is equal to or greater than the class threshold. In step ST224, if the class score 110D is less than the class threshold, the determination is negative and the imaging support process proceeds to step ST206. In step ST224, if the class score 110D is equal to or greater than the class threshold, the determination is positive and the imaging support process proceeds to step ST226 shown in FIG. 31B.
[0172] In step ST226 shown in FIG. 31B, the derivation unit 86G acquires the class 110C from the subject identification information 110 in the memory 90. The class 110C acquired here is the class 110C having the largest class score 110D among the classes 110C included in the subject identification information 110 in the memory 90.
[0173] In step ST228, the derivation unit 86G refers to the division method derivation table 128 in the storage 88 and derives a division method corresponding to the class 110C acquired in step ST226. That is, the derivation unit 86G derives a division method corresponding to the class 110C acquired in step ST226 from the division method derivation table 128.
[0174] In step ST230, the generation unit 86D acquires the bounding box position information 110A from the subject identification information 110 in the memory 90. After the process of step ST230 is executed, the imaging support process proceeds to step ST232.
[0175] In step ST232, the generation unit 86D generates the AF area frame 134 based on the bounding box position information 110A acquired in step ST230. After the process of step ST232 is executed, the imaging support process proceeds to step ST234.
[0176] In step ST234, the division unit 86F divides the AF area frame 134 generated in step ST232 by the division method derived in step ST228 (see FIGS. 25 to 27).
[0177] In step ST236, the generation unit 86D generates the live view image 136 with the AF area frame using the live view image 130 received in step ST200 and the AF area frame 134 obtained in step ST234 (see FIG. 28). After the process of step ST236 is executed, the imaging support process proceeds to step ST238.
[0178] In step ST238, the transmission unit 86H transmits the live view image 136 with the AF area frame generated in step ST236 to the imaging device 12 via the communication I / F 84 (see FIG. 4). After the process of step ST238 is executed, the imaging support process proceeds to step ST240.
[0179] In step ST240, the determination unit 86C determines whether or not a condition for ending the imaging support process (hereinafter also referred to as the "imaging support process end condition") is satisfied. As an example of the imaging support process end condition, there are conditions such as the imaging mode set for the imaging device 12 being canceled, or an instruction to end the imaging support process being received by the reception device 76. In step ST240, if the imaging support process end condition is not satisfied, the determination is negative and the imaging support process proceeds to step ST200 shown in FIG. 31A. In step ST240, if the imaging support process end condition is satisfied, the determination is affirmative and the imaging support process ends.
[0180] As described above, in the imaging support device 14, the type of the subject is acquired by the derivation unit 86G as the class 110C based on the captured image 108. The AF area frame 134 surrounds a division area that can distinguish the subject from other areas in the imaging range. The AF area frame 134 is divided by the division unit 86F according to the division method derived from the division method derivation table 128. This means that the division area surrounded by the AF area frame 134 is divided according to the division method derived from the division method derivation table 128. The division method for dividing the AF area frame 134 is derived from the division method derivation table 128 according to the class 110C acquired by the derivation unit 86G. The AF area frame 134 divided by the division method derived from the division method derivation table 128 is included in the live view image 136 with the AF area frame, and the live view image 136 with the AF area frame is transmitted to the imaging device 12 by the transmission unit 86H. Since the division method can be specified from the AF area frame 134 included in the live view image 136 with the AF area frame, the transmission of the live view image 136 with the AF area frame to the imaging device 12 means that the information indicating the division method derived from the division method derivation table 128 is also transmitted to the imaging device 12. Therefore, according to this configuration, compared with the case where the AF area frame 134 is always divided by a fixed division method regardless of the type of the subject, the control (for example, AF control) regarding the imaging of the subject by the image sensor 20 of the imaging device 12 can be performed with high accuracy.
[0181] Also, in the imaging support device 14, the class 110C is inferred based on the subject identification information 110, which is the output result output from the learned model 93 when the live view image 130 is given to the learned model 93. Therefore, according to this configuration, compared with the case where the class 110C is acquired only by the template matching method without using the learned model 93, the class 110C to which the subject belongs can be accurately grasped.
[0182] In addition, in the imaging support device 14, the learned model 93 assigns an object (for example, a subject image such as an airliner image 132, a face image 120A, or an automobile image 124A) within the bounding box 116 applied to the live view image 130 to the corresponding class 110C. The subject identification information 110, which is the output result from the learned model 93, includes a class score 110D that is a value based on the probability that the object within the bounding box 116 applied to the live view image 130 belongs to a specific class 110C. Therefore, according to this configuration, compared with the case where the class 110C is inferred only by the template matching method without using the class score 110D obtained from the learned model 93, the class 110C to which the subject belongs can be accurately grasped.
[0183] Also, in the imaging support device 14, the class score 110D is obtained as a value based on the probability that the object within the bounding box 116 belongs to a specific class 110C when the objectness score 110B is equal to or greater than the objectness threshold. Then, the class 110C to which the subject belongs is inferred based on the obtained class score 110D. Therefore, according to this configuration, compared with the case where the class 110C is inferred based on the class score 110D obtained as a value based on the probability that the object within the bounding box 116 belongs to a specific class 110C regardless of whether an object exists within the bounding box 116, the class 110C to which the subject belongs, that is, the type of the subject, can be accurately specified.
[0184] Also, in the imaging support device 14, the class 110C to which the subject belongs is inferred based on the class score 110D that is equal to or greater than the class threshold. Therefore, according to this configuration, compared with the case where the class 110C to which the subject belongs is inferred based on the class score 110D that is less than the class threshold, the class 110C to which the subject belongs, that is, the type of the subject, can be accurately specified.
[0185] In addition, in the imaging support device 14, when the objectness score 110B is less than the objectness threshold value, the AF area frame 134 is expanded. Therefore, according to this configuration, the probability that an object exists within the AF area frame 134 can be increased as compared with the case where the AF area frame 134 always has a constant size.
[0186] In addition, in the imaging support device 14, the number of divisions for dividing the AF area frame 134 by the division method derived from the division method derivation table 128 by the derivation unit 86G is defined. Therefore, according to this configuration, as compared with the case where the number of divisions of the AF area frame 134 is always constant, control regarding imaging by the image sensor 20 can be accurately performed for the location within the subject intended by the user or the like.
[0187] In addition, in the imaging support device 14, the number of vertical divisions and the number of horizontal divisions of the AF area frame 134 by the division method derived from the division method derivation table 128 by the derivation unit 86G are defined. Therefore, according to this configuration, as compared with the case where only the number of divisions in one of the vertical direction and the horizontal direction of the AF area frame 134 is defined, control regarding imaging by the image sensor 20 can be accurately performed for the location within the subject intended by the user or the like.
[0188] In addition, in the imaging support device 14, the number of vertical divisions and the number of horizontal divisions of the AF area frame 134 divided by the division method derived from the division method derivation table 128 by the derivation unit 86G are defined based on the composition of the subject image (for example, the airliner image 132 shown in FIG. 25) in the live view image 130. Therefore, according to this configuration, as compared with the case where the number of vertical divisions and the number of horizontal divisions of the AF area frame 134 are defined regardless of the composition of the subject image in the live view image 130, control regarding imaging by the image sensor 20 can be accurately performed for the location within the subject intended by the user or the like.
[0189] Further, in the imaging support device 14, a live view image 136 with an AF area frame is transmitted to the imaging device 12. Then, the live view image 136 with the AF area frame is displayed on the display 28 of the imaging device 12. Therefore, according to this configuration, the user or the like can recognize the subject to be controlled regarding imaging by the image sensor 20 of the imaging device 12.
[0190] In addition, in the above embodiment, a form example has been described in which the live view image 136 including the AF area frame 134 divided according to the division method is transmitted to the imaging device 12 by the transmission unit 86H. However, the technology of the present disclosure is not limited to this. For example, the CPU 86 of the imaging support device 14 may transmit information contributing to moving the focus lens 40B to the focusing position according to the class 110C acquired from the subject identification information 110 to the imaging device 12 via the communication I / F 84. Specifically, the CPU 86 of the imaging support device 14 may transmit, via the communication I / F 84, information contributing to moving the focus lens 40B to the focusing position corresponding to the divided area 134A corresponding to the class 110C acquired from the subject identification information 110 among the plurality of divided areas 134A to the imaging device 12.
[0191] In this case, as an example, as shown in FIG. 32, a face image 120A obtained by imaging the face of a person 120 (see FIG. 13) from the front side is surrounded by an AF area frame 134. When the face image 120A is located at a first predetermined position of the AF area frame 134, similarly to the above-described embodiment, the dividing unit 86F divides the AF area frame 134 into an even number of equal parts in the horizontal direction. Then, the dividing unit 86F, similarly to the deriving unit 86G, acquires the class 110C from the subject identification information 110 in the memory 90. Then, the dividing unit 86F detects a pupil image 120A2 indicating a pupil from the face image 120A, and assigns focus priority position information 150 to a divided area 134A including the detected pupil image 120A2. The focus priority position information 150 is information that contributes to moving the focus lens 40B to a focusing position that is in focus with the pupil indicated by the pupil image 120A2. The focus priority position information 150 includes position information that can specify the relative position of the divided area 134A within the AF area frame 134. Note that the focus priority position information 150 is an example of "information that contributes to moving the focus lens to a focusing position corresponding to a divided region according to the type" and "information that contributes to moving the focus lens to a focusing position according to the type" according to the technology of the present disclosure.
[0192] In the imaging support process shown in FIG. 33, step ST300 is inserted between step ST234 and step ST236. In step ST300, the dividing unit 86F detects a pupil image 120A2 indicating a pupil from the face image 120A, and assigns focus priority position information 150 to a divided area 134A including the detected pupil image 120A2 among the plurality of divided areas 134A included in the AF area frame 134 divided in step ST234.
[0193] As a result, the live view image 136 with the AF area frame includes the AF area frame 134 to which the focus priority position information 150 is assigned. The live view image 136 with the AF area frame is transmitted to the imaging device 12 by the transmitting unit 86H (see step ST238 shown in FIG. 33).
[0194] The CPU 62 of the imaging device 12 receives the live view image 136 with the AF area frame transmitted by the transmission unit 86H, and performs AF control according to the focus priority position information 150 of the AF area frame 134 included in the received live view image 136 with the AF area frame. That is, the CPU 62 moves the focus lens 40B to the in-focus position that is in focus with respect to the position of the pupil specified from the focus priority position information 150. In this case, compared with the case where the focus lens 40B is moved only to the in-focus position corresponding to the same location in the imaging range all the time, it is possible to perform high-precision focusing on the important position of the subject.
[0195] Note that the form in which the AF control is performed according to the focus priority position information 150 is only an example, and the important position of the subject may be notified to the user or the like using the focus priority position information 150. In this case, for example, the divided area 134A to which the focus priority position information 150 is given in the AF area frame 134 displayed on the display 28 of the imaging device 12 may be displayed in a manner distinguishable from the other divided areas 134A (for example, a manner of emphasizing the frame of the divided area 134A). Thereby, it becomes possible for the user or the like to visually recognize which position in the AF area frame 134 is the important position.
[0196] In the above embodiment, a form example in which the bounding box 116 is directly used as the AF area frame 134 has been described, but the technology of the present disclosure is not limited to this. For example, the CPU 62 may change the size of the AF area frame 134 according to the class score 110D obtained from the subject identification information 110.
[0197] In this case, as an example shown in FIG. 34, in the imaging support process, step ST400 is inserted between step ST232 and step ST234. In step ST400, the generation unit 86D acquires the class score 110D from the subject identification information 110 in the memory 90, and changes the size of the AF area frame 134 generated in step ST232 according to the acquired class score 110D. In the example shown in FIG. 35, a morphological example is shown in which the size of the AF area frame 134 is reduced according to the class score 110D. In this case, among the class scores 110D (see FIG. 9) given to each cell 114 (see FIG. 9) in the AF area frame 134 before reduction, it may be reduced to a size that surrounds the area where the class scores 110D equal to or higher than the reference value are distributed. Note that the reference value may be a variable value that is changed according to an instruction given to the imaging support device 14 and / or various conditions, or may be a fixed value.
[0198] According to the examples shown in FIGS. 34 and 35, since the size of the AF area frame 134 is changed according to the class score 110D acquired from the subject identification information 110, the accuracy of control related to imaging by the image sensor 20 for the subject can be improved as compared with the case where the size of the AF area frame 134 is always constant.
[0199] In the above embodiment, an example has been described in which the imaging device 12 performs AF control using the AF area frame 134 included in the live view image 136 with the AF area frame. However, the technology of the present disclosure is not limited to this, and the imaging device 12 may perform control related to imaging other than AF control using the AF area frame 134 included in the live view image 136 with the AF area frame. For example, as shown in FIG. 36, control related to imaging other than AF control may include custom control. Custom control recommends changing the control content (hereinafter, also simply referred to as "control content") of control related to imaging according to the subject, and is control for changing the control content according to a given instruction. The CPU 86 acquires the class 110C from the subject identification information 110 in the memory 90, and outputs information that contributes to changing the control content according to the acquired class 110C.
[0200] For example, custom control is control including at least one of post-wait focus control, subject speed tolerance range setting control, and focus adjustment priority area setting control. The post-wait focus control is control that waits for a predetermined time (e.g., 10 seconds) when the position of the focus lens 40B is out of the in-focus position where the subject is in focus, and then moves the focus lens 40B toward the in-focus position. The subject speed tolerance range setting control is control for setting the tolerance range of the speed of the subject to be focused. The focus adjustment priority area setting control is control for setting which of the plurality of divided areas 134A within the AF area frame 134 to prioritize focus adjustment for.
[0201] When the AF area frame 134 is used for custom control in this way, for example, the imaging support process shown in FIG. 37 is performed by the CPU 86. The imaging support process shown in FIG. 37 is different from the imaging support process shown in FIG. 34 in that it has step ST500 between step ST300 and step ST236, and has steps ST502 and ST504 between step ST236 and step ST238.
[0202] In the imaging support process shown in FIG. 37, at step ST500, the CPU 86 acquires class 110C from the subject identification information 110 in the memory 90, and generates change instruction information for instructing a change in the control content of the custom control according to the acquired class 110C, and stores it in the memory 90. Examples of the control content include, for example, the waiting time used in the subject speed tolerance range setting control, the tolerance range used in the subject speed tolerance range setting control, and information capable of specifying the relative position within the AF area frame 134 of the focus priority divided area (i.e., the divided area 134A for which focus adjustment is prioritized) used in the focus adjustment priority area setting control. Further, the change instruction information may include specific change contents. The change contents may be, for example, contents predetermined for each class 110C. Note that the change instruction information is an example of the "information contributing to the change of the control content" according to the technology of the present disclosure.
[0203] In step ST502, the CPU 86 determines whether change instruction information is stored in the memory 90. In step ST502, if the change instruction information is not stored in the memory 90, the determination is negative, and the imaging support process proceeds to step ST238. In step ST502, if the change instruction information is stored in the memory 90, the determination is positive, and the imaging support process proceeds to step ST504.
[0204] In step ST504, the CPU 86 adds the change instruction information in the memory 90 to the live view image 136 with the AF area frame generated in step ST236. Then, the change instruction information is erased from the memory 90.
[0205] When the live view image 136 with the AF area frame is transmitted to the imaging device 12 by executing the process of step ST238, the CPU 62 of the imaging device 12 receives the live view image 136 with the AF area frame. Then, the CPU 62 causes the display 28 to display an alert prompting a change in the control content of the custom control according to the change instruction information added to the live view image 136 with the AF area frame.
[0206] Here, an example of the form in which an alert is displayed on the display 28 has been described. However, the technology of the present disclosure is not limited to this. The CPU 62 may change the control content of the custom control according to the change instruction information. Further, the CPU 62 may store the history of receiving the change instruction information in the NVM 64. When storing the history of receiving the change instruction information in the NVM 64, the live view image 130 included in the AF area frame-added live view image 136 to which the change instruction information is added, the live view image 130 included in the AF area frame-added live view image 136 to which the change instruction information is added, the thumbnail image of the AF area frame-added live view image 136 to which the change instruction information is added, or the thumbnail image of the live view image 130 included in the AF area frame-added live view image 136 to which the change instruction information is added may be associated with the history and stored in the NVM 64. In this case, it is possible to make the user or the like understand what kind of scene the custom control should be performed in.
[0207] According to the examples shown in FIGS. 36 and 37, the custom control is included in the control related to imaging other than the AF control. The CPU 86 acquires the class 110C from the subject identification information 110 in the memory 90 and outputs information that contributes to the change of the control content according to the acquired class 110C. Therefore, compared with the case where the control content of the custom control is changed according to an instruction given from the user or the like relying only on one's own intuition, imaging using the custom control suitable for the class 110C can be realized.
[0208] In the example shown in FIGS. 36 and 37, the CPU 86 obtains the class 110C from the subject identification information 110 in the memory 90, and outputs information that contributes to the change of the control content according to the obtained class 110C. However, the technology of the present disclosure is not limited to this. For example, in addition to the class 110C, the subject identification information 110 may have the state of the subject (for example, the subject is moving, the subject is stationary, the speed of the movement of the subject, and the trajectory of the movement of the subject, etc.) as subclasses. In this case, the CPU 86 obtains the class 110C and the subclasses from the subject identification information 110 in the memory 90, and outputs information that contributes to the change of the control content according to the obtained class 110C and subclasses. Thereby, imaging using custom control suitable for the class 110C and the subclasses can be realized as compared with the case where the control content of the custom control is changed according to an instruction given from a user or the like relying only on one's own intuition.
[0209] In the above embodiment, for the sake of convenience of explanation, an example has been described in which one bounding box 116 is applied to one frame of the live view image 130. However, when a plurality of subjects are included in the imaging range, a plurality of bounding boxes 116 appear for one frame of the live view image 130. And it is also conceivable that a plurality of bounding boxes 116 overlap. For example, when two bounding boxes 116 overlap, as shown in FIG. 38 as an example, one bounding box 116 is generated as the AF area frame 152 and the other bounding box 116 is generated as the AF area frame 154 by the generation unit 86D, and the AF area frame 152 and the AF area frame 154 overlap. In this case, the person image 156 which is an object in the AF area frame 152 and the person image 158 which is an object in the AF area frame 154 overlap. In the example shown in FIG. 38, since the person image 158 overlaps behind the person image 156, if the AF area frame 154 surrounding the person image 158 is used for AF control, there is a possibility that the person indicated by the person image 156 will be focused.
[0210] Therefore, the generation unit 86D obtains the class 110C to which the person image 156 within the AF area frame 152 belongs and the class 110C to which the person image 158 within the AF area frame 154 belongs from the subject identification information 110 (see FIG. 18) described in the above embodiment. Here, the class 110C to which the person image 156 belongs and the class 110C to which the person image 158 belongs are examples of the "object-by-object class information" according to the technology of the present disclosure.
[0211] Based on the class 110C to which the person image 156 belongs and the class 110C to which the person image 158 belongs, the generation unit 86D narrows down the images surrounded by frames from the person images 156 and 158. For example, different priorities are assigned in advance to the class 110C to which the person image 156 belongs and the class 110C to which the person image 158 belongs, and the images of the class 110C with a higher priority are narrowed down as the objects to be narrowed down by a frame. In the example shown in FIG. 38, since the priority of the class 110C to which the person image 158 belongs is higher than the priority of the class 110C to which the person image 156 belongs, the range of the person image 158 surrounded by the AF area frame 154 is narrowed down to the range excluding the overlapping area with the person image 156 (in the example shown in FIG. 38, the overlapping area between the AF area frame 152 and the AF area frame 154). In this way, the AF area frame 154 with the range of the person image 158 narrowed down is divided by the division method described in the above embodiment. Then, the live view image 136 with the AF area frame including the AF area frame 154 is transmitted to the imaging device 12 by the transmission unit 86H. As a result, the CPU 62 of the imaging device 12 performs AF control and the like using the AF area frame 154 with the range of the person image 158 narrowed down (that is, the AF area frame 154 narrowed down to the range excluding the area where the range of the person image 158 overlaps with the person image 156). In this case, compared with the case where the AF area frames 154 and 156 are directly used for AF control and the like, even if the person indicated by the person image 156 and the person indicated by the person image 158 overlap in the depth direction, it is possible to more easily focus on the person intended by the user or the like. Here, a person is exemplified as the subject, but this is merely an example, and it goes without saying that the subject may be other than a person.
[0212] Also, as an example, as shown in FIG. 39, for the portion of the AF area frame 154 that overlaps with the AF area frame 152, the out-of-focus target information 160 may be given by the dividing unit 86F. The out-of-focus target information 160 is information indicating an area that is excluded from the target to be focused. Here, the out-of-focus target information 160 is exemplified, but not limited thereto, and imaging-related control target-excluding information may be applied instead of the out-of-focus target information 160. The imaging-related control target-excluding information is information indicating an area that is excluded from the target of the control related to the imaging described above.
[0213] In the above embodiment, an example form in which the division method reflecting the composition of the airliner image 132 is derived from the division method derivation table 128 has been described, but the technology of the present disclosure is not limited to this. For example, a division method reflecting the composition of a high-rise building image obtained by imaging the high-rise building from an obliquely lower side or an obliquely upper side may be derived from the division method derivation table 128. In this case, the AF area frame 134 surrounding the high-rise building image is formed vertically long, and the number of divisions in the vertical direction is larger than the number of divisions in the horizontal direction. Note that not limited to the airliner image 132 and the high-rise building image, for an image obtained by imaging a subject with a composition for giving a three-dimensional effect, a division method corresponding to the composition may be derived from the division method derivation table 128.
[0214] In the above embodiment, the AF area frame 134 is exemplified, but the technology of the present disclosure is not limited to this, and instead of the AF area frame 134, or together with the AF area frame 134, an area frame for restricting targets such as exposure control, white balance control, and / or gradation control may be used. Also in this case, the area frame is generated by the imaging support device 14 in the same manner as in the above embodiment, and the generated area frame is transmitted from the imaging support device 14 to the imaging device 12 and used by the imaging device 12.
[0215] In the above-described embodiment, an example of the form in which a subject is recognized by the AI subject recognition method has been described. However, the technology of the present disclosure is not limited to this, and the subject may be recognized by other subject recognition methods such as the template matching method.
[0216] In the above-described embodiment, the live view image 130 has been exemplified. However, the technology of the present disclosure is not limited to this. For example, a post-view image may be used instead of the live view image 130. That is, the subject recognition process (see step ST202 shown in FIG. 31A) may be executed based on the post-view image. Further, the subject recognition process may be executed based on the captured image 108. Further, the subject recognition process may be executed based on a phase difference image including a plurality of phase difference pixels. In this case, the imaging support device 14 can provide the imaging device 12 with a plurality of phase difference pixels used for distance measurement together with the AF area frame 134.
[0217] In the above-described embodiment, an example of the form in which the imaging device 12 and the imaging support device 14 are separate bodies has been described. However, the technology of the present disclosure is not limited to this, and the imaging device 12 and the imaging support device 14 may be integrated. In this case, for example, as shown in FIG. 40, in addition to the display control processing program 80, a learned model 93, an imaging support processing program 126, and a division method derivation table 128 are stored in the NVM 64 of the imaging device main body 16, and the CPU 62 uses the learned model 93, the imaging support processing program 126, and the division method derivation table 128 in addition to the display control processing program 80.
[0218] Further, when the imaging support device 14 functions for the imaging device 12 in this way, at least one other CPU, at least one GPU, and / or at least one TPU may be used instead of or together with the CPU 62.
[0219] In the above-described embodiment, an example has been described in which the imaging support processing program 126 is stored in the storage 88. However, the technology of the present disclosure is not limited to this. For example, the imaging support processing program 126 may be stored in a portable non-temporary storage medium such as an SSD or a USB memory. The imaging support processing program 126 stored in the non-temporary storage medium is installed in the computer 82 of the imaging support device 14. The CPU 86 executes the imaging support processing according to the imaging support processing program 126.
[0220] Also, the imaging support processing program 126 may be stored in a storage device such as another computer or a server device connected to the imaging support device 14 via the network 34, and the imaging support processing program 126 may be downloaded and installed in the computer 82 in response to a request from the imaging support device 14.
[0221] Note that it is not necessary to store all of the imaging support processing program 126 in a storage device such as another computer or a server device connected to the imaging support device 14, or in the storage 88, and a part of the imaging support processing program 126 may be stored.
[0222] Also, although the imaging device 12 shown in FIG. 2 has a built-in controller 44, the technology of the present disclosure is not limited to this. For example, the controller 44 may be provided outside the imaging device 12.
[0223] In the above-described embodiment, the computer 82 is illustrated. However, the technology of the present disclosure is not limited to this, and a device including an ASIC, an FPGA, and / or a PLD may be applied instead of the computer 82. Also, instead of the computer 82, a combination of a hardware configuration and a software configuration may be used.
[0224] As the hardware resources for executing the imaging support process described in the above embodiment, various types of processors shown below can be used. As the processor, for example, a general-purpose processor such as a CPU that functions as a hardware resource for executing the imaging support process by executing software, that is, a program, can be mentioned. Further, as the processor, for example, a dedicated electric circuit which is a processor having a circuit configuration specifically designed for executing a specific process such as an FPGA, a PLD, or an ASIC can be mentioned. A memory is built in or connected to any of these processors, and any of these processors executes the imaging support process by using the memory.
[0225] The hardware resources for executing the imaging support process may be constituted by one of these various processors, or may be constituted by a combination of two or more processors of the same type or different types (for example, a combination of a plurality of FPGAs, or a combination of a CPU and an FPGA). Further, the hardware resources for executing the imaging support process may be one processor.
[0226] As an example of constituting with one processor, first, there is a form in which one processor is constituted by a combination of one or more CPUs and software, and this processor functions as a hardware resource for executing the imaging support process. Second, as represented by an SoC or the like, there is a form in which a processor that realizes the functions of the entire system including a plurality of hardware resources for executing the imaging support process with one IC chip is used. Thus, the imaging support process is realized as a hardware resource by using one or more of the above various processors.
[0227] Furthermore, as the hardware structure of these various processors, more specifically, an electric circuit combining circuit elements such as semiconductor elements can be used. Also, the above imaging support process is merely an example. Therefore, it goes without saying that unnecessary steps may be deleted, new steps may be added, or the processing order may be changed within the scope of not departing from the gist.
[0228] The description and illustration shown above are detailed descriptions of the part related to the technology of the present disclosure and are merely examples of the technology of the present disclosure. For example, the descriptions regarding the above configurations, functions, operations, and effects are descriptions of examples of the configurations, functions, operations, and effects of the part related to the technology of the present disclosure. Therefore, it goes without saying that within the scope not departing from the gist of the technology of the present disclosure, the description and illustration shown above may be modified by deleting unnecessary parts, adding new elements, or making replacements. Also, in order to avoid complication and facilitate the understanding of the part related to the technology of the present disclosure, in the description and illustration shown above, descriptions regarding common technical knowledge and the like that do not particularly require explanation for implementing the technology of the present disclosure are omitted.
[0229] In this specification, "A and / or B" is synonymous with "at least one of A and B". That is, "A and / or B" means that it may be only A, only B, or a combination of A and B. Also, in this specification, when expressing three or more matters connected by "and / or", the same concept as that of "A and / or B" is applicable.
[0230] All documents, patent applications, and technical standards described in this specification are incorporated herein by reference to the same extent as if each individual document, patent application, and technical standard were specifically and individually stated to be incorporated by reference.
Claims
1. A processor, a memory connected to or incorporated in the processor, and includes: The processor is obtaining the type of the subject based on an image obtained by imaging an imaging range including the subject by an image sensor, outputting information indicating a division method for dividing a division area that can distinguish the subject from other areas in the imaging range according to the obtained type An imaging support device.
2. The processor obtains information regarding a person, an animal, or a vehicle as the type The imaging support device according to claim 1.
3. The processor obtains information regarding a specific person, the face of a specific person, a specific automobile, a specific passenger aircraft, a specific bird, or a specific train as the type The imaging support device according to claim 1.
4. The processor obtains the type based on an output result output from the learned model by providing the image to the learned model on which machine learning has been performed The imaging support device according to claim 1.
5. The learned model assigns an object within a bounding box applied to the image to a corresponding class, The output result includes a value based on the probability that the object within the bounding box applied to the image belongs to a specific class The imaging support device according to claim 4.
6. The output result is based on the probability that an object exists within the bounding box When the value is equal to or greater than a first threshold, the value includes a value based on the probability that the object within the bounding box belongs to a specific class The imaging support device according to claim 5.
7. The output result includes a value equal to or greater than a second threshold among the values based on the probability that the object belongs to the specific class The imaging support device according to claim 5 or claim 6.
8. The processor expands the bounding box when the value based on the probability that the object exists within the bounding box is less than a third threshold The imaging support device according to any one of claims 5 to 7.
9. The processor changes the size of the division area according to the value based on the probability that the object belongs to the specific class The imaging support device according to any one of claims 5 to 8.
10. The imaging range includes a plurality of subjects, The learned model assigns each of a plurality of objects within a plurality of bounding boxes applied to the image to a corresponding class. The output result includes object-by-object class information indicating each class to which the plurality of objects within the plurality of bounding boxes applied to the image belong. The processor narrows down at least one subject surrounded by the divided region from the plurality of subjects based on the object-by-object class information. The imaging support device according to any one of claims 4 to 9.
11. The division method defines the number of divisions for dividing the divided region. The imaging support device according to any one of claims 1 to 10.
12. The divided region is defined by a first direction and a second direction intersecting the first direction. The division method defines the number of divisions in the first direction and the number of divisions in the second direction. The imaging support device according to claim 11.
13. The number of divisions in the first direction and the number of divisions in the second direction are defined based on the composition of the subject image indicating the subject within the image. The imaging support device according to claim 12.
14. The number of divisions in the first direction and the number of divisions in the second direction are defined based on whether the subject image indicating the subject within the image gives a sense of depth. The imaging support device according to claim 12.
15. When the focus of the focus lens is adjustable by moving the focus lens along the optical axis that guides incident light to the image sensor, The processor outputs information that contributes to moving the focus lens to a focusing position corresponding to the divided region corresponding to the acquired type among the plurality of divided regions obtained by dividing the divided region by the number of divisions. The imaging support device according to any one of claims 11 to 13.
16. When the focus of the focus lens is adjustable by moving the focus lens along the optical axis that guides incident light to the image sensor, The processor moves the focus lens to the focusing position corresponding to the acquired type and outputs information that contributes to the movement. The imaging support device according to any one of claims 1 to 15.
17. The divided region is used for control related to imaging of the subject by the image sensor. The imaging support device according to any one of claims 1 to 16.
18. The control related to the imaging includes custom control, The custom control is a control including at least one of post-waiting focus control, subject speed tolerance range setting control, and focus adjustment priority area setting control, The processor outputs information that contributes to the change of the custom control according to the acquired type. The imaging support device according to claim 17.
19. The processor further acquires the state of the subject based on the image, Outputs information that contributes to the change of the custom control according to the acquired state and the type. The imaging support device according to claim 18.
20. The state is that the subject is moving, the subject is stationary, the speed of movement of the subject, and / or the trajectory of movement of the subject. The imaging support device according to claim 19.
21. In the case where the focus of the focus lens can be adjusted by moving the focus lens that guides incident light to the image sensor along the optical axis, The post-waiting focus control is a control including at least one of controls that wait for a predetermined time when the position of the focus lens is deviated from the in-focus position where the focus is on the subject and then move the focus lens toward the in-focus position. The imaging support device according to claim 18 or claim 19.
22. In the case where the focus of the focus lens can be adjusted by moving the focus lens that guides incident light to the image sensor along the optical axis, The subject speed tolerance range setting control is a control for setting the allowable range of the speed of the subject to be focused. The imaging support device according to claim 18.
23. In the case where the focus of the focus lens can be adjusted by moving the focus lens that guides incident light to the image sensor along the optical axis, The focus adjustment priority area setting control is a control for setting which area among a plurality of areas within the divided area is prioritized for focus adjustment. The imaging support device according to claim 18.
24. The divided area is a frame surrounding the subject. The imaging support device according to any one of claims 1 to 21.
25. The processor outputs information for causing the display to display a live view image based on the image and to display the frame within the live view image. The imaging support apparatus according to claim 24. **Claim 26** The processor adjusts the focus of the focus lens by moving the focus lens that guides incident light to the image sensor along the optical axis, wherein the frame is a focus frame that defines an area that is a candidate for focusing. The imaging support apparatus according to claim 24 or claim 25. **Claim 27** An imaging apparatus comprising: a processor; a memory connected to or incorporated in the processor; and an image sensor, wherein the processor acquires the type of the subject based on an image obtained by imaging an imaging range including the subject with the image sensor, and outputs information indicating a division method for dividing a division area that discriminately divides the subject from other areas in the imaging range according to the acquired type. **Claim 28** acquiring the type of the subject based on an image obtained by imaging an imaging range including the subject with an image sensor; and outputting information indicating a division method for dividing a division area that discriminately divides the subject from other areas in the imaging range according to the acquired type. An imaging support method. **Claim 29** A program for causing a computer to acquire the type of the subject based on an image obtained by imaging an imaging range including the subject with an image sensor; and output information indicating a division method for dividing a division area that discriminately divides the subject from other areas in the imaging range according to the acquired type.
Citation Information
Patent Citations
Vehicle detecting system
JP2009175846A
Segmenting spatiotemporal data based on user gaze data
JP2013045445A
Object recognition device
JP2015215868A
Photographing device, photographing method, and program
JP2016061884A
Subject detection device, imaging apparatus, method of detection and program
JP2020057871A