Imaging support device, imaging device, imaging support method, and storage medium

CN116685904BActive Publication Date: 2026-08-11FUJIFILM CORP
View PDF 4 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-12-21
Publication Date
2026-08-11

AI Technical Summary

Benefits of technology

[0028]本发明的技术所涉及的第23方式为程序,其用于使计算机执行包括如下步骤的处理:根据通过由图像传感器拍摄包括被摄体的拍摄范围而得的图像来获取被摄体的类型;及输出表示根据获取到的类型来分割划分区域的分割方式的信息,该划分区域将被摄体划分成能够与拍摄范围的其他区域区分。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116685904B_ABST
    Figure CN116685904B_ABST
Patent Text Reader

Abstract

This invention relates to a camera support device, a camera apparatus, a camera support method, and a program storage medium. The camera support device of this invention includes: a processor; and a memory connected to or integrated into the processor. The processor acquires the type of the subject based on an image obtained by capturing a shooting range including the subject using an image sensor, and outputs information indicating a segmentation method for dividing a region according to the acquired type, wherein the segmented region divides the subject into areas distinguishable from other areas of the shooting range.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to a camera support device, a camera device, a camera support method, and a procedure. Background Technology

[0002] Japanese Patent Application Publication No. 2015-046917 discloses a photographic apparatus characterized by comprising: an imaging unit that captures an image of a subject imaged by an imaging lens and outputs image data; a focus determination unit that performs a weighting based on at least one of the focal position of the subject and characteristics of the subject region, and determines the focus sequence from which region to which region within the captured image to focus based on the weighting result; a focus control unit that drives the imaging lens according to the focus sequence determined by the focus determination unit; and a recording unit that records image data of moving images and still images based on the image data. The focus determination unit determines a main subject, a secondary subject, or a main region and a secondary region based on the weighting result, and determines the focus sequence with one of the main subject, secondary subject, or main region and secondary region as the starting point and the other as the ending point. The recording unit records image data of moving images while the imaging lens is driven by the focus control unit, and records image data of still images after the driving of the imaging lens is stopped by the focus control unit.

[0003] Japanese Patent Application Publication No. 2012-128287 discloses a focus detection device, characterized by comprising: a face detection mechanism for detecting the position and size of a human face in an image; a setting mechanism for setting a first focus detection area for the presence of a human face and a second focus detection area for predicting the location of the human face when observed from the position of the human face; and a focus adjustment mechanism for adjusting the focus by moving the camera optical system based on the signal output of the focus detection area, wherein if the size of the face detected by the face detection mechanism is smaller than a predetermined size, the setting mechanism sets the size of the second focus detection area to be larger than the size of the first focus detection area. Summary of the Invention

[0004] One embodiment of the present invention provides a camera support device, camera apparatus, camera support method, and program capable of precisely controlling the image sensor's capture of a subject.

[0005] means for solving technical problems

[0006] The first aspect of the technology of the present invention is a camera support device, which includes: a processor; and a memory connected to or built into the processor, the processor performing the following processing: obtaining the type of the subject based on an image obtained by capturing a shooting range including the subject by an image sensor; and outputting information indicating a segmentation method for dividing a region according to the obtained type, the segmentation region dividing the subject into regions that can be distinguished from other regions of the shooting range.

[0007] In the second aspect of the technology of the present invention, in the camera support device of the first aspect, the processor obtains the type based on the output result of the learned model output by assigning an image to the learned model through machine learning.

[0008] The third aspect of the present invention relates to the camera support device of the second aspect, wherein a learned model classifies objects within a bounding box of an image into corresponding categories, and the output includes a value based on the probability that an object within a bounding box of an image belongs to a specific category.

[0009] In the camera support device of the third method, the fourth aspect of the present invention outputs a value based on the probability that the object exists within the bounding box, when the value is greater than or equal to a first threshold. This includes the probability that the object within the bounding box belongs to a specific category.

[0010] In the fifth aspect of the present invention, in the camera support device involved in the third or fourth aspect, the output result includes a value above the second threshold among the values ​​of the probability that the object belongs to a specific category.

[0011] In the camera support device involved in any of the third to fifth embodiments of the present invention, the sixth aspect of the technology relates to an expanding bounding box when the value of the probability based on the object existing within the bounding box is less than a third threshold.

[0012] In the seventh aspect of the present invention, in the camera support device involved in any of the third to sixth aspects, the processor changes the size of the partitioned area based on a value of the probability that the object belongs to a specific category.

[0013] In the camera support device involved in any of the second to seventh embodiments of the present invention, the shooting range includes multiple subjects, the learned model assigns multiple objects within multiple bounding boxes applicable to the image to corresponding categories, the output results include object category information representing the categories to which the multiple objects within the multiple bounding boxes applicable to the image belong, and the processor retrieves at least one subject surrounded by a divided region from the multiple subjects based on the object category information.

[0014] In the ninth aspect of the present invention, in the camera support device involved in any of the first to eighth aspects, the segmentation method specifies the number of segments for dividing the region.

[0015] In the 10th aspect of the present invention, in the camera support device according to the 9th aspect, the area is defined by a first direction and a second direction intersecting the first direction, and the segmentation method defines the number of segments in the first direction and the number of segments in the second direction.

[0016] In the 11th aspect of the present invention, in the camera support device involved in any of the 1st to 10th aspects, the number of divisions in the first direction and the number of divisions in the second direction are determined according to the composition of the subject image representing the subject in the image.

[0017] In the camera support device according to any one of the 9th to 11th embodiments of the present invention, when the focus of the focusing lens can be adjusted by moving the focusing lens that guides the incident light to the image sensor along the optical axis, the processor outputs information that helps to move the focusing lens to a focus position corresponding to a segmented region in a plurality of segmented regions that corresponds to the type obtained, the plurality of segmented regions being obtained by dividing the region by a number of segments.

[0018] In the camera support device according to any one of the first to 12th embodiments, the 13th aspect of the present invention, when the focus of the focusing lens can be adjusted by moving the focusing lens that guides the incident light to the image sensor along the optical axis, the processor outputs information that helps to move the focusing lens to a focus position corresponding to the acquired type.

[0019] In the 14th aspect of the present invention, in the camera support device involved in any of the 1st to 13th aspects, a region is divided for control related to the image sensor's capture of the subject.

[0020] In the 15th aspect of the present invention, in the camera support device of the 14th aspect, the shooting-related control includes custom control, which is as follows: it is suggested to change the control content of the shooting-related control according to the subject, and the control content is changed according to the instruction given, and the processor outputs information that helps to change the control content according to the type obtained.

[0021] In the 16th aspect of the present invention, in the camera support device of the 15th aspect, the processor also acquires the state of the subject based on the image, and outputs information that helps to change the control content based on the acquired state and type.

[0022] In the camera support device according to the 15th or 16th aspect of the present invention, the 17th aspect of the present invention, in the case where the focus of the focusing lens can be adjusted by moving the focusing lens that guides the incident light to the image sensor along the optical axis, is a custom control that includes at least one of the following controls: when the position of the focusing lens deviates from the focus position aligned with the subject, after a predetermined standby time, the focusing lens is moved toward the focus position; an allowable range for the speed of the subject to be aligned with the focus is set; and a setting is made for which of the multiple regions within the divided area is prioritized for focus adjustment.

[0023] In the 18th aspect of the present invention, in the camera support device involved in any of the 1st to 17th aspects, the area is divided into a frame surrounding the subject.

[0024] In the 19th aspect of the present invention, in the camera support device of the 18th aspect, the processor outputs information for causing the display to show an image-based instant preview image and displaying a frame within the instant preview image.

[0025] In the 20th aspect of the present invention, in the camera support device of the 18th or 19th aspect, the processor adjusts the focus of the focusing lens by moving the focusing lens that guides incident light to the image sensor along the optical axis, and the frame is a focusing frame that defines a candidate area for focusing.

[0026] The 21st aspect of the present invention is a camera device comprising: a processor; a memory connected to or built into the processor; and an image sensor, wherein the processor performs the following processing: obtaining the type of the subject based on an image obtained by capturing a shooting range including the subject by the image sensor; and outputting information representing a segmentation method for dividing a region according to the obtained type, the segmented region dividing the subject into regions that can be distinguished from other regions of the shooting range.

[0027] The 22nd aspect of the present invention is a camera support method, which includes the following steps: obtaining the type of a subject based on an image obtained by capturing a shooting range including the subject using an image sensor; and outputting information representing a segmentation method for dividing a region according to the obtained type, wherein the segmented region divides the subject into regions that can be distinguished from other regions of the shooting range.

[0028] The 23rd aspect of the technology of the present invention is a program that causes a computer to perform a process including the following steps: obtaining the type of a subject based on an image obtained by capturing a shooting range including the subject by an image sensor; and outputting information representing a segmentation method for dividing a region according to the obtained type, the segmentation region dividing the subject into regions that can be distinguished from other regions of the shooting range. Attached Figure Description

[0029] Figure 1 This is a schematic diagram illustrating an example of the overall structure of a camera system.

[0030] Figure 2 This is a schematic structural diagram illustrating an example of the hardware structure of the optical and electrical systems of the camera device included in a camera system.

[0031] Figure 3 This is a block diagram illustrating an example of the functionality of the main components of the CPU included in the main body of the camera device.

[0032] Figure 4 This is a schematic structural diagram illustrating an example of the hardware structure of an electrical system including a camera support device in a camera system.

[0033] Figure 5 This is a block diagram representing an example of the content processed during CNN machine learning.

[0034] Figure 6 This is a block diagram illustrating an example of constructing a learned model by optimizing a CNN.

[0035] Figure 7 This is a block diagram illustrating an example of how subject-specific information is extracted from a learned model when a captured image is assigned to that model.

[0036] Figure 8 This is a concept diagram illustrating an example of how anchor frames are applied to captured images.

[0037] Figure 9 This is a concept diagram representing an example of an inference method using bounding boxes.

[0038] Figure 10 This is a concept diagram illustrating an example of a subject recognition method that utilizes template matching.

[0039] Figure 11 This is a concept diagram representing an example of the recognition result when identifying a subject using template matching.

[0040] Figure 12This is a conceptual diagram illustrating the difference between the recognition results obtained when subject recognition is performed using template matching and the recognition results obtained when subject recognition is performed using AI subject recognition.

[0041] Figure 13 This is a conceptual diagram illustrating the differences in depth length when the subject is a human, when the subject is a bird, and when the subject is a car.

[0042] Figure 14 This is a block diagram representing an example of the stored content of the storage unit of a camera support device.

[0043] Figure 15 This is a block diagram illustrating an example of the functionality of the main components of the CPU included in a camera support device.

[0044] Figure 16 This is a block diagram illustrating an example of the processing content of the acquisition unit, generation unit, and transmission unit of a camera device.

[0045] Figure 17 This is a block diagram illustrating an example of the processing content of the receiving and executing sections of a camera support device.

[0046] Figure 18 This is a block diagram illustrating an example of the processing content of the execution section and the decision section of a camera support device.

[0047] Figure 19 This is a block diagram representing an example of the content processed in the first decision step.

[0048] Figure 20 This is a block diagram representing an example of the content processed in the second decision step.

[0049] Figure 21 This is a block diagram representing an example of the content processed in the first segment.

[0050] Figure 22 This is a conceptual diagram illustrating an example of how AF region boxes are divided based on their area.

[0051] Figure 23 This is a block diagram representing an example of the content processed in the second segmentation.

[0052] Figure 24 This is a block diagram representing an example of the content processed in the third segmentation.

[0053] Figure 25 This is a conceptual diagram illustrating an example of how an AF region box is segmented based on a segmentation method derived from the category determined by the subject.

[0054] Figure 26This is a conceptual diagram illustrating an example of how an AF region box is segmented based on a segmentation method derived from the category obtained from the subject determination information, when the AF region box surrounds an image of a person.

[0055] Figure 27 This is a conceptual diagram illustrating an example of how an AF region box is segmented based on a segmentation method derived from the category obtained from the subject determination information, when the AF region box surrounds a car image.

[0056] Figure 28 This is a block diagram representing an example of the content generated, sent, and processed from an image with an AF region bounding box.

[0057] Figure 29 This is a conceptual diagram illustrating an example of the processing content within a camera device when an image with an AF region frame is displayed on the camera device's monitor.

[0058] Figure 30 This is a flowchart illustrating an example of the display control processing flow.

[0059] Figure 31A This is a flowchart illustrating an example of the camera support processing flow.

[0060] Figure 31B yes Figure 31A The following diagram is a follow-up to the flowchart shown.

[0061] Figure 32 This is a conceptual diagram illustrating an example of how focus priority position information is assigned to the AF region bounding box.

[0062] Figure 33 This is a flowchart representing the first variation of the camera support processing flow.

[0063] Figure 34 This is a flowchart of the second variation of the camera support processing flow.

[0064] Figure 35 This is a conceptual diagram illustrating an example of how an AF region box is sized based on category scores.

[0065] Figure 36 This is a block diagram representing an example of custom control content.

[0066] Figure 37 This is a flowchart representing the third variation of the camera support processing flow.

[0067] Figure 38 This is a concept diagram illustrating an example of retrieving the processing content of a subject surrounded by multiple overlapping AF region frames.

[0068] Figure 39This is a conceptual diagram illustrating an example of how information about the outside of the focus object is added to an AF region box when multiple AF region boxes overlap.

[0069] Figure 40 This is a block diagram illustrating a structural example of the main body of a camera device when the camera device assumes the function of a camera support device. Detailed Implementation

[0070] Hereinafter, an example of an embodiment of the camera support device, camera apparatus, camera support method and program related to the present invention will be described with reference to the accompanying drawings.

[0071] First, let me explain the words and phrases used in the following description.

[0072] CPU stands for Central Processing Unit. GPU stands for Graphics Processing Unit. TPU stands for Tensor Processing Unit. NVM stands for Non-volatile memory. RAM stands for Random Access Memory. IC stands for Integrated Circuit. ASIC stands for Application Specific Integrated Circuit. PLD stands for Programmable Logic Device. FPGA stands for Field-Programmable Gate Array. SoC stands for System-on-a-chip. SSD stands for Solid State Drive. USB stands for Universal Serial Bus. HDD stands for Hard Disk Drive. EEPROM is an abbreviation for "Electrically Erasable and Programmable Read Only Memory". EL is an abbreviation for "Electro-Luminescence". I / F is an abbreviation for "Interface". UI is an abbreviation for "User Interface". fps is an abbreviation for "Frames Per Second". MF is an abbreviation for "Manual Focus". AF is an abbreviation for "Auto Focus". CMOS is an abbreviation for "Complementary Metal Oxide Semiconductor". LAN is an abbreviation for "Local Area Network". WAN is an abbreviation for "Wide Area Network". CNN is an abbreviation for "Convolutional Neural Network".AI is an abbreviation for "Artificial Intelligence". TOF is an abbreviation for "Time of Flight".

[0073] In this specification, "perpendicular" means not only perfectly perpendicular, but also perpendicular to a degree of error that is generally permissible in the technical field to which this invention pertains and does not depart from the spirit of this invention. In this specification, "orthogonal" means not only perfectly orthogonal, but also orthogonal to a degree of error that is generally permissible in the technical field to which this invention pertains and does not depart from the spirit of this invention. In this specification, "parallel" means not only perfectly parallel, but also parallel to a degree of error that is generally permissible in the technical field to which this invention pertains and does not depart from the spirit of this invention. In this specification, "coinciding" means not only perfectly coincident, but also coincident to a degree of error that is generally permissible in the technical field to which this invention pertains and does not depart from the spirit of this invention.

[0074] As an example, such as Figure 1 As shown, the camera system 10 includes a camera device 12 and a camera support device 14. The camera device 12 is a device for photographing a subject. Figure 1 In the example shown, a digital camera with an interchangeable lens is illustrated as an example of a camera device 12. The camera device 12 includes a camera device body 16 and an interchangeable lens 18. The interchangeable lens 18 is interchangeably mounted to the camera device body 16. The interchangeable lens 18 is provided with a focus ring 18A. The focus ring 18A is operated by the user of the camera device 12 (hereinafter referred to as "user") or others when manually adjusting the focus on the subject through the camera device 12.

[0075] In addition, in this embodiment, an interchangeable-lens digital camera is exemplified as the camera device 12, but this is only one example. It can also be a fixed-lens digital camera, or a digital camera built into various electronic devices such as smart devices, wearable terminals, cell observation devices, ophthalmic observation devices or surgical microscopes.

[0076] The main body 16 of the imaging device is equipped with an image sensor 20. The image sensor 20 is a CMOS image sensor. The image sensor 20 captures a shooting range including at least one subject. When the interchangeable lens 18 is mounted on the main body 16 of the imaging device, subject light representing the subject is transmitted through the interchangeable lens 18 and imaged onto the image sensor 20, and image data representing the subject is generated by the image sensor 20. Subject light is an example of "incident light" as described in the technology of this invention.

[0077] In addition, in this embodiment, a CMOS image sensor is exemplified as the image sensor 20, but the technology of the present invention is not limited to this and other image sensors may also be used.

[0078] The upper surface of the main body 16 of the camera device is provided with a release button 22 and a dial 24. The dial 24 is operated when setting the operation mode of the camera system and the operation mode of the playback system. By operating the dial 24, the shooting mode and playback mode can be selectively set as the operation mode in the camera device 12.

[0079] The release button 22 functions as both a shooting preparation indicator and a shooting indicator, detecting press operations in both the shooting preparation indicator state and the shooting indicator state. The shooting preparation indicator state refers, for example, to the state from the standby position to the middle position (half-pressed position), and the shooting indicator state refers to the state from the standby position to the final pressed position (fully pressed position). Furthermore, hereinafter, the state from the standby position to the half-pressed position will be referred to as the "half-pressed state," and the state from the standby position to the fully pressed position will be referred to as the "fully pressed state." Depending on the structure of the camera device 12, the shooting preparation indicator state can also refer to the state where the user's finger touches the release button 22, and the shooting indicator state can also refer to the state where the user's finger moves from the state of touching the release button 22 to the state of releasing it.

[0080] The back of the camera device body 16 is equipped with a touch screen display 32 and indicator keys 26.

[0081] Touchscreen display 32 includes display 28 and touch panel 30 (see also...) Figure 2 As an example of display 28, an EL display (e.g., an organic EL display or an inorganic EL display) can be cited. Display 28 may also be other types of displays, such as liquid crystal displays, instead of EL displays.

[0082] The display 28 displays images and / or character information, etc. When the camera device 12 is in shooting mode, the display 28 is used to display a real-time preview image 130 (see reference). Figure 16 The instant preview image 130 is obtained by taking instant preview images (i.e., continuous shooting). To obtain the instant preview image 130 (refer to...), Figure 16 The shooting (hereinafter also referred to as "shooting for real-time preview images") is performed, for example, at a frame rate of 60 fps. 60 fps is just one example; it can also be a frame rate less than 60 fps or a frame rate greater than 60 fps.

[0083] Here, "live preview image" refers to a dynamic image for display based on image data captured by the image sensor 20. Live preview images are also commonly referred to as live view images.

[0084] When the camera device 12 is instructed to take a still image via the release button 22, the display 28 is also used to display a still image obtained by taking a still image. Furthermore, the display 28 is also used to display playback images and menu screens when the camera device 12 is in playback mode.

[0085] The touch panel 30 is a transmissive touch panel that is superimposed on the surface of the display area of ​​the display 28. The touch panel 30 receives user instructions by detecting contact from a finger or stylus. Additionally, for ease of explanation, the "full press state" mentioned below also includes the state where the user presses the soft key for starting photography via the touch panel 30.

[0086] Furthermore, in this embodiment, as an example of a touchscreen display 32, an external touchscreen display in which the touch panel 30 is superimposed on the surface of the display area of ​​the display 28 is provided, but this is only one example. For example, as a touchscreen display 32, an embedded or in-wall touchscreen display may also be used.

[0087] Indicator key 26 handles various instructions. Here, "various instructions" refers to, for example, instructions for displaying menu screens with various menu options, instructions for selecting one or more menus, instructions for confirming selections, instructions for deleting selections, zooming in, zooming out, and frame forward, etc. Furthermore, these instructions can also be made via touch panel 30.

[0088] The camera device body 16 is connected to the camera support device 14 via a network 34, the details of which will be described later. The network 34 is, for example, the Internet. The network 34 is not limited to the Internet; it can also be a WAN and / or a LAN, etc. Furthermore, in this embodiment, the camera support device 14 is a server that provides services corresponding to requests from the camera device 12. Alternatively, the server can be a mainframe computer used locally with the camera device 12, or it can be an external server implemented through cloud computing. Furthermore, the server can also be an external server implemented through network computing such as fog computing, edge computing, or grid computing. Here, a server is listed as an example of the camera support device 14, but this is only one example; at least one personal computer or the like can also be used as the camera support device 14 instead of a server.

[0089] As an example, such as Figure 2As shown, the image sensor 20 includes a photoelectric conversion element 72. The photoelectric conversion element 72 has a light-receiving surface 72A. The photoelectric conversion element 72 is disposed within the camera device body 16 such that the center of the light-receiving surface 72A coincides with the optical axis OA (see also...). Figure 1 The photoelectric conversion element 72 has a plurality of photosensitive pixels arranged in a matrix, and the light-receiving surface 72A is formed by the plurality of photosensitive pixels. The photosensitive pixel is a physical pixel having a photodiode (not shown), which performs photoelectric conversion on the received light and outputs an electrical signal corresponding to the amount of light received.

[0090] The interchangeable lens 18 includes an imaging lens 40. The imaging lens 40 includes an objective lens 40A, a focusing lens 40B, a zoom lens 40C, and an aperture 40D. The objective lens 40A, focusing lens 40B, zoom lens 40C, and aperture 40D are arranged sequentially along the optical axis OA from the subject side (object side) to the imaging device body 16 side (image side).

[0091] Furthermore, the interchangeable lens 18 includes a control device 36, a first actuator 37, a second actuator 38, and a third actuator 39. The control device 36 controls the entire interchangeable lens 18 according to instructions from the camera device main body 16. The control device 36 is, for example, a computer device including a CPU, NVM, and RAM. While a computer is illustrated here, this is only one example; devices including ASICs, FPGAs, and / or PLDs can also be used. Moreover, the control device 36 can also be a device implemented using a combination of hardware and software structures.

[0092] The first actuator 37 includes a focusing sliding mechanism (not shown) and a focusing motor (not shown). A focusing lens 40B is mounted on the focusing sliding mechanism in a manner that allows it to slide along the optical axis OA. Furthermore, the focusing sliding mechanism is connected to the focusing motor, and the focusing sliding mechanism operates by receiving power from the focusing motor, thereby moving the focusing lens 40B along the optical axis OA.

[0093] The second actuator 38 includes a zoom sliding mechanism (not shown) and a zoom motor (not shown). The zoom lens 40C is mounted on the zoom sliding mechanism in a manner that allows it to slide along the optical axis OA. Furthermore, the zoom motor is connected to the zoom sliding mechanism, and the zoom sliding mechanism operates by receiving power from the zoom motor, thereby moving the zoom lens 40C along the optical axis OA.

[0094] The third actuator 39 includes a power transmission mechanism (not shown) and an aperture motor (not shown). The aperture 40D has an opening 40D1, the size of which is variable. The opening 40D1 is formed by multiple aperture blades 40D2. The multiple aperture blades 40D2 are connected to the power transmission mechanism. Furthermore, the aperture motor is connected to the power transmission mechanism, which transmits power from the aperture motor to the multiple aperture blades 40D2. The multiple aperture blades 40D2 operate by receiving power from the power transmission mechanism, thereby changing the size of the opening 40D1. The aperture 40D adjusts the exposure by changing the size of the opening 40D1.

[0095] The focusing motor, zoom motor, and aperture motor are connected to the control device 36, which controls their respective drives. In this embodiment, a stepper motor is used as an example of the focusing motor, zoom motor, and aperture motor. Therefore, the focusing motor, zoom motor, and aperture motor operate synchronously according to commands and pulse signals from the control device 36. Here, an example is shown where the focusing motor, zoom motor, and aperture motor are located on the interchangeable lens 18; however, this is only one example, and at least one of the focusing motor, zoom motor, and aperture motor may be located on the imaging device body 16. Furthermore, the structure and / or operation method of the interchangeable lens 18 can be changed as needed.

[0096] In the imaging device 12, when in shooting mode, the MF mode and AF mode can be selectively set according to the instructions given to the imaging device body 16. MF mode is the manual focus mode. In MF mode, for example, by the user operating the focus ring 18A, the focusing lens 40B moves along the optical axis OA by a movement corresponding to the operation amount of the focus ring 18A, thereby adjusting the focus.

[0097] In AF mode, the camera body 16 calculates the focus position corresponding to the distance to the subject and moves the focusing lens 40B toward the calculated focus position, thereby adjusting the focus. Here, the focus position refers to the position of the focusing lens 40B on the optical axis OA in the focusing state. Furthermore, for ease of explanation, the control of aligning the focusing lens 40B with the focus position will also be referred to as "AF control" below.

[0098] The main body 16 of the camera device includes an image sensor 20, a controller 44, an image memory 46, a UI device 48, an external I / F 50, a communication I / F 52, a photoelectric conversion element driver 54, a mechanical shutter driver 56, a mechanical shutter actuator 58, a mechanical shutter 60, and an input / output interface 70. Furthermore, the image sensor 20 includes a photoelectric conversion element 72 and a signal processing circuit 74.

[0099] The input / output interface 70 is connected to a controller 44, an image memory 46, a UI device 48, an external I / F 50, a photoelectric conversion element driver 54, a mechanical shutter driver 56, and a signal processing circuit 74. Furthermore, the input / output interface 70 is also connected to a control device 36 for the interchangeable lens 18.

[0100] The controller 44 includes a CPU 62, an NVM 64, and a RAM 66. The CPU 62, NVM 64, and RAM 66 are interconnected via a bus 68, which is connected to an input / output interface 70.

[0101] In addition, Figure 2 In the example shown, for ease of illustration, bus 68 is depicted as a single bus, but multiple buses are also possible. Bus 68 can be a serial bus or a parallel bus that includes a data bus, an address bus, and a control bus.

[0102] NVM64 is a non-transitory storage medium that stores various parameters and programs. For example, NVM64 is an EEPROM. However, this is just one example; it can also replace EEPROM or be used in conjunction with HDDs and / or SSDs as NVM64. Furthermore, RAM66 temporarily stores various information and is used as working memory.

[0103] CPU 62 reads the necessary programs from NVM 64 and executes the read programs on RAM 66. CPU 62 controls the entire camera device 12 according to the programs executed on RAM 66. Figure 2 In the example shown, the image memory 46, UI device 48, external I / F 50, communication I / F 52, photoelectric conversion element driver 54, mechanical shutter driver 56, and control device 36 are controlled by CPU 62.

[0104] A photoelectric conversion element driver 54 is connected to the photoelectric conversion element 72. The photoelectric conversion element driver 54 supplies a shooting timing signal to the photoelectric conversion element 72 according to an instruction from the CPU 62. This shooting timing signal specifies the timing for shooting performed by the photoelectric conversion element 72. The photoelectric conversion element 72 performs reset, exposure, and outputs electrical signals based on the shooting timing signal supplied from the photoelectric conversion element driver 54. Examples of shooting timing signals include, for example, a vertical synchronization signal and a horizontal synchronization signal.

[0105] With the interchangeable lens 18 mounted on the main body 16 of the imaging device, the subject light incident on the imaging lens 40 is imaged onto the light-receiving surface 72A by the imaging lens 40. Under the control of the photoelectric conversion element driver 54, the photoelectric conversion element 72 performs photoelectric conversion on the subject light received by the light-receiving surface 72A and outputs an electrical signal corresponding to the amount of subject light to the signal processing circuit 74 as analog image data representing the subject light. Specifically, the signal processing circuit 74 reads analog image data from the photoelectric conversion element 72 in exposure sequence for each horizontal line, frame by frame.

[0106] The signal processing circuit 74 generates digital image data by digitizing analog image data. Furthermore, for ease of explanation, the term "captured image 108" will be used hereafter without distinguishing between the digital image data, which is the object of internal processing within the camera device body 16, and the image represented by the digital image data (i.e., the image displayed on the display 28, etc., based on the visualization of the digital image data).

[0107] Mechanical shutter 60 is a focal plane shutter, positioned between aperture 40D and light-receiving surface 72A. Mechanical shutter 60 has a front curtain (not shown) and a rear curtain (not shown). Both the front and rear curtains have multiple blades. The front curtain is positioned closer to the subject than the rear curtain.

[0108] The mechanical shutter actuator 58 is an actuator comprising a linkage mechanism (not shown), a front curtain solenoid (not shown), and a rear curtain solenoid (not shown). The front curtain solenoid is the drive source for the front curtain and is mechanically connected to the front curtain via the linkage mechanism. The rear curtain solenoid is the drive source for the rear curtain and is mechanically connected to the rear curtain via the linkage mechanism. The mechanical shutter driver 56 controls the mechanical shutter actuator 58 according to instructions from the CPU 62.

[0109] The front curtain, powered by a solenoid under the control of the mechanical shutter driver 56, selectively raises and lowers the front curtain. Similarly, the rear curtain, also powered by a solenoid under the control of the mechanical shutter driver 56, selectively raises and lowers the rear curtain. In the imaging device 12, the exposure to the photoelectric conversion element 72 is controlled by the opening and closing of the front and rear curtains, controlled by the CPU 62.

[0110] In the imaging device 12, real-time preview images are captured in exposure sequence readout mode (rolling shutter mode), and recorded images are captured for recording still images and / or moving images. The image sensor 20 has an electronic shutter function, and the real-time preview image capture is achieved by activating the electronic shutter function while keeping the mechanical shutter 60 fully open and not operating it.

[0111] In contrast, shooting with formal exposure (i.e., shooting still images) is achieved by activating the electronic shutter function and using the mechanical shutter 60 to transition the mechanical shutter 60 from the front curtain closed state to the rear curtain closed state.

[0112] The captured image 108 generated by the signal processing circuit 74 is stored in the image memory 46. That is, the signal processing circuit 74 causes the image memory 46 to store the captured image 108. The CPU 62 retrieves the captured image 108 from the image memory 46 and uses the retrieved captured image 108 to perform various processes.

[0113] The UI device 48 includes a display 28, and the CPU 62 enables the display 28 to display various information. Furthermore, the UI device 48 includes a receiving device 76. The receiving device 76 includes a touch panel 30 and a hard key unit 78. The hard key unit 78 includes indicator keys 26 (see reference). Figure 1 The CPU 62 operates according to various instructions received via the touch panel 30. Furthermore, while the hard key unit 78 is included in the UI system device 48, the technology of the present invention is not limited thereto; for example, the hard key unit 78 may also be connected to an external I / F 50.

[0114] The external I / F50 controls the exchange of various information with devices located outside the camera device 12 (hereinafter also referred to as "external devices"). An example of an external I / F50 is a USB interface. External devices (not shown) such as smart devices, personal computers, servers, USB storage devices, memory cards, and / or printers are connected directly or indirectly to the USB interface.

[0115] Communication I / F52 via network 34 (reference) Figure 1 Control CPU62 and camera support device 14 (reference) Figure 1 The communication I / F 52 exchanges information with the camera support device 14 via the network 34, corresponding to a request from the CPU 62. Furthermore, the communication I / F 52 receives information from the camera support device 14 and outputs the received information to the CPU 62 via the input / output interface 70.

[0116] As an example, such as Figure 3 As shown, the NVM64 stores a display control processing program 80. The CPU62 reads the display control processing program 80 from the NVM64 and executes the read display control processing program 80 on the RAM66. The CPU62 performs display control processing according to the display control processing program 80 executed on the RAM66 (see reference). Figure 30 ).

[0117] CPU 62 operates as the acquisition unit 62A, generation unit 62B, transmission unit 62C, receiving unit 62D, and display control unit 62E by executing display control processing program 80. The specific processing details of the acquisition unit 62A, generation unit 62B, transmission unit 62C, receiving unit 62D, and display control unit 62E will be discussed later. Figure 16 , Figure 17 , Figure 28 , Figure 29 and Figure 30 To narrate.

[0118] As an example, such as Figure 4 As shown, the camera support device 14 includes a computer 82 and a communication I / F 84. The computer 82 includes a CPU 86, a storage unit 88, and a memory 90. Here, the computer 82 is an example of a "computer" according to the technology of the present invention, the CPU 86 is an example of a "processor" according to the technology of the present invention, and the memory 90 is an example of a "memory" according to the technology of the present invention.

[0119] The CPU 86, storage unit 88, memory 90, and communication I / F 84 are connected to the bus 92. Additionally, in Figure 4 In the example shown, for ease of illustration, bus 92 is depicted as a single bus, but multiple buses are also possible. Bus 92 can be a serial bus or a parallel bus that includes a data bus, an address bus, and a control bus.

[0120] CPU 86 controls the entire camera support device 14. Storage unit 88 is a non-temporary storage medium, a non-volatile storage device that stores various programs and parameters. Examples of storage units 88 include EEPROM, SSD, and / or HDD. Memory 90 is a temporary storage memory that is used as working memory by CPU 86. Examples of memory 90 include RAM.

[0121] Communication I / F 84 is connected to communication I / F 52 of camera device 12 via network 34. Communication I / F 84 controls the exchange of information between CPU 86 and camera device 12. For example, communication I / F 84 receives information sent from camera device 12 and outputs the received information to CPU 86. Furthermore, it sends information corresponding to requests from CPU 62 to camera device 12 via network 34.

[0122] Storage unit 88 stores the learned model 93. The learned model 93 is obtained by using a convolutional neural network (i.e., CNN96, reference 96). Figure 5 The model is generated through learning. The CPU 86 uses the learned model 93 to identify the subject based on the captured image 108.

[0123] Here, for reference Figure 5 and Figure 6 An example of how to create the learned model 93 (i.e., an example of the learning phase) will be explained.

[0124] As an example, such as Figure 5 As shown, the storage unit 88 stores the learning execution processing program 94 and CN N96. The CPU 86 reads the learning execution processing program 94 from the storage unit 88 and executes the read learning execution processing program 94, thereby operating as the learning stage calculation unit 87A, the error calculation unit 87B, and the adjustment value calculation unit 87C.

[0125] During the learning phase, a training data supply device 98 is used. The training data supply device 98 stores training data 100 and supplies training data 100 to the CPU 86. The training data 100 includes multiple learning images 100A and multiple correct answer data 100B. A one-to-one correspondence is established between the correct answer data 100B and the multiple learning images 100A.

[0126] The learning phase calculation unit 87A acquires the learning image 100A. The error calculation unit 87B acquires the correct answer data 100B corresponding to the learning image 100A acquired by the learning phase calculation unit 87A.

[0127] CNN96 has an input layer 96A, multiple intermediate layers 96B, and an output layer 96C. The learning phase computation unit 87A extracts feature data representing the characteristics of a subject determined based on the learning image 100A by passing it through the input layer 96A, multiple intermediate layers 96B, and the output layer 96C. Here, the features of the subject refer to, for example, comprehensive features related to contours, hues, surface texture, part features, and overall characteristics. The output layer 96C determines which cluster the multiple feature data extracted through the input layer 96A and multiple intermediate layers 96B belong to and outputs a CNN signal 102 representing the determination result. Here, a cluster refers to the collective feature data of each of the multiple subjects. The determination result is, for example, a category 110C (referencing the cluster) representing the feature data and the category determined based on the cluster. Figure 7 Information regarding the probability correspondence between the feature data and category 110C (e.g., information indicating the probability that the correspondence between the feature data and category 110C is the correct answer). Category 110C refers to the type of subject. Furthermore, category 110C is an example of the "type of subject" in the technology of this invention.

[0128] The correct answer data 100B is data that is pre-determined to be equivalent to the ideal CNN signal 102 output from output layer 96C. Correct answer data 100B includes information that establishes a correspondence between feature data and category 110C.

[0129] Error calculation unit 87B calculates the error 104 between the CNN signal 102 and the correct answer data 100B. Adjustment value calculation unit 87C calculates multiple adjustment values ​​106 that minimize the error 104 calculated by error calculation unit 87B. Learning phase operation unit 87A uses the multiple adjustment values ​​106 to adjust multiple optimization variables within CNN96 to minimize the error 104. Here, multiple optimization variables refer to, for example, multiple connection weights and multiple offset values ​​included in CNN96.

[0130] During the learning phase, the computation unit 87A uses multiple adjustment values ​​106 calculated by the adjustment value calculation unit 87C to adjust multiple optimization variables within the CNN96 for each of the multi-frame learning images 100A, thereby minimizing the error 104 and optimizing the CNN96. Furthermore, as an example, such as... Figure 6 As shown, CNN96 is optimized by adjusting multiple optimization variables, thereby constructing a learned model 93.

[0131] As an example, such as Figure 7 As shown, the learned model 93 has an input layer 93A, multiple intermediate layers 93B, and an output layer 93C. The input layer 93A is an optimized... Figure 5 The layer shown is derived from the input layer 96A, and the multiple intermediate layers 93B are obtained through optimization. Figure 5 The layer shown is obtained by multiple intermediate layers 96B, and the output layer 93C is obtained by optimization. Figure 5 The layer shown is obtained from the output layer 96C.

[0132] In the camera support device 14, the captured image 108 is applied to the input layer 93A of the learned model 93, and subject determination information 110 is output from the output layer 93C. The subject determination information 110 includes bounding box position information 110A, objectivity score 110B, category 110C, and category score 110D. The bounding box position information 110A is a bounding box 116 (refer to) that can determine the area within the captured image 108. Figure 9 The relative position information of the image. In addition, the captured image 108 is shown here, but this is only one example, and it can also be an image based on the captured image 108 (e.g., real-time preview image 130).

[0133] The objectivity score 110B represents the probability that an object exists within bounding box 116. This example illustrates the probability of an object existing within bounding box 116, but it is merely an example and can be a value obtained by fine-tuning the probability of an object existing within bounding box 116, as long as it is based on the probability of an object existing within bounding box 116.

[0134] Category 110C represents the type of the subject. Category score 110D represents the probability that an object within bounding box 116 belongs to a specific category 110C. Here, the probability of an object within bounding box 116 belonging to a specific category 110C is illustrated, but this is only one example. It can also be a value obtained by fine-tuning the probability of an object within bounding box 116 belonging to a specific category 110C, as long as it is based on the probability of an object within bounding box 116 belonging to a specific category 110C.

[0135] Furthermore, as an example of a specific category 110C, examples include a specific person, a specific person's face, a specific car, a specific passenger plane, a specific bird, and a specific tram. Additionally, there are multiple specific categories 110C, each assigned a category score 110D.

[0136] If the captured image 108 is input into the learned model 93, then as an example, Figure 8 As shown, an anchor frame 112 is applied to the captured image 108. The anchor frame 112 is a collection of multiple virtual frames, each with a pre-defined height and width. Figure 8 In the example shown, as an example of anchor frame 112, a first virtual frame 112A, a second virtual frame 112B, and a third virtual frame 112C are shown. The first virtual frame 112A and the third virtual frame 112C have the same shape as each other (in...). Figure 8 In the example shown, the first virtual frame 112A, the second virtual frame 112B, and the third virtual frame 112C are rectangular and the same size. The second virtual frame 112B is square. Within the captured image 108, the centers of the first virtual frame 112A, the second virtual frame 112B, and the third virtual frame 112C coincide with each other. Furthermore, the outer frame of the captured image 108 is rectangular, and the first virtual frame 112A, the second virtual frame 112B, and the third virtual frame 112C are arranged within the captured image 108 with the short side of the first virtual frame 112A, a specific side of the second virtual frame 112B, and the long side of the third virtual frame 112C parallel to a specific side of the outer frame of the captured image 108.

[0137] exist Figure 8 In the example shown, the captured image 108 includes a person image 109, and the first virtual frame 112A, the second virtual frame 112B, and the third virtual frame 112C each include a portion of the person image 109. Furthermore, for ease of explanation, the first virtual frame 112A, the second virtual frame 112B, and the third virtual frame 112C will be referred to as "virtual frames" without further distinction.

[0138] The captured image 108 is divided into multiple virtual units 114. If the captured image 108 is input into the learned model 93, the CPU 86 sets the center of each unit 114 to the center of the anchor frame 112 and calculates the anchor frame score and the anchor category score.

[0139] Anchor frame score refers to the probability that the image of person 109 is within anchor frame 112. The anchor frame score is a probability derived from the probability that the image of person 109 is within the first virtual frame 112A (hereinafter referred to as the "first virtual frame probability"), the probability that the image of person 109 is within the second virtual frame 112B (hereinafter referred to as the "second virtual frame probability"), and the probability that the image of person 109 is within the third virtual frame 112C (hereinafter referred to as the "third virtual frame probability"). For example, the anchor frame score is the average of the first virtual frame probability, the second virtual frame probability, and the third virtual frame probability.

[0140] Anchor category score refers to the image within the anchor frame 112 (in Figure 8 In the example shown, the probability that the subject (e.g., a person) represented by the image 109 belongs to a specific category 110C (e.g., a specific person) is an anchor category score. The anchor category score is a probability derived from the probability that the subject represented by the image within the first virtual frame 112A belongs to a specific category 110C (hereinafter referred to as the "first virtual frame category probability"), the probability that the subject represented by the image within the second virtual frame 112B belongs to a specific category 110C (hereinafter referred to as the "second virtual frame category probability"), and the probability that the subject represented by the image within the third virtual frame 112C belongs to a specific category 110C (hereinafter referred to as the "third virtual frame category probability"). For example, the anchor category score is the average of the first virtual frame category probability, the second virtual frame category probability, and the third virtual frame category probability.

[0141] CPU86 calculates the anchor frame value when applying anchor frame 112 to the captured image 108. The anchor frame value is obtained based on cell determination information, anchor frame score, anchor category score, and anchor frame constant. Here, cell determination information refers to information that determines the width, height, and position of cell 114 within the captured image 108 (e.g., two-dimensional coordinates that can determine the position within the captured image 108). Anchor frame constant is a constant that is preset for the type of anchor frame 112.

[0142] The anchor frame value is calculated according to the formula "(anchor frame value) = (number of all units 114 present in the captured image 108) × {(unit determination information), (anchor frame score), (anchor category score)} × (anchor frame constant)".

[0143] The CPU 86 calculates anchor frame values ​​by applying anchor frame 112 to each unit 114. Furthermore, the CPU 86 determines anchor frames 112 with anchor frame values ​​exceeding an anchor frame threshold. The anchor frame threshold can be a fixed value or a variable value that changes based on instructions given to the camera support device 14 and / or various conditions.

[0144] CPU86 determines the maximum virtual frame probability based on the first, second, and third virtual frame probabilities of anchor frame 112 with an anchor frame value exceeding the anchor frame threshold. The maximum virtual frame probability is the one with the highest value among the first, second, and third virtual frame probabilities.

[0145] As an example, such as Figure 9 As shown, CPU86 identifies the virtual box with the highest virtual box probability as bounding box 116 and deletes the remaining virtual boxes. Figure 9 In the example shown, within the captured image 108, the first virtual frame 112A is determined as the bounding box 116, and the second virtual frame 112B and the third virtual frame 112C are deleted. The bounding box 116 is used as the AF region box 134. The AF region box 134 is an example of the "box" and "focus box" involved in the technology of this invention. The AF region box 134 is a box that surrounds the subject. Within the captured image 108, the subject image (in the image representing a specific subject (e.g., a specific person) is defined as the frame that surrounds the subject (e.g., a specific person). Figure 8 The example shown is the frame for the image 109. The AF region frame 134 refers to the focus frame that defines the candidate area (hereinafter also referred to as the "focus candidate area") for focusing the lens 40B in AF mode. Furthermore, while AF region frame 134 is shown here, it is only one example and can also be used as the focus frame for defining the focus candidate area in MF mode.

[0146] CPU 86 infers the category 110C and category score 110D for each unit 114 within the bounding box 116, and establishes a correspondence between the category 110C and category score 110D and each unit 114. Then, CPU 86 extracts information from the learned model 93, including the inference results inferred for each unit 114 within the bounding box 116, as subject determination information 110.

[0147] exist Figure 9 In the example shown, the subject determination information 110 includes bounding box position information 110A regarding the bounding box 116 within the captured image 108, objectivity score 110B regarding the bounding box 116 within the captured image 108, multiple categories 110C obtained based on categories 110C corresponding to each unit 114 within the bounding box 116, and multiple category scores 110D obtained based on category scores 110D corresponding to each unit 114 within the bounding box 116.

[0148] The objectivity score 110B included in the subject determination information 110 is the probability that the subject exists within the bounding box 116 determined according to the bounding box position information 110A included in the subject determination information 110. For example, an objectivity score 110B of 0% indicates that there is no subject image representing the subject (e.g., person image 109, etc.) within the bounding box 116, while an objectivity score 110B of 100% indicates that there is definitely a subject image representing the subject within the bounding box 116.

[0149] In the subject determination information 110, each category 110C has a category score 110D. That is, the multiple categories 110C included in the subject determination information 110 each have an independent category score 110D. The category score 110D included in the subject determination information 110 is, for example, the average of all category scores 110D that establish a correspondence between the corresponding category 110C and all units 114 within the bounding box 116.

[0150] The CPU86 determines that the category 110C with the highest category score 110D among multiple categories 110C is most likely the category 110C to which the subject represented by the subject image within the bounding box 116 belongs. The most likely category 110C is the category 110C corresponding to the highest overall category score 110D among all category scores 110D that correspond to each unit 114 within the bounding box 116. The highest overall category score 110D, for example, refers to the category 110C with the highest average of all category scores 110D corresponding to all units 114 within the bounding box 116.

[0151] The method of identifying a subject using subject identification information 110 is a subject recognition method utilizing AI (hereinafter also referred to as "AI subject recognition method"). As another subject recognition method, a method of identifying a subject through template matching (hereinafter also referred to as "template matching method") can be cited. In the template matching method, for example, such as... Figure 10 As shown, subject recognition template 118 is used. In Figure 10 In the example shown, the CPU 86 uses a subject recognition template 118, which is an image representing a reference subject (e.g., a face created through machine learning, etc., as a typical human face). The CPU 86 performs raster scanning within the captured image 108 using the subject recognition template 118 while changing its size. Figure 10In the example shown, a scan is performed along the direction of the straight arrow to detect the person image 109 included in the captured image 108. In the template matching mode, the CPU 86 measures the difference between the entire subject recognition template 118 and a portion of the captured image 108, and calculates the matching degree (hereinafter also referred to as "matching degree") between the subject recognition template 118 and a portion of the captured image 108 based on the measured difference.

[0152] As an example, such as Figure 11 As shown, the matching degree is radially distributed from the pixel with the highest matching degree within the captured image 108. Figure 11 In the example shown, the outer frame 118A of the subject recognition template 118 is radially distributed from the center. In this case, the outer frame 118A, which is configured with the pixel having the maximum matching degree within the captured image 108 as the AF region frame 134, is used as the AF region frame 134.

[0153] Thus, in the template matching method, the matching degree between the template and the subject recognition template 118 is calculated. In contrast, in the AI ​​subject recognition method, the probability of a specific subject (e.g., a specific category 110C) is inferred. Therefore, the metrics used to evaluate subjects in the template matching and AI subject recognition methods are completely different. In the template matching method, a single value similar to the category score 110D (i.e., a single value that mixes object score 110B and category score 110D and is closer to category score 110D than object score 110B) is used as the matching degree for subject evaluation. In contrast, in the AI ​​subject recognition method, both object score 110B and category score 110D are used for subject evaluation.

[0154] In template matching, even if the matching degree changes when the conditions of the subject change, the matching degree is radially distributed from the pixel with the highest matching degree within the captured image 108. Therefore, for example, if the matching degree within the outer frame 118A is above a threshold equivalent to the category threshold used for comparison with category score 110D when determining the reliability of category score 110D, it may sometimes be identified as a specific category 110C or as a category 110C other than a specific category 110C. That is, as an example, such as... Figure 12As shown, category 110C is sometimes correctly identified, but sometimes incorrectly identified. If the captured image 108 obtained by the camera device 12 is used in a template matching manner under a constant environment (e.g., an environment where conditions such as the subject and lighting are constant), the subject is considered to be almost correctly identified. However, if the captured image 108 obtained by the camera device 12 is used in a template matching manner under a changing environment (e.g., an environment where conditions such as the subject and lighting are changing), the subject is more likely to be incorrectly identified compared to the case when the image was captured under a constant environment.

[0155] In contrast, in AI subject recognition methods, as an example, such as Figure 12 As shown, the object threshold used for comparison with the object score 110B is used when determining the reliability of the object score 110B. In the AI ​​subject recognition method, if the object score 110B is above the object threshold, the subject image is determined to be within the bounding box 116. Furthermore, in the AI ​​subject recognition method, assuming the subject image has been determined to be within the bounding box 116, if the category score 110D is above the category threshold, the subject is determined to belong to a specific category 110C. If the category score 110D is less than the category threshold, even if the subject image has been determined to be within the bounding box 116, the subject is determined not to belong to the specific category 110C.

[0156] Furthermore, when a horse image representing a horse is used as a subject recognition template 118 and a zebra image representing a zebra is used for subject recognition in a template-matching manner, the zebra image is not recognized. In contrast, when a learned model 93, obtained by machine learning using a horse image as a learning image 100A, is used for subject recognition of the zebra image in an AI subject recognition manner, the objectivity score 110B is above the objectivity threshold, and the category score 110D is below the category threshold, thus determining that the zebra represented by the zebra image does not belong to the specific category 110C (here, as an example, a horse). That is, the zebra represented by the zebra image is identified as a foreign object other than a specific subject.

[0157] Thus, the difference in subject recognition results between template matching and AI subject recognition is because, in template matching, the difference in stripes between a horse and a zebra affects the matching degree, and the matching degree is a unique value used for subject recognition. In contrast, in AI subject recognition, the determination of whether a subject is a specific subject is made probabilistically based on a set of local feature data such as the horse's overall shape, face, four legs, and tail.

[0158] In this embodiment, the AI ​​subject recognition method is used to control the AF region frame 134. Examples of controlling the AF region frame 134 include adjusting the size of the AF region frame 134 according to the subject and selecting a segmentation method for dividing the AF region frame 134 according to the subject.

[0159] In the AI ​​subject recognition method, bounding box 116 is used as AF region box 134. Therefore, it can be said that the reliability of the size of AF region box 134 when the objectivity score 110B is high (e.g., when the objectivity score 110B is above the objectivity threshold) is higher than the reliability of the size of AF region box 134 when the objectivity score 110B is low (e.g., when the objectivity score 110B is less than the objectivity threshold). Conversely, it can be said that the reliability of the size of AF region box 134 when the objectivity score 110B is low is lower than the reliability of the size of AF region box 134 when the objectivity score 110B is high, thus increasing the likelihood that the subject is outside the focus candidate region. In this case, for example, AF region box 134 may be enlarged or distorted.

[0160] In conventional techniques, when segmenting the AF region frame 134, the number of segments is changed according to the size of the subject image enclosed by the AF region frame 134. However, if the number of segments is changed according to the size of the subject image, then the segmentation line 134A1 (reference) Figure 22 and Figures 25-27 (That is, multiple segmented regions 134A obtained by segmentation (reference)) Figure 22 and Figures 25-27 The boundary line between the two points () sometimes overlaps with the part corresponding to the important position of the subject. It is predicted that the important position of the subject is generally more likely to be the focus of attention than a position that is completely unrelated to the important position of the subject.

[0161] For example, such as Figure 13 As shown, when the subject is a person 120, the eyes of the person 120 can be considered as one of the important positions of the subject. Similarly, when the subject is a bird 122, the eyes of the bird 122 can be considered as one of the important positions of the subject. Furthermore, when the subject is a car 124, the position of the car logo on the front of the car 124 can be considered as one of the important positions of the subject.

[0162] Furthermore, when the camera device 12 photographs the subject from the front side, the length of the subject in the depth direction (i.e., the length from the reference position to the important position of the subject) varies depending on the subject. Regarding the reference position of the subject when the camera device 12 photographs the subject from the front side, for example, it might be the tip of the nose of person 120, the tip of the beak of bird 122, or the area between the roof and windshield of car 124. Thus, although the length from the reference position to the important position of the subject varies depending on the subject, if the focus is on the reference position of the subject during the photograph, the image representing the important position of the subject within the captured image 108 will become blurry. In this case, as long as the part corresponding to the important position of the subject is completely included in the segmented area 134A of the AF area frame 134, it is easier to focus on the important position of the subject in the captured image 108 compared to the case where the part corresponding to the important position of the subject does not enter the segmented area 134A of the AF area frame 134, but the part corresponding to the reference position of the subject enters the segmented area 134A of the AF area frame 134.

[0163] Therefore, in order to ensure that the portion corresponding to the important position of the subject is completely included within the segmented area of ​​the AF region frame 134, as an example, such as Figure 14 As shown, in addition to the learned model 93, the storage unit 88 of the camera support device 14 also stores a camera support processing program 126 and a segmentation method export table 128. The camera support processing program 126 is an example of a "program" involved in the technology of this invention.

[0164] As an example, such as Figure 15 As shown, the CPU 86 reads the camera support processing program 126 from the storage unit 88 and executes the read camera support processing program 126 on the memory 90. The CPU 86 performs camera support processing based on the camera support processing program 126 executed on the memory 90 (see also...). Figure 31A and Figure 31B ).

[0165] By performing camera support processing, the CPU 86 obtains category 110C based on the captured image 108 and outputs information indicating the segmentation method for dividing the region according to the obtained category 110C. This segmentation divides the subject into regions that can be distinguished from other regions within the shooting range. Furthermore, the CPU 86 obtains category 110C based on subject determination information 110 output from the learned model 93 by applying it to the captured image 108. Here, subject determination information 110 is an example of the "output result" involved in the technology of this invention.

[0166] CPU 86 operates as a receiving unit 86A, an execution unit 86B, a determination unit 86C, a generation unit 86D, an amplification unit 86E, a segmentation unit 86F, an output unit 86G, and a transmitting unit 86H by executing camera support processing program 126. The specific processing details of the receiving unit 86A, execution unit 86B, determination unit 86C, generation unit 86D, amplification unit 86E, segmentation unit 86F, output unit 86G, and transmitting unit 86H will be discussed later. Figures 17-28 , Figure 31A and Figure 31B To narrate.

[0167] As an example, such as Figure 16 As shown, in the imaging device 12, when the captured image 108 has already been stored in the image memory 46, the acquisition unit 62A acquires the captured image 108 from the image memory 46. The generation unit 62B generates a real-time preview image 130 based on the captured image 108 acquired by the acquisition unit 62A. The real-time preview image 130 is, for example, an image obtained by periodically removing pixels from the captured image 108 according to predetermined rules. Figure 16 In the example shown, as a live preview image 130, an image including an image 132 representing an aircraft is displayed. The transmitting unit 62C transmits the live preview image 130 to the camera support device 14 via the communication I / F 52.

[0168] As an example, such as Figure 17 As shown, in the camera support device 14, the receiving unit 86A receives the real-time preview image 130 sent from the transmitting unit 62C of the camera device 12. The execution unit 86B performs subject recognition processing based on the real-time preview image 130 received by the receiving unit 86A. Here, subject recognition processing refers to processing that includes detecting a subject using AI subject recognition (i.e., determining whether a subject exists) and determining the type of the subject using AI subject recognition.

[0169] The execution unit 86B uses the learned model 93 in the storage unit 88 to perform subject recognition processing. The execution unit 86B extracts subject determination information 110 from the learned model 93 by assigning an instant preview image 130 to the learned model 93.

[0170] As an example, such as Figure 18As shown, the execution unit 86B stores the subject determination information 110 in the memory 90. The determination unit 86C retrieves the object score 110B from the subject determination information 110 in the memory 90 and determines whether the subject exists by referring to the retrieved object score 110B. For example, if the object score 110B is greater than "0", the subject is determined to exist; if the object score 110B is "0", the subject is determined not to exist. If the determination unit 86C determines that the subject exists, the CPU 86 performs the first determination process. If the determination unit 86C determines that the subject does not exist, the CPU 86 performs the first segmentation process.

[0171] As an example, such as Figure 19 As shown, in the first determination process, the determination unit 86C obtains the object score 110B from the subject determination information 110 in the memory 90 and determines whether the obtained object score 110B is above or above the object threshold. If the object score 110B is above or above the object threshold, the CPU 86 performs the second determination process. If the object score 110B is less than the object threshold, the CPU 86 performs the second segmentation process. Furthermore, the object score 110B above the object threshold is an example of the "output result" and "value based on the probability that an object within a bounding box applicable to an image belongs to a specific category" according to the technology of this invention. Also, the object threshold when the camera support process enters the second segmentation process is an example of the "first threshold" according to the technology of this invention.

[0172] As an example, such as Figure 20 As shown, in the second determination process, the determination unit 86C obtains the category score 110D from the subject determination information 110 in the memory 90 and determines whether the obtained category score 110D is above or above the category threshold. If the category score 110D is above or above the category threshold, the CPU 86 performs the third segmentation process. If the category score 110D is less than the category threshold, the CPU 86 performs the first segmentation process. Furthermore, the category score 110D above the category threshold is an example of the "output result" and "value above the second threshold among the values ​​of the probability that the object belongs to a specific category" according to the technology of this invention. And the category threshold is an example of the "second threshold" according to the technology of this invention.

[0173] As an example, such as Figure 21As shown, in the first segmentation process, the generation unit 86D obtains the bounding box position information 110A from the subject determination information 110 in the memory 90, and generates the AF region box 134 based on the obtained bounding box position information 110A. The segmentation unit 86F calculates the area of ​​the AF region box 134 generated by the generation unit 86D, and segments the AF region box 134 according to the segmentation method corresponding to the calculated area. After segmenting the AF region box 134 according to the segmentation method corresponding to the area, the CPU 86 performs the image generation and transmission process with the AF region box.

[0174] The segmentation method specifies the number of segments into which the AF region frame 134 is divided, forming a region (hereinafter also simply referred to as a "divided region") that can be distinguished from other regions within the captured image 108, enclosed by the AF region frame 134. The divided region is defined by a vertical direction and a horizontal direction orthogonal to the vertical direction. The vertical direction corresponds to the row direction in the defined row and column directions of the captured image 108, and the horizontal direction corresponds to the column direction. The segmentation method specifies the number of segments in the vertical direction and the number of segments in the horizontal direction. Furthermore, the row direction is an example of the "first direction" according to the technology of this invention, and the column direction is an example of the "second direction" according to the technology of this invention.

[0175] As an example, such as Figure 22 As shown, in the first segmentation process, the segmentation unit 86F changes the number of segments (i.e., the number of segmented regions 134A) of the AF region frame 134 within a predetermined range (e.g., within a range where the lower limit is set to 2 and the upper limit is set to 30) based on the area of ​​the AF region frame 134. That is, the larger the area of ​​the segmented region 134A, the more the segmentation unit 86F increases the number of segmented regions 134A within the predetermined range. Specifically, the larger the area of ​​the segmented region 134A, the more the segmentation unit 86F increases the number of segments in both the vertical and horizontal directions.

[0176] Furthermore, the fewer the number of segmented regions 134A, the less the number of segmented regions 134A the segmenting part 86F reduces within a predetermined range. Specifically, the fewer the number of segmented regions 134A, the less the number of segments in the longitudinal and transverse directions the segmenting part 86F reduces.

[0177] As an example, such as Figure 23 As shown, in the second segmentation process, the generation unit 86D obtains the bounding box position information 110A from the subject determination information 110 in the memory 90, and generates the AF region frame 134 based on the obtained bounding box position information 110A. The enlargement unit 86E enlarges the AF region frame 134 at a predetermined magnification (e.g., 1.25x). Furthermore, the predetermined magnification can be a variable value that changes according to instructions given to the camera support device 14 and / or various conditions, or it can be a fixed value.

[0178] In the second segmentation process, the segmentation unit 86F calculates the area of ​​the AF region box 134 expanded by the enlargement unit 86E, and segments the AF region box 134 according to the segmentation method corresponding to the calculated area. After segmenting the AF region box 134 according to the segmentation method corresponding to the area, the CPU 86 performs image generation and transmission processing with the AF region box.

[0179] As an example, such as Figure 24 As shown, the partitioning method derivation table 128 within storage unit 88 establishes a correspondence between category 110C and partitioning methods. There are multiple categories 110C, and the partitioning methods are inherent to each of these categories. Figure 24 In the example shown, segmentation method A corresponds to category A, segmentation method B corresponds to category B, and segmentation method C corresponds to category C.

[0180] The exporting unit 86G retrieves the category 110C corresponding to the category score 110D above the category threshold from the subject determination information 110 in the memory 90, and exports the segmentation method corresponding to the retrieved category 110C from the segmentation method export table 128. The generating unit 86D retrieves the bounding box position information 110A from the subject determination information 110 in the memory 90, and generates the AF region box 134 based on the retrieved bounding box position information 110A. The segmentation unit 86F segments the AF region box 134 generated by the generating unit 86D according to the segmentation method exported by the exporting unit 86G. Segmenting the AF region box 134 according to the segmentation method exported by the exporting unit 86G means dividing the subject into a segmented region that can be distinguished from other areas of the shooting range according to the segmentation method exported by the exporting unit 86G.

[0181] Generally, when shooting a subject from an angle rather than from the front, a subject image with a sense of depth can be obtained compared to shooting it from the front. Thus, when shooting with a composition that gives a sense of depth to a person (hereinafter also referred to as "the viewer") viewing the instant preview image 130, it is important to make the number of vertical divisions of the AF area frame 134 greater than the number of horizontal divisions, or vice versa, according to category 110C.

[0182] Therefore, the number of vertical and horizontal divisions of the AF region frame 134 based on the division method exported by the export unit 86G is determined according to the composition of the subject image surrounded by the AF region frame 134 in the instant preview image 130.

[0183] As an example, such as Figure 25As shown, when the aircraft is photographed from a lower angle using the camera device 12, the segmentation unit 86F divides the AF region frame 134 into a frame that surrounds the aircraft image 132 within the real-time preview image 130. The number of horizontal divisions of the AF region frame 134 is greater than the number of vertical divisions. Figure 25 In the example shown, the AF region box 134 is divided into 24 horizontally and 10 vertically. Furthermore, the AF region box 134 is equally divided both horizontally and vertically. Thus, the region enclosed by the AF region box 134 is divided into 240 (=24×10) segmented regions 134A. Figure 25 In the example shown, the 240 segmented regions 134A are an example of the "multiple segmented regions" involved in the technology of this invention.

[0184] Furthermore, as an example, such as Figure 26 As shown, the person 120 (reference) is photographed from the front side, surrounded by the AF area frame 134. Figure 13 When the facial image 120A is obtained from the face of the AF region box 134 and the facial image 120A is located at the first predetermined position of the AF region box 134, the segmentation unit 86F divides the AF region box 134 into an even number of equal parts in the horizontal direction (in the case of the face image 120A obtained from the face image 120A and the face image 120A being located at the first predetermined position of the AF region box Figure 26 In the example shown, there are "8". The first predetermined position refers to the position where the part corresponding to the nose of person 120 is located at the horizontal center of the AF region box 134. Thus, by dividing the AF region box 134 into an even number of equal parts horizontally, it is easier to ensure that the part 120A1 corresponding to the eyes in the facial image 120A within the AF region box 134 is completely contained within a segmented region 134A, compared to dividing the AF region box 134 into an odd number of equal parts horizontally. In other words, this means that compared to dividing the AF region box 134 into an odd number of equal parts horizontally, the part 120A1 corresponding to the eyes in the facial image 120A is less likely to overlap with the segmentation line 134A1. In addition, the AF region box 134 is also divided equally vertically. The number of vertical segments is less than the number of horizontal segments (in Figure 26 (The example shown is "4"). Figure 26 In the example shown, the AF region box 134 is also segmented vertically, but this is only one example, and it is also possible not to segment vertically. The size of the segmented region 134A is a size that is pre-derived by a simulator or the like, so that the portion 120A1 corresponding to the eye in the face image 120A is completely included in a segmented region 134A.

[0185] Furthermore, as an example, such as Figure 27 As shown, the car 124 is photographed from the front side, surrounded by the AF area frame 134 (reference). Figure 13Given a car image 124A obtained and located at the second predetermined position of the AF region box 134, the segmentation unit 86F divides the AF region box 134 horizontally into an odd number of segments (in...). Figure 27 (In the example shown, "7"). The second predetermined position refers to the position where the center of the car's front view (e.g., the car logo) is located in the horizontal center of the AF region box 134. Thus, by dividing the AF region box 134 into an odd number of equal parts horizontally, it is easier to completely include the portion 124A1 of the car image 124A corresponding to the car logo within a segmented region 134A compared to dividing it into an even number of equal parts horizontally.

[0186] In other words, this means that, compared to the case where the AF region box 134 is divided into an even number of equal parts horizontally, the part 124A1 corresponding to the car logo in the car image 124A is less likely to overlap with the dividing line 134A1.

[0187] Furthermore, the AF region box 134 is also equally divided vertically. The number of vertical divisions is less than the number of horizontal divisions (in... Figure 27 (The example shown is "3"). Figure 27 In the example shown, the AF region box 134 is also segmented vertically, but this is only one example, and it is also possible not to segment it vertically. The size of the segmented region 134A is a size that is pre-derived by a simulator or the like, so that the part 124A1 corresponding to the car logo in the car image 124A can be completely included in a segmented region 134A.

[0188] As an example, such as Figure 28 As shown, in the image generation and transmission process with AF region frames, the generation unit 86D uses the real-time preview image 130 received by the receiving unit 86A and the AF region frames 134 divided into multiple segmented regions 134A by the segmentation unit 86F to generate a real-time preview image 136 with AF region frames. The real-time preview image 136 with AF region frames is an image in which the AF region frames 134 are superimposed on the real-time preview image 130. The AF region frames 134 are positioned within the real-time preview image 130, surrounding the aircraft image 132. The transmission unit 86H causes the display 28 of the imaging device 12 to display the real-time preview image 130 based on the captured image 108, and transmits it via communication I / F 84 (reference). Figure 4The generator 86D sends a real-time preview image 136 with AF region frames generated by the generator 86D to the camera device 12 as information for displaying the AF region frames 134 within the real-time preview image 130. The segmentation method derived by the generator 86G can be determined based on the AF region frames 134, which are divided into multiple segmented regions 134A by the segmentation unit 86F. Therefore, sending the real-time preview image 136 with AF region frames to the camera device 12 indicates outputting information representing the segmentation method to the camera device 12. Furthermore, the AF region frames 134, divided into multiple segmented regions 134A by the segmentation unit 86F, are an example of the "information representing the segmentation method" involved in the technology of this invention.

[0189] In the camera device 12, the receiver 62D is connected via communication I / F52 (reference). Figure 2 ) Receive the real-time preview image 136 with AF region frame sent from the transmitting unit 86H.

[0190] As an example, such as Figure 29 As shown, in the imaging device 12, the display control unit 62E causes the display 28 to display a real-time preview image 136 with an AF area frame received by the receiving unit 62D. For example, the real-time preview image 136 with an AF area frame displays an AF area frame 134, which is divided into multiple segmented regions 134A. For example, multiple segmented regions 134A are selectively designated according to an instruction received via the touch panel 30 (e.g., a touch operation). The CPU 62 performs AF control to focus on the real space region corresponding to the designated segmented region 134A. The CPU 62 performs AF control targeting the real space region corresponding to the designated segmented region 134A. That is, the CPU 62 calculates the focus position relative to the real space region corresponding to the designated segmented region 134A and moves the focusing lens 40B to the calculated focus position. In addition, the AF control can be either phase-difference AF control or contrast-difference AF control. Furthermore, an AF method based on the parallax of a pair of images obtained from a stereo camera or an AF method using the ranging results of a TOF method based on a laser beam, etc., can also be used.

[0191] Next, refer to Figures 30-31B The function of the camera system 10 will be explained.

[0192] Figure 30 An example of the display control processing flow performed by the CPU 62 of the camera device 12 is shown. Figure 30In the display control process shown, firstly, in step ST100, the acquisition unit 62A determines whether the captured image 108 has been stored in the image memory 46. If, in step ST100, the captured image 108 is not stored in the image memory 46, the determination is negative, and the display control process proceeds to step ST112. If, in step ST100, the captured image 108 is stored in the image memory 46, the determination is positive, and the display control process proceeds to step ST102.

[0193] In step ST102, the acquisition unit 62A acquires the captured image 108 from the image memory 46. After the processing in step ST102 is performed, the display control processing proceeds to step ST104.

[0194] In step ST104, the generation unit 62B generates a real-time preview image 130 based on the captured image 108 obtained in step ST102. After the processing in step ST104 is performed, the display control processing proceeds to step ST106.

[0195] In step ST106, the transmitting unit 62C transmits the real-time preview image 130 generated in step ST104 to the camera support device 14 via the communication I / F 52. After the processing in step ST106 is performed, the display control processing proceeds to step ST108.

[0196] In step ST108, the receiving unit 62D determines whether it has received a message via communication I / F52 through the execution of the included... Figure 31B The real-time preview image 136 with AF region frame is sent from the camera support device 14 during step ST238 of the camera support processing shown. In step ST108, if the real-time preview image 136 with AF region frame is not received via communication I / F 52, the determination is negative, and step ST108 is repeated. In step ST108, if the real-time preview image 136 with AF region frame is received via communication I / F 52, the determination is positive, and the display control processing proceeds to step ST110.

[0197] In step ST110, the display control unit 62E causes the display 28 to display the real-time preview image 136 with the AF region frame received in step ST108. After the processing in step ST110 is performed, the display control processing proceeds to step ST112.

[0198] In step ST112, the display control unit 62E determines whether the conditions for ending the display control process (hereinafter also referred to as "display control process end conditions") are met. Examples of display control process end conditions include the condition that the shooting mode set for the camera device 12 has been deactivated, or the condition that an instruction to end the display control process has been received by the receiving device 76. In step ST112, if the display control process end condition is not met, the determination is negative, and the display control process proceeds to step ST100. In step ST112, if the display control process end condition is met, the determination is positive, and the display control process ends.

[0199] Figure 31A and Figure 31B An example of the camera support processing flow performed by the CPU 86 of the camera support device 14 is shown. Additionally, Figure 31A and Figure 31B The illustrated camera support processing flow is an example of the "camera support method" involved in the technology of this invention.

[0200] exist Figure 31A In the camera support processing shown, in step ST200, the receiving unit 86A determines whether it has received a signal via communication I / F84 after execution. Figure 30 The live preview image 130 is sent during the processing in step ST106. In step ST200, if the live preview image 130 is not received via communication I / F84, the determination is rejected, and the camera support processing begins. Figure 31B The step ST240 is shown. In step ST200, if the real-time preview image 130 is received via communication I / F84, the determination is positive, and the camera support processing proceeds to step ST202.

[0201] In step ST202, the execution unit 86B performs subject recognition processing based on the real-time preview image 130 received in step ST200. The execution unit 86B uses the learned model 93 to perform the subject recognition processing. That is, the execution unit 86B extracts subject determination information 110 from the learned model 93 by assigning the real-time preview image 130 to the learned model 93, and stores the extracted subject determination information 110 in the memory 90. After performing the processing in step ST202, the camera support processing proceeds to step ST204.

[0202] In step ST204, the determination unit 86C determines whether a subject has been detected by the subject identification process. Specifically, the determination unit 86C retrieves an objectivity score 110B from the subject determination information 110 in the memory 90 and uses the retrieved objectivity score 110B to determine whether a subject exists. If no subject exists in step ST204, the determination is rejected, and the camera support process proceeds to step ST206. If a subject exists in step ST204, the determination is affirmative, and the camera support process proceeds to step ST212.

[0203] In step ST206, the generation unit 86D obtains bounding box position information 110A from the subject determination information 110 stored in the memory 90. After the processing in step ST206 is performed, the camera support processing proceeds to step ST208.

[0204] In step ST208, the generation unit 86D generates the AF region box 134 based on the bounding box position information 110A obtained in step ST206. After the processing in step ST208 is performed, the camera support processing proceeds to step ST210.

[0205] In step ST210, the segmentation unit 86F calculates the area of ​​the AF region frame 134 generated in step ST208 or step ST220, and segments the AF region frame in a segmentation method corresponding to the calculated area (see reference). Figure 22 After the processing in step ST210, the camera support processing begins. Figure 31B Step ST236 is shown.

[0206] In step ST212, the determination unit 86C obtains the objectivity score 110B from the subject determination information 110 stored in the memory 90. After the processing in step ST212 is performed, the camera support processing proceeds to step ST214.

[0207] In step ST214, the determination unit 86C determines whether the objectivity score 110B obtained in step ST212 is above or above the objectivity threshold. If, in step ST214, the objectivity score 110B is less than the objectivity threshold, the determination is rejected, and the camera support processing proceeds to step ST216. If, in step ST214, the objectivity score 110B is above or above the objectivity threshold, the determination is affirmative, and the camera support processing proceeds to step ST222. Furthermore, the objectivity threshold at which the camera support processing proceeds to step ST216 is an example of the "third threshold" involved in the technology of this invention.

[0208] In step ST216, the generation unit 86D obtains the bounding box position information 110A from the subject determination information 110 stored in the memory 90. After the processing in step ST216 is performed, the camera support processing proceeds to step ST218.

[0209] In step ST218, the generation unit 86D generates the AF region box 134 based on the bounding box position information 110A obtained in step ST216. After the processing in step ST218 is performed, the camera support processing proceeds to step ST220.

[0210] In step ST220, the enlargement unit 86E enlarges the AF area frame 134 by a predetermined magnification (e.g., 1.25x). After the processing in step ST220 is performed, the camera support processing proceeds to step ST210.

[0211] In step ST222, the determination unit 86C retrieves the category score 110D from the subject determination information 110 stored in the memory 90. After the processing in step ST222 is performed, the camera support processing proceeds to step ST224.

[0212] In step ST224, the determination unit 86C determines whether the category score 110D obtained in step ST222 is above the category threshold. In step ST224, if the category score 110D is less than the category threshold, the determination is rejected, and the camera support processing proceeds to step ST206. In step ST224, if the category score 110D is above the category threshold, the determination is affirmative, and the camera support processing proceeds to step ST206. Figure 31B Step ST226 is shown.

[0213] exist Figure 31B In step ST226 shown, the export unit 86G retrieves category 110C from the subject determination information 110 stored in the memory 90. The category 110C retrieved here is the category 110C with the largest category score 110D among the categories 110C included in the subject determination information 110 stored in the memory 90.

[0214] In step ST228, the export unit 86G refers to the partitioning method export table 128 in the storage unit 88 to export the partitioning method corresponding to the category 110C obtained in step ST226. That is, the export unit 86G exports the partitioning method corresponding to the category 110C obtained in step ST226 from the partitioning method export table 128.

[0215] In step ST230, the generation unit 86D obtains bounding box position information 110A from the subject determination information 110 stored in the memory 90. After performing the processing in step ST230, the camera support processing proceeds to step ST232.

[0216] In step ST232, the generation unit 86D generates the AF region box 134 based on the bounding box position information 110A obtained in step ST230. After the processing in step ST232 is performed, the camera support processing proceeds to step ST234.

[0217] In step ST234, the segmentation unit 86F segments the AF region bounding box 134 generated in step ST232 (reference) according to the segmentation method derived in step ST228. Figures 25-27 ).

[0218] In step ST236, the generation unit 86D uses the real-time preview image 130 received in step ST200 and the AF region bounding box 134 obtained in step ST234 to generate a real-time preview image 136 with an AF region bounding box (see reference). Figure 28 After the processing in step ST236 is performed, the camera support processing proceeds to step ST238.

[0219] In step ST238, the transmitting unit 86H transmits via communication I / F84 (reference). Figure 4 The real-time preview image 136 with the AF region bounding box generated in step ST236 is sent to the camera device 12. After the processing in step ST238 is performed, the camera support processing proceeds to step ST240.

[0220] In step ST240, the determination unit 86C determines whether the conditions for ending the camera support process (hereinafter also referred to as the "camera support process ending condition") are met. Examples of camera support process ending conditions include the condition that the shooting mode set for the camera device 12 has been deactivated, or the condition that an instruction to end the camera support process has been received by the receiving device 76. In step ST240, if the camera support process ending condition is not met, the determination is rejected, and the camera support process continues. Figure 31A The step ST200 is shown. In step ST240, if the camera support processing termination condition is met, a positive determination is made, and the camera support processing ends.

[0221] As described above, in the camera support device 14, the export unit 86G obtains the type of the subject as category 110C based on the captured image 108. The AF region frame 134 surrounds a segmented area that divides the subject into regions distinguishable from other areas of the shooting range. The segmentation unit 86F segments the AF region frame 134 according to a segmentation method derived from the segmentation method export table 128. This means that the segmented area surrounded by the AF region frame 134 is segmented according to the segmentation method derived from the segmentation method export table 128. The export unit 86G exports the segmentation method for segmenting the AF region frame 134 from the segmentation method export table 128 corresponding to the obtained category 110C. The AF region frame 134 segmented according to the segmentation method derived from the segmentation method export table 128 is included in the real-time preview image 136 with the AF region frame, and the transmission unit 86H transmits the real-time preview image 136 with the AF region frame to the camera device 12. The segmentation method can be determined based on the AF region box 134 included in the live preview image 136 with the AF region box. Therefore, sending the live preview image 136 with the AF region box to the imaging device 12 also means sending information to the imaging device 12 indicating the segmentation method derived from the segmentation method derivation table 128. Therefore, according to this structure, compared to the case where the AF region box 134 is segmented in a always constant segmentation method regardless of the type of the subject, precise control (e.g., AF control) related to the image sensor 20 of the imaging device 12 capturing the subject can be performed.

[0222] Furthermore, in the camera support device 14, the category 110C is inferred based on the subject determination information 110, which is the output result from the learned model 93 by applying the instant preview image 130 to the learned model 93. Therefore, according to this structure, compared to obtaining the category 110C solely through template matching without using the learned model 93, the category 110C to which the subject belongs can be accurately determined.

[0223] Furthermore, in the camera support device 14, the learned model 93 assigns objects (e.g., subject images such as passenger plane image 132, face image 120A, or car image 124A) within the bounding box 116 of the instant preview image 130 to the corresponding category 110C. The subject determination information 110, as an output from the learned model 93, includes a category score 110D, which is a value based on the probability that an object within the bounding box 116 of the instant preview image 130 belongs to a specific category 110C. Therefore, according to this structure, compared to inferring the category 110C solely through template matching without using the category score 110D obtained from the learned model 93, the category 110C to which the subject belongs can be accurately determined.

[0224] Furthermore, in the camera support device 14, a category score 110D is acquired as the probability value that an object within the bounding box 116 belongs to a specific category 110C when the objectivity score 110B is above the objectivity threshold. Then, the category 110C to which the subject belongs is inferred based on the acquired category score 110D. Therefore, according to this structure, compared to the case where the category 110C is inferred based on the category score 110D, which is acquired as the probability value that an object within the bounding box 116 belongs to a specific category 110C, regardless of whether the object exists within the bounding box 116, the category 110C to which the subject belongs (i.e., the type of the subject) can be determined more accurately.

[0225] Furthermore, in the camera support device 14, the category 110C to which the subject belongs is inferred based on the category score 110D above the category threshold. Therefore, according to this structure, compared to the case where the category 110C to which the subject belongs is inferred based on the category score 110D below the category threshold, the category 110C to which the subject belongs (i.e., the type of the subject) can be determined more accurately.

[0226] Furthermore, in the camera support device 14, when the objectivity score 110B is less than the objectivity threshold, the AF region box 134 is expanded. Therefore, according to this structure, compared to the case where the AF region box 134 is always a constant size, the probability that an object exists within the AF region box 134 can be increased.

[0227] Furthermore, in the camera support device 14, the number of divisions of the AF region frame 134 is specified according to the division method derived from the division method export table 128 by the export unit 86G. Therefore, according to this structure, compared to the case where the number of divisions of the AF region frame 134 is always constant, it is possible to accurately control the position of the subject desired by the user or others in relation to the shooting of the image sensor 20.

[0228] Furthermore, in the camera support device 14, the number of vertical and horizontal divisions of the AF region frame 134 is specified according to the division method derived from the division method export table 128 by the export unit 86G. Therefore, according to this structure, compared to the case where only the number of divisions in one direction of the AF region frame 134 is specified, the position of the subject desired by the user or others can be accurately controlled in relation to the shooting of the image sensor 20.

[0229] Furthermore, in the camera support device 14, the vertical and horizontal division numbers of the AF region frame 134, divided according to the division method exported from the division method export table 128 by the export unit 86G, are determined based on the subject image within the real-time preview image 130 (e.g., ...). Figure 25The composition of the passenger aircraft image 132 shown is defined. Therefore, according to this structure, compared to the case where the number of vertical and horizontal divisions of the AF area frame 134 is defined regardless of the composition of the subject image in the real-time preview image 130, precise control related to the shooting of the image sensor 20 can be performed on the position of the subject as desired by the user, etc.

[0230] Furthermore, the camera support device 14 sends a real-time preview image 136 with an AF region frame to the camera device 12. Then, the real-time preview image 136 with the AF region frame is displayed on the display 28 of the camera device 12. Therefore, according to this structure, users can identify the subject as the object to be controlled in relation to the image sensor 20 of the camera device 12.

[0231] Furthermore, in the above embodiment, an example of sending an AF region frame real-time preview image 136, including AF region frames 134 divided in a segmentation manner, to the imaging device 12 via the transmission unit 86H has been described, but the technology of the present invention is not limited thereto. For example, the CPU 86 of the imaging support device 14 may also send information to the imaging device 12 via communication I / F 84 that helps to move the focusing lens 40B to a focus position corresponding to the category 110C obtained from the subject determination information 110. Specifically, the CPU 86 of the imaging support device 14 may send information to the imaging device 12 via communication I / F 84 that helps to move the focusing lens 40B to a focus position corresponding to the segmented region 134A of the plurality of segmented regions 134A that corresponds to the category 110C obtained from the subject determination information 110.

[0232] In this case, as an example, such as Figure 32 As shown, the person 120 (reference) is photographed from the front side, surrounded by the AF area frame 134. Figure 13When the face image 120A obtained from the face is located at the first predetermined position of the AF region frame 134, the segmentation unit 86F divides the AF region frame 134 into an even number of equal parts in the horizontal direction, similar to the embodiment described above. Then, the segmentation unit 86F retrieves the category 110C from the subject determination information 110 stored in the memory 90, similar to the export unit 86G. Then, the segmentation unit 86F detects the eye image 120A2 representing the eye from the face image 120A and assigns focus priority position information 150 to the segmented region 134A including the detected eye image 120A2. The focus priority position information 150 is information that helps to move the focusing lens 40B to a focus position that is aligned with the eye represented by the eye image 120A2. The focus priority position information 150 includes position information that can determine the relative position of the segmented region 134A within the AF region frame 134. Furthermore, the focus priority position information 150 is an example of "information that helps to move the focusing lens to a focus position corresponding to a segmented region corresponding to the type" and "information that helps to move the focusing lens to a focus position corresponding to the type" involved in the technology of the present invention.

[0233] exist Figure 33 In the camera support processing shown, step ST300 is inserted between step ST234 and step ST236. In step ST300, the segmentation unit 86F detects an eye image 120A2 representing the eye from the face image 120A, and assigns focus priority position information 150 to the segmented region 134A that includes the detected eye image 120A2 among the multiple segmented regions 134A included in the AF region box 134 segmented in step ST234.

[0234] Therefore, the real-time preview image 136 with AF region frame includes an AF region frame 134 assigned focus priority position information 150. The real-time preview image 136 with AF region frame is transmitted to the camera device 12 (reference) via the transmitting unit 86H. Figure 33 Step ST238 is shown.

[0235] The CPU 62 of the imaging device 12 receives the real-time preview image 136 with an AF area frame sent by the transmitting unit 86H, and performs AF control based on the focus priority position information 150 of the AF area frame 134 included in the received real-time preview image 136 with an AF area frame. That is, the CPU 62 moves the focusing lens 40B to a focus position that aligns with the eye position determined according to the focus priority position information 150. In this case, compared to the case where the focusing lens 40B is only moved to a focus position corresponding to the same position within the shooting range, precise focusing on important positions of the subject is possible.

[0236] Furthermore, the method of AF control based on focus priority position information 150 is only one example. Focus priority position information 150 can also be used to notify users of important positions of the subject. In this case, for example, the segmented region 134A within the AF area frame 134 displayed on the display 28 of the imaging device 12, which is assigned focus priority position information 150, can be displayed in a way that distinguishes it from other segmented regions 134A (e.g., by emphasizing the frame of segmented region 134A). This allows users to visually identify which position within the AF area frame 134 is important.

[0237] In the above embodiments, examples of using the bounding box 116 directly as the AF region box 134 have been described, but the technology of the present invention is not limited thereto. For example, the CPU 62 may also change the size of the AF region box 134 based on the category score 110D obtained from the subject determination information 110.

[0238] In this case, as an example, such as Figure 34 As shown, in the camera support processing, step ST400 is inserted between step ST232 and step ST234. In step ST400, the generation unit 86D obtains the category score 110D from the subject determination information 110 in the memory 90, and changes the size of the AF region frame 134 generated in step ST232 according to the obtained category score 110D. Figure 35 The example shown illustrates a form in which the size of the AF region box 134 is reduced based on the category score 110D. In this case, it can be reduced to enclose the individual cells 114 distributed within the original AF region box 134 (see reference). Figure 9 The assigned category score is 110D (reference). Figure 9 The size of the area of ​​category score 110D above the reference value. In addition, the reference value can be a variable value that changes according to the instructions made to the camera support device 14 and / or various conditions, or it can be a fixed value.

[0239] according to Figure 34 and Figure 35 In the example shown, the size of the AF region box 134 is changed according to the category score 110D obtained from the subject determination information 110, thus improving the accuracy of the control related to subject shooting by the image sensor 20 compared to the case where the size of the AF region box 134 is always constant.

[0240] In the above embodiments, an example of AF control using an AF region frame 134 included in a real-time preview image 136 with an AF region frame has been described. However, the technology of the present invention is not limited thereto, and the camera device 12 may also use the AF region frame 134 included in a real-time preview image 136 with an AF region frame to perform shooting-related controls other than AF control. For example, such as Figure 36 As shown, shooting-related controls other than AF control can include custom controls. Custom controls are controls that suggest changing the control content of shooting-related controls (hereinafter, also referred to as "control content") according to the subject, and changing the control content according to the instructions given. The CPU 86 retrieves the category 110C from the subject determination information 110 in the memory 90, and outputs information that helps to change the control content based on the retrieved category 110C.

[0241] For example, the custom control includes at least one of the following: post-standby focus control, subject speed allowable range setting control, and focus adjustment priority area setting control. Post-standby focus control is configured to move the focusing lens 40B toward the focus position after a predetermined standby time (e.g., 10 seconds) if the position of the focusing lens 40B deviates from the focus position aligned with the subject. Subject speed allowable range setting control sets the allowable range of speed for the subject to be focused. Focus adjustment priority area setting control prioritizes which of the multiple segmented areas 134A within the AF area frame 134 for which focus is adjusted.

[0242] Thus, when the AF region box 134 is used for custom control, for example, by CPU 86... Figure 37 The camera support processing shown. Figure 37 The camera support processing shown Figure 34 The difference in the camera support processing shown is that step ST500 is included between step ST300 and step ST236, and steps ST502 and ST504 are included between step ST236 and step ST238.

[0243] exist Figure 37In the camera support processing shown, in step ST500, the CPU 86 retrieves category 110C from the subject determination information 110 in the memory 90, generates change instruction information indicating changes to the control content of the custom control based on the retrieved category 110C, and stores it in the memory 90. Examples of control content include, for instance, the standby time used in the subject speed allowable range setting control, the allowable range used in the subject speed allowable range setting control, and information used in the focus adjustment priority area setting control that determines the relative position of the focus priority segmentation area (i.e., the segmentation area 134A for priority focus adjustment) within the AF area frame 134. Furthermore, the change instruction information may include specific change content. The change content may, for example, be content pre-set for each category 110C. Additionally, the change instruction information is an example of "information that helps change the control content" as described in the technology of this invention.

[0244] In step ST502, CPU86 determines whether change indication information is stored in memory 90. If, in step ST502, change indication information is not stored in memory 90, the determination is negative, and the camera support process proceeds to step ST238. If, in step ST502, change indication information is stored in memory 90, the determination is positive, and the camera support process proceeds to step ST504.

[0245] In step ST504, CPU 86 attaches the change indication information in memory 90 to the real-time preview image 136 with AF region bounding box generated in step ST236. Then, the change indication information is deleted from memory 90.

[0246] If the real-time preview image 136 with AF region frame is sent to the camera device 12 through the processing of step ST238, the CPU 62 of the camera device 12 receives the real-time preview image 136 with AF region frame. Then, the CPU 62 causes the display 28 to display an alarm prompting the change of the control content of the custom control based on the change instruction information attached to the real-time preview image 136 with AF region frame.

[0247] Furthermore, while the example of displaying an alarm on the display 28 has been given here, the technology of the present invention is not limited thereto. The CPU 62 can also change the control content of the custom control based on the change instruction information. Moreover, the CPU 62 can also store the history of received change instruction information in the NVM 64. When storing the history of received change instruction information in the NVM 64, the real-time preview image 136 with the attached change instruction information, the real-time preview image 130 included in the real-time preview image 136 with the attached change instruction information, the thumbnail image of the real-time preview image 136 with the attached change instruction information, or the thumbnail image of the real-time preview image 130 included in the real-time preview image 136 with the attached change instruction information, and the thumbnail image of the real-time preview image 130 included in the real-time preview image 136 with the attached change instruction information, can be associated with the history and stored in the NVM 64. In this case, users can understand in what scenarios the custom control should be performed.

[0248] according to Figure 36 and Figure 37 In the example shown, shooting-related controls other than AF control include custom controls. The CPU 86 retrieves the category 110C from the subject determination information 110 in the memory 90 and outputs information that helps to change the control content based on the retrieved category 110C. Therefore, compared to the situation where the control content of custom controls is changed based solely on one's own intuition according to instructions given by the user, it is possible to shoot using custom controls suitable for category 110C.

[0249] In addition, Figure 36 and Figure 37 The example shown illustrates how the CPU 86 retrieves category 110C from the subject determination information 110 in the memory 90 and outputs information that helps change the control content based on the retrieved category 110C. However, the technology of the present invention is not limited to this. For example, in addition to category 110C, the subject determination information 110 may also include the state of the subject (e.g., the subject is moving, the subject is stopped, the subject's moving speed, and the subject's moving trajectory, etc.) as a subcategory. In this case, the CPU 86 retrieves category 110C and subcategories from the subject determination information 110 in the memory 90 and outputs information that helps change the control content based on the retrieved category 110C and subcategories. Thus, compared to the situation where the control content of custom controls is changed solely based on the user's instructions, it is possible to achieve shooting using custom controls suitable for category 110C and subcategories.

[0250] In the above embodiments, for ease of explanation, an example of applying one bounding box 116 to a single frame of real-time preview image 130 has been given. However, when the shooting range includes multiple subjects, multiple bounding boxes 116 may appear for a single frame of real-time preview image 130. Furthermore, it is also conceivable that multiple bounding boxes 116 may overlap. For example, in the case of two overlapping bounding boxes 116, as an example, ... Figure 38 As shown, the generation unit 86D generates one bounding box 116 as the AF region box 152 and another bounding box 116 as the AF region box 154, and the AF region box 152 overlaps with the AF region box 154. In this case, the person image 156, which is an object within the AF region box 152, overlaps with the person image 158, which is an object within the AF region box 154. Figure 38 In the example shown, the person image 158 is superimposed on the inside of the person image 156. Therefore, if the AF region box 154 surrounding the person image 158 is used for AF control, the focus may be on the person represented by the person image 156.

[0251] Therefore, the generation unit 86D uses the subject determination information 110 (reference) described in the above embodiment. Figure 18 The category 110C to which the person image 156 within the AF region frame 152 belongs and the category 110C to which the person image 158 within the AF region frame 154 belongs are obtained. Here, the category 110C to which the person image 156 belongs and the category 110C to which the person image 158 belongs are examples of "object category information" involved in the technology of the present invention.

[0252] The generation unit 86D retrieves bounding boxes from person images 156 and 158 based on the category 110C to which person image 156 and person image 158 belong. For example, different priorities are pre-assigned to the categories 110C to which person image 156 and person image 158 belong, and images in category 110C with higher bounding box retrieval priority are used as the retrieval targets. Figure 38 In the example shown, the category 110C to which person image 158 belongs has a higher priority than the category 110C to which person image 156 belongs. Therefore, the range of person image 158 enclosed by AF region box 154 is retrieved by excluding the region that overlaps with person image 156 (in... Figure 38In the example shown, the area where AF region frame 152 and AF region frame 154 overlap is the range. The AF region frame 154, obtained by retrieving the area surrounding the person image 158, is divided using the segmentation method described in the above embodiment. Then, the transmitting unit 86H sends a real-time preview image 136, including the AF region frame 154, to the imaging device 12. Thus, the CPU 62 of the imaging device 12 uses the AF region frame 154 (i.e., the AF region frame 154 obtained by retrieving the area surrounding the person image 158, excluding the area overlapping with the person image 156) to perform AF control, etc. In this case, compared to directly using AF region frames 154 and 156 for AF control, even if the person represented by person image 156 overlaps with the person represented by person image 158 in the depth direction, it is easier to focus on the person desired by the user or others. Furthermore, while a person is exemplified as the subject here, this is only one example; other subjects besides people can also be used.

[0253] Furthermore, as an example, such as Figure 39 As shown, the portion of the AF region frame 154 overlapping with the AF region frame 152 can be assigned focus object out-of-focus information 160 by the dividing portion 86F. Focus object out-of-focus information 160 indicates that this portion is an area excluded from the object to be focused on. Furthermore, while focus object out-of-focus information 160 is illustrated here, it is not limited to this; shooting-related control object out-of-focus information can also be used instead. Shooting-related control object out-of-focus information indicates that this portion is an area excluded from the aforementioned shooting-related controlled objects.

[0254] In the above embodiments, examples of deriving a segmentation method reflecting the composition of the aircraft image 132 from the segmentation method derivation table 128 have been described, but the technology of the present invention is not limited thereto. For example, a segmentation method reflecting the composition of a skyscraper image obtained by photographing a skyscraper from a lower or upper angle can also be derived from the segmentation method derivation table 128. In this case, the AF region frame 134 surrounding the skyscraper image is formed longitudinally, and the number of segments in the longitudinal direction is greater than the number of segments in the horizontal direction. Furthermore, it is not limited to the aircraft image 132 and the skyscraper image; any segmentation method corresponding to the composition can be derived from the segmentation method derivation table 128 for images obtained by photographing a subject with a composition used to impart a sense of depth.

[0255] In the above embodiment, an AF region frame 134 is exemplified, but the technology of the present invention is not limited thereto. Region frames that limit exposure control, white balance control, and / or grayscale control, etc., may also be used in conjunction with the AF region frame 134. In this case, similarly to the above embodiment, the region frame is generated by the camera support device 14, and the generated region frame is sent from the camera support device 14 to the camera device 12 for use by the camera device 12.

[0256] In the above embodiments, examples of identifying subjects using AI subject recognition methods have been given for illustration. However, the technology of the present invention is not limited to this, and subjects can also be identified using other subject recognition methods such as template matching.

[0257] In the above embodiment, an instant preview image 130 is illustrated, but the technology of the present invention is not limited thereto. For example, a later-viewed image may be used instead of the instant preview image 130. That is, subject recognition processing may also be performed based on the later-viewed image (see reference). Figure 31A (See step ST202 shown). Furthermore, subject recognition processing can also be performed based on the captured image 108. Additionally, subject recognition processing can also be performed based on a phase difference image including multiple phase difference pixels. In this case, the camera support device 14 can provide the multiple phase difference pixels used for ranging along with the AF region frame 134 to the camera device 12.

[0258] In the above embodiments, examples of separate camera devices 12 and camera support devices 14 have been described, but the technology of the present invention is not limited thereto, and the camera device 12 and camera support devices 14 may also be integrated. In this case, for example, as Figure 40 As shown, the NVM64 of the camera device body 16 can store the learned model 93, the camera support processing program 126 and the segmentation method export table 128 in addition to the display control processing program 80, and the CPU 62 can use the learned model 93, the camera support processing program 126 and the segmentation method export table 128 in addition to the display control processing program 80.

[0259] Furthermore, in such a case where the camera device 12 is responsible for the functions of the camera support device 14, it may replace the CPU 62 or be used in conjunction with the CPU 62 with at least one other CPU, at least one GPU and / or at least one TPU.

[0260] In the above embodiments, examples of the camera support processing program 126 being stored in the storage unit 88 have been described, but the technology of the present invention is not limited thereto. For example, the camera support processing program 126 may also be stored in a portable non-transitory storage medium such as an SSD or a USB memory. The camera support processing program 126 stored in the non-transitory storage medium is installed in the computer 82 of the camera support device 14. The CPU 86 executes camera support processing according to the camera support processing program 126.

[0261] Alternatively, the camera support processing program 126 can be stored in the storage device of another computer or server device connected to the camera support device 14 via the network 34, and the camera support processing program 126 can be downloaded and installed on the computer 82 according to the request of the camera support device 14.

[0262] In addition, it is not necessary to store the entire camera support processing program 126 in the storage device or storage unit 88 of other computers or server devices connected to the camera support device 14; a portion of the camera support processing program 126 may also be stored.

[0263] and, Figure 2 The camera device 12 shown has a built-in controller 44, but the technology of the present invention is not limited thereto. For example, the controller 44 can also be set outside the camera device 12.

[0264] In the above embodiments, a computer 82 is exemplified, but the technology of the present invention is not limited thereto, and devices including ASICs, FPGAs, and / or PLDs can be used instead of computer 82. Furthermore, a combination of hardware and software structures can be used instead of computer 82.

[0265] As the hardware resource for performing the camera support processing described in the above embodiments, various processors as shown below can be used. For example, a CPU can be used as a processor; this CPU is a general-purpose processor that performs the function of executing the camera support processing hardware resource by executing software (i.e., a program). Furthermore, a dedicated circuit can be used as a processor, such as an FPGA, PLD, or ASIC, which has a circuit structure specifically designed for performing a particular process. Regardless of the type of processor, it has built-in or connected memory, and regardless of the type of processor, it performs the camera support processing using the memory.

[0266] The hardware resources for performing camera support processing can consist of one of these various processors, or a combination of two or more processors of the same or different types (e.g., a combination of multiple FPGAs or a combination of a CPU and an FPGA). Furthermore, the hardware resources for performing camera support processing can also be a single processor.

[0267] As an example of a single processor, firstly, there is a form where a processor is a combination of one or more CPUs and software, and this processor performs the functions of hardware resources for camera support processing. Secondly, represented by SoCs (System-on-a-Chip), there is a form where a processor uses a single IC chip to implement the overall functions of a system including multiple hardware resources for performing camera support processing. Thus, camera support processing is implemented by using one or more of the aforementioned processors as hardware resources.

[0268] Furthermore, the hardware structure of these various processors, more specifically, can use circuits composed of combined semiconductor components and other circuit elements. Moreover, the aforementioned camera support processing is merely one example. Therefore, it is certainly possible to delete unnecessary steps, add new steps, or change the processing order without departing from the main point.

[0269] The above description and illustrations are detailed explanations of the parts related to the technology of this invention, and are merely one example of the technology of this invention. For example, the descriptions of the structure, function, effect, and effect described above are examples of the structure, function, effect, and effect of the parts related to the technology of this invention. Therefore, it is of course possible to delete unnecessary parts, add new elements, or replace the above description and illustrations without departing from the spirit of the technology of this invention. Furthermore, in order to avoid trouble and facilitate understanding of the parts related to the technology of this invention, explanations of technical common sense that do not require special explanation when implementing the technology of this invention have been omitted in the above description and illustrations.

[0270] In this specification, "A and / or B" has the same meaning as "at least one of A and B". That is, "A and / or B" means that it can be only A, only B, or a combination of A and B. Furthermore, in this specification, when three or more items are expressed by associating them with "and / or", the same meaning as "A and / or B" applies.

[0271] All documents, patent applications and technical standards described in this specification may be referenced in this specification to the same extent as the specific and individually described instances of reference to each document, patent application and technical standard.

Claims

1. A camera support device, comprising: Processor; and The memory is connected to or built into the processor. The processor performs the following processing: The type of the subject is obtained based on an image captured by an image sensor, including the subject's field of view; and The output represents information about the segmentation method used to divide the subject into regions based on the acquired type. These regions divide the subject into areas that can be distinguished from other areas within the shooting range. The defined region is determined by a first direction and a second direction intersecting the first direction. The segmentation method specifies that the number of segments in the first direction and the number of segments in the second direction are used as the number of segments to divide the partitioned region. The number of segments in the first direction and the number of segments in the second direction are determined based on whether the subject image representing the subject in the image is a composition that gives a sense of depth.

2. The camera support device according to claim 1, wherein, The processor acquires information about people, animals, or vehicles as the type.

3. The camera support device according to claim 1, wherein, The processor acquires information about a specific person, a specific person's face, a specific car, a specific passenger plane, a specific bird, or a specific tram as the type.

4. The camera support device according to claim 1, wherein, The processor obtains the type based on the output of the learned model, which is applied to the image by the learned model learned through machine learning.

5. The camera support device according to claim 1, wherein, The segmentation method specifies the number of divisions into which the region is equally divided according to the type.

6. The camera support device according to claim 4, wherein, The learned model assigns objects within the bounding boxes applied to the image to their corresponding categories. The output includes a value representing the probability that an object within the bounding box applied to the image belongs to a specific category.

7. The camera support device according to claim 6, wherein, The output includes the probability value of the object within the bounding box belonging to a specific category based on the following condition: the probability value of the object existing within the bounding box is above a first threshold.

8. The camera support device according to claim 6 or 7, wherein, The output includes values ​​above the second threshold based on the probability that the object belongs to the specific category.

9. The camera support device according to claim 6 or 7, wherein, If the probability that the object exists within the bounding box is less than a third threshold, the processor expands the bounding box.

10. The camera support device according to claim 6 or 7, wherein, The processor changes the size of the partitioned region based on a value representing the probability that the object belongs to the specific category.

11. The camera support device according to claim 4, wherein, The shooting range includes multiple subjects. The learned model assigns multiple objects within multiple bounding boxes applied to the image to their respective categories. The output includes: object category information representing the categories to which the multiple objects within the multiple bounding boxes applied to the image belong. The processor retrieves at least one subject surrounded by the defined region from the plurality of subjects based on the object category information.

12. A camera support device, comprising: Processor; and The memory is connected to or built into the processor. The processor performs the following processing: The type of the subject is obtained based on an image captured by an image sensor, including the subject's field of view; and The output represents information about the segmentation method used to divide the subject into regions based on the acquired type. These regions divide the subject into areas that can be distinguished from other areas within the shooting range. The segmentation method specifies the number of segments to divide the region. In cases where the focus of the focusing lens can be adjusted by moving the focusing lens that guides the incident light to the image sensor along the optical axis, The processor outputs information that helps move the focusing lens to a focus position corresponding to a segmented region of the acquired type within a plurality of segmented regions, the plurality of segmented regions being obtained by dividing the segmented regions with the number of segments.

13. A camera support device, comprising: Processor; and The memory is connected to or built into the processor. The processor performs the following processing: The type of the subject is obtained based on an image captured by an image sensor, including the subject's field of view; and The output represents information about the segmentation method used to divide the subject into regions based on the acquired type. These regions divide the subject into areas that can be distinguished from other areas within the shooting range. In cases where the focus of the focusing lens can be adjusted by moving the focusing lens that guides the incident light to the image sensor along the optical axis, The processor outputs information that helps move the focusing lens to a focus position corresponding to the acquired type.

14. A camera support device comprising: Processor; and The memory is connected to or built into the processor. The processor performs the following processing: The type of the subject is obtained based on an image captured by an image sensor, including the subject's field of view; and The output represents information about the segmentation method used to divide the subject into regions based on the acquired type. These regions divide the subject into areas that can be distinguished from other areas within the shooting range. The defined region is used for shooting-related control, where the shooting is the image sensor capturing the subject. The shooting-related controls include custom controls. The custom control includes at least one of the following: standby focus control, subject speed allowable range setting control, and focus adjustment priority area setting control. The processor outputs information that helps to modify the custom control based on the obtained type.

15. The camera support device according to claim 14, wherein, The processor performs the following processing: The state of the subject is also obtained from the image. Based on the obtained state and type, output information that helps to change the custom control.

16. The camera support device according to claim 15, wherein, The state refers to the subject moving, the subject stopping, the moving speed of the subject, and / or the moving trajectory of the subject.

17. The camera support device according to claim 14 or 15, wherein, In cases where the focus of the focusing lens can be adjusted by moving the focusing lens that guides the incident light to the image sensor along the optical axis, The standby focus control is a control that includes at least one of the following controls: If the position of the focusing lens deviates from the focus position of the subject, after a predetermined standby time, the focusing lens is moved toward the focus position.

18. The camera support device according to claim 14, wherein, In cases where the focus of the focusing lens can be adjusted by moving the focusing lens that guides the incident light to the image sensor along the optical axis, The subject speed allowable range setting control sets the allowable range of the subject speed for focusing.

19. The camera support device according to claim 14, wherein, In cases where the focus of the focusing lens can be adjusted by moving the focusing lens that guides the incident light to the image sensor along the optical axis, The focus adjustment priority area setting control determines whether to prioritize the adjustment of the focus of any one of the multiple areas within the divided area.

20. A camera support device, comprising: Processor; and The memory is connected to or built into the processor. The processor performs the following processing: The type of the subject is obtained based on an image captured by an image sensor, including the subject's field of view; and The output represents information about the segmentation method used to divide the subject into regions based on the acquired type. These regions divide the subject into areas that can be distinguished from other areas within the shooting range. The defined region is a frame that surrounds the subject. The processor adjusts the focus of the focusing lens by moving the focusing lens that guides the incident light to the image sensor along the optical axis. The frame is a focus frame that defines the region as a candidate area for focusing.

21. The camera support device according to claim 20, wherein, The processor outputs information for causing the display to show a real-time preview image based on the image and to display the frame within the real-time preview image.

22. A camera support device, comprising: Processor; and The memory is connected to or built into the processor. The processor performs the following processing: The type of the subject is obtained based on an image captured by an image sensor, including the subject's field of view; and The output represents information about the segmentation method used to divide the subject into regions based on the acquired type. These regions divide the subject into areas that can be distinguished from other areas within the shooting range. The defined region is determined by a first direction and a second direction intersecting the first direction. The segmentation method specifies that the number of segments in the first direction and the number of segments in the second direction are used as the number of segments to divide the partitioned region. Based on the type and the location of the specific part of the subject within the divided area, it is determined whether the number of divisions in the first direction and / or the number of divisions in the second direction is set to an even number or an odd number.

23. The camera support device according to claim 22, wherein, The number of divisions for the partitioned region is determined based on the area of ​​the partitioned region.

24. The camera support device according to claim 22 or 23, wherein, The size of each of the multiple segmented regions obtained by dividing the region according to the segmentation method is determined to be the size at which the specific part of the subject captured in the image completely enters the segmented region.

25. A camera device comprising: The camera support device according to any one of claims 1 to 24; and The image sensor.

26. A camera support method, comprising the following steps: The type of the subject is obtained based on an image captured by an image sensor, including the subject's field of view; and The output represents information about the segmentation method used to divide the subject into regions based on the acquired type. These regions divide the subject into areas that can be distinguished from other areas within the shooting range. The defined region is determined by a first direction and a second direction intersecting the first direction. The segmentation method specifies that the number of segments in the first direction and the number of segments in the second direction are used as the number of segments to divide the partitioned region. The number of segments in the first direction and the number of segments in the second direction are determined based on whether the subject image representing the subject in the image is a composition that gives a sense of depth.

27. A camera support method, comprising the following steps: The type of the subject is obtained based on an image captured by an image sensor, including the subject's field of view; and The output represents information about the segmentation method used to divide the subject into regions based on the acquired type. These regions divide the subject into areas that can be distinguished from other areas within the shooting range. The segmentation method specifies the number of segments to divide the region. The camera support method further includes the following steps: when the focus of the focusing lens can be adjusted by moving the focusing lens that guides incident light to the image sensor along the optical axis, outputting information that helps move the focusing lens to a focus position corresponding to a segmented region of the acquired type in a plurality of segmented regions, wherein the plurality of segmented regions are obtained by dividing the segmented regions by the number of segments.

28. A camera support method, comprising the following steps: The type of the subject is determined by capturing an image of the subject's range using an image sensor. The output represents information about the segmentation method for dividing the area according to the acquired type, wherein the segmented area divides the subject into areas that can be distinguished from other areas of the shooting range; and When the focus of the focusing lens can be adjusted by moving the focusing lens that guides the incident light to the image sensor along the optical axis, information is output that helps to move the focusing lens to a focus position corresponding to the acquired type.

29. A camera support method, comprising the following steps: The type of the subject is obtained based on an image captured by an image sensor, including the subject's field of view; and The output represents information about the segmentation method used to divide the subject into regions based on the acquired type. These regions divide the subject into areas that can be distinguished from other areas within the shooting range. The defined region is used for shooting-related control, where the shooting is the image sensor capturing the subject. The shooting-related controls include custom controls. The custom control includes at least one of the following: standby focus control, subject speed allowable range setting control, and focus adjustment priority area setting control. The camera support method further includes the following step: outputting information that helps to change the custom control based on the obtained type.

30. A camera support method, comprising the following steps: The type of the subject is obtained based on an image captured by an image sensor, including the subject's field of view; and The output represents information about the segmentation method used to divide the subject into regions based on the acquired type. These regions divide the subject into areas that can be distinguished from other areas within the shooting range. The defined region is a frame that surrounds the subject. The camera support method further includes the step of adjusting the focus of the focusing lens by moving the focusing lens that guides incident light to the image sensor along the optical axis. The frame is a focus frame that defines the region as a candidate area for focusing.

31. A camera support method, comprising the following steps: The type of the subject is obtained based on an image captured by an image sensor, including the subject's field of view; and The output represents information about the segmentation method used to divide the subject into regions based on the acquired type. These regions divide the subject into areas that can be distinguished from other areas within the shooting range. The defined region is determined by a first direction and a second direction intersecting the first direction. The segmentation method specifies that the number of segments in the first direction and the number of segments in the second direction are used as the number of segments to divide the partitioned region. Based on the type and the location of the specific part of the subject within the divided area, it is determined whether the number of divisions in the first direction and / or the number of divisions in the second direction is set to an even number or an odd number.

32. A storage medium storing a program for causing a computer to perform a process comprising the following steps: The type of the subject is obtained based on an image captured by an image sensor, including the subject's field of view; and The output represents information about the segmentation method used to divide the subject into regions based on the acquired type. These regions divide the subject into areas that can be distinguished from other areas within the shooting range. The defined region is determined by a first direction and a second direction intersecting the first direction. The segmentation method specifies that the number of segments in the first direction and the number of segments in the second direction are used as the number of segments to divide the partitioned region. The number of segments in the first direction and the number of segments in the second direction are determined based on whether the subject image representing the subject in the image is a composition that gives a sense of depth.

33. A storage medium storing a program for causing a computer to perform a process comprising the following steps: The type of the subject is obtained based on an image captured by an image sensor, including the subject's field of view; and The output represents information about the segmentation method used to divide the subject into regions based on the acquired type. These regions divide the subject into areas that can be distinguished from other areas within the shooting range. The segmentation method specifies the number of segments to divide the region. The processing further includes the following steps: when the focus of the focusing lens can be adjusted by moving the focusing lens that guides the incident light to the image sensor along the optical axis, the processor outputs information that helps move the focusing lens to a focus position corresponding to a segmented region of the acquired type within a plurality of segmented regions, the plurality of segmented regions being obtained by dividing the segmented regions by the number of segments.

34. A storage medium storing a program for causing a computer to perform a process comprising the following steps: The type of the subject is determined by capturing an image of the subject's range using an image sensor. The output represents information about the segmentation method for dividing the area according to the acquired type, wherein the segmented area divides the subject into areas that can be distinguished from other areas of the shooting range; and When the focus of the focusing lens can be adjusted by moving the focusing lens that guides the incident light to the image sensor along the optical axis, information is output that helps to move the focusing lens to a focus position corresponding to the acquired type.

35. A storage medium storing a program for causing a computer to perform a process comprising the following steps: The type of the subject is obtained based on an image captured by an image sensor, including the subject's field of view; and The output represents information about the segmentation method used to divide the subject into regions based on the acquired type. These regions divide the subject into areas that can be distinguished from other areas within the shooting range. The defined region is used for shooting-related control, where the shooting is the image sensor capturing the subject. The shooting-related controls include custom controls. The custom control includes at least one of the following: standby focus control, subject speed allowable range setting control, and focus adjustment priority area setting control. The process further includes the following steps: outputting information that helps to change the custom control based on the obtained type.

36. A storage medium storing a program for causing a computer to perform a process comprising the following steps: The type of the subject is obtained based on an image captured by an image sensor, including the subject's field of view; and The output represents information about the segmentation method used to divide the subject into regions based on the acquired type. These regions divide the subject into areas that can be distinguished from other areas within the shooting range. The defined region is a frame that surrounds the subject. The process further includes the step of adjusting the focus of the focusing lens by moving the focusing lens that guides the incident light to the image sensor along the optical axis. The frame is a focus frame that defines the region as a candidate area for focusing.

37. A storage medium storing a program for causing a computer to perform a process comprising the following steps: The type of the subject is obtained based on an image captured by an image sensor, including the subject's field of view; and The output represents information about the segmentation method used to divide the subject into regions based on the acquired type. These regions divide the subject into areas that can be distinguished from other areas within the shooting range. The defined region is determined by a first direction and a second direction intersecting the first direction. The segmentation method specifies that the number of segments in the first direction and the number of segments in the second direction are used as the number of segments to divide the partitioned region. Based on the type and the location of the specific part of the subject within the divided area, it is determined whether the number of divisions in the first direction and / or the number of divisions in the second direction is set to an even number or an odd number.

Citation Information

Patent Citations

  • Focus detection device and focus detection method

    JP2012128287A

  • Imaging device and imaging method

    JP2015046917A

  • Classification device, classification method, attribute recognition device, and machine learning device

    JP2019109843A

  • Electronic device and area selection method

    JP2020053720A