Image processing apparatus and control method thereof

The image processing apparatus addresses the issue of overlapping detection results by using detection confidence calculation and priority settings to accurately determine the main subject, enhancing the reliability of subject detection.

JP2026081301APending Publication Date: 2026-05-18CANON KK
View PDF 1 Cites 0 Cited by

Patent Information

Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
CANON KK
Filing Date
2026-02-03
Publication Date
2026-05-18

Smart Images

  • Figure 2026081301000001_ABST
    Figure 2026081301000001_ABST
Patent Text Reader

Abstract

When multiple detection results are obtained for the same subject using multiple dictionaries, the type of subject may not be selected correctly. [Solution] The system includes a detection means for detecting multiple types of subjects in an input image, a detection reliability calculation means for calculating a detection reliability for the detected subjects, a priority subject setting means for setting a subject type to be prioritized as the main subject, and a main subject determination means for determining the detection result that will be the main subject from the detected subjects based on the set priority subject and the detection reliability. The main subject determination means is characterized in that, when multiple types of subject detection results exist in the same area, it determines one type of subject in the area based on the set priority subject, the detection reliability and the detected subject type.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to an image processing apparatus having a subject detection function and a control method thereof.

Background Art

[0002] There is a technique for detecting a plurality of types of subjects based on a learned model obtained by performing machine learning for each subject type in order to detect a plurality of types of subjects from image data captured by an imaging device such as a digital camera. In order to adjust focus, brightness, and color to a suitable state based on the detected subject, it is necessary to determine one main subject from among the plurality of obtained subjects. Patent Document 1 discloses a method of determining a main subject based on a stable presence degree indicating whether a plurality of detected subjects are stably detected over a plurality of frames.

Prior Art Documents

Patent Documents

[0003]

Patent Document 1

Summary of the Invention

Problems to be Solved by the Invention

[0004] However, there is no description on how to determine and output the main subject when detection results of a plurality of types overlap for the same subject.

[0005] In view of the above problems, an object of the image processing apparatus of the present invention is to provide an image processing apparatus and a control method of the image processing apparatus capable of appropriately detecting a subject even when there are a plurality of detection results by a plurality of dictionaries for the same subject.

Means for Solving the Problems

[0006] Subject detection means for detecting a plurality of types of subjects for an input image, A detection confidence calculation means capable of calculating the detection confidence level for the detected subject, A priority subject setting means that allows setting the subject type to be prioritized as the main subject, The system includes a main subject determination means that determines the detection result that will be the main subject from the detected subjects based on the set priority subjects and the detection confidence level, The main subject determination means, when multiple types of subject detection results exist in the same area, The method is characterized by determining a single type of subject in the region based on the set priority subject, the detection confidence level, and the detected subject type. [Effects of the Invention]

[0007] According to the present invention, even when there are multiple detection results from multiple dictionaries for the same subject, it is possible to select the correct detection type. [Brief explanation of the drawing]

[0008] [Figure 1] External view of an imaging device including an image processing device. [Figure 2] Block diagram showing the configuration of an imaging system including an image processing device. [Figure 3] This diagram illustrates an example of how users can set priority target subjects for detection. [Figure 4] Flowchart of the overall process [Figure 5] A diagram showing an example of a switching sequence for multiple dictionary data. [Figure 6] Flowchart for the decision process to determine the type of subject within the same area. [Figure 7] This figure shows an example of a decision process for determining the type of subject within the same area. [Figure 8] Flowchart for determining the main subject [Figure 9] This diagram shows an example of the main subject determination process. [Figure 10] This diagram shows an example of a sequence for switching between multiple dictionary data sets when specified by the user. [Modes for carrying out the invention]

[0009] Figures 1(a) and 1(b) show external views of an imaging device 100, including an image processing device, as an example of an apparatus to which the present invention can be applied. Figure 1(a) is a front perspective view of the imaging device 100, and Figure 1(b) is a rear perspective view of the imaging device 100.

[0010] In Figure 1, the display unit 28 is a display unit located on the back of the camera that displays images and various information. The touch panel 70a can detect touch operations on the display surface (operating surface) of the display unit 28. The viewfinder-external display unit 43 is a display unit located on the top of the camera that displays various camera settings, including shutter speed and aperture. The shutter button 61 is an operation unit for issuing shooting instructions. The mode selector switch 60 is an operation unit for switching between various modes. The terminal cover 40 is a cover that protects connectors (not shown) such as connection cables that connect external devices to the imaging device 100.

[0011] The main electronic dial 71 is a rotary operating element included in the control unit 70, and by rotating this main electronic dial 71, settings such as shutter speed and aperture can be changed. The power switch 72 is an operating element that switches the power of the imaging device 100 ON and OFF. The sub electronic dial 73 is also included in the control unit 70 and is a rotary operating element included in the control unit 70, which can be used to move the selection frame and advance images. The cross key 74 is included in the control unit 70 and is a cross key (4-way key) with up, down, left, and right sections that can be pressed. Operations can be performed according to the part of the cross key 74 that is pressed. The SET button 75 is included in the control unit 70 and is a push button, mainly used to confirm selection items.

[0012] The video button 76 is used to start and stop video recording. The AE lock button 77 is included in the control unit 70 and, when pressed in shooting standby mode, locks the exposure state. The zoom button 78 is included in the control unit 70 and is used to turn the zoom mode ON and OFF in the live view display in shooting mode. By turning the zoom mode ON and then operating the main electronic dial 71, the LV image can be enlarged or reduced. In playback mode, it functions as a zoom button to enlarge the playback image and increase the magnification ratio. The playback button 79 is included in the control unit 70 and is used to switch between shooting mode and playback mode. By pressing the playback button 79 in shooting mode, the camera switches to playback mode and displays the latest image from the images recorded on the recording medium 200 on the display unit 28. The menu button 81 is included in the control unit 70 and, when pressed, displays various configurable menu screens on the display unit 28. The user can intuitively make various settings using the menu screen displayed on the display unit 28, the directional keys 74, and the SET button 75.

[0013] The touch bar 82 is a line-shaped touch sensor that can accept touch operations and is positioned to be operated by the thumb of the right hand holding the grip 90. The touch bar 82 can accept tap operations (touching and releasing without moving within a predetermined period of time), left and right slide operations (touching and then moving the touch position while keeping the touch on the screen), and so on. The touch bar 82 is a different operating component from the touch panel 70a and does not have a display function.

[0014] The communication terminal 10 is a communication terminal for the imaging device 100 to communicate with the lens side (detachable). The eyepiece part 16 is the eyepiece part of an eyepiece finder (a peeping type finder), and the user can visually recognize the video displayed on the internal EVF 29 through the eyepiece part 16. The eyepiece detection part 57 is an eyepiece detection sensor that detects whether the photographer is looking through the eyepiece part 16. The lid 202 is the lid of the slot that stores the recording medium 200. The grip part 90 is a holding part with a shape that is easy for the user to hold with the right hand when holding the imaging device 100. With the grip part 90 held by the little finger, ring finger, and middle finger of the right hand to hold the digital camera, the shutter button 61 and the main electronic dial 71 are arranged at positions operable by the index finger of the right hand. Also, in the same state, the sub - electronic dial 73 and the touch bar 82 are arranged at positions operable by the thumb of the right hand.

[0015] (Configuration of Imaging Device) FIG. 2 is a block diagram showing a configuration example of the imaging device 100 according to the present embodiment. In FIG. 2, the lens unit 150 is a lens unit equipped with an interchangeable photographing lens. The lens 103 is usually composed of a plurality of lenses, but here it is shown simply as a single lens for simplicity. The communication terminal 6 is a communication terminal for the lens unit 150 to communicate with the imaging device 100 side, and the communication terminal 10 is a communication terminal for the imaging device 100 to communicate with the lens unit 150 side. The lens unit 150 communicates with the system control unit 50 via these communication terminals 6 and 10, and controls the aperture 1 via the aperture drive circuit 2 by the internal lens system control circuit 4, and focuses by displacing the position of the lens 103 via the AF drive circuit 3.

[0016] The shutter 101 is a focal - plane shutter that can freely control the exposure time of the imaging unit 22 under the control of the system control unit 50.

[0017] The imaging unit 22 is an image sensor composed of a CCD or CMOS element, etc., which converts an optical image into an electrical signal. The imaging unit 22 may have an image plane phase difference sensor that outputs defocus amount information to the system control unit 50. The A / D converter 23 converts an analog signal into a digital signal. The A / D converter 23 is used to convert the analog signal output from the imaging unit 22 into a digital signal.

[0018] The image processing unit 24 performs resizing and color conversion processing, such as predetermined pixel interpolation and reduction, on the data from the A / D converter 23 or the data from the memory control unit 15. The image processing unit 24 also performs predetermined calculations using the captured image data. Based on the calculation results obtained by the image processing unit 24, the system control unit 50 performs exposure control and distance measurement control. This enables TTL (through-the-lens) AF (autofocus), AE (automatic exposure), and EF (flash pre-flash) processing. The image processing unit 24 further performs predetermined calculations using the captured image data and performs TTL AWB (auto white balance) processing based on the calculation results obtained.

[0019] The output data from the A / D converter 23 is written to the memory 32 via the image processing unit 24 and the memory control unit 15, or directly via the memory control unit 15. The memory 32 stores image data obtained by the imaging unit 22 and converted into digital data by the A / D converter 23, as well as image data for display on the display unit 28 and EVF 29. The memory 32 has sufficient storage capacity to store a predetermined number of still images, a predetermined amount of video footage, and audio.

[0020] Furthermore, memory 32 also serves as memory for image display (video memory). The D / A converter 19 converts the image display data stored in memory 32 into an analog signal and supplies it to the display unit 28 and EVF 29. In this way, the display image data written to memory 32 is displayed by the display unit 28 and EVF 29 via the D / A converter 19. The display unit 28 and EVF 29 display the image on a display device such as an LCD or organic EL according to the analog signal from the D / A converter 19. The digital signal, which has been A / D converted once by the A / D converter 23 and stored in memory 32, is converted to analog by the D / A converter 19 and sequentially transferred to the display unit 28 or EVF 29 for display, thereby enabling live view display (LV display). Hereinafter, the image displayed in live view will be referred to as a live view image (LV image).

[0021] The external LCD display unit 43 displays various camera settings, including shutter speed and aperture, via the external display unit drive circuit 44.

[0022] The non-volatile memory 56 is an electrically erasable and recordable memory, such as an EEPROM. The non-volatile memory 56 stores constants for the operation of the system control unit 50, programs, etc. The program referred to here is a program for executing various flowcharts described later in this embodiment.

[0023] The system control unit 50 is a control unit consisting of at least one processor or circuit, and controls the entire imaging device 100. It realizes each of the processes of this embodiment, which will be described later, by executing the program recorded in the non-volatile memory 56 mentioned above. For example, RAM is used in the system memory 52, and constants, variables for the operation of the system control unit 50, the program read from the non-volatile memory 56, etc. are stored there. The system control unit 50 also performs display control by controlling the memory 32, the D / A converter 19, the display unit 28, etc.

[0024] The system timer 53 is a timekeeping unit that measures the time used for various controls and the time of the built-in clock.

[0025] The operation unit 70 is an operating means for inputting various operation instructions to the system control unit 50. The mode switching switch 60 is an operating component included in the operation unit 70 and switches the operating mode of the system control unit 50 to one of the following: still image shooting mode, video shooting mode, playback mode, etc. Modes included in the still image shooting mode include auto shooting mode, auto scene detection mode, manual mode, aperture priority mode (Av mode), shutter speed priority mode (Tv mode), and program AE mode (P mode). There are also various scene modes and custom modes that provide shooting settings for different shooting scenes. The user can switch directly to any of these modes using the mode switching switch 60. Alternatively, the user can switch to a list screen of shooting modes using the mode switching switch 60, select one of the displayed modes, and then switch using another operating component. Similarly, the video shooting mode may also include multiple modes.

[0026] The first shutter switch 62 turns ON during the operation of the shutter button 61 on the imaging device 100, specifically when it is half-pressed (instruction to prepare for shooting), and generates the first shutter switch signal SW1. The first shutter switch signal SW1 initiates shooting preparation operations such as AF (autofocus), AE (automatic exposure), AWB (auto white balance), and EF (flash pre-flash).

[0027] The second shutter switch 64 turns ON when the shutter button 61 is fully pressed (shooting instruction), generating the second shutter switch signal SW2. The system control unit 50 starts a series of shooting processes, from reading the signal from the imaging unit 22 to writing the captured image to the recording medium 200 as an image file, in response to the second shutter switch signal SW2.

[0028] The operation unit 70 consists of various operating components that act as an input unit for receiving user input. The operation unit 70 includes at least the following operating components: shutter button 61, main electronic dial 71, power switch 72, sub electronic dial 73, cross key 74, SET button 75, video button 76, AF lock button 77, zoom button 78, playback button 79, menu button 81, and touch bar 82. Other operating components 70b represent a group of operating components that are not individually described in the block diagram.

[0029] The power control unit 80 consists of a battery detection circuit, a DC-DC converter, a switch circuit for switching which blocks are energized, and detects whether a battery is installed, the type of battery, and the remaining battery level. The power control unit 80 also controls the DC-DC converter based on the detection results and instructions from the system control unit 50, supplying the necessary voltage to each part, including the recording medium 200, for the required period. The power supply unit 30 consists of primary batteries such as alkaline batteries and lithium batteries, secondary batteries such as NiCd batteries, NiMH batteries, and Li batteries, and an AC adapter.

[0030] The recording medium I / F18 is an interface to the recording medium 200, such as a memory card or hard disk. The recording medium 200 is a recording medium such as a memory card for recording captured images, and is composed of semiconductor memory, magnetic disks, etc.

[0031] The communication unit 54 is connected wirelessly or via a wired cable and transmits and receives video and audio signals. The communication unit 54 can also connect to a wireless LAN (Local Area Network) or the internet. Furthermore, the communication unit 54 can communicate with external devices using Bluetooth® or Bluetooth Low Energy. The communication unit 54 can transmit images (including LV images) captured by the imaging unit 22 and images recorded on the recording medium 200, and can also receive images and other various information from external devices.

[0032] The attitude detection unit 55 detects the attitude of the imaging device 100 relative to the direction of gravity. Based on the attitude detected by the attitude detection unit 55, it is possible to determine whether the image captured by the imaging device 100 was taken with the imaging device 100 held horizontally or vertically. The system control unit 50 can add orientation information corresponding to the attitude detected by the attitude detection unit 55 to the image file of the image captured by the imaging device 22, or rotate the image before recording. An acceleration sensor or gyro sensor can be used as the attitude detection unit 55. It is also possible to detect the movement of the imaging device 100 (pan, tilt, lift, whether it is stationary or not, etc.) using the acceleration sensor or gyro sensor in the attitude detection unit 55.

[0033] (Image processing unit configuration) Figure 2(b) illustrates the characteristic configuration of the image processing unit 24 in this embodiment. The image processing unit 24 includes a subject detection unit 201, a detection history storage unit 202, a dictionary data storage unit 203, a dictionary data selection unit 204, and a main subject determination unit 205. In this embodiment, it is described as part of the image processing unit 24, but it may be part of the system control unit 50, or it may be provided separately from the image processing unit 24 and the system control unit 50. The image processing device may be mounted on, for example, a smartphone or a tablet terminal.

[0034] The image processing unit 24 sends the image data generated based on the data output from the A / D converter 23 to the subject detection unit 201 within the image processing unit 24.

[0035] In this embodiment, the subject detection unit 201 is composed of a machine learning (deep learning) CNN (convolutional neural network) and performs detection of a specific subject. The types of detectable subjects are based on dictionary data stored in the dictionary data storage unit 203. In this embodiment, the subject detection unit 201 is composed of different CNNs (different network parameters) depending on the type of detectable subject. The subject detection unit 201 may be implemented using a GPU (graphics processing unit) or a circuit specialized for estimation processing by CNN.

[0036] Machine learning of a CNN can be performed using any method. For example, a predetermined computer, such as a server, may perform machine learning of the CNN, and the imaging device 100 may acquire the trained CNN from the predetermined computer. In this embodiment, the predetermined computer takes training image data as input and performs supervised learning using the position information of subjects corresponding to the training image data as training data (annotation), thereby training the CNN of the subject detection unit 201. As a result, a trained CNN is generated. The CNN training may be performed by the imaging device 100 or the image processing device described above.

[0037] As described above, the subject detection unit 201 includes a CNN (trained model) trained by machine learning. The subject detection unit 201 takes image data as input, estimates the position, size, confidence level, etc., of the subject, and outputs the estimated information. The CNN may be, for example, a network in which a fully connected layer and an output layer are connected to a layer structure in which convolutional layers and pooling layers are stacked alternately. In this case, for example, backpropagation may be applied to train the CNN. Alternatively, the CNN may be a neocognitive CNN, which consists of a feature detection layer (S layer) and a feature integration layer (C layer) as a set. In this case, a training method called "Add-if Silent" may be applied to train the CNN.

[0038] The subject detection unit 201 may use any pre-trained model other than a pre-trained CNN. For example, a pre-trained model generated by machine learning, such as a support vector machine or a decision tree, may be applied to the subject detection unit 201. Furthermore, the subject detection unit 201 does not have to be a pre-trained model generated by machine learning. For example, any subject detection method that does not use machine learning may be applied to the subject detection unit 201.

[0039] The detection history storage unit 202 stores the subject detection history detected by the subject detection unit 201 in the image data, and the system control unit 50 sends the history to the dictionary data selection unit 204. In this embodiment, the subject detection history stores the dictionary data used for detection, the position of the detected subject, the size of the detected subject, and the confidence level of the detected subject. Other data such as the number of detections and the identifier of the image data containing the detected subject may also be stored.

[0040] The dictionary data storage unit 203 stores dictionary data for detecting specific subjects, and the system control unit 50 reads the dictionary data selected by the dictionary data selection unit 204 from the dictionary data storage unit 203 and sends it to the subject detection unit 201. Each dictionary data is, for example, data in which the characteristics of each region of a specific subject are registered. Furthermore, in order to detect multiple types of subjects, dictionary data for each subject and for each region of the subject may be used. The dictionary data storage unit 203 stores multiple types of subject detection dictionary data, such as dictionary data for detecting "people," dictionary data for detecting "animals," and dictionary data for detecting "vehicles." Furthermore, in addition to "animals," it may also store dictionary data for detecting "birds," which have a special shape and are in high demand for detection. Furthermore, the dictionary data for detecting "vehicles" can be further subdivided into "cars," "motorcycles," "railways," "airplanes," etc., and stored individually.

[0041] The area of ​​the subject detected by multiple dictionary data stored in the dictionary data storage unit 203 can be used as a focus detection area. For example, in a composition where there is an obstacle in the foreground and a subject in the background, it is possible to focus on the desired subject by focusing on the detected area.

[0042] Furthermore, in this embodiment, the multiple dictionary data used for detection by the subject detection unit 201 are generated by machine learning, but rule-based dictionary data may be used or used in combination. Rule-based dictionary data, for example, stores images of the subject to be detected or features specific to the subject, as predetermined by the designer. By comparing the images or features of the dictionary data with the images or features of the image data obtained by capturing, the subject can be detected. Rule-based dictionary data is less complex than the models set by machine learning-trained models, so it requires less data, and subject detection using rule-based dictionary data is faster (and has a lower processing load) than that of trained models.

[0043] The dictionary data selection unit 204 selects the next dictionary data to be used based on the detection history stored in the detection history storage unit 202, a predetermined order or rule, or instructions from the user, and notifies the dictionary data storage unit 203.

[0044] In this embodiment, the dictionary data storage unit 203 individually stores dictionary data for multiple types of subjects and for each region of a subject, and subject detection is performed multiple times by switching between multiple dictionary data for the same image data. The dictionary data selection unit 204 determines the dictionary data switching sequence and determines the dictionary data to be used according to the determined sequence. An example of the dictionary data switching sequence will be described later.

[0045] The type determination unit 205 determines the type of subject in an area when multiple subjects are detected within the same area. Based on the subject detection priority setting set by the user via the operation unit 70 from among the multiple detection histories stored in the detection history storage unit 202, one detection result is determined. The determination method will be described later.

[0046] Figure 3 shows an example of how to set the subject to be detected with priority, where the user selects whether or not to prioritize the type of subject to be detected from the menu screen displayed on the display unit 28. Figure 3 shows the detection subject selection setting screen displayed on the display unit 28, where the user selects the subject to be detected with priority from specific detectable subjects (e.g., vehicles, animals, people) by operating the operation unit 70. In Figure 3, "Vehicles" is selected. Also, "None" in Figure 3 is a mode in which no subjects are detected, and "Automatic" is a mode in which all specific detectable subjects are detected without priority.

[0047] The main subject determination unit 206 determines the main subject based on multiple detection histories stored in the detection history storage unit 202, the subject settings to be prioritized for detection set by the user via the operation unit 70, and the subject determined by the type determination unit 205. The method for determining the main subject will be described later.

[0048] (Processing flow of the imaging device) Figure 4 is a flowchart showing the characteristic processing flow of the present invention performed by the imaging device 100 of this embodiment. Each step in this flowchart is executed by the system control unit 50 or by instructions from the system control unit 50. At the start of this flowchart, the imaging device 100 is assumed to be powered on, in live view imaging mode, and in a state where it can be instructed to start capturing (recording) still images or videos via the operation unit 70.

[0049] The series of processes from step S401 to step S409 in Figure 4 are assumed to be the processes performed when the imaging unit 22 of the imaging device 100 captures one frame (one image data). However, the series of processes from step S401 to step S409 may be performed across multiple frames. That is, the result of subject detection in the first frame may be reflected in any frame from the second frame onward.

[0050] In step S401, the system control unit 50 acquires the image data captured by the imaging unit 22 and output by the A / D conversion unit 23.

[0051] In step S402, the image processing unit 24 resizes the image data to a size that is easy to process (e.g., QVGA) and sends the resized image data to the image data generation unit 201.

[0052] In step S403, the dictionary data selection unit 204 selects dictionary data generated by machine learning to be used for subject detection, and sends selection information about what the selected dictionary data is to the dictionary data storage unit 203.

[0053] Here, dictionary data generated by machine learning can be created by extracting common features of a specific subject from a large amount of image data containing that subject. Common features include, for example, the size, position, and color of the subject, as well as areas outside the specific subject, such as the background. Therefore, the more limited the background in which the detected subject exists, the easier it is to improve detection performance (detection accuracy) with less training. On the other hand, if the model is trained to detect a specific subject regardless of the background, it will have high versatility for various shooting scenes, but the detection accuracy will be lower. The more diverse and numerous the image data used to generate the dictionary data, the higher the detection performance tends to be. However, by limiting the size and position of the detection area of ​​the subject to be detected in the image data used for detection, it is possible to improve detection performance even if the number and diversity of image data required for dictionary data generation are reduced. Also, if part of the subject is cut off outside the image data, some of the subject's features are lost, resulting in lower detection performance.

[0054] Furthermore, generally speaking, the larger the area of ​​the subject, the more features it contains. In detection using the machine learning-generated dictionary data mentioned earlier, there is a possibility of misidentifying objects with similar features to the specific subject being detected, in addition to the object itself. A region defined as a local region is a smaller area compared to the overall area. The smaller the region, the fewer features it contains, and the fewer features it contains, the more objects with similar features it will contain, leading to an increase in false detections.

[0055] Figure 5 illustrates the switching sequence for multiple dictionary data for one frame (one image data) in step S403. When multiple dictionary data are stored in the dictionary data storage unit 203, it is possible to perform detection with multiple dictionaries for one frame. On the other hand, in live view mode, where images are captured sequentially and processed, and in video data during video recording, the number of times a subject can be detected per frame is limited.

[0056] At this time, the type and order of dictionary data to be used should be determined based on factors such as whether or not subjects have been detected in the past, the type of dictionary data used at that time, and the type of subject to be prioritized for detection. Depending on the dictionary data switching sequence, the dictionary data for detecting a specific subject may not be selected when that subject is included in the frame, potentially causing a missed detection opportunity. Therefore, the dictionary data switching sequence also needs to be changed according to the settings and scene.

[0057] As an example, Figure 5 shows an example of a dictionary data switching sequence when a vehicle is selected as the preferred subject for detection in a structure that can perform subject detection up to three times per frame (or has three detectors that can process in parallel). V0 and V1 each represent the vertical synchronization period for one frame, and the rectangular blocks such as the person's head, vehicle 1 (motorcycle), and vehicle 2 (car) indicate that subject detection can be performed using three dictionary data sets (trained models) in a time series within one vertical synchronization period.

[0058] Figure 5(a) shows an example of dictionary data switching when no detected subject is present. In the first frame, the dictionary data is switched in the order of human head, vehicle 1 (motorcycle), and vehicle 2 (car), and in the second frame, the dictionary data is switched in the order of animal (dog / cat), vehicle 1 (motorcycle), and vehicle 2 (car). For example, without a switching sequence, the system always uses dictionary data that can detect subjects selected by the user from the menu screen, as shown in Figure 3. In this case, it is necessary to switch the priority detection subject setting for each scene, such as vehicles when vehicles are present, and people and animals otherwise. Also, if it is unclear when a vehicle will appear, there is a risk that the shooting will not be completed in time if the priority detection subject setting is switched after noticing the vehicle. In contrast, in this embodiment, as shown in Figure 5(a), all dictionary data is switched across multiple frames during periods when no specific subject is detected, making it possible to shoot without worrying about the priority detection subject setting. By selecting dictionary data according to the priority detection subject setting in either the first or second frame while switching all dictionary data, it is possible to detect all detectable subjects while also improving the detection accuracy of the priority detection subject. This reduces the number of times the priority detection subject setting needs to be switched. Additionally, there may be a separate mode that always cycles through a specific dictionary (or group of dictionaries) according to user-specified settings.

[0059] Figure 5(b) shows an example of switching dictionary data in the next frame after detecting a motorcycle in the previous frame, switching the dictionary data in the order of vehicle 1 (motorcycle), person's head, vehicle 1 (motorcycle). The order of detection when switching dictionary data is not necessarily as described above. For example, the dictionary data for the person's head in the above example of dictionary data switching could be changed to match the scene, for example, in a scene where a motorcycle is being photographed, to dictionary data that is more likely to be selected as a subject other than a motorcycle. At this time, detection using the dictionary data for "animals," which is less likely to be detected in parallel with vehicles, could be controlled mutually. Depending on the texture (pattern) and color of the vehicle, it may be detected as an animal, and controlling it mutually in this way can ultimately improve the detection accuracy of the desired subject.

[0060] In step S404, the subject detection unit 201 uses dictionary data for detecting specific subjects (objects) stored in the dictionary data storage unit 203 to detect subjects (or the subject area in which subjects exist) from image data captured by the imaging unit 22 and input to the image processing unit 24. The position and size of the detected subject, the calculated confidence level, and other information, as well as the type of dictionary data used and the identifier of the image data used for detection, are stored in the detection history storage unit 202.

[0061] In step S405, it is determined from the detection history stored in the detection history storage unit 203 whether detection has been performed using all necessary dictionary data for image data with the same identifier (image data from the same frame). If the result of the determination is Yes, the process proceeds to S406; otherwise, the process returns to step S403 and the dictionary data to be used next is selected.

[0062] In step S406, it is determined from the detection history stored in the detection history storage unit 203 whether detection has been performed with all dictionary data. If the result of the determination is Yes, the process proceeds to S407; otherwise, the process proceeds to the next frame. For example, in Figure 5(a), it takes two frames to perform detection with all the necessary dictionary data, so the subsequent processing is skipped for the first frame, and the process proceeds to the next frame, and then to S407 after two frames. In this embodiment, the subsequent processing is skipped until detection has been performed with all the necessary dictionary data. However, this is not limited to this, and for processes that require immediate response, such as autofocus, the subsequent processing may be performed only with the subjects detected in each frame without waiting for detection with all the dictionary data. Also, for example, if it is possible to cycle through all the currently set dictionary data in two frames, as in this embodiment, the subsequent processing from step S407 onwards may be performed based on the detection results of two frames, always including the previous past frame.

[0063] In step S407, the settings for selecting which subjects to prioritize for detection from among the specific detectable subjects that the user has previously set via the control unit 70 are read.

[0064] In step S408, it is determined from the detection history of image data with the same identifier stored in the detection history storage unit 203 whether or not multiple detection results exist in the same area. If multiple results exist in the same area, the process proceeds to S409; otherwise, the process proceeds to step S410. Whether or not multiple detection results exist in the same area can be determined, for example, if the detection center coordinates are within different detection result areas. Alternatively, it can be determined that multiple detection results exist in the same area if the detection areas overlap by a percentage greater than a threshold; the method is not limited.

[0065] In step S409, the type determination unit 205 determines one region detection result from the priority subject setting set in step S407, the detection result saved in step S405, and the result determined to exist in the same region in step S408. The determination method will be described later.

[0066] In step S410, the main subject determination unit 206 uses the priority subject setting set in step S407 to determine the main subject from among multiple detection results of image data with the same identifier from the detection history stored in the detection history storage unit 203. If it is determined in step S408 that there are multiple detection results in the same area, the result from step S409 is also used. At this time, the system control unit 50 may display some or all of the information output by the main subject determination unit 206 on the display unit 28. The determination method will be described later.

[0067] (Flowchart for determining the type of subject based on the detection results of multiple subjects within the same area) The type determination process in S409 will be explained using the flowchart in Figure 6, Figure 7, and Table 1. Each step in this flowchart is executed by the system control unit 50 or by the unit instructed by the system control unit 50.

[0068] Figure 7 illustrates an example of the type determination process. Figure 7(a) is the input image, and it is assumed that motorcycle 701 is the subject. In Figure 7(b), the person dictionary is selected in step S403 and person 702 is detected. In Figure 7(c), the motorcycle dictionary is selected in step S403 and motorcycle 703 is detected. In Figure 7(d), the automobile dictionary is selected in step S403 and automobile 704 is falsely detected. In Figure 7(e), the dog dictionary is selected in step S403 and dog 705 is falsely detected. In Figure 7(f), the cat dictionary is selected in step S403, and the result of the process is that no detection result was found.

[0069] In step S601, priority is assigned to the detected subject types according to the priority settings set in step S407.

[0070] Table 1 shows an example of priority classification based on priority settings and subject type. The vertical axis in Table 1 represents the settable priority settings, which are "People," "Animals," "Vehicles," "None," and "Automatic," as in Figure 3. The horizontal axis represents the detected subject types, which are "People," "Dogs," "Cats," "Cars," and "Motorcycles," as in Figure 7. In Table 1, a smaller priority number indicates a higher priority, and "No priority" indicates subjects that will not be detected.

[0071] In this embodiment, we have used three values: priority subjects (priority 1 in the table), non-priority subjects (priority 2 in the table), and subjects not to be used (no priority in the table), but we are not limited to these. For example, we could use two values: subjects to be used and subjects not to be used. Alternatively, we could use four values: highest priority subjects, priority subjects, non-priority subjects, and subjects not to be used, and this can be changed depending on the number of subject types that can be detected and the priority settings that can be set. Also, in Table 1, when vehicles are prioritized, cars and motorcycles are classified as priority subjects, people as non-priority subjects, and dogs and cats as subjects not to be used. However, the classification method is not limited to this, and for example, if you do not want to detect any subject types other than those set as priority, you can classify people as subjects not to be used. Also, if you want to detect subjects other than priority subjects, you can classify dogs and cats as non-priority subjects.

[0072] [Table 1]

[0073] In step S602, the process of determining the type of subject within the same area based on the priority determined in step S601 is performed.

[0074] The specific method will be explained using Figure 7. For example, if people are set as the priority, according to Table 1, the subject with priority 1 is a person, so we check if there is a person detection result. As Figure 7(b) Person 702 exists, we select it and end the type determination process as the subject type within the region. If vehicles are set as the priority, according to Table 1, the subjects with priority 1 are cars and motorcycles, so we check if there is a car or motorcycle detection result with priority 1. As both Figure 7(c) Motorcycle 702 and Figure 7(d) Car exist, we proceed to step S603. If neither Figure 7(c) Motorcycle 702 nor Figure 7(d) Car exist at this time, we check if there is a person detection result with priority 2. If there is no person result, since dogs and cats have no priority according to Table 1, we conclude that there are no subjects within the region and end the type determination process.

[0075] In step S603, the confidence level of the detection results saved in step S405 is normalized for each subject. The purpose of normalization is that the maximum value and the confidence threshold for a subject that can be trusted vary depending on the dictionary used. Normalization makes it possible to compare the confidence levels of subjects from different dictionaries in subsequent processing. In this embodiment, the minimum possible confidence level for each dictionary is normalized to 0 and the maximum value to 1, thereby reducing the confidence level of all subjects from 0 to 1 and enabling comparison by confidence level. The normalization method is not limited to this; for example, the confidence threshold for a subject that can be trusted may be set to 1, and the minimum possible confidence level may be set to 0, and the method is not limited.

[0076] In step S604, if step S602 confirms that there are multiple subject types with the same priority, step S603 determines the subject with the highest confidence level, which has been normalized, as the subject within the region, and terminates the type determination process. In this embodiment, the subject within the region is determined by confidence level, but the method is not limited. For example, the detection results of past frames may be referred to, and the subject type that has been detected most frequently across multiple frames may be determined as the subject within the region.

[0077] In Figure 7, since step S602 has determined that Figure 7(c) motorcycle 703 and Figure 7(d) automobile 704 are priority subjects in the same region, we compare them. In this embodiment, since the input subject is motorcycle 701, we assume that the confidence level of Figure 7(c) motorcycle 703 is the highest, and motorcycle 703 is determined to be the subject in the region.

[0078] Before comparing confidence levels in step S604, subject selection is performed in step S602 based on priority. The purpose of this is that, for example, if subjects have similar common characteristics, such as both dogs and cats being quadrupedal, inputting a cat image into the dog dictionary is likely to result in a false positive, where the cat is mistakenly identified as a dog. However, if subjects do not have similar common characteristics, such as dogs and motorcycles, inputting a motorcycle image into the dog dictionary is less likely to result in a false positive, where the motorcycle is mistakenly identified as a dog. However, in cases of false positives, such as dog 705 in Figure 7(e), it can be difficult to determine which feature of the image triggered the reaction, resulting in a high confidence level, and in such cases, it is difficult to prevent the final output from being a dog. Therefore, by first selecting subjects according to the set priority, false positives of undesired subjects are eliminated.

[0079] (Flow of the process for determining the main subject) The process for determining the main subject in S410 will be explained using the flowchart in Figure 8 and Figure 9. Each step in this flowchart is executed by the system control unit 50 or by the instruction of the system control unit 50.

[0080] Figure 9 shows an example of primary subject determination when multiple subjects are detected in the same frame. Figure 9(a) shows the detection of human face 901, cat 902, and cat 903. Figure 9(b) shows that human face 904 is selected as the primary subject from among human face 901, cat 902, and cat 903. Figure 9(c) shows that cat 905 is selected as the primary subject from among human face 901, cat 902, and cat 903.

[0081] In step S801, a primary subject candidate is selected according to the priority settings established in step S407. If a primary subject candidate is uniquely determined, that candidate is designated as the primary subject and the primary subject determination is terminated. If no candidate is found, the primary subject determination is terminated, and no primary subject is found. If multiple candidates are found, the process proceeds to step S802.

[0082] A specific example will be explained using Figure 9.

[0083] If "person" is set in step S407, then, among the human face 901, cat 902, and cat 903 in Figure 9(a), the human face will be designated as the main subject as shown in Figure 9(b) 904 according to the priority setting, and the main subject determination will be completed.

[0084] If "animal" is set in step S407, then since there are multiple detection results for cats among the human face 901, cat 902, and cat 903 in Figure 9(a), the process proceeds to step S802.

[0085] If "Automatic" is set in step S407 as shown in Figure 3, there are no subjects to prioritize for detection, and since there are multiple detection results for people and cats, the process proceeds to step S802.

[0086] If the "vehicle" in Figure 3 is set in step S407, then none of the human face 901, cat 902, and cat 903 in Figure 9(a) will be used as the subject, so there is no main subject, and the main subject determination is terminated.

[0087] In step S802, the primary subject is selected from among the multiple candidate subjects determined in step S801 based on the position, size, and reliability of the subjects detected in step S404. For example, if the primary subject is determined to be the subject closest to the center of the field of view, and three candidates remain in step S801—a human face 901, a cat 902, and a cat 903—then the human face 901 is the closest to the center, so the human face is selected as the primary subject, as shown in 904 of Figure 9(b).

[0088] If cats 902 and 903 remain as candidates, cat 902 is closest to the center, so cat 905 in Figure 9(c) will be chosen as the main subject.

[0089] In this embodiment, the subject closest to the center of the field of view was selected as the primary subject from among the candidate subjects, but this is not limited to this. For example, the subject closest to the center of the autofocusable area may be selected as the primary subject, or a larger subject may be selected as the primary subject, or a subject with high detection reliability may be selected as the primary subject, or the primary subject may be determined by a combination of these factors.

[0090] (An embodiment in which the user performs an operation specified on the screen) The above-described embodiment is an example in which the imaging device automatically detects a subject, determines the type of subject within the same area, and determines the main subject. In this embodiment, we will describe an example in which, when the user arbitrarily specifies an area within the live view screen displayed on the display unit 28, the dictionary switching sequence is changed to determine the type of subject within the same area and determine the main subject.

[0091] Figure 10 illustrates the dictionary switching sequence when the user specifies an arbitrary area within the live view screen in the dictionary data selection unit 204 in step S403.

[0092] In Figure 5, the switching sequence was changed based on already detected subjects and priority detection subject settings. However, in this embodiment, if the user specifies an area within the live view screen, all detectable dictionaries are switched regardless of detected subjects and priority detection subject settings. This is because, regardless of already detected subjects, all detectable dictionaries are switched to accurately detect subjects in the specified area in order to accurately reflect the user's area specification.

[0093] Figure 10 shows an example of dictionary data switching. In the first frame, the dictionary data is switched in the order of human head, vehicle 1 (motorcycle), and vehicle 2 (car). In the second frame, the dictionary data is switched in the order of human head, animal (dog / cat), and animal (bird), and so on, switching dictionary data across multiple frames. In this embodiment, human heads are switched in both the first and second frames, but one of them may be assigned to a different dictionary depending on the priority detection subject setting. For example, if vehicle priority is set, the dictionary for any vehicle may be assigned to the second frame, or if animal priority is set, the dictionary for any animal may be assigned to the second frame.

[0094] The characteristic processing of the type determination process in S409 of this embodiment will be described. In this embodiment, processing is performed when multiple types of subjects are detected within the area specified by the user.

[0095] The characteristic processing of the main subject determination process in S410 of this embodiment will be described. In this embodiment, a subject within the area specified by the user is determined to be the main subject. If no subject is detected within the specified area, the specified area is determined to be the main subject, but in the dictionary data switching sequence of step S403 in the next frame, all dictionaries are switched until a detectable subject is detected within the specified area.

[0096] Here, you can restrict the types of subjects within the designated area that will be designated as the primary subject by setting the priority detection subject. For example, when prioritizing people, all subjects can be designated as the primary subject, but when prioritizing animals, vehicles will not be designated as the primary subject even if they are detected in the designated area, and similar restrictions can be imposed when prioritizing vehicles, even if animals are detected in the designated area, they will not be designated as the primary subject. When restricting the type of primary subject, you can either designate the designated area as the primary subject, as in the case where no subjects are detected in the designated area as described above, or you can use only the position and size of the subject from the detection results.

[0097] If a restricted subject is detected, the system may control the system by not switching to the restricted subject dictionary in subsequent frames, but instead using the priority dictionary. For example, when animals are prioritized, even if a vehicle is specified, the system will not switch to the vehicle dictionary in subsequent frames to avoid vehicle detection. Instead, it will switch to the animal dictionary more frequently to make animal detection easier. This control makes it easier to switch to the priority subject.

[0098] This embodiment describes how to specify an area within the display screen of the display unit in live view shooting, where images are sequentially input from the image sensor and displayed sequentially on the display unit. However, the method of specifying an area is not limited; it may also be done by eye-tracking on the screen displayed in the viewfinder, or by displaying a pointer on the live view screen or the screen displayed in the viewfinder and operating the pointer with a controller.

[0099] The object of the present invention can also be achieved as follows: a storage medium containing program code for software describing the procedures for realizing the functions of each embodiment described above is supplied to a system or device. The computer (or CPU, MPU, etc.) of that system or device then reads and executes the program code stored on the storage medium.

[0100] In this case, the program code read from the storage medium itself realizes the novel function of the present invention, and the storage medium and program that store that program code constitute the present invention.

[0101] Furthermore, storage media for supplying program code include, for example, flexible disks, hard disks, optical disks, and magneto-optical disks. CD-ROMs, CD-Rs, CD-RWs, DVD-ROMs, DVD-RAMs, DVD-RWs, DVD-Rs, magnetic tapes, non-volatile memory cards, and ROMs can also be used.

[0102] Furthermore, the functions of each of the embodiments described above are realized by making the program code read by the computer executable. In addition, this also includes cases in which the OS (operating system) running on the computer performs some or all of the actual processing based on the instructions of the program code, and the functions of each of the embodiments described above are realized through that processing.

[0103] Furthermore, the following cases are also included: First, program code read from a storage medium is written to the memory of a function expansion board inserted into a computer or a function expansion unit connected to a computer. Then, based on the instructions of that program code, the CPU or other components of that function expansion board or unit perform some or all of the actual processing. Although preferred embodiments of the present invention have been described above, the present invention is not limited to these embodiments, and various modifications and changes are possible within the scope of its gist. [Explanation of Symbols]

[0104] 100 Imaging device 28 Display section 50 System Control Unit 201 Subject detection unit 202 Detection history storage unit 203 Dictionary data storage unit 204 Dictionary Data Selection Section 205 Category Decision Department 206 Subject-Subject Determination Section

Claims

1. A detection means that detects multiple types of subjects in an input image, A setting means for setting a subject classification to be prioritized in order to determine the main subject from among a plurality of subject classifications, in which at least a portion of the plurality of subject types are classified, It has a means for determining the type of subject that will be the main subject, The setting means sets the preferred subject classification based on one of the following modes selected by the user: a plurality of priority modes in which each of the plurality of subject classifications is set as the preferred subject classification, and a mode in which the user does not specify a preferred subject classification. The main subject determination means is characterized in that it determines the subject type that will be the main subject from the priority subject classification set by the setting means and the subject type detected by the detection means.

2. The system includes a calculation means for calculating the reliability of detecting an object detected by the aforementioned detection means, The image processing apparatus according to claim 1, wherein the main subject determination means determines the type of subject in the same region based on the classification of preferred subjects set by the setting means, the type of subject detected by the detection means, and the confidence level calculated by the calculation means.

3. The image processing apparatus according to claim 1 or 2, characterized in that the detection means has dictionary data learned based on a neural network for each type of subject, and each of the dictionary data has different network parameters.

4. The image processing apparatus according to claim 3, characterized in that it has a control means for switching between the plurality of dictionary data based on predetermined settings.

5. The image processing apparatus according to any one of claims 1 to 3, characterized in that the main subject determination means performs the main subject determination process after obtaining the detection results of a plurality of pre-set types of subjects.

6. The image processing apparatus according to any one of claims 1 to 5, wherein the setting means sets a priority for each type of subject in accordance with the setting of the classification of the preferred subject.

7. The main subject determination means, when the detection results of the multiple types of subjects by the detection means exist in the same area, The image processing apparatus according to claim 6, characterized in that the subject with the highest priority is used as the main subject.

8. The image processing apparatus according to claim 6, characterized in that, when multiple types of subject detection results with the same priority exist in the same area, the subject with the highest reliability of detection by the detection means is designated as the main subject.

9. The image processing apparatus according to claim 2, characterized in that the main subject determination means normalizes the reliability of detection by the detection means according to the type of subject, and determines the main subject using the normalized reliability.

10. The image processing apparatus according to claim 4, characterized in that the control means changes to a switching sequence that switches all detectable dictionaries when an arbitrary region of the input image is specified by a user's region specification operation.

11. The image processing apparatus according to any one of claims 1 to 10, characterized in that the subject type includes at least one of birds, dogs, cats, motorcycles, and automobiles, wherein the birds, dogs, and cats are classified as animals as subject classifications, and the motorcycles and automobiles are classified as vehicles as subject classifications.

12. The image processing apparatus according to any one of claims 1 to 10, characterized in that the mode in which there is no setting of subject classification to be preferred by the user is a mode in which all subject classifications are set with the same priority.

13. The image processing apparatus according to any one of claims 1 to 10, characterized in that the mode in which the user does not have a preferred subject classification setting is a mode in which all subject classifications are set with the same priority.

14. A detection means that detects multiple types of subjects in an input image, A setting means for setting a subject classification to be prioritized in order to determine the main subject from among a plurality of subject classifications, in which at least a portion of the plurality of subject types are classified, It has a means for determining the type of subject that will be the main subject, The main subject determination means is characterized in that, when regions corresponding to multiple subject types detected by the detection means overlap, it determines one subject type for the overlapping region based on the preferred subject classification set by the setting means.

15. The image processing apparatus according to claim 14, characterized in that the detection means has dictionary data learned based on a neural network for each type of subject, and each of the dictionary data has different network parameters.

16. The image processing apparatus according to claim 14, characterized in that it has a control means for switching between the plurality of dictionary data based on predetermined settings.

17. The image processing apparatus according to claim 14, characterized in that the main subject determination means performs the main subject determination process after obtaining the detection results of a plurality of pre-set types of subjects.

18. The image processing apparatus according to claim 14, wherein the setting means sets a priority for each type of subject in accordance with the setting of the classification of the preferred subject.

19. The image processing apparatus according to any one of claims 14 to 18, characterized in that the subject type includes at least one of birds, dogs, cats, motorcycles, and automobiles, wherein the birds, dogs, and cats are classified as animals as subject classifications, and the motorcycles and automobiles are classified as vehicles as subject classifications.

20. A detection step that detects multiple subject types in an input image, A setting step of setting a subject classification to be prioritized for determining the main subject from among a plurality of subject classifications, in which at least a portion of the plurality of subject types are classified, A control method for an image processing apparatus, comprising a main subject determination step of determining one subject type to be the main subject from the classification of preferred subjects set in the setting step and the plurality of subject types detected in the detection step.

21. A detection step that detects multiple subject types in an input image, A setting step of setting a subject classification to be prioritized for determining the main subject from among a plurality of subject classifications, in which at least a portion of the plurality of subject types are classified, It includes a main subject determination step, which determines the type of subject that will be the main subject, The control method for an image processing apparatus is characterized in that, in the main subject determination step, if there is an overlap in the regions corresponding to multiple subject types detected in the detection step, one subject type is determined for the overlapping regions based on the priority subject classification set in the setting step.

22. A computer-executable program describing a procedure for controlling an image processing apparatus according to claim 20 or 21.

23. A computer-readable storage medium in which a program is stored that causes the computer to perform each step of the control method for the image processing apparatus described in claim 20 or 21.