Image processing device and control method thereof

The image processing device addresses overlapping detection results by using a subject detection unit, reliability calculation, and priority setting to determine the main subject accurately, enhancing subject identification clarity.

JP7814848B2Active Publication Date: 2026-02-17CANON KK
View PDF 7 Cites 0 Cited by

Patent Information

Application Number
JP2021065015
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Filing Date
2021-04-06
Publication Date
2026-02-17
Estimated Expiration
2041-04-06

AI Technical Summary

Technical Problem

Existing image processing systems struggle to determine a main subject when multiple types of detection results overlap for the same subject, leading to ambiguity in subject identification.

Method used

An image processing device and method that includes a subject detection unit, detection reliability calculation, priority subject setting, and main subject determination to select the correct detection type even when multiple detection results exist for the same subject.

Benefits of technology

Enables accurate selection of the main subject by prioritizing detection results based on reliability and user-defined priorities, effectively resolving overlapping detection issues.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007814848000002
    Figure 0007814848000002
  • Figure 0007814848000003
    Figure 0007814848000003
  • Figure 0007814848000004
    Figure 0007814848000004
Patent Text Reader

Abstract

To solve a problem that in a case where a plurality of detection results by a plurality of dictionaries exists for the same subject, a subject type may not be correctly selected.SOLUTION: A image processing apparatus include: detection means configured to detect a plurality of types of subjects for an input image; detection reliability calculation means configured to calculate detection reliability for the detected subjects; priority subject setting means configured to set a type of a subject as a priority main subject; and main subject determination means configured to determine a detection result as a main subject from among the detected subjects based on the set priority subject and the detection reliability. In a case where detection results of a plurality of types of subjects exist in a same region, the main subject determination means determines one subject type in the same region based on the set priority subject and the types of the detected subjects.SELECTED DRAWING: Figure 4
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] The present invention relates to an image processing device having a subject detection function and a control method thereof. [Background technology]

[0002] There is a technology for detecting multiple types of subjects from image data captured by an imaging device such as a digital camera, based on a trained model that has undergone machine learning for each type of subject. To capture an image with optimal focus, brightness, and color based on the detected subjects, it is necessary to determine one main subject from the multiple subjects. Patent Document 1 discloses a method for determining the main subject based on a stable presence index that indicates whether multiple detected subjects are stably detected across multiple frames. [Prior art documents] [Patent documents]

[0003] [Patent Document 1] Japanese Patent Application Laid-Open No. 2017-5738 Summary of the Invention [Problem to be solved by the invention]

[0004] However, there is no description as to how to determine the main subject and output the results when multiple types of detection results are output for the same subject in an overlapping manner.

[0005] In view of the above problems, an object of the image processing device of the present invention is to provide an image processing device and a control method for an image processing device that can appropriately detect a subject even when there are multiple detection results for the same subject using multiple dictionaries. [Means for solving the problem]

[0006] a subject detection means for detecting a plurality of types of subjects from an input image; from a plurality of classifications each including at least some of the types of the plurality of types of subjects,The present invention includes a detection reliability calculation means for calculating a detection reliability for the detected subject, a priority subject setting means for setting a subject type to be prioritized as a main subject, and a main subject determination means for determining a detection result to be a main subject from the detected subjects based on the set priority subject and the detection reliability, and the main subject determination means determines the detection result to be a main subject from the detected subjects based on the set priority subject and the detection reliability when a plurality of types of subject detection results exist in the same area. Set by means Priority given do subject Classification of and 、 The aforementioned By detection means The feature of this method is that the type of subject in the area is determined to be one. [Effects of the Invention]

[0007] According to the present invention, even when there are multiple detection results from multiple dictionaries for the same subject, it is possible to select the correct detection type. [Brief explanation of the drawings]

[0008] [Figure 1] External view of an imaging device including an image processing device [Figure 2] A block diagram showing the configuration of an imaging system including an image processing device. [Figure 3] FIG. 10 is a diagram showing an example of a method by which a user sets a target subject to be detected with priority. [Figure 4] Overall process flowchart [Figure 5] FIG. 10 is a diagram showing an example of a sequence for switching between multiple dictionary data sets. [Figure 6] Flowchart of a process for determining the type of subject within the same area [Figure 7] FIG. 10 is a diagram showing an example of a process for determining the type of subject in the same area; [Figure 8] Flowchart of main subject determination process [Figure 9] FIG. 10 is a diagram showing an example of main subject determination processing; [Figure 10] A diagram showing an example of a sequence for switching between multiple dictionary data sets at the user's discretion. DETAILED DESCRIPTION OF THE INVENTION

[0009] 1(a) and 1(b) show external views of an imaging device 100 including an image processing device as an example of an apparatus to which the present invention can be applied. Fig. 1(a) is a front perspective view of the imaging device 100, and Fig. 1(b) is a rear perspective view of the imaging device 100.

[0010] In FIG. 1, display unit 28 is a display unit provided on the back of the camera that displays images and various information. Touch panel 70a can detect touch operations on the display surface (operation surface) of display unit 28. Outside viewfinder display unit 43 is a display unit provided on the top surface of the camera, and displays various camera setting values ​​such as shutter speed and aperture. Shutter button 61 is an operation unit for issuing shooting instructions. Mode selector switch 60 is an operation unit for switching between various modes. Terminal cover 40 is a cover that protects connectors (not shown) such as connection cables that connect external devices to imaging device 100.

[0011] The main electronic dial 71 is a rotary operation member included in the operation unit 70, and by turning this main electronic dial 71, settings such as shutter speed and aperture can be changed. The power switch 72 is an operation member that switches the power of the imaging device 100 on and off. The sub electronic dial 73 is a rotary operation member included in the operation unit 70, and is used to move the selection frame, advance images, etc. The cross key 74 is included in the operation unit 70, and is a cross key (four-way key) that can be pressed up, down, left, and right. Operations can be performed according to the part of the cross key 74 that is pressed. The SET button 75 is included in the operation unit 70, and is a push button that is mainly used to confirm selections, etc.

[0012] The video button 76 is used to start and stop video shooting (recording). The AE lock button 77 is included in the operation unit 70 and can be pressed to lock the exposure state while the camera is ready to shoot. The enlarge button 78 is included in the operation unit 70 and is an operation button for turning the enlargement mode on and off in the live view display in shooting mode. Turning the enlargement mode on and operating the main electronic dial 71 allows the user to enlarge or reduce the LV image. In playback mode, the button functions as an enlargement button for enlarging the playback image and increasing the magnification. The playback button 79 is included in the operation unit 70 and is an operation button for switching between shooting mode and playback mode. Pressing the playback button 79 in shooting mode switches to playback mode, allowing the most recent image recorded on the recording medium 200 to be displayed on the display unit 28. The menu button 81 is included in the operation unit 70 and, when pressed, displays a menu screen on the display unit 28 that allows various settings to be made. The user can intuitively configure various settings using the menu screen displayed on the display unit 28, the cross key 74, and the SET button 75.

[0013] The touch bar 82 is a line-shaped touch operation member (line touch sensor) that can receive touch operations, and is located in a position that can be operated with the thumb of the right hand that is gripping the grip portion 90. The touch bar 82 can receive tap operations (operations in which the user touches and then releases the touch bar without moving it within a predetermined period of time), left and right slide operations (operations in which the user touches and then moves the touched position while still touching the bar), and the like. The touch bar 82 is an operation member that is different from the touch panel 70a, and does not have a display function.

[0014] The communication terminal 10 is a communication terminal that enables the imaging device 100 to communicate with the lens side (detachable). The eyepiece 16 is the eyepiece of an eyepiece finder (a peer-type finder), and the user can view the image displayed on the internal EVF 29 through the eyepiece 16. The eyepiece detection unit 57 is an eyepiece detection sensor that detects whether the photographer has placed their eye on the eyepiece 16. The lid 202 is a lid for a slot that stores a recording medium 200. The grip unit 90 is a holding unit shaped to be easily held in the right hand when the user holds the imaging device 100. When the digital camera is held by gripping the grip unit 90 with the little finger, ring finger, and middle finger of the right hand, the shutter button 61 and main electronic dial 71 are located in positions that can be operated with the index finger of the right hand. In the same state, the sub electronic dial 73 and touch bar 82 are located in positions that can be operated with the thumb of the right hand.

[0015] (Configuration of imaging device) FIG. 2 is a block diagram showing an example of the configuration of an image capture device 100 according to this embodiment. In FIG. 2, lens unit 150 is a lens unit equipped with an interchangeable photographic lens. Lens 103 is usually composed of multiple lenses, but for simplicity, only a single lens is shown here. Communication terminal 6 is a communication terminal through which lens unit 150 communicates with image capture device 100, and communication terminal 10 is a communication terminal through which image capture device 100 communicates with lens unit 150. Lens unit 150 communicates with system controller 50 via communication terminals 6 and 10, controls aperture 1 via aperture drive circuit 2 using an internal lens system control circuit 4, and adjusts focus by displacing the position of lens 103 via AF drive circuit 3.

[0016] The shutter 101 is a focal plane shutter that can freely control the exposure time of the imaging unit 22 under the control of the system control unit 50.

[0017] The imaging unit 22 is an imaging element configured with a CCD, CMOS element, or the like that converts an optical image into an electrical signal. The imaging unit 22 may have an imaging surface phase difference sensor that outputs defocus amount information to the system control unit 50. The A / D converter 23 converts an analog signal into a digital signal. The A / D converter 23 is used to convert the analog signal output from the imaging unit 22 into a digital signal.

[0018] The image processing unit 24 performs predetermined pixel interpolation, resizing such as reduction, and color conversion processing on the data from the A / D converter 23 or the data from the memory control unit 15. The image processing unit 24 also performs predetermined arithmetic processing using the captured image data. The system control unit 50 performs exposure control and distance measurement control based on the arithmetic results obtained by the image processing unit 24. This allows TTL (through-the-lens) type AF (autofocus) processing, AE (autoexposure) processing, and EF (flash pre-flash) processing to be performed. The image processing unit 24 also performs predetermined arithmetic processing using the captured image data, and performs TTL type AWB (auto white balance) processing based on the arithmetic results obtained.

[0019] The output data from the A / D converter 23 is written into the memory 32 via the image processing unit 24 and the memory control unit 15, or directly via the memory control unit 15. The memory 32 stores image data obtained by the imaging unit 22 and converted into digital data by the A / D converter 23, as well as image data to be displayed on the display unit 28 and the EVF 29. The memory 32 has a storage capacity sufficient to store a predetermined number of still images and a predetermined period of moving images and audio.

[0020] The memory 32 also serves as a memory (video memory) for image display. The D / A converter 19 converts the image display data stored in the memory 32 into an analog signal and supplies it to the display unit 28 and the EVF 29. In this way, the display image data written to the memory 32 is displayed on the display unit 28 and the EVF 29 via the D / A converter 19. The display unit 28 and the EVF 29 perform display according to the analog signal from the D / A converter 19 on a display device such as an LCD or an organic EL. The digital signal that has been A / D converted once by the A / D converter 23 and stored in the memory 32 is converted to analog in the D / A converter 19 and then sequentially transferred and displayed on the display unit 28 or the EVF 29, thereby performing a live view display (LV display). Hereinafter, an image displayed in live view will be referred to as a live view image (LV image).

[0021] The outside-finder liquid crystal display 43 displays various camera settings such as shutter speed and aperture via an outside-finder display drive circuit 44 .

[0022] The nonvolatile memory 56 is an electrically erasable and recordable memory, such as an EEPROM. The nonvolatile memory 56 stores constants, programs, etc. for the operation of the system control unit 50. The programs referred to here are programs for executing various flowcharts described later in this embodiment.

[0023] The system control unit 50 is a control unit made up of at least one processor or circuit, and controls the entire imaging device 100. Each process of this embodiment, which will be described later, is realized by executing a program recorded in the nonvolatile memory 56 described above. The system memory 52, for example, is a RAM, and constants and variables for the operation of the system control unit 50, programs read from the nonvolatile memory 56, and the like are loaded into the system memory 52. ​​The system control unit 50 also performs display control by controlling the memory 32, D / A converter 19, display unit 28, etc.

[0024] The system timer 53 is a timekeeping unit that measures the time used for various controls and the time of a built-in clock.

[0025] The operation unit 70 is an operating means for inputting various operational instructions to the system control unit 50. The mode selector switch 60 is an operating member included in the operation unit 70 and switches the operation mode of the system control unit 50 to one of still image capture mode, video capture mode, playback mode, etc. Modes included in the still image capture mode include auto capture mode, auto scene determination mode, manual mode, aperture priority mode (Av mode), shutter speed priority mode (Tv mode), and program AE mode (P mode). There are also various scene modes and custom modes that provide capture settings for specific shooting scenes. The mode selector switch 60 allows the user to directly switch to one of these modes. Alternatively, the user may first switch to a list screen of shooting modes using the mode selector switch 60, then select one of the displayed modes and switch using other operating members. Similarly, the video capture mode may also include multiple modes.

[0026] The first shutter switch 62 is turned on and generates a first shutter switch signal SW1 when the shutter button 61 provided on the imaging device 100 is pressed halfway (a shooting preparation instruction) during operation. The first shutter switch signal SW1 starts shooting preparation operations such as AF (autofocus) processing, AE (auto exposure) processing, AWB (auto white balance) processing, and EF (pre-flash) processing.

[0027] The second shutter switch 64 is turned on when the shutter button 61 is fully pressed (photographing instruction) and generates a second shutter switch signal SW2. The system control unit 50 starts a series of photographing processing operations, from reading out a signal from the imaging unit 22 to writing the captured image to the recording medium 200 as an image file, in response to the second shutter switch signal SW2.

[0028] The operation unit 70 is a variety of operation members that serve as an input unit for accepting operations from the user. The operation unit 70 includes at least the following operation members: the shutter button 61, the main electronic dial 71, the power switch 72, the sub electronic dial 73, the cross key 74, the SET button 75, the movie button 76, the AF lock button 77, the magnification button 78, the playback button 79, the menu button 81, and the touch bar 82. Other operation members 70b collectively represent operation members that are not individually shown in the block diagram.

[0029] The power supply control unit 80 is composed of a battery detection circuit, a DC-DC converter, a switch circuit for switching between powered blocks, etc., and detects whether a battery is installed, the battery type, and the remaining battery charge. The power supply control unit 80 also controls the DC-DC converter based on the detection results and instructions from the system control unit 50, and supplies the required voltage for the required period to each unit, including the recording medium 200. The power supply unit 30 is composed of primary batteries such as alkaline batteries or lithium batteries, secondary batteries such as NiCd batteries, NiMH batteries, or Li batteries, an AC adapter, etc.

[0030] The recording medium I / F 18 is an interface with a recording medium 200 such as a memory card or a hard disk. The recording medium 200 is a recording medium such as a memory card for recording captured images, and is composed of a semiconductor memory, a magnetic disk, or the like.

[0031] The communication unit 54 is connected wirelessly or via a wired cable and transmits and receives video signals and audio signals. The communication unit 54 can also be connected to a wireless LAN (Local Area Network) or the Internet. The communication unit 54 can also communicate with external devices via Bluetooth (registered trademark) or Bluetooth Low Energy. The communication unit 54 can transmit images (including LV images) captured by the imaging unit 22 and images recorded on the recording medium 200, and can also receive images and various other information from external devices.

[0032] The orientation detection unit 55 detects the orientation of the image capture device 100 with respect to the direction of gravity. Based on the orientation detected by the orientation detection unit 55, it is possible to determine whether an image captured by the image capture unit 22 was captured with the image capture device 100 held horizontally or vertically. The system control unit 50 can add orientation information corresponding to the orientation detected by the orientation detection unit 55 to the image file of the image captured by the image capture unit 22, or rotate and record the image. An acceleration sensor, a gyro sensor, or the like can be used as the orientation detection unit 55. The acceleration sensor or gyro sensor of the orientation detection unit 55 can also be used to detect movement of the image capture device 100 (panning, tilting, lifting, whether it is stationary, etc.).

[0033] (Configuration of image processing unit) 2(b) illustrates the configuration characteristic of the present embodiment of the image processing unit 24. The image processing unit 24 has a subject detection unit 201, a detection history storage unit 202, a dictionary data storage unit 203, a dictionary data selection unit 204, and a main subject determination unit 205. Although described as part of the image processing unit 24 in this embodiment, these may be part of the system control unit 50, or may be provided separately from the image processing unit 24 and the system control unit 50. The image data generation, detection history storage, and image processing device may be installed in, for example, a smartphone, a tablet terminal, or the like.

[0034] The image processing unit 24 generates image data based on the data output from the A / D converter 23 and sends the image data to a subject detection unit 201 within the image processing unit 24 .

[0035] In this embodiment, the object detection unit 201 is configured with a CNN (convolutional neural network) that has undergone machine learning (deep learning) and detects specific objects. The types of detectable objects are based on dictionary data stored in the dictionary data storage unit 203. In this embodiment, the object detection unit 201 is configured with different CNNs (different network parameters) depending on the types of detectable objects. The object detection unit 201 may be realized by a GPU (graphics processing unit) or a circuit specialized for estimation processing using CNN.

[0036] The machine learning of the CNN may be performed by any method. For example, a predetermined computer such as a server may perform the machine learning of the CNN, and the imaging device 100 may acquire the trained CNN from the predetermined computer. In this embodiment, the predetermined computer receives training image data as input and performs supervised learning using position information and the like of the subject corresponding to the training image data as training data (annotation), thereby training the CNN of the subject detection unit 201. In this manner, a trained CNN is generated. The training of the CNN may be performed by the imaging device 100 or the image processing device described above.

[0037] As described above, the object detection unit 201 includes a CNN (trained model) trained by machine learning. The object detection unit 201 receives image data as input, estimates the position, size, reliability, etc. of the object, and outputs the estimated information. The CNN may be, for example, a network in which a fully connected layer and an output layer are connected to a layer structure in which convolutional layers and pooling layers are alternately stacked. In this case, for example, backpropagation may be applied to train the CNN. The CNN may also be a neocognitron CNN that includes a feature detection layer (S layer) and a feature integration layer (C layer). In this case, a learning method called "Add-if Silent" may be applied to train the CNN.

[0038] The object detection unit 201 may use any trained model other than a trained CNN. For example, a trained model generated by machine learning such as a support vector machine or a decision tree may be applied to the object detection unit 201. Furthermore, the object detection unit 201 does not need to use a trained model generated by machine learning. For example, the object detection unit 201 may use any object detection method that does not use machine learning.

[0039] The detection history storage unit 202 stores the subject detection history detected by the subject detection unit 201 in the image data, and the system control unit 50 sends the history to the dictionary data selection unit 204. In this embodiment, the subject detection history stores the dictionary data used for detection, the position of the detected subject, the size of the detected subject, and the reliability of the detected subject. Other data such as the number of detections and the identifier of the image data containing the detected subject may also be stored.

[0040] The dictionary data storage unit 203 stores dictionary data for detecting specific subjects, and the system control unit 50 reads dictionary data selected by the dictionary data selection unit 204 from the dictionary data storage unit 203 and sends it to the subject detection unit 201. Each dictionary data is, for example, data in which the characteristics of each area of ​​a specific subject are registered. Furthermore, to detect multiple types of subjects, dictionary data for each subject and each area of ​​the subject may be used. The dictionary data storage unit 203 stores multiple types of subject detection dictionary data, such as dictionary data for detecting "people," dictionary data for detecting "animals," and dictionary data for detecting "vehicles." Furthermore, in addition to "animals," dictionary data for detecting "birds," which have unique shapes and are in high demand for detection, may also be stored. Furthermore, the dictionary data for detecting "vehicles" can be further subdivided into "cars," "motorcycles," "trains," "airplanes," and other categories and stored separately.

[0041] The area of ​​the subject detected by the multiple dictionary data stored in the dictionary data storage unit 203 can be used as a focus detection area. For example, in a composition where an obstacle is present in the foreground and the subject is present in the background, it is possible to focus on the desired subject by focusing within the detected area.

[0042] Furthermore, in this embodiment, the multiple dictionary data used for detection by the object detection unit 201 are generated by machine learning, but dictionary data generated by a rule base may also be used or used in combination. Rule-based dictionary data is, for example, dictionary data that stores images of objects to be detected or feature quantities specific to the objects, as determined in advance by a designer. The objects can be detected by comparing the images or feature quantities of the dictionary data with the images or feature quantities of captured image data. Rule-based dictionary data is less complicated than models set by machine learning trained models, and therefore requires less data volume. Object detection using rule-based dictionary data also has faster processing speeds (lower processing loads) than trained models.

[0043] The dictionary data selection unit 204 selects the dictionary data to be used next based on the detection history stored in the detection history storage unit 202, a predetermined order / rule, or instructions from the user, and notifies the dictionary data storage unit 203 of the selection.

[0044] In this embodiment, dictionary data for each of a plurality of types of objects and each region of the objects is stored separately in the dictionary data storage unit 203, and object detection is performed multiple times for the same image data by switching between the plurality of dictionary data. The dictionary data selection unit 204 determines a dictionary data switching sequence and determines the dictionary data to be used in accordance with the determined sequence. An example of the dictionary data switching sequence will be described later.

[0045] When multiple subjects are detected in the same area, the type determination unit 205 determines the type of subject in that area. One detection result is determined from the multiple detection histories stored in the detection history storage unit 202 based on the subject detection priority setting set by the user via the operation unit 70. The determination method will be described later.

[0046] Regarding the method of setting the subject to be detected with priority, an example in which the user selects whether or not to prioritize the type of subject to be detected from a menu screen displayed on the display unit 28 is shown in Fig. 3. Fig. 3 shows the detection subject selection setting screen displayed on the display unit 28, and the user selects the subject to be detected with priority from among specific detectable subjects (for example, vehicles, animals, people) by operating the operation unit 70. In Fig. 3, "vehicle" is selected. Also, "none" in Fig. 3 is a mode in which no subjects are detected, and "auto" is a mode in which all specific detectable subjects are detected without assigning any priority.

[0047] Main subject determination unit 206 determines the main subject based on multiple detection histories stored in detection history storage unit 202, subject settings to be preferentially detected set by the user via operation unit 70, and the subject determined by type determination unit 205. The main subject determination method will be described later.

[0048] (Processing flow of the imaging device) 4 is a flowchart showing the flow of characteristic processing of the present invention performed by the imaging device 100 of this embodiment. Each step of this flowchart is executed by the system control unit 50 or by each unit in response to an instruction from the system control unit 50. When this flowchart starts, it is assumed that the imaging device 100 is powered on and in live view imaging mode, and that an instruction to start capturing (recording) still images or videos can be given by operating the operation unit 70.

[0049] 4 is assumed to be processing performed when one frame (one piece of image data) is captured by the imaging unit 22 of the imaging device 100. However, this is not limiting, and the processing from step S401 to step S409 may be performed over multiple frames. In other words, the result of subject detection in the first frame may be reflected in any of the frames from the second frame onwards.

[0050] In step S401, the system control unit 50 acquires the captured image data captured by the imaging unit 22 and output by the A / D conversion unit .

[0051] In step S 402 , the image processing unit 24 resizes the image data to an image size that is easy to process (for example, QVGA), and sends the resized image data to the image data generation unit 201 .

[0052] In step S403, the dictionary data selection unit 204 selects dictionary data generated by machine learning to be used for subject detection, and sends selection information indicating which dictionary data has been selected to the dictionary data storage unit 203.

[0053] Here, dictionary data generated by machine learning can be generated by extracting common features of specific subjects from a large amount of image data containing those subjects. Common features include, for example, the size, position, and color of the subject, as well as areas outside the specific subject, such as the background. Therefore, the more limited the background in which the subject is detected, the easier it is to improve detection performance (detection accuracy) with less training. On the other hand, training to detect specific subjects regardless of the background provides high versatility across shooting scenes but is less likely to improve detection accuracy. The more image data and diversity are used to generate the dictionary data, the higher the detection performance tends to be. On the other hand, by limiting the size and position of the detection area of ​​the subject to be detected in the image data used for detection, it is possible to improve detection performance even if the number and diversity of image data required to generate the dictionary data is reduced. Furthermore, if part of the subject is outside the image data, some of the subject's features are lost, resulting in lower detection performance.

[0054] In general, the larger the area of ​​a subject, the more features it contains. In detection using the aforementioned machine-learned dictionary data, there is a possibility that an object with similar features other than the specific subject to be detected using the dictionary data may be mistakenly detected as the specific subject. The area defined as a local area is narrower than the entire area. The narrower the area, the fewer features it contains, and the fewer features it contains, the more objects with similar features there will be, leading to an increase in false detections.

[0055] The switching sequence for multiple dictionary data for one frame (one piece of image data) in step S403 will be described using Figure 5. When multiple dictionary data are stored in the dictionary data storage unit 203, it is possible to perform detection using multiple dictionaries for one frame. However, due to issues with imaging speed and processing speed, images captured sequentially are output and processed in live view mode, and the number of times subject detection can be performed for one frame is considered to be limited for video data when recording a video.

[0056] In this case, the type and order of dictionary data to be used can be determined based on, for example, whether or not a subject has been detected in the past, the type of dictionary data used at that time, and the type of subject to be detected with priority. Depending on the dictionary data switching sequence, when a specific subject is included in a frame, the corresponding dictionary data for subject detection may not be selected, resulting in a missed detection opportunity. For this reason, the dictionary data switching sequence must also be changed according to the settings and scene.

[0057] For example, in a structure that can perform subject detection up to three times per frame (or has three detectors capable of processing in parallel), an example of a dictionary data switching sequence when a vehicle is selected as the subject to be detected first is shown in Figure 5. V0 and V1 each indicate the vertical synchronization period for one frame, and the rectangular blocks such as a person's head, vehicle 1 (motorcycle), and vehicle 2 (car) indicate that subject detection can be performed using three dictionary data (trained models) in chronological order within one vertical synchronization period.

[0058] FIG. 5(a) shows an example of dictionary data switching when no detected subject is present. In the first frame, the dictionary data is switched in the order of a person's head, vehicle 1 (motorcycle), and vehicle 2 (car). In the second frame, the dictionary data is switched in the order of an animal (dog / cat), vehicle 1 (motorcycle), and vehicle 2 (car). For example, without a switching sequence, dictionary data capable of detecting the subject selected by the user from the menu screen, as shown in FIG. 3, is always used. In this case, the priority detection subject setting is switched for each scene, with vehicles selected when a vehicle is present and people and animals selected otherwise, resulting in the time-consuming task of switching the priority detection subject setting for each scene. Furthermore, if it is unclear when a vehicle will arrive, switching the priority detection subject setting after noticing the vehicle may result in missing the shot. In contrast, in this embodiment, as shown in FIG. 5(a), all dictionary data is switched across multiple frames during periods when no specific subject is detected, allowing for capture without worrying about the priority detection subject setting. By switching all dictionary data while selecting dictionary data corresponding to the priority detection subject setting in both the first and second frames, it is possible to detect all detectable subjects while improving the detection accuracy of the priority detection subject. This reduces the number of times the priority detection subject setting needs to be switched. Also, a separate mode may be provided in which only a specific dictionary (group) is always rotated according to the settings specified by the user.

[0059] FIG. 5(b) shows an example of dictionary data switching in the next frame when a motorcycle is detected in the previous frame, with the dictionary data switching in the order Vehicle 1 (motorcycle), person's head, and Vehicle 1 (motorcycle). The order of dictionary data switching does not necessarily have to be as described above. For example, the person's head dictionary data in the example of dictionary data switching described above may be switched to dictionary data that is more likely to be selected as a subject other than a motorcycle in a scene capturing a motorcycle, depending on the scene. Furthermore, exclusive control may be used to prevent detection using dictionary data for "animals," which are unlikely to be detected alongside vehicles. Depending on the texture (pattern) or color of a vehicle, it may be detected as an animal. This exclusive control ultimately improves the accuracy of detecting the desired subject.

[0060] In step S404, the subject detection unit 201 uses dictionary data for detecting specific subjects (objects) stored in the dictionary data storage unit 203 to detect a subject (or a subject area in which the subject exists) from image data captured by the imaging unit 22 and input to the image processing unit 24. Information such as the position and size of the detected subject, the calculated reliability, the type of dictionary data used, and the identifier of the image data used for detection are stored in the detection history storage unit 202.

[0061] In step S405, it is determined whether detection has been performed using all necessary dictionary data for image data having the same identifier (image data of the same frame) from the detection history stored in the detection history storage unit 203. If the determination result is Yes, the process proceeds to S406, and if No, the process returns to step S403 to select the next dictionary data to be used.

[0062] In step S406, it is determined whether detection has been performed using all dictionary data from the detection history stored in the detection history storage unit 203. If the determination result is Yes, the process proceeds to S407; if No, the process proceeds to the next frame. For example, in FIG. 5A, since it takes two frames to perform detection using all necessary dictionary data, the subsequent processing is skipped for the first frame, the process proceeds to the next frame, and the process proceeds to S407 for the second frame. In this embodiment, subsequent processing is skipped until detection has been performed using all necessary dictionary data. However, this is not limited to this. For processing that requires quick response, such as autofocus, subsequent processing may be performed using only the subjects detected in each frame without waiting for detection using all dictionary data. Furthermore, for example, if all currently set dictionary data can be processed in two frames, as in this embodiment, subsequent processing from step S407 onwards may be performed based on the detection results for two frames, including the previous frame.

[0063] In step S407, a setting for selecting a subject to be preferentially detected from among specific detectable subjects set in advance by the user via the operation unit 70 is read out.

[0064] In step S408, it is determined whether or not multiple detection results exist in the same region from the detection history of the detection results of image data with the same identifier stored in the detection history storage unit 203. If multiple results exist in the same region, proceed to S409; if not, proceed to step S410. Whether multiple detection results exist in the same region may be determined, for example, if the detection center coordinates are within different detection result regions. Alternatively, it may be determined that multiple detection results exist in the same region if the detection regions overlap at a rate equal to or greater than a threshold; the method is not limited thereto.

[0065] In step S409, the type determination unit 205 determines one area detection result from the priority subject setting set in step S407, the detection result saved in step S405, and the result determined to exist in the same area in step S408. The determination method will be described later.

[0066] In step S410, main subject determination unit 206 uses the priority subject setting set in step S407 to determine a main subject from among multiple detection results of image data having the same identifier based on the detection history stored in detection history storage unit 203. At this time, if it is determined in step S408 that there are multiple detection results for the same region, the result of step S409 is also used. At this time, system control unit 50 may cause display unit 28 to display some or all of the information output by main subject determination unit 206. The determination method will be described later.

[0067] (Flow of type determination process to determine the type of subject from multiple subject detection results within the same area) The type determination process in S409 will be described with reference to the flowchart in Fig. 6, Fig. 7, and Table 1. Each step in this flowchart is executed by the system control unit 50 or by each unit in response to an instruction from the system control unit 50.

[0068] An example of the type determination process will be explained with reference to Figure 7. Figure 7(a) shows an input image in which a motorcycle 701 is shown as the subject. Figure 7(b) shows that the person dictionary was selected in step S403 and a person 702 was detected. Figure 7(c) shows that the motorcycle dictionary was selected in step S403 and a motorcycle 703 was detected. Figure 7(d) shows that the car dictionary was selected in step S403 and a car 704 was erroneously detected. Figure 7(e) shows that the dog dictionary was selected in step S403 and a dog 705 was erroneously detected. Figure 7(f) shows that the cat dictionary was selected in step S403 and no detection results were found as a result of the process.

[0069] In step S601, priorities are assigned to the detected subject types in accordance with the priority settings made in step S407.

[0070] Table 1 shows an example of priority classification based on priority settings and subject type. The vertical axis in Table 1 represents the priority settings that can be set, with "People," "Animals," "Vehicles," "None," and "Automatic" in line with Figure 3. The horizontal axis represents the subject type that can be detected, with "People," "Dogs," "Cats," "Cars," and "Motorcycles" in line with Figure 7. The smaller the priority number in Table 1, the higher the priority, and "No priority" indicates that the subject will not be used.

[0071] In this embodiment, three values ​​are used: priority subject (priority 1 in the table), non-priority subject (priority 2 in the table), and non-selected subject (no priority in the table), but this is not limited to this. For example, two values, selected subject and non-selected subject, may be used. Furthermore, four values, top priority subject, priority subject, non-priority subject, and non-selected subject, may be used. This can be changed depending on the number of subject types that can be detected and the priority settings that can be set. Furthermore, in Table 1, when priority is given to vehicles, cars and motorcycles are classified as priority subjects, people as non-priority subjects, and dogs and cats as non-selected subjects. However, the classification method is not limited to this. For example, if it is desired to only detect subject types that have been set as priority, people may also be classified as non-selected subjects. Furthermore, if it is desired to detect subjects other than those set as priority subjects, dogs and cats may also be classified as non-priority subjects.

[0072] [Table 1]

[0073] In step S602, a subject type determination process is performed within the same area based on the priority determined in step S601.

[0074] A specific method will be explained using FIG. 7. For example, if a person is set as the priority setting, the subject with priority 1 from Table 1 is a person, so a check is made to see if there are any human detection results. Since a person 702 is present in FIG. 7(b), this is adopted and this is determined to be the subject type within the area, and the type determination process is terminated. If a vehicle is set as the priority setting, the subjects with priority 1 from Table 1 are a car and a motorcycle, so a check is made to see if there are any car or motorcycle detection results with priority 1. Since both a motorcycle 702 in FIG. 7(c) and a car in FIG. 7(d) are present, the process proceeds to step S603. In this case, if neither a motorcycle 702 in FIG. 7(c) nor a car in FIG. 7(d) is selected, a check is made to see if there are any human detection results with priority 2. If there are no human results, since dogs and cats are not priorities from Table 1, it is determined that there are no subjects within the area, and the type determination process is terminated.

[0075] In step S603, the reliability of the detection results saved in step S405 is normalized for each subject. The purpose of normalization is that the maximum reliability of the detection results and the reliability threshold for a subject vary depending on the dictionary used. Normalization makes it possible to compare the reliability of subjects in different dictionaries in subsequent processing. In this embodiment, the minimum possible reliability value for each dictionary is normalized to 0 and the maximum possible reliability value to 1, thereby reducing the reliability of all subjects between 0 and 1 and enabling comparison based on reliability. The normalization method is not limited to this; for example, the reliability threshold for a subject may be set to 1 and the minimum possible reliability value may be set to 0, and any other method may be used.

[0076] In step S604, if it is confirmed in step S602 that there are multiple object types with the same priority, the result with the highest reliability normalized in step S603 is determined as the object within the region, and the type determination process ends. In this embodiment, the object within the region is determined based on reliability, but the method is not limited to this. For example, the detection results of past frames may be referenced, and the object type that is most frequently detected in multiple frames may be determined as the object within the region.

[0077] In Fig. 7, the motorcycle 703 in Fig. 7(c) and the car 704 in Fig. 7(d) are determined to be priority subjects in the same area in step S602, so they are compared. In this embodiment, since the input subject is the motorcycle 701, it is assumed that the reliability of the motorcycle 703 in Fig. 7(c) is the highest, and the motorcycle 703 is determined to be the subject in the area.

[0078] Prior to comparing the reliability in step S604, object selection is performed based on priority in step S602. The purpose of this is that, for example, when objects share similar common characteristics, such as when both dogs and cats walk on all fours, there is a high possibility that a cat will be mistakenly detected as a dog even if a cat image is input into the dog dictionary. However, when objects share dissimilar common characteristics, such as when a dog and a motorcycle are input into the dog dictionary, there is a low possibility that the motorcycle will be mistakenly detected as a dog even if an image of a motorcycle is input into the dog dictionary. However, when a false detection occurs, such as with dog 705 in Figure 7(e), it can be difficult to determine which image feature was reacted to, resulting in a high reliability. In such cases, it is difficult to prevent the final output from being a dog. Therefore, by first selecting objects based on the set priority, false detection of undesired objects is eliminated.

[0079] (Flow of main subject determination process) The main subject determination process in S410 will be described with reference to the flowchart in Figure 8 and Figure 9. Each step in this flowchart is executed by the system control unit 50 or by each unit in response to an instruction from the system control unit 50.

[0080] Figure 9 shows an example of main subject determination when multiple subjects are detected in the same frame. Figure 9(a) shows the detection of a human face 901, and the detection of cats 902 and 903. Figure 9(b) shows the selection of human face 904 as the main subject among human face 901, cat 902, and cat 903. Figure 9(c) shows the selection of cat 905 as the main subject among human face 901, cat 902, and cat 903.

[0081] In step S801, a main subject candidate is selected according to the priority setting established in step S407. If a unique main subject candidate is determined, the main subject candidate is set as the main subject and main subject determination ends. If no candidate is found, it is determined that there is no main subject and main subject determination ends. If multiple candidates are found, proceed to step S802.

[0082] A specific example will be described with reference to FIG.

[0083] If "person" in FIG. 3 is set in step S407, the human face 901, cat 902, and cat 903 in FIG. 9(a) are selected as the main subject as shown in 904 in FIG. 9(b) in accordance with the priority setting, and main subject determination is completed.

[0084] When "animal" in FIG. 3 is set in step S407, there are multiple cats detected among the human face 901, cat 902, and cat 903 in FIG. 9(a), so the process proceeds to step S802.

[0085] If "Auto" in FIG. 3 is set in step S407, there is no subject to be detected with priority, and there are multiple detection results for people and cats, so the process proceeds to step S802.

[0086] If "vehicle" in FIG. 3 is set in step S407, none of the human face 901, cat 902, and cat 903 in FIG. 9A will be used as the subject, so there is no main subject and main subject determination ends.

[0087] In step S802, a main subject is selected from the multiple candidate subjects determined in step S801 based on the position, size, reliability, etc. of the multiple subjects detected in step S404. For example, when determining that a subject close to the center of the angle of view is the main subject, if three candidates, a human face 901, a cat 902, and a cat 903, remain in step S801, the human face 901 is closest to the center and is therefore determined to be the main subject, as shown in 904 in Figure 9(b).

[0088] When two cats, cat 902 and cat 903, remain as candidates, cat 902 is closest to the center, so cat 905 in FIG. 9(c) is selected as the main subject.

[0089] In this embodiment, the subject closest to the center of the angle of view among the candidate subjects is determined as the main subject, but this is not limited to this. For example, the main subject may be the subject closest to the center of the area where autofocus is possible, or another large subject, or a subject with a high degree of reliability as the detected subject, or the main subject may be determined by a combination of these factors.

[0090] (Example in which the user performs a designation operation on the screen) In the above-described embodiment, the imaging device automatically detects a subject, determines the type of subject within the same area, and determines the main subject. In this embodiment, when the user arbitrarily specifies an area within the live view screen displayed on the display unit 28, the dictionary switching sequence is changed, and the type of subject within the same area and the main subject are determined.

[0091] The dictionary switching sequence in step S403 when the user designates an arbitrary area on the live view screen in the dictionary data selection unit 204 will be described with reference to FIG.

[0092] In Figure 5, the switching sequence was changed depending on the subjects that had already been detected and the priority detection subject setting. However, in this embodiment, when the user specifies an area within the live view screen, all detectable dictionaries are switched regardless of the detected subjects and the priority detection subject setting. This is because, in order to accurately reflect the area specification by the user, regardless of the subjects that have already been detected, all detectable dictionaries are switched to accurately detect the subject in the specified area.

[0093] FIG. 10 shows an example of dictionary data switching. In the first frame, dictionary data is switched in the order of person's head, vehicle 1 (motorcycle), and vehicle 2 (car), and in the second frame, dictionary data is switched in the order of person's head, animal (dog / cat), and animal (bird), so that dictionary data is switched across multiple frames. In this embodiment, the person's head is switched in both the first and second frames, but either one may be a different dictionary depending on the priority detection subject setting. For example, when vehicle priority is set, one of the vehicle dictionaries may be used in the second frame, and when animal priority is set, one of the animal dictionaries may be used in the second frame.

[0094] The type determination process in S409, which is a characteristic process of this embodiment, will be described below. In this embodiment, the process is performed when multiple types of subjects are detected within the area specified by the user.

[0095] The main subject determination process in S410 is a characteristic process of this embodiment. In this embodiment, a subject within a region designated by the user is determined to be the main subject. If no subject is detected within the designated region, the designated region is determined to be the main subject. However, in the dictionary data switching sequence in step S403 in the next frame, all dictionaries are switched until a detectable subject is detected within the designated region.

[0096] Here, the type of subject within the specified area that is to be the main subject may be limited by setting a priority detection subject. For example, when people are prioritized, all subjects can be designated as the main subject by designation, but when animals are prioritized, a vehicle may not be designated as the main subject even if it is detected in the specified area, and when vehicles are prioritized, an animal may not be designated as the main subject even if it is detected in the specified area. When limiting the type of main subject, the specified area may be designated as the main subject, as in the case when no subject is detected within the specified area described above, or only the position and size of the subject from the detection results may be used.

[0097] If it is determined that a restricted subject has been specified, control may be performed so that the dictionary of the priority setting is used instead of switching to the dictionary of the restricted subject in the next frame onwards. For example, even if a vehicle subject is specified when animal priority is set, vehicle detection is not performed by not switching to the dictionary of vehicles in the following frames, and instead the dictionary of animals is switched to more frequently, making it easier to detect animals. Control in this way makes it easier to switch to the priority setting subject.

[0098] This embodiment has described area designation within the display screen of the display unit in live view photography, in which images input sequentially from the image sensor are displayed sequentially on the display unit. However, the area designation method is not limited to this; it is also possible to designate an area by directing the line of sight on the screen displayed in the viewfinder, or by displaying a pointer on the live view screen or the screen displayed in the viewfinder and operating the pointer with a controller.

[0099] The object of the present invention can also be achieved as follows: A storage medium storing software program code describing procedures for realizing the functions of each of the above-described embodiments is supplied to a system or device, and the computer (or CPU, MPU, etc.) of the system or device reads and executes the program code stored in the storage medium.

[0100] In this case, the program code itself read from the storage medium will realize the novel functions of the present invention, and the storage medium storing the program code and the program will constitute the present invention.

[0101] Furthermore, examples of storage media for supplying the program code include flexible disks, hard disks, optical disks, magneto-optical disks, etc. Also usable are CD-ROMs, CD-Rs, CD-RWs, DVD-ROMs, DVD-RAMs, DVD-RWs, DVD-Rs, magnetic tapes, non-volatile memory cards, ROMs, etc.

[0102] The functions of the above-described embodiments are realized by making the computer executable the read program code. Furthermore, the functions of the above-described embodiments may be realized by an operating system (OS) or the like running on the computer performing some or all of the actual processing based on the instructions of the program code.

[0103] The following cases are also included: First, program code is read from a storage medium and written into the memory of an expansion board inserted into a computer or an expansion unit connected to the computer. Then, based on the instructions of the program code, the CPU or other device in the expansion board or expansion unit performs some or all of the actual processing. Although the preferred embodiments of the present invention have been described above, the present invention is not limited to these embodiments, and various modifications and changes are possible within the scope of the gist of the present invention. [Explanation of symbols]

[0104] 100 Imaging device 28 Display section 50 System control section 201 Subject detection unit 202 Detection history memory unit 203 Dictionary data storage unit 204 Dictionary Data Selection Unit 205 Category Decision Department 206 Subject-Subject Determination Section

Claims

1. a detection means for detecting a plurality of subject types from an input image; a setting means for setting a priority object classification for determining a main object from a plurality of object classifications into which at least some of the plurality of object types are classified; a main subject determining means for determining a subject type to be a main subject; When there are detection results of a plurality of types of subjects in the same area, the main subject determination means 10. An image processing apparatus comprising: a processor for determining one type of subject in the same area based on the prioritized subject classification set by said setting means and the type of subject detected by said detecting means;

2. a calculation unit for calculating a reliability of detection of the object detected by the detection unit; 2. The image processing device according to claim 1, wherein the main subject determining means determines the type of subject in the same area based on the classification of the prioritized subject set by the setting means, the type of subject detected by the detection means, and the reliability calculated by the calculation means.

3. 3. The image processing apparatus according to claim 1, wherein said detecting means has dictionary data for each type of subject that has been trained based on a neural network, and the dictionary data has different network parameters for each type of subject.

4. 4. The image processing apparatus according to claim 3, further comprising control means for switching between said plurality of dictionary data based on a predetermined setting.

5. 4. The image processing device according to claim 1, wherein the main subject determining means performs the process of determining the main subject after obtaining detection results of a plurality of types of subjects set in advance.

6. 6. The image processing apparatus according to claim 1, wherein the setting unit sets a priority for each type of subject in accordance with a setting of the priority subject classification.

7. When the detection results of the plurality of types of subjects by the detection means exist in the same area, the main subject determination means 7. The image processing apparatus according to claim 6, wherein the subject with the highest priority is set as a main subject.

8. 7. The image processing device according to claim 6, wherein when a plurality of types of subject detection results having the same priority exist in the same area, the main subject determination means determines the subject detected by the detection means with the highest reliability as the main subject.

9. 3. The image processing apparatus according to claim 2, wherein said main subject determining means normalizes the reliability of detection by said detecting means in accordance with the type of subject, and determines the main subject using said normalized reliability.

10. 5. The image processing device according to claim 4, wherein the control means changes the switching sequence to one that switches all detectable dictionaries when a user specifies an arbitrary area of ​​the input image by a region specifying operation.

11. 11. The image processing device according to claim 1, wherein the subject types include at least one of birds, dogs, cats, motorcycles, and automobiles, and the birds, dogs, and cats are classified as animals, and the motorcycles and automobiles are classified as vehicles.

12. a detection means for detecting a plurality of subject types from an input image; a setting means for setting a priority object classification for determining a main object from a plurality of object classifications into which at least some of the plurality of object types are classified; a main subject determining means for determining a subject type to be a main subject, When subject detection results of a plurality of subject types detected by the detection means are obtained, the main subject determination means determines one subject type to be the main subject from the prioritized subject classification set by the setting means and the plurality of subject types detected by the detection means.

13. 13. The image processing apparatus according to claim 12, wherein said detecting means has dictionary data for each type of subject that has been trained based on a neural network, and the dictionary data has different network parameters for each type of subject.

14. 14. The image processing device according to claim 12, wherein when subject detection results of a plurality of subject types detected by the detection means overlap, the main subject determination means determines the subject type in the overlapping area as the subject type detected by the detection means with a higher reliability.

15. 15. The image processing device according to claim 12, wherein the subject types include at least one of birds, dogs, cats, motorcycles, and automobiles, and the birds, dogs, and cats are classified as animals in the subject classification, and the motorcycles and automobiles are classified as vehicles in the subject classification.

16. a detection step of detecting a plurality of types of subjects from an input image; a setting step of setting a subject classification to be prioritized as a main subject from a plurality of classifications each including at least some of the plurality of types of subjects; a main subject determination step of determining a main subject from the plurality of types of subjects detected in the detection step, In the main subject determination step, if there are detection results of a plurality of types of subjects in the same area, A control method for an image processing device, characterized in that the type of subject in the same area is determined to be one based on the classification of the prioritized subject set in the setting step and the type of subject detected in the detection step.

17. a detection step of detecting a plurality of subject types from an input image; a setting step of setting a priority object classification for determining a main object from a plurality of object classifications into which at least some of the plurality of object types are classified; a main subject determination step of determining one subject type to be a main subject from the classification of prioritized subjects set in the setting step and the plurality of subject types detected in the detection step.

18. 18. A computer-executable program in which the procedure of the control method for the image processing apparatus according to claim 16 or 17 is written.

19. 18. A computer-readable storage medium storing a program for causing a computer to execute each step of the method for controlling an image processing apparatus according to claim 16 or 17.

Citation Information

Patent Citations

  • Image data generating apparatus and image data generating method

    JP2007281533A

  • Image processing apparatus, imaging apparatus, image processing apparatus control method, image processing apparatus control program, and storage medium

    JP2015103852A

  • Image processor, image processing method, and program

    JP2017005738A

  • Imaging device and imaging method

    JP2019146022A

  • Image processing device, image processing method, program and storage medium

    JP2020008899A