Imaging device

JP2026148175APending Publication Date: 2026-09-17PANASONIC INTELLECTUAL PROPERTY MANAGEMENT CO LTD
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
JP2025036586
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Filing Date
2025-03-07
Publication Date
2026-09-17

AI Technical Summary

Benefits of technology

【0007】 本開示に係る撮像装置によると、画像においてユーザの意図に沿った被写体の領域を認識し易くすることができる。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2026148175000001_ABST
    Figure 2026148175000001_ABST
Patent Text Reader

Abstract

The present invention provides an imaging device that makes it easier to recognize areas of subjects in images that align with the user's intent. [Solution] The imaging device comprises an imaging unit, a first recognition unit, a second recognition unit, and a control unit. The imaging unit captures an image of a subject and generates image data. The first recognition unit recognizes the region of the first subject in the image indicated by the image data based on the image data generated by the imaging unit. The second recognition unit recognizes the region of the second subject in the image that is included in or separate from the first subject, based on the image data. The control unit determines the subject region of the recognition result in the image data based on the region of the first subject recognized by the first recognition unit and the region of the second subject recognized by the second recognition unit.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to an imaging apparatus that recognizes a region of a subject in an image.

Background Art

[0002] Patent Document 1 discloses an imaging apparatus aimed at appropriately selecting a subject desired by a user when selecting a main subject from an image. This imaging apparatus determines a main subject from an image during live view display using a trained CNN, and further selects a data set to be used for additional learning of the CNN in accordance with a user operation to change the main subject from the determination result. In this way, from the data set of the image displayed in live view and the main subject region, respective data sets are selected for positive learning, which learns regions that should be determined as the main subject, and negative learning, which learns regions that should not be determined as the main subject. The imaging apparatus of Patent Document 1 performs additional learning if the number of images in the selected data set is equal to or greater than a threshold value.

Prior Art Literature

Patent Literature

[0003]

Patent Document 1

Summary of Invention

Problem to be Solved by the Invention

[0004] The present disclosure provides an imaging apparatus that can make it easier to recognize a region of a subject in an image that aligns with the user's intention.

Means for Solving the Problem

[0005] An imaging device in one aspect of the present disclosure comprises an imaging unit, a first recognition unit, a second recognition unit, and a control unit. The imaging unit captures an image of a subject and generates image data. The first recognition unit recognizes the region of a first subject in the image indicated by the image data based on the image data generated by the imaging unit. The second recognition unit recognizes the region of a second subject in the image that is included in or separate from the first subject, based on the image data. The control unit determines the recognized subject region in the image data based on the region of the first subject recognized by the first recognition unit and the region of the second subject recognized by the second recognition unit.

[0006] An imaging device in another aspect of the present disclosure comprises an imaging unit, an operation unit, and a control unit. The imaging unit captures an image of a subject and generates image data. The operation unit receives user input. The control unit controls image recognition processing for a first subject on the image data generated by the imaging unit in response to user input in the operation unit. The control unit receives user input in the operation unit to set a second subject that is included in or separate from the first subject, and controls image recognition processing based on the second subject set in the operation unit to determine the subject area of ​​the recognition result in the image data. [Effects of the Invention]

[0007] The imaging device described herein makes it easier to recognize areas of subjects in images that align with the user's intent. [Brief explanation of the drawing]

[0008] [Figure 1] A diagram showing the configuration of a digital camera according to Embodiment 1 of this disclosure. [Figure 2] A diagram illustrating the basic operation of a digital camera. [Figure 3] A flowchart illustrating the filtering settings process in a digital camera. [Figure 4] This figure shows an example of the display of the filtering settings screen in the digital camera of Embodiment 1. [Figure 5]A flowchart illustrating the subject recognition process in the digital camera of Embodiment 1. [Figure 6] A diagram illustrating the subject recognition process in the digital camera of Embodiment 1. [Figure 7] A flowchart illustrating subject filtering in subject recognition processing. [Figure 8] A flowchart illustrating the exclusion determination process in subject filtering. [Figure 9] Diagram to explain the exclusion determination process [Figure 10] A flowchart illustrating the aperture determination process in subject filtering. [Figure 11] Diagram to explain the filtering and judgment process. [Figure 12] A diagram illustrating the operation of the digital camera according to Embodiment 2. [Figure 13] Flowchart illustrating subject recognition processing in the digital camera of Embodiment 2 [Figure 14] A diagram illustrating subject filtering in a digital camera of Embodiment 2. [Modes for carrying out the invention]

[0009] The embodiments will be described in detail below, with reference to the drawings as appropriate. However, unnecessary details may be omitted. For example, detailed explanations of already well-known matters or redundant explanations of substantially identical configurations may be omitted. This is to avoid the following explanation becoming unnecessarily verbose and to facilitate understanding by those skilled in the art. The inventors provide the accompanying drawings and the following explanation so that those skilled in the art can fully understand this disclosure, and do not intend to limit the subject matter described in the claims by means of these.

[0010] (Embodiment 1) Embodiment 1 describes a digital camera that recognizes the area of ​​a subject in an captured image as an example of an imaging device according to the present disclosure.

[0011] 1. Configuration FIG. 1 is a diagram showing the configuration of a digital camera 100 according to the present embodiment. The digital camera 100 of the present embodiment includes an image sensor 115, an image processing engine 120, a display monitor 130, and a controller 135. Furthermore, the digital camera 100 includes a buffer memory 125, a card slot 140, a flash memory 145, an operation unit 150, and a communication module 155. The digital camera 100 also includes, for example, an optical system 110 and a lens driving unit 112.

[0012] The optical system 110 includes a focus lens, a zoom lens, an optical image stabilization lens (OIS), an aperture, a shutter, and the like. The focus lens is a lens for changing the focus state of a subject image formed on the image sensor 115. The zoom lens is a lens for changing the magnification of a subject image formed by the optical system. Each of the focus lens and the like is composed of one or more lenses.

[0013] The lens driving unit 112 drives the focus lens and the like in the optical system 110. The lens driving unit 112 includes a motor, and moves the focus lens along the optical axis of the optical system 110 under the control of the controller 135. The configuration for driving the focus lens in the lens driving unit 112 can be implemented by a DC motor, a stepping motor, a servo motor, an ultrasonic motor, or the like.

[0014] The image sensor 115 captures a subject image formed via the optical system 110 to generate captured data. The captured data constitutes image data representing an image captured by the image sensor 115. The image sensor 115 generates new frame image data at a predetermined frame rate (for example, 30 frames per second). The generation timing of captured data and the electronic shutter operation in the image sensor 115 are controlled by the controller 135. Various image sensors such as a CMOS image sensor, a CCD image sensor, or an NMOS image sensor can be used as the image sensor 115.

[0015] The image sensor 115 performs imaging operations such as still image capturing, through image capturing, and the like. A through image is mainly a moving image, and is displayed on the display monitor 130 for a user to determine a composition for, for example, still image capturing. The image sensor 115 is an example of the imaging unit in the present embodiment.

[0016] The image processing engine 120 performs various types of processing on captured data output from the image sensor 115 to generate image data, or performs various types of processing on image data to generate an image to be displayed on the display monitor 130. Examples of the various types of processing include, but are not limited to, white balance correction, gamma correction, YC conversion processing, electronic zoom processing, compression processing, decompression processing, and the like. The image processing engine 120 may be configured with a hard-wired electronic circuit, or may be configured with a microcomputer, a processor, or the like using a program.

[0017] The image processing engine 120 in the present embodiment includes a human recognition unit 122 and a plurality of specific object recognition units 124 that detect corresponding regions by using image recognition technology based on image data representing a captured image, with subjects such as people and objects in the captured image set as recognition targets. Each specific object recognition unit 124 performs detection processing through image recognition of such a subject region for a specific type of subject, respectively. In the digital camera 100 of the present embodiment, for example, in response to a user operation for selecting a type of subject via the operation unit 150, the specific object recognition unit 124 corresponding to the selected subject is used together with the human recognition unit 122 to detect the subject region.

[0018] Each recognition unit 122, 124 is configured as a pre-trained model using machine learning, such as a CNN or other type of DNN (Deep Neural Network). The person recognition unit 122 is trained to detect regions corresponding to people in an image. Each specific object recognition unit 124 is trained to detect regions corresponding to specific types of objects in an image. The person recognition unit 122 and the specific object recognition unit 124 output, for example, the position and size on the image indicating the detected region as the detection results for the person region and the specific object region, respectively.

[0019] The trained model for the person recognition unit 122 is generated by supervised learning based on training data that associates multiple images with regions corresponding to the entire body of a person in each image. The trained model for each specific object recognition unit 124 is generated, for example, based on training data that associates multiple images with regions corresponding to a specific type of object in each image, similar to the person recognition unit 122. The object to be recognized by the specific object recognition unit 124 is not limited to subjects other than those recognized by the person recognition unit 122, but may also be subjects included in the object to be recognized by the person recognition unit 122, such as a specific type of person.

[0020] The trained models of each recognition unit 122,124 are not limited to the DNN described above, but may be various machine learning models for image recognition. Furthermore, each recognition unit 122,124 is not limited to a machine learning model, but may detect the area of ​​the subject using rule-based image recognition techniques such as template matching. The person recognition unit 122 and each specific object recognition unit 124 are examples of the first recognition unit and the second recognition unit in this embodiment, respectively.

[0021] The display monitor 130 is an example of a display unit that displays various information. For example, the display monitor 130 displays an image (through image) represented by image data captured by the image sensor 115 and processed by the image processing engine 120. The display monitor 130 also displays a menu screen or the like for the user to make various settings for the digital camera 100. The display monitor 130 can be made of, for example, a liquid crystal display device or an organic EL device.

[0022] The operation unit 150 is a collective term for at least one operation interface, such as an operation button or operation dial, provided on the exterior of the digital camera 100, and accepts operations from the user. The operation unit 150 includes, for example, a shutter release button, a mode dial, a touch panel on the display monitor 130, a joystick, etc. When the operation unit 150 accepts an operation from the user, it transmits an operation signal corresponding to the user operation to the controller 135. The operation unit 150 may also be various GUIs such as virtual buttons, icons, cursors, a software keyboard, and / or objects displayed on the display monitor 130.

[0023] The controller 135 is a hardware controller that provides overall control over the operation of the digital camera 100. The controller 135 includes a CPU, and the CPU executes programs (software) to realize predetermined functions. Instead of a CPU, the controller 135 may include a processor consisting of dedicated electronic circuits designed to realize predetermined functions. In other words, the controller 135 can be realized with various processors such as a CPU, MPU, GPU, DSU, FPGA, and ASIC. The controller 135 may consist of one or more processors. Alternatively, the controller 135 may be configured on a single semiconductor chip together with the image processing engine 120, etc.

[0024] The buffer memory 125 is a recording medium that functions as work memory for the image processing engine 120 and the controller 135. The buffer memory 125 is implemented using DRAM (Dynamic Random Access Memory) or the like. The flash memory 145 is a non-volatile recording medium.

[0025] Although not shown in the diagram, the controller 135 may have various internal memories, such as a built-in ROM. The ROM stores various programs that the controller 135 executes. The controller 135 may also have a built-in RAM that functions as a workspace for the CPU.

[0026] The card slot 140 is a means for inserting a removable memory card 142. The card slot 140 can electrically and mechanically connect to the memory card 142. The memory card 142 is an external memory equipped with recording elements such as flash memory. The memory card 142 can store data such as image data generated by the image processing engine 120.

[0027] The communication module 155 is a communication module (circuit) that performs communication compliant with the communication standard IEEE 802.11 or the Wi-Fi standard, etc. The digital camera 100 can communicate with other devices via the communication module 155. The digital camera 100 may communicate directly with other devices via the communication module 155, or it may communicate via an access point. The communication module 155 may be connectable to a communication network such as the Internet. The communication module 155 is an example of a communication unit in this embodiment.

[0028] 2. Operation The operation of the digital camera 100, configured as described above, will be explained below.

[0029] In this embodiment, the digital camera 100 determines, for example, the area to be used for autofocus (AF) control based on the detection results from two types of recognition units 122 and 124 in the process of recognizing the area of ​​a subject in the captured image. As the subject area of ​​the recognition result from this subject recognition process, the digital camera 100 of this embodiment extracts a portion of the human area from the human area detection result by

[0030] Figure 2 shows an example of filtering settings in the digital camera 100 of this embodiment. In the digital camera 100 of this embodiment, filtering settings can be configured according to user operation on a menu screen displayed on the display monitor 130, as shown in Figure 2. In the example in Figure 2, the filtering function in the digital camera 100 is turned on, i.e., enabled, and filtering conditions corresponding to the "Filtering Mode" and "Filter Target" menu items are set. Filtering conditions include, for example, the method for extracting the human area and the recognition target of the specific object recognition unit 124 used for filtering.

[0031] In this embodiment, the digital camera 100 determines the human area of ​​the recognition result from the subject recognition process according to the filtering mode, which is set as the filtering operation mode, for example. In the example in Figure 2, an exclusion mode is set to exclude the human area from the recognition result for a specific subject. In exclusion mode, the digital camera 100 excludes the human area from the recognition result for subjects of a type set as the filter target, for example.

[0032] In the example shown in Figure 2, assuming that the user wants to focus on the players using AF control but not on the referee in a shooting scene such as a sports match, the filter target is set to the referee. In this case, a specific object recognition unit 124 is used to detect the area in the image corresponding to the referee. After this setting, the digital camera 100 can determine the person area of ​​the recognition result by filtering the detection result from the person recognition unit 122, for example, by excluding the person area corresponding to the area of ​​the referee detected by the specific object recognition unit 124 in the captured image. The operation of the digital camera 100 in this embodiment will be described in detail below.

[0033] 2-1. Filtering settings process The digital camera 100 of this embodiment sets the filtering settings for subject recognition processing in response to user operations on a menu screen as shown in Figure 2. The filtering setting process in this digital camera 100 will be explained using Figures 3 and 4.

[0034] Figure 3 is a flowchart illustrating the filtering setting process in the digital camera 100. The process shown in Figure 3 is initiated, for example, in response to user operations on a predetermined menu screen for configuring the digital camera 100.

[0035] First, the controller 135 of the digital camera 100 displays the filtering settings screen on the display monitor 130 as a menu screen for making filtering-related settings (S41).

[0036] Figure 4 shows an example of the display of the filtering settings screen in the digital camera 100 of this embodiment. Figures 4(A) and 4(B) illustrate the transitions of the filtering screen until the filtering conditions shown in the setting example of Figure 2 are set. The filtering settings screen in this example includes "ON / OFF setting", "filtering mode", and "filter target" as menu items that the user can select, as well as a back button 54. "ON / OFF setting" displays options to switch between enabling and disabling filtering in the subject recognition process, for example, according to the user's selection operation. The back button 54 accepts a user operation to return the screen transition on the display monitor 130 of the digital camera 100 by one screen.

[0037] The controller 135 determines whether or not a user operation on the filtering settings screen has been input via the operation unit 150 (S42). If no user operation has been input (NO in S42), the controller 135 repeats the process from step S41 onwards.

[0038] If a user operation is input (YES in S42), the controller 135 determines whether or not the filtering conditions have been determined by that user operation (S43). For example, when the user inputs the back button 54 on the filtering settings screen, the controller 135 determines that the filtering conditions have been determined (YES in S43).

[0039] If the filtering conditions have not been determined (NO in S43), the controller 135 repeats the processing from step S41 onwards. Figure 4(A) shows an example of the display when a user operation is input on the filtering settings screen (YES in S42) and the "Filtering Mode" menu item is selected. In response to the user operation, the controller 135 displays the filtering mode options for the digital camera 100 on the filtering settings screen (NO in S43, S41). In the digital camera 100 of this embodiment, for example, in addition to the exclusion mode, a narrowing mode can be used as a filtering mode.

[0040] Figure 4(B) shows an example where the exclusion mode is selected in the example of Figure 4(A), and a user operation is input in step S42 to select a menu item to be filtered. In response to this user operation, the controller 135 displays the selection of subject types that can be selected as filters in the exclusion mode on the filtering settings screen, for example as shown in Figure 4(B) (NO in S43, S41). The selection of filters is displayed, for example, in the digital camera 100, corresponding to the recognition targets of the specific object recognition unit 124 that are available in each filtering mode.

[0041] Figure 4(C) shows an example where the filtering mode is changed to the refinement mode in the example of Figure 4(A). The refinement mode is a filtering mode that narrows down the recognition results from the subject recognition process to a specific person area based on the detection results from each recognition unit 122, 124. Figure 4(D) shows an example from the example of Figure 4(C) in which the selection of the filter target in the refinement mode is displayed on the filtering settings screen in response to the user operation in step S42. For example, in the refinement mode, the filter target is set to a specific object area corresponding to the subject used for refinement.

[0042] When the filtering conditions are determined in response to user operation on the filtering settings screen (YES in S43), the controller 135 sets the filtering conditions for the subject recognition process (S44). For example, the controller 135 generates setting information including whether filtering is enabled or disabled, as well as the determined filtering conditions, and stores it in the buffer memory 125. After that, the controller 135 terminates the processing of this flowchart.

[0043] According to the filtering setting process described above, filtering conditions in the subject recognition process are set according to user operations (S41-S44). As a result, the recognition results can be filtered in the subject recognition process based on the set filtering conditions.

[0044] The filtering setting process and filtering setting screen are not limited to the examples above. For example, in response to user operation in step S42, step S41 may transition from the filtering setting screen to another screen displaying options for filtering conditions. Such a setting screen may display text or other information, such as explanations for each option. Furthermore, the decision in step S43 is not limited to the examples above. For example, the filtering setting screen may be provided with a completion button to accept user operation to complete the setting of filtering conditions, and the determination of the filtering conditions may be made in response to the user operation of pressing the completion button.

[0045] 2-2. Subject Recognition Processing In the digital camera 100 of this embodiment, the subject recognition process using the filtering conditions set as described above will be explained with reference to Figures 5 and 6.

[0046] Figure 5 is a flowchart illustrating the subject recognition process in the digital camera 100 of this embodiment. The process shown in Figure 5 starts with the setting information, such as filtering conditions set by the filtering setting process, being held in the buffer memory 125. The process in Figure 5 is performed, for example, at predetermined intervals, such as the frame period, while the image sensor 115 is performing the image acquisition operation of various moving images, such as through images.

[0047] First, the controller 135 of the digital camera 100 acquires the captured image generated by the imaging operation from the image sensor 115 (S1). An example of such a captured image Im is shown in Figure 6(A). Figure 6 is a diagram for explaining the subject recognition process in the digital camera 100 of this embodiment. Below, we will explain an example in which the filtering conditions shown in Figure 2 are set, assuming a shooting scene during a sports match as shown in the captured image Im in Figure 6(A).

[0048] The controller 135 instructs the person recognition unit 122 to perform detection processing on the acquired image Im and obtains detection information indicating the detection result of the person region (S2). Figure 6(B) illustrates the person regions R11 to R13 detected in the image Im of Figure 6(A). In this example, person regions R11 and R12 corresponding to two players and person region R13 corresponding to one referee have been detected.

[0049] The controller 135 determines whether or not a human region has been detected in the captured image Im based on the detection information obtained from the person recognition unit 122 (S3). If no human region has been detected (NO in S3), the controller 135 terminates the processing of this flowchart and, for example, while the imaging operation is in progress, repeats the processing from step S1 onwards for the next frame.

[0050] If a person area is detected (YES in S3), the controller 135 determines whether filtering is enabled, i.e., turned on, by referring to the configuration information held in the buffer memory 125, for example (S4).

[0051] If filtering is enabled (YES in S4), the controller 135 causes the specific object recognition unit 124, which corresponds to the object to be filtered in the filtering conditions of the setting information, to perform detection processing on the captured image Im acquired in step S1 (S5). By performing this detection processing, the controller 135 in this embodiment acquires detection information of the specific object region detected by the specific object recognition unit 124 in the captured image Im (S5). In this example, the specific object recognition unit 124 is used to detect the corresponding region with the referee, which is set as the object to be filtered in the filtering conditions of Figure 2, as the specific subject. Figure 6(C) illustrates the specific object region R23 detected in the captured image Im of Figure 6(A).

[0052] The controller 135 determines whether or not a specific object region has been detected in the captured image Im based on the detection information obtained from the specific object recognition unit 124 (S6).

[0053] If a specific object region is detected (YES in S6), the controller 135 determines the person region as the recognition result of the subject recognition process in the captured image Im based on the detection information of the person region and the specific object region acquired in steps S2 and S5, respectively (S7). In step S7, the controller 135 determines the recognition result by filtering multiple detected person regions according to the detection result of the specific object region. Figure 6(D) shows an example in which person regions R11 and R12 are determined as the recognition result in the captured image Im from the person regions R11 to R13 and specific object region R23 shown in Figures 6(B) and (C) through this subject filtering process (S7).

[0054] In the example in Figure 6(D), the person region R13 corresponding to the specific object region R23 is excluded from the person regions R11 to R13. In step S7, the controller 135 determines, for example, based on detection information from each recognition unit 122, 124, whether each person region R11 to R13 corresponds to the specific object region R23, and decides to exclude the person region R13. The excluded area is determined, for example, in the captured image Im, as the person region corresponding to the subject to be filtered in exclusion mode. After executing this subject filtering process (S7), the controller 135 terminates the processing of this flowchart. Details of the subject filtering process (S7) will be described later.

[0055] On the other hand, if no specific object region is detected (NO in S6), the controller 135 does not perform subject filtering (S7) and terminates the processing of this flowchart. Furthermore, if filtering is off (NO in S4), the controller 135 does not perform filtering-related processing (S5-S7), such as referencing the specific object region, and terminates the processing of this flowchart. In these cases, the controller 135 may determine the recognition result based on the detection information of the person region before filtering. For example, the detection information of the person region acquired in step S2 may be used as the recognition result of the subject recognition processing for the frame in which the captured image Im was acquired in step S1.

[0056] According to the subject recognition process described above, subject filtering is performed based on the detection information from the person recognition unit 122 and the specific object recognition unit 124, according to the filtering settings (S4) (S2-S3, S5-S6, S7). For example, from the person regions R11-R13 detected in the captured image Im shown in Figure 6, some person regions R11 and R12 are extracted by referencing the specific object region R23 according to the set filtering conditions, and these are determined as the recognition result (S7). In this way, for example, in addition to the person recognition unit 122, the specific object recognition unit 124 for a specific type of person set as the filtering target can be used to determine the recognition result, making it easier to obtain subject recognition results that match the user's intentions.

[0057] In the digital camera 100, the subject recognition process is not limited to the above example; for example, it may be performed only when filtering is turned on. In this case, it is not necessary to determine whether filtering is enabled or disabled in step S4.

[0058] 2-2-1. Subject Filtering The details of the subject filtering process in step S7 of Figure 5 will be explained using Figure 7.

[0059] Figure 7 is a flowchart illustrating the subject filtering process (S7) in the subject recognition process (Figure 5). The process in this flowchart starts with the detection information obtained by the person recognition unit 122 and the specific object recognition unit 124, respectively, acquired in steps S2 and S5 of Figure 5, in addition to setting information such as filtering conditions, stored in the buffer memory 125.

[0060] The controller 135 of the digital camera 100 determines, for example, whether the filtering mode is set to exclusion mode by referring to the setting information (S11).

[0061] If the filtering mode is set to exclusion mode (YES in S11), the controller 135 determines whether each person region corresponds to a subject to be filtered in exclusion mode based on the detection information from each recognition unit 122, 124 (S12). For example, for each detected person region, the controller 135 determines the object to be excluded by comparing its position and size on the captured image Im with the detected specific object region, according to the filter target setting in exclusion mode. Details of this exclusion determination process (S12) will be described later.

[0062] Next, the controller 135 determines whether or not an exclusion target has been determined in the exclusion determination process (S12) (S13).

[0063] If the exclusion targets are determined (YES in S13), the controller 135 removes the person regions determined to be excluded from the pre-filtering recognition result, which includes all person regions detected in the captured image Im, and determines the post-filtering recognition result (S14). After that, the controller 135 terminates the processing in this flowchart and repeats the processing from step S1 onwards in Figure 5 for, for example, the captured image Im of the next frame.

[0064] If the exclusion target has not been determined (NO in S13), the controller 135 will not perform the processing in step S14 and will terminate the processing of this flowchart.

[0065] If the filtering mode is set to the narrowing mode instead of the exclusion mode (NO in S13), the controller 135 determines whether each person region is a target for narrowing based on the detection information from each recognition unit 122, 124 (S15). The targets for narrowing are determined, for example, as person regions in the captured image Im that correspond to subjects that are filtered in the narrowing mode. In step S15, the controller 135 determines the targets for narrowing by, for example, comparing the positions on the image between each detected person region and a specific object region. This narrowing determination process (S15) will be described later.

[0066] The controller 135 determines whether or not the target for filtering has been determined in the filtering determination process (S15) (S16).

[0067] If the target for narrowing down is determined (YES in S16), the controller 135 restricts the person region to be included in the recognition result from the detection information of the person recognition unit 122 to the person region targeted for narrowing down (S17). After determining the recognition result filtered by this narrowing down (S17), the controller 135 terminates the processing of this flowchart.

[0068] If the target for filtering has not been determined (NO in S16), the controller 135 will not perform the processing in step S17 and will terminate the processing of this flowchart. If the target for exclusion or filtering has not been determined (NO in S13 or NO in S16) and the processing in steps S14 and S17 is not performed, for example, detection information of all human areas in the captured image Im may be used as the recognition result.

[0069] According to the subject filtering process (S7) described above, an exclusion determination process (S12) or a narrowing determination process (S15) is performed according to the set filtering mode (S11), and the recognition result by the subject recognition process is determined (S14, S17). For example, from all the person regions detected by the person recognition unit 122, the recognition result is determined by excluding the excluded targets according to the result of the exclusion determination process (S12~S14). Also, according to the result of the narrowing determination process, the recognition result is determined that is limited to the narrowed targets (S15~S17). In this way, by determining the recognition result according to the setting of the filter targets in each filtering mode, it becomes easier to obtain a subject recognition result that matches the user's intention.

[0070] (1) Exclusion determination process The details of the exclusion determination process in step S12 of Figure 7 will be explained using Figures 8 and 9.

[0071] Figure 8 is a flowchart illustrating the exclusion determination process (S12) in the subject filtering process (S7 in Figures 7 and 5). Figure 9 is a diagram illustrating the exclusion determination process (S12). Below, we will explain an example in which the exclusion determination process (S12) is performed based on detection information of human areas and specific object areas in the captured image Im, similar to Figure 6(A).

[0072] The controller 135 selects one person region from the detection information by the person recognition unit 122, for example, stored in the buffer memory 125 (S21). Figure 9(A) is an explanatory diagram showing the state in which person region R11 is selected from person regions R11 to R13, similar to Figure 6(B).

[0073] Furthermore, the controller 135 selects one specific object region from the detection information by the specific object recognition unit 124 (S22). In this example, the same specific object region R23 as in Figure 6(C) is selected.

[0074] The controller 135 determines whether the centers of each region in the captured image Im are located within a predetermined distance between the person region and the specific object region selected in steps S21 and S22 (S23). For example, in the state shown in Figure 9(A), it is determined whether the distance d1 between the center positions of each region R11 and R23 is within a predetermined distance. The predetermined distance is defined, for example, by the ratio to the size of the person region in the captured image Im, and is set to a value small enough that the person region and the specific object region can be considered to correspond to the same subject (for example, 3% of the diagonal length of the person region).

[0075] If the distance between the center positions of the two selected regions is within a predetermined distance (YES in S23), the controller 135 determines whether the difference in size between the two regions in the captured image Im is within a predetermined value (S24). The predetermined value of the size difference is defined, for example, by the ratio of the size of one region, and is set to a value small enough that the width and height of each rectangular region in the captured image Im can be considered to correspond to the same subject (for example, 5% of the width or height of the person region).

[0076] On the other hand, if the distance between the center positions in the two selected regions is not within a predetermined distance (NO in S23), the controller 135 proceeds to step S26. In step S26, the controller 135 determines whether the detection information by the specific object recognition unit 124 includes other specific object regions not selected in step S22. If the difference in size between the two selected regions is not within a predetermined value (NO in S24), the controller 135 also proceeds to the determination in step S26.

[0077] If other specific object regions are included in the detection information (YES in S26), the controller 135 selects the next specific object region in the detection information (S22) and repeats the subsequent processing. Whether or not each specific object region has been selected may be managed by a predetermined flag, for example, and this flag may be reset when proceeding to the next step S27.

[0078] If the detection information by the specific object recognition unit 124 does not include any other specific object regions (NO in S26), the controller 135 determines whether the detection information by the person recognition unit 122 includes any other person regions not selected in step S21 (S27). For example, in the state shown in Figure 9(A), if the distance d1 is not within the predetermined distance (NO in S23) and the detection information does not include any other specific object regions (NO in S26), the determination in step S27 is made.

[0079] If other person regions are included in the detection information (YES in S27), the controller 135 selects the next person region in the detection information (S21) and repeats the subsequent processing. Figure 9(B) illustrates the state in which the next person region R12 is selected from the state in Figure 9(A), and the specific object region R23 is selected in step S22. Figure 9(C) illustrates the state in which, from the state in Figure 9(B), for example, the next person region R13 is selected in place of person region R12, depending on whether the distance d2 between regions R12 and R23 is not within a predetermined distance (NO in S23).

[0080] If the distance between the selected person region and the specific object region is within a predetermined distance (YES in S23), and the difference in size between the two regions is within a predetermined value (YES in S24), the controller 135 decides to exclude the person region (S25). For example, in the state shown in Figure 9(C), the person region R13 is excluded because the distance d3 between the person region R13 and the specific object region R23 is within a predetermined distance (YES in S23), and the difference in size is within a predetermined value (YES in S24).

[0081] After determining the exclusion targets (S25), the controller 135 repeats the subsequent processing according to the decision in step S26. Comparisons are performed on all person regions and specific object regions in the detection information, such as in step S23. If no other person regions are included (NO in S27), the controller 135 terminates the processing in this flowchart. After that, the controller 135 returns to the next processing in the subject filtering process (S7) (S13 in Figure 7).

[0082] According to the exclusion determination process (S12) described above, based on a comparison of the position and size between each person region and the specific object region detected in the captured image Im (S23, S24), the person region to be excluded is determined (S25). For example, as shown in Figures 6 and 9, from the detection information of each person region R11 to R13, the person region R13, which has a position and size that can be considered to correspond to the same subject as the specific object region R23, is determined to be excluded. In this way, the exclusion determination process (S12) allows for the determination of the exclusion target with simple calculations, thereby reducing the processing load on, for example, the digital camera 100.

[0083] Furthermore, without having to construct a new dataset for the machine learning model or perform additional training, as in Patent Document 1, for example, determining the recognition result after excluding the excluded items (S14 in Figure 7) makes it easier to obtain the recognition result intended by the user.

[0084] (2) Filtering and judgment process Next, the details of the filtering decision process in step S15 of Figure 7 will be explained using Figures 10 and 11.

[0085] Figure 10 is a flowchart illustrating the narrowing-down determination process (S15) in the subject filtering process (S7, Figure 7). Figure 11 is a diagram illustrating the narrowing-down determination process (S15). In this example, we will explain the case where, through the filtering setting process (Figure 3), the filtering mode is set to the narrowing-down mode in the filtering setting screen illustrated in Figures 4(C) and (D), and the filter target is set to "bicycle". In the subject recognition process of this example, in addition to the person recognition unit 122, a specific object recognition unit 124 is used that recognizes bicycles as subjects related to people.

[0086] Figure 11(A) illustrates an image Im captured in a different scene from Figure 6(A). Figure 11(B) illustrates the person regions R14~R17 detected by the person recognition unit 122 and the specific object regions R25, R27 detected by the specific object recognition unit 124 in the image Im captured in Figure 11(A). In this example, the person regions R14, R16 of spectators and the person regions R15, R17 of a cyclist are detected. Figure 11(C) shows an example in which the recognition result is determined by the narrowing-down determination process (S15) in the image Im captured in Figure 11(A).

[0087] In the example shown in Figure 11, based on the assumption that there is a one-to-one correspondence between the athlete and the bicycle, the filtering conditions are set so that the person regions R15 and R17 are selected from the person regions R14 to R17, and the recognition results are limited to these selected regions.

[0088] The following describes an example in which the refinement judgment process (S15) shown in Figure 10 is performed on the captured image Im shown in Figure 11(A), with the detection information of the human region R14~R17 and the detection information of the specific object region R25,R27 shown in Figure 11(B) acquired (S2,S5 in Figure 5).

[0089] The controller 135 selects, for example, one specific object region from the detection information by the specific object recognition unit 124 (S31), and further selects one person region from the detection information by the person recognition unit 122 (S32).

[0090] The controller 135 determines whether the centers of the regions in the captured image Im are located within a predetermined distance between the person region and the specific object region selected in steps S31 and S32 (S33). The predetermined distance for the narrowing-down determination process may be set to a larger value than the predetermined distance in step S23 of the exclusion determination process (for example, 50% of the diagonal length of the person region), from the viewpoint of identifying person regions located near the specific object region in the captured image Im. In step S33, the controller 135 may calculate the distance between the centers of the two regions based on the detection information and store the calculation result in association with the two regions.

[0091] If the distance between the center positions of the two selected regions is within a predetermined distance (YES in S33), the controller 135 stores the detection information of the selected person region in a predetermined buffer region, for example, as a candidate for determining the target to narrow down (S34). Such a buffer region for narrowing down candidates may be provided in a buffer memory 125, for example.

[0092] After buffering the narrowing candidates (S34), the controller 135 determines whether the detected information includes other person regions, similar to step S27 of the exclusion determination process (S35). If the distance between the two selected regions is not within a predetermined distance (NO in S33), the controller 135 proceeds to the determination in step S35 without performing the processing in step S34.

[0093] If the detection information includes other person regions (YES in S35), the controller 135 selects the next person region (S32) and repeats the subsequent processing.

[0094] If the detection information does not include other human regions (NO in S35), the controller 135 determines whether or not there are candidates for narrowing down the candidates stored in the buffer area (S36). For example, if the specific object region R25 shown in Figure 11(B) is selected (S31), the distance to the specific object region R25 is compared for each human region R14 to R17, and candidates for narrowing down the candidates are buffered according to the comparison results (S32 to S35).

[0095] If a candidate for narrowing down the area is stored in the buffer area (YES in S36), the controller 135 determines the candidate that is closest to the specific object area being selected in the captured image Im from among the stored candidates to be narrowed down (S37). This closest candidate can be determined, for example, based on the distance between the two areas calculated in step S33. For example, for the specific object area R25 in Figure 11(B), if the human areas R14 and R15 are candidates for narrowing down the area, the human area R15, which is closer to the specific object area R25, is determined to be the target for narrowing down the area.

[0096] The controller 135 determines the target for narrowing down from the candidate person regions for the selected specific object region (S37), and then determines whether the detection information includes other specific object regions (S38). If no narrowing down candidates are stored in the buffer region (YES in S36), the controller 135 skips the processing in step S37 and proceeds to the determination in step S38.

[0097] If the detection information includes other specific object regions (YES in S38), the controller 135 selects the next specific object region (S31) and repeats the subsequent processing.

[0098] If the detection information does not include any other specific object regions (NO in S38), the controller 135 terminates the processing in this flowchart and returns to the next step in the subject filtering process (S7) (S16 in Figure 7).

[0099] According to the filtering determination process (S15) described above, the target for filtering is determined from each person region based on a comparison of the positions between each person region and the specific object region detected in the captured image Im (S31-S38). For example, in the captured image Im shown in Figure 11, from the detection information of person regions R14-R17, the person regions R15 and R17 that are closest to the specific object regions R25 and R27, respectively, are determined to be the target for filtering. In this way, the target for filtering can be determined with a simple calculation process, and this also allows the recognition results of the subject recognition process to be narrowed down to the target person region according to the setting of the filtering conditions.

[0100] The above describes an example in which each selected object region is compared with each person region (S31-S33, S35, S38). The execution order of each process is not limited to the above example. For example, each time a person region is selected, the candidate for narrowing down the region may be buffered by comparing it with each specific object region, or after comparing all person regions and specific object regions in the detection information, the person region closest to each specific object region may be determined as the target for narrowing down the region. Alternatively, for example, for each specific object region, two or more person regions may be determined as targets for narrowing down the region from the candidate for narrowing down the region, depending on the distance to that specific object region in the captured image.

[0101] 3. Summary As described above, the digital camera 100 of this embodiment includes an image sensor 115, which is an example of an imaging unit; a person recognition unit 122, which is an example of a first recognition unit; a specific object recognition unit 124, which is an example of a second recognition unit; and a controller 135, which is an example of a control unit. The image sensor 115 captures an image of a subject and generates image data. The person recognition unit 122 detects a person region as an example of recognizing the region of a first subject in the captured image, which is an example of an image indicated by the image data, based on the image data generated by the image sensor 115. The specific object recognition unit 124 detects a specific object region as an example of recognizing the region of a second subject that is included in the first subject or is separate from the first subject, based on the image data. The controller 135 determines the person region as an example of the subject region of the recognition result in the image data (S7), based on the person region detected by the person recognition unit 122 and the specific object region detected by the specific object recognition unit 124 (see S2, S5).

[0102] According to the digital camera 100 described above, the person recognition unit 122 and the specific object recognition unit 124 for objects included in or other recognized objects of the person recognition unit 122 are used to determine the person region as the subject region of the recognition result (S7). This allows the recognition result to be determined according to the relationship between the objects recognized by the two recognition units 122 and 124, making it easier to recognize the subject region in the image in accordance with the user's intentions. For example, using the subject region of the recognition result, AF control and the like can be performed with high accuracy for the subject desired by the user.

[0103] In this embodiment, the controller 135 refers to the specific object region detected by the specific object recognition unit 124 in the captured image, extracts a portion of the region from the person region detected by the person recognition unit 122, and determines the extracted region to be the person region of the recognition result (S7, Figure 7). For example, as shown in Figures 6 and 11, the person region corresponding to the specific object region can be excluded from one or more person regions (S14), or the region can be narrowed down to the person region related to the specific object region (S17), and the extracted specific person region can be determined to be the recognition result.

[0104] In this embodiment, for example, the filtering target for an example of a second subject in exclusion mode is a specific type of subject included in the person in the example of a first subject. The controller 135, for example, in the captured image Im shown in Figure 6, excludes region R13 which has a predetermined correspondence with the specific object region R23 detected by the specific object recognition unit 124 from the person regions R11 to R13 detected by the person recognition unit 122, and determines the excluded regions R11 and R12 as the person regions of the recognition result (S14). In this embodiment, for the person region and the specific object region, if the center positions of each region in the captured image Im are within a predetermined distance (YES in S23) and the size difference is within a predetermined value (YES in S24), it is determined that the two regions have a predetermined correspondence. According to this filtering process in exclusion mode (S7, S12 to S14), for example, using the detection results of the two recognition units 122 and 124, a specific subject can be excluded from the recognition result by simple calculation processing.

[0105] In this embodiment, for example, the filtering target for an example of a second subject in the narrowing mode is another subject related to the person in the example of a first subject. The controller 135 narrows down the person regions R14 to R17 detected by the person recognition unit 122 in the captured image Im, for example as shown in Figure 11, to a portion of the region where the specific object regions R25 and R27 detected by the specific object recognition unit 124 are located within a predetermined range, and determines the narrowed-down regions R15 and R17 as the recognized person regions R15 and R17 (S7, S15 to S17). In this embodiment, if the center positions of each region are within a predetermined distance (YES in S33), it is determined that the two regions are located within a predetermined range. Even with this narrowing mode filtering process (S7, S15 to S17), for example, the recognition results for some subjects can be narrowed down by simple calculation processing using the detection results from the two recognition units 122 and 124.

[0106] In this embodiment, the digital camera 100 further includes an operation unit 150 that receives user input. The controller 135 receives user input from the operation unit 150 to set a filter target for an example of a second subject (see S42, Figures 4(A) and (B)), and causes the specific object recognition unit 124 to detect the specific object region of the set filter target (see S5). As a result, in response to user input, for example, the recognition target of the specific object recognition unit 124 is set as the filter target, and filtering processing using the specific object recognition unit 124 (S7) makes it easier to recognize the subject region that aligns with the user's intentions.

[0107] In this embodiment, when no filter target is set, such as when filtering is off (NO in S4), the controller 135 determines the person region of the recognition result based on the person region detected by the person recognition unit 122 without referring to the specific object region. When a filter target is set, such as when filtering is on (YES in S4), the controller 135 determines the person region of the recognition result based on both the person region detected by the person recognition unit 122 and the specific object region recognized by the specific object recognition unit 124 (S5-S7). In this way, the person region determined in the recognition result changes depending on the filtering setting, for example, and the digital camera 100 can utilize recognition results that correspond to settings made by the user.

[0108] In this embodiment, the digital camera 100 includes an image sensor 115, which is an example of an imaging unit that captures an image of a subject and generates image data; an operation unit 150 that receives user operations; and a controller 135, which is an example of a control unit. The controller 135 controls subject recognition processing for a person as an example of image recognition processing for a first subject on the image data generated by the image sensor 115, in response to user operations on the operation unit 150 (S42-S44) (S1-S7). The controller 135 receives a user operation on the operation unit 150 to set a filter target for an example of a second subject that is included in the first subject or is different from the first subject (S42), and controls the subject recognition processing based on the filter target set on the operation unit 150 to determine the person region, which is an example of the subject region of the recognition result in the image data (S7).

[0109] According to the digital camera 100 described above, the subject recognition process can be controlled according to the filter target settings made by the user, making it easier to recognize the area of ​​the subject in accordance with the user's intentions.

[0110] (Embodiment 2) Embodiment 2 of this disclosure will be described below with reference to Figures 12 to 14. Embodiment 1 described a digital camera 100 that determines the recognition result in an captured image using a person recognition unit 122 and one specific object recognition unit 124 in the subject recognition process. Embodiment 2 describes a digital camera 100 that determines the recognition result by using multiple specific object recognition units 124 in combination.

[0111] Hereinafter, descriptions of the configuration and operation similar to that of the digital camera 100 according to Embodiment 1 will be omitted as appropriate, and the digital camera according to this embodiment will be described.

[0112] 1. Filtering settings Figure 12 is a diagram illustrating the operation of the digital camera 100 according to Embodiment 2. In this embodiment, the digital camera 100 sets filtering conditions for each of the multiple specific object recognition units 124 used in combination. In the filtering setting process of this embodiment, the controller 135 of the digital camera 100 first displays a screen on the display monitor 130 as a filtering setting screen, for example, as shown in Figure 12(A) (S41 in Figure 3).

[0113] In the filtering settings screen illustrated in Figure 12(A), options such as "Filtering 1" and "Filtering 2" are displayed for setting filtering conditions by multiple specific object recognition units 124. In the example in Figure 12(A), the enable or disable of filtering based on the settings of each option is displayed as "ON" / "OFF". The controller 135 displays a screen for setting the filtering conditions for a given option in response to a user operation, for example, selecting one option (S41~S43). Figure 12(B) illustrates the filtering settings screen that transitions from the screen in Figure 12(A) in response to the selection of "Filtering 2".

[0114] In the digital camera 100 of this embodiment, each filtering condition can be set in a filtering setting screen similar to that in Figure 4, for example, as shown in Figure 12(B). The controller 135 determines that the filtering conditions have been determined in response to user operations on the screen shown in Figure 12(A) (YES in S43), and sets the filtering conditions for the subject recognition process (S44). In this way, in this embodiment, filtering conditions using multiple specific object recognition units 124 can be set.

[0115] 2. Subject Recognition Processing Figure 13 is a flowchart illustrating the subject recognition process in the digital camera 100 of Embodiment 2. Below, we will describe an example in which filtering conditions are set by two specific object recognition units 124 in a filtering setting screen, such as the one shown in Figure 12, and these conditions are stored in the setting information as filtering conditions for the first filtering and the second filtering.

[0116] In this embodiment, the digital camera 100 performs the processing shown in the flowchart of Figure 13 instead of the flowchart of Figure 5 during the imaging operation, similar to Embodiment 1.

[0117] The controller 135, for example, acquires an image Im (S1), obtains detection information for a person's region in the image Im (S2), and then determines whether or not a person's region has been detected in the image Im based on the detection information (S3).

[0118] In the digital camera 100 of this embodiment, if a person area is detected (YES in S3), the controller 135 determines whether the first filtering setting is ON or OFF (S4A), instead of, for example, step S4 relating to one filtering setting in Embodiment 1.

[0119] If the first filtering setting is turned on (YES in S4A), the controller 135, instead of step S5 in Embodiment 1, for example, causes the specific object recognition unit 124 corresponding to the filtering target of the first filtering to perform detection processing on the captured image Im (S5A). In step S5A, the controller 135 acquires the detection information of the first specific object region output from the specific object recognition unit 124 of the first filtering.

[0120] Based on the detection information acquired in step S5A, the controller 135 determines whether or not a first specific object region has been detected in the captured image Im (S6A). If a first specific object region is detected (YES in S6A), the controller 135 performs subject filtering in the same manner as in Embodiment 1, based on, for example, the detection information of the person region acquired in step S1 and the first specific object region acquired in step S5A (S7).

[0121] In this embodiment, the controller 135 performs subject filtering (S7) based on the detection information of the person region and the first specific object region, and then determines whether the setting for the second filtering is turned on or off (S4B).

[0122] If the second filtering setting is turned on (YES in S4B), the controller 135, for example, instead of the first filtering in step S5A, causes the specific object recognition unit 124 corresponding to the filtering target of the second filtering to perform detection processing on the captured image Im (S5B). The controller 135 obtains detection information of the second specific object region from the specific object recognition unit 124 of the second filtering (S5B).

[0123] Based on the detection information acquired in step S5B, the controller 135 determines whether or not a second specific object region has been detected in the captured image Im (S6B).

[0124] If a second specific object region is detected (YES in S6B), the controller 135 of this embodiment performs further subject filtering (S7A). In step S7B, the controller 135 performs subject filtering based on the processing result of step S7 and the detection information of the second specific object region, respectively, instead of the detection information of the person region and the detection information of the first specific object region in step S7.

[0125] Figure 14 is a diagram illustrating the subject filtering process (S7A) in the digital camera of Embodiment 2. Figure 14(A) shows the result of subject filtering in step S7 of the captured image Im, similar to the example in Figure 6, in which the person region R11 corresponding to the first specific object region R23 shown in Figure 6(C) is excluded from the person regions R11 to R13 in Figure 6(B). Below, we will describe an example in which the filtering mode is set to the narrowing mode and the filter target is set to the jersey number "10" as shown in Figure 4(D). In the second filtering of this example, the specific object recognition unit 124 detects the corresponding region with the jersey number "10" set as the filter target as a specific subject.

[0126] Figure 14(B) illustrates a second specific object region R32 detected in the captured image Im of this example. For example, in step S5B, the controller 135 acquires detection information for this second specific object region R32. In the subject filtering process (S7A) of this example, the controller 135 determines the target for filtering from the human regions R12 and R13 in Figure 14(A) by the filtering determination process (S15) of Figure 10 (S37). In this example, the human region R12, which is closest to the specific object region R32 on the captured image Im, is determined to be the target for filtering. Figure 14(C) illustrates the recognition result determined by this subject filtering process (S7A).

[0127] As described above, in the subject recognition process, the digital camera 100 of this embodiment determines the recognition result in the captured image Im by a two-stage subject filtering process (S7, S7A) using two specific object recognition units 124. By combining filtering conditions using multiple specific object recognition units 124 in this way, it is possible to make it easier to recognize the area of ​​the subject that matches the user's intention. The combination of filtering conditions is not limited to the example in Figure 14. For example, in a scene of shooting a ball game, two specific object recognition units 124 may be used, one for the referee and the other for the ball, and in the filtering process, after excluding the area of ​​the person corresponding to the referee, the area may be narrowed down to the area of ​​the person that appears near the ball.

[0128] Furthermore, the combination of filtering conditions is not limited to exclusion and narrowing, but may also be a two-stage exclusion or a two-stage narrowing. In addition, as a filtering mode, an additional mode may be provided that allows specifying a person combined with a specific object, such as "a player holding a ball," as the filtering target, in addition to the exclusion mode and the narrowing mode. In this case, in the additional mode, for example, a specific object recognition unit 124 that recognizes a ball as the recognition target may be used to determine the recognition result as a person region captured near the ball through a process similar to the narrowing determination process.

[0129] The above describes an example in which the digital camera 100 performs subject filtering in two stages using two specific object recognition units 124. In the digital camera 100 of this embodiment, subject filtering may be performed in stages using three or more specific object recognition units 124. For example, from the above example, subject filtering may be performed further based on the processing result of step S7A and the detection information from the third specific object recognition unit 124 to determine the recognition result.

[0130] (Other embodiments) As described above, Embodiments 1 and 2 have been explained as examples of the technology disclosed in this application. However, the technology in this disclosure is not limited thereto and can be applied to embodiments that are modified, replaced, added to, or omitted as appropriate. Furthermore, it is possible to create new embodiments by combining the components described in each of the above embodiments.

[0131] In each of the embodiments described above, a digital camera 100 having both an exclusion mode and a narrowing mode as filtering modes was described. The digital camera 100 of this embodiment may have only one of either the exclusion mode or the narrowing mode as filtering modes.

[0132] In each of the embodiments described above, an example was described in which the image processing engine 120 includes a person recognition unit 122 and a plurality of specific object recognition units 124. In the digital camera 100, each recognition unit 122, 124 may be configured separately from the image processing engine 120. Furthermore, the digital camera 100 of this embodiment may have only one specific object recognition unit 124.

[0133] In each of the embodiments described above, an example was described in which the digital camera 100 is equipped with a trained model that functions as a person recognition unit 122 and a specific object recognition unit 124. In this embodiment, the digital camera 100 does not need to be equipped with such a trained model in advance. In this embodiment, the trained models that function as the respective recognition units 122 and 124 may be stored in an information processing device, such as a cloud server that can communicate data with the digital camera 100. The digital camera 100 in this embodiment may, for example, acquire the trained model from such an external information processing device via a communication module 155, and make the acquired trained model function as the person recognition unit 122 and / or the specific object recognition unit 124.

[0134] For example, in the above embodiment, the digital camera 100 may further implement at least one of the functions of the person recognition unit 122 and the specific object recognition unit 124 by cooperating with an external information processing device as described above. For example, the digital camera 100 in this embodiment may transmit image data of the captured image to the information processing device via the communication module 155 to detect the person region and / or the specific object region in the captured image. The controller 135 may obtain detection information from each recognition unit 122 and 124 by having the person recognition unit 122 and / or the specific object recognition unit 124 acquire the regions recognized by the information processing device. In this case, the digital camera 100 does not need to have pre-trained models that implement the functions of each recognition unit 122 and 124, nor does it need to acquire them from an external source.

[0135] As described above, in this embodiment, the digital camera 100 further includes a communication module 155, which is an example of a communication unit that communicates data with an external information processing unit. In this embodiment, the controller 135, via the communication module 155, causes the information processing unit to recognize the area of ​​a subject in the image shown by the image data, and causes at least one of the person recognition unit 122, which is an example of a first recognition unit, and the specific object recognition unit 124, which is an example of a second recognition unit, to acquire the recognized area. This further reduces the load on the image recognition processing in the digital camera 100.

[0136] In the embodiments described above, an example was explained in which a person recognition unit 122 detects a person region as an example of a first recognition unit in the subject recognition process. However, the subject to be recognized by the first recognition unit is not limited to people. In the digital camera 100 of this embodiment, for example, a first recognition unit may be used that recognizes a region corresponding to the entire body of a specific type of animal, such as a dog or a horse, in an image. The subject to be recognized may also be a specific type of insect, plant, or vehicle. In this embodiment, a specific object recognition unit 124 that recognizes a specific object region corresponding to the object recognized by the first recognition unit may be used as a second recognition unit.

[0137] For example, in a shooting scene such as a horse race or a car race, the first recognition unit may recognize the area of ​​a horse or car, and the second recognition unit may recognize a specific number, thereby narrowing down the recognition result to a horse or car with a specific number. Furthermore, the digital camera 100 of this embodiment may also be equipped with an eye recognition function by image recognition. For example, when the subject is a person or an animal, the eye may be recognized from the peripheral area of ​​the area recognized by the first recognition unit in the captured image, and AF control may be performed to focus on the eye area within the recognized area.

[0138] In each of the embodiments described above, a digital camera 100 that uses the recognition result of the subject recognition process for AF control was described. In this embodiment, the digital camera 100 is not limited to AF control, but may also perform white balance (WB) control and / or automatic exposure (AE) based on the recognition result.

[0139] In each of the embodiments described above, a digital camera 100 was described that determines the subject of recognition results by subject recognition processing for captured images Im such as through images. In this embodiment, the digital camera 100 may perform subject recognition processing on images shown in image data stored in the memory card 142.

[0140] In each of the embodiments described above, a digital camera 100 comprising an optical system 110 and a lens drive unit 112 was illustrated. In this embodiment, the imaging device does not necessarily have to include an optical system 110 and a lens drive unit 112; for example, it may be a camera with interchangeable lenses.

[0141] In the embodiments described above, a digital camera was used as an example of an imaging device, but the invention is not limited to this. The imaging device of this disclosure may be any electronic device having an image capture function (e.g., a video camera, smartphone, tablet terminal, etc.).

[0142] (Summary of characteristics) The various aspects of this disclosure are listed below.

[0143] A first aspect of the present disclosure is an imaging device comprising: an imaging unit that captures an image of a subject and generates image data; a first recognition unit that recognizes a region of a first subject in an image indicated by the image data based on the image data generated by the imaging unit; a second recognition unit that recognizes a region of a second subject in an image that is included in or separate from the first subject based on the image data; and a control unit that determines the recognized subject region in the image data based on the region of the first subject recognized by the first recognition unit and the region of the second subject recognized by the second recognition unit.

[0144] The second embodiment is an imaging apparatus as described in the first embodiment, wherein the control unit refers to the region of the second subject recognized by the second recognition unit in the image, extracts a portion of the region of the first subject recognized by the first recognition unit, and determines the extracted region to be the subject region.

[0145] In the second embodiment, the control unit refers to the region of the second subject recognized by the second recognition unit in the image, extracts a specific region from the region of one or more first subjects recognized by the first recognition unit, and determines the extracted region to be the subject region.

[0146] A third embodiment is an imaging apparatus according to the first or second embodiment, wherein the second subject is a specific type of subject included in the first subject, and the control unit excludes from the region of the first subject recognized by the first recognition unit a region that has a predetermined correspondence with the region of the second subject recognized by the second recognition unit, and determines the remaining region as the subject region.

[0147] A fourth embodiment is an imaging device according to the first or second embodiment, wherein the second subject is another subject related to the first subject, and the control unit narrows down the region of the first subject recognized by the first recognition unit in the image to a portion of the region of the second subject recognized by the second recognition unit that is located within a predetermined range, and determines the narrowed-down region to be the subject region.

[0148] The fifth embodiment is an imaging device according to any of the first to fourth embodiments, further comprising an operating unit for receiving user input, wherein the pre-control unit receives a user input from the operating unit to set a second subject, and causes the second recognition unit to recognize the area of ​​the set second subject.

[0149] The sixth aspect is an imaging apparatus according to any of the first to fifth aspects, wherein the control unit is If a second subject is not set, the subject area is determined based on the area of ​​the first subject recognized by the first recognition unit, without referring to the area of ​​the second subject. If a second subject is set, the subject area is determined based on both the area of ​​the first subject recognized by the first recognition unit and the area of ​​the second subject recognized by the second recognition unit.

[0150] The seventh embodiment is an imaging device according to any of the first to sixth embodiments, further comprising a communication unit that communicates data with an external information processing device, wherein the control unit causes the information processing device to recognize the area of ​​a subject in the image shown by the image data via the communication unit, and causes at least one of the first recognition unit and the second recognition unit to acquire the recognized area of ​​the subject.

[0151] The eighth aspect is an imaging device comprising: an imaging unit that captures an image of a subject and generates image data; an operation unit that receives user input; and a control unit that controls image recognition processing for a first subject on the image data generated by the imaging unit in response to user input in the operation unit. The control unit receives user input in the operation unit to set a second subject that is included in or separate from the first subject, and controls image recognition processing based on the second subject set in the operation unit to determine the subject area of ​​the recognition result in the image data.

[0152] As described above, embodiments have been explained as examples of the technology in this disclosure. For this purpose, accompanying drawings and a detailed description have been provided.

[0153] Therefore, the components described in the attached drawings and detailed descriptions may include not only components essential for solving the problem, but also components that are not essential for solving the problem, provided that they illustrate the technology described above. For this reason, the mere presence of these non-essential components in the attached drawings and detailed descriptions should not be immediately assumed to mean that they are essential.

[0154] Furthermore, since the embodiments described above are for illustrative purposes of the technology described herein, various modifications, substitutions, additions, omissions, etc., can be made within the claims or their equivalents. [Industrial applicability]

[0155] This disclosure is applicable to imaging devices that recognize the region of a subject in an image. [Explanation of Symbols]

[0156] 100 Digital Cameras 115 Image Sensor 120 Image Processing Engines 122 Person recognition section 124 Specific object recognition unit 125 buffer memory 130 Display Monitor 135 Controller 145 Flash Memory 150 Operation section 155 Communication Module

Claims

1. An imaging unit that captures an image of the subject and generates image data, A first recognition unit recognizes the region of the first subject in the image shown by the image data, based on the image data generated by the imaging unit. A second recognition unit recognizes, based on the image data, a region of a second subject that is included in the first subject or is separate from the first subject in the image, The system includes a control unit that determines the subject region of the recognition result in the image data based on the region of the first subject recognized by the first recognition unit and the region of the second subject recognized by the second recognition unit. Imaging device.

2. The control unit refers to the region of the second subject recognized by the second recognition unit in the image, extracts a portion of the region of the first subject recognized by the first recognition unit, and determines the extracted region to be the subject region. The imaging apparatus according to claim 1.

3. The second subject is a specific type of subject included in the first subject, The control unit, in the image, excludes from the region of the first subject recognized by the first recognition unit the region of the second subject recognized by the second recognition unit a region that has a predetermined correspondence with the region of the second subject recognized by the second recognition unit, and determines the remaining region to be the subject region. The imaging apparatus according to claim 2.

4. The second subject is another subject related to the first subject, The control unit narrows down the region of the first subject recognized by the first recognition unit in the image to a portion of the region of the second subject recognized by the second recognition unit that is located within a predetermined range, and determines the narrowed-down region to be the subject region. The imaging apparatus according to claim 2.

5. It also includes an operating section that accepts user input, The control unit receives a user operation to set the second subject in the operation unit and causes the second recognition unit to recognize the area of ​​the set second subject. The imaging apparatus according to claim 1.

6. The control unit, If the second subject is not set, the subject area is determined based on the area of ​​the first subject recognized by the first recognition unit, without referring to the area of ​​the second subject. If the second subject is set, the subject area is determined based on both the area of ​​the first subject recognized by the first recognition unit and the area of ​​the second subject recognized by the second recognition unit. The imaging apparatus according to claim 5.

7. The imaging device further comprises a communication unit that communicates data with an external information processing device, The control unit, via the communication unit, causes the information processing device to recognize the area of ​​the subject in the image shown by the image data, and causes at least one of the first recognition unit and the second recognition unit to acquire the recognized area of ​​the subject. The imaging apparatus according to claim 1.

8. An imaging unit that captures an image of the subject and generates image data, An operating unit that accepts user input, The system comprises a control unit that controls image recognition processing for a first subject on image data generated by the imaging unit in response to user operation in the operation unit, The control unit, The operation unit receives a user operation to set a second subject that is included in the first subject or is different from the first subject, Based on the second subject set in the operation unit, the image recognition process is controlled to determine the subject area of ​​the recognition result in the image data. Imaging device.

Citation Information

Patent Citations

  • Image processing device, control method thereof, program, and storage media

    JP2020126434A