Image processing apparatus and image processing method
The image processing apparatus uses a machine learning model to reliably detect the positional relationship and main subject in a single frame, addressing the challenge of varying photographer-subject distances and angles in multi-subject scenarios.
Patent Information
- Application Number
- JP2024007544
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2024-01-22
- Publication Date
- 2025-08-01
AI Technical Summary
In scenarios where multiple subjects are moving in the same direction, such as track and field events or races, it is challenging to consistently determine the moving direction and positional relationship of subjects due to varying photographer-subject distances and angles of view.
An image processing apparatus and method using a machine learning model to detect the positional relationship and determine a main subject region based on the reliability of the estimated positional information, without requiring motion vectors and minimizing erroneous detections.
Enables accurate detection of the positional relationship and main subject in a single frame, reducing errors caused by changing photographer-subject distances and angles, and ensuring consistent focus and exposure on the main subject.
Smart Images

Figure 2025112958000001_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to an image processing apparatus and an image processing method, and particularly to a technique for detecting the positional relationship of a plurality of subjects from a photographed image.
Background Art
[0002] In photographing a scene where a plurality of subjects move in the same direction, such as track and field events or races, there are cases where it is desired to track not a specific subject but a subject at a specific position (for example, the leading position). Therefore, in Patent Document 1, the leading subject is detected based on the moving direction from among a plurality of subjects moving in the same direction.
Prior Art Documents
Patent Documents
[0003]
Patent Document 1
Summary of the Invention
Problems to be Solved by the Invention
[0004] However, the positional relationship between the photographer and a plurality of subjects is not constant, and the photographing angle of view is not always constant. Therefore, it may not be easy to specify the moving direction of a plurality of subjects. In one aspect of the present invention, there is provided an image processing apparatus and an image processing method capable of detecting the positional relationship in the moving direction of a plurality of subjects from an image of one frame.
Means for Solving the Problems
[0005] In one aspect of the present invention, there is provided an image processing apparatus having: acquisition means for acquiring, for each of a plurality of subjects, information regarding a positional relationship in a direction from an input image based on image data obtained by photographing a scene in which the plurality of subjects are moving in the same direction, using a machine learning model; and determination means for determining a main subject region among regions of the plurality of subjects included in the image data based on the information, wherein the determination means determines whether or not to determine the main subject region based on the reliability of the information.
Advantages of the Invention
[0006] According to the present invention, it is possible to provide an image processing apparatus and an image processing method capable of detecting a positional relationship in a moving direction of a plurality of subjects from an image of one frame.
Brief Description of the Drawings
[0007]
Figure 1
Figure 2
Figure 3
Figure 4
Figure 5
Figure 6
Figure 7
Figure 8
Figure 9
Figure 10
Figure 11
Embodiments for Carrying Out the Invention
[0008] Hereinafter, the present invention will be described in detail based on its exemplary embodiments with reference to the accompanying drawings. Note that the following embodiments do not limit the invention according to the claims. Also, although a plurality of features are described in the embodiments, not all of them are essential to the invention, and a plurality of features may be arbitrarily combined. Further, in the accompanying drawings, the same or similar configurations are denoted by the same reference numerals, and redundant descriptions are omitted.
[0009] In the following embodiments, the case where the present invention is implemented in a digital single-lens reflex camera will be described. Although the present invention can be suitably used for controlling a photographing operation, a photographing function is not essential for the device for implementing the present invention. The present invention can be implemented with any electronic device capable of handling image data. Such electronic devices include video cameras, computer devices (personal computers, tablet computers, media players, PDAs, etc.), mobile phones, smartphones, game machines, robots, drones, and drive recorders. These are merely examples, and the present invention can also be implemented with other electronic devices.
[0010] ●(First Embodiment) FIG. 1 is a vertical view schematically showing an example of the main components of a digital single-lens reflex camera (hereinafter simply referred to as a camera) 100 as an image processing apparatus according to the first embodiment of the present invention and an example of their arrangement. The camera 100 has a main body 101 and a detachable lens unit 120. The lens unit 120 has a plurality of lenses including a focus lens 121, a diaphragm 122, and a mechanism for driving a movable member.
[0011] The lens unit 120 and the main body 101 are detachably connected by a mount portion that engages with each other. The mount portion is provided with a contact portion 123 having a plurality of terminals, and power is supplied from the main body 101 to the lens unit 120 through the contact portion 123. Also, the operations of the focus lens 121 and the aperture 122 can be controlled from the main body 101 (the arithmetic unit 102) through the contact portion 123.
[0012] The shutter 103 provided between the lens unit 120 and the imaging element 104 opens and closes under the control of the arithmetic unit 102 to control the exposure period of the imaging element 104. Note that when controlling the exposure period of the imaging element 104 by a so-called electronic shutter, the shutter 103 that physically opens and closes may not be necessary.
[0013] The imaging element 104 may be, for example, a known CCD or CMOS color image sensor having a color filter in a primary color Bayer array. The imaging element 104 has a pixel array in which a plurality of pixels are two-dimensionally arranged and a peripheral circuit for reading signals from each pixel. Each pixel accumulates electric charges corresponding to the incident light amount by photoelectric conversion. By reading out a signal having a voltage corresponding to the amount of electric charges accumulated during the exposure period from each pixel, a pixel signal group (analog image signal) representing the subject image formed by the lens unit 120 on the imaging surface is obtained.
[0014] The display unit 105 is provided on the surface of the housing of the camera 100 and displays a live view video, a recorded image, a menu screen, and the like. The display unit 105 may be, for example, a liquid crystal display device (LCD). Also, the display unit 105 may be a touch display.
[0015] The operation unit 106 is a general term for a plurality of input devices (buttons, switches, dials, etc.) provided on the camera 100 for the user to input various instructions to the camera 100. The input devices constituting the operation unit 106 have names corresponding to the assigned functions. For example, a release switch, a video recording switch, a shooting mode selection dial for selecting a shooting mode, a menu button, direction keys, a determination key, etc. The release switch is a switch for still image recording, and the arithmetic unit 102 recognizes the half-pressed state of the release switch as an instruction for shooting preparation and the fully pressed state as an instruction for starting shooting. Also, when the video recording switch is pressed in the shooting standby state, the arithmetic unit 102 recognizes it as an instruction for starting video recording, and when it is pressed during video recording, it recognizes it as an instruction for stopping recording. Note that the functions assigned to the same input device may be variable. Also, the input device may be a software button or key using a touch display. Further, the operation unit 106 may include an input device corresponding to a non-contact input method such as voice input or gaze input.
[0016] The arithmetic unit 102 has one or more processors (hereinafter referred to as CPUs) capable of executing programs, and reads, for example, a program stored in the ROM 112 (FIG. 2) into the RAM 111 (FIG. 2) and executes it with the CPU. The arithmetic unit 102 realizes the functions of the camera 100 by controlling the operations of the components of the camera 100 according to the program. Also, when the arithmetic unit 102 detects an operation on the operation unit 106, it executes an operation corresponding to the detected operation.
[0017] FIG. 2 is a block diagram showing an example of the functional configuration of the camera 100. The same components as those in FIG. 1 are given the same reference numerals as in FIG. 1.
[0018] Each of the functional blocks 201 to 205 included in the arithmetic unit 102 schematically shows the main functions realized by one or more CPUs included in the arithmetic unit 102 by executing a program. Note that one or more of the functional blocks may be implemented by a CPU using another hardware circuit. For example, a GPU (Graphics Processing Unit) or an ASIC (Application Specific Integrated Circuit) can be used for processing related to image processing. Also, an NPU (Neural Processing Unit) can be used for processing related to a machine learning model. These hardware circuits may be included in the arithmetic unit 102 or may be external circuits of the arithmetic unit 102.
[0019] Further, the functional blocks 201 to 205 are part of the operations performed by the arithmetic unit 102. Therefore, in the following description, there are cases where the operating entity is the arithmetic unit 102 and cases where the functional blocks 201 to 205 are involved.
[0020] The RAM 111 is used to load a program executed by the CPU of the arithmetic unit 102 and to store values necessary during the execution of the program. Also, a part of the RAM 111 is used as a buffer for temporarily storing captured image data or as a video memory for the display unit 105.
[0021] The ROM 112 is an electrically rewritable non-volatile memory. The ROM 112 stores programs executable by the CPU of the arithmetic unit 102, setting values of the camera 100, GUI data, and the like.
[0022] Note that FIG. 2 shows only a part of the components of the camera 100 in order to simplify the description of the embodiment. In reality, the camera 100 has components such as a power source, a recording medium, and a communication interface that are common to general cameras.
[0023] The control unit 201 controls the operations of the main body 101 and the lens unit 120. For example, the control unit 201 controls the operation timing of the imaging device 104 and reads out an analog image signal from the imaging device 104. Further, the control unit 201 applies A / D conversion or predetermined image processing to the analog image signal to generate a signal or image data suitable for the application, and acquires and / or generates various types of information. For example, the control unit 201 generates image data for display or recording, and generates evaluation values or signals used for autofocus detection (AF) and automatic exposure control (AE).
[0024] The main subject calculation unit 202 detects a subject area included in the image, and determines a main subject area to be focused by the camera 100 from among the detected subject areas. The main subject calculation unit 202 determines, based on an image of one frame that has been captured, as the main subject area, a subject area that is estimated to have the highest probability of being located in a specific order (here, the top) in the moving direction among a plurality of subject areas. Details of the operation of the main subject calculation unit 202 will be described later.
[0025] The tracking calculation unit 203 searches for the main subject area determined by the main subject calculation unit 202 in subsequent frames. There is no limitation on the search method, and for example, any known method such as pattern matching using the main subject area as a template or a method using feature amounts extracted from the main subject area can be used. The tracking calculation unit 203 outputs the position information of the searched main subject area and generates information to be used for the next search according to the search method. For example, when using template matching, the tracking calculation unit 203 uses the determined main subject area as a template for searching for the next main subject area.
[0026] The focus calculation unit 204 calculates the driving amount and driving direction of the focus lens 121 such that the focus detection area set in the main subject area is in focus. The focus calculation unit 204 can calculate the driving amount and driving direction of the focus lens 121 by a known method such as the contrast method or the phase difference detection method.
[0027] The exposure calculation unit 205 determines exposure parameters (aperture value, shutter speed, shooting sensitivity) such that the main subject area has proper exposure. The exposure calculation unit 205 can determine the exposure parameters, for example, from the evaluation value generated by the control unit 201 and the program diagram.
[0028] Next, the operation of the camera 100 in this embodiment will be described using the flowchart shown in FIG. 3. This operation starts in the standby state of the still image shooting mode. In the standby state, video shooting for performing live view display on the display unit 105 is being executed. Note that since a scene in which a plurality of subjects are moving in substantially the same direction is assumed, the operation described below may be executed when, for example, a mode for shooting such a scene is set.
[0029] In S301, the control unit 201 reads an image signal for one frame from the imaging device 104 and generates image data for display on the display unit 105 from the image signal. The control unit 201 stores the generated image data in the RAM 111.
[0030] In S302, the arithmetic unit 102 determines whether the currently processed frame is the first frame. If it is determined to be the first frame, S306 is executed; otherwise, S303 is executed.
[0031] In S303, the tracking calculation unit 203 uses the tracking reference information generated in S310 executed for the previous frame to search for the position of the main subject area set in the previous frame in the current frame. The tracking calculation unit 203 stores the information (position, size, etc.) of the main subject area searched in the current frame in the RAM 111 as the tracking result.
[0032] In S304, the focus calculation unit 204 sets a focus detection area in the main subject area based on the tracking result generated in S303. The focus calculation unit 204 uses the evaluation value and signal generated by the control unit 201 in S301 to determine the movement amount and movement direction of the focus lens 121 required for the focus detection area to be in focus. The focus calculation unit 204 notifies the control unit 201 of the determined movement amount and movement direction. When receiving the notification, the control unit 201 executes control to drive the focus lens 121 in the notified movement amount and movement direction.
[0033] In S305, the exposure calculation unit 205 uses the evaluation value generated in S301 to determine exposure parameters such that the main subject area based on the tracking result has proper exposure. The exposure calculation unit 205 notifies the control unit 201 of the determined exposure parameters. When receiving the notification, the control unit 201 controls the shutter speed, shooting sensitivity, and aperture 122 at the time of the next frame shooting according to the exposure parameters.
[0034] In S306, the control unit 201 determines whether a still image shooting start instruction has been detected through the operation unit 106. If it is determined that the shooting start instruction has been detected, the control unit 201 executes S307; otherwise, it executes S308.
[0035] In S307, the control unit 201 executes still image shooting processing. The control unit 201 drives the shutter 103 based on the exposure parameters notified in S305 to expose the image sensor 104. Then, still image data for recording is generated from the signal read from the image sensor 104. The control unit 201 records the generated image data on a recording medium (not shown) such as a memory card.
[0036] In S308, the control unit 201 instructs the main subject calculation unit 202 to determine the main subject area. In response to the instruction, the main subject calculation unit 202 determines the main subject area based on the image data generated in S301. The main subject calculation unit 202 notifies the control unit 201 and the tracking calculation unit 203 of the information on the determined main subject area. Details of the main subject area determination process will be described later.
[0037] In S309, based on the information of the main subject area notified in S308 and the image data generated in S301, the tracking calculation unit 203 generates tracking information for use in the tracking process (S302) for the next frame. The tracking information can vary depending on the tracking method as described above.
[0038] In S310, based on the information of the main subject area notified in S308, the control unit 201 superimposes an index (e.g., a frame) indicating the main subject area on the image data generated in S301 and displays it on the display unit 105.
[0039] The above operations are executed each time one frame of the live view image is captured. While the camera 100 is operating in the still image shooting mode, the series of processes shown in FIG. 3 are repeatedly executed.
[0040] Next, the determination process of the main subject area executed by the main subject calculation unit 202 in S308 will be described in more detail with reference to the flowchart shown in FIG. 4. In S401, the main subject calculation unit 202 executes a subject detection process on the image data generated in S301. Here, it is assumed that a human subject is detected, but other types of subjects such as animals and vehicles may also be detected. Any known method can be used for detection. For example, a machine learning model such as a convolutional neural network (CNN) trained using human images can be used. Also, a method such as AdaBoost that combines multiple machine learning models may be used.
[0041] In S402, the main subject operation unit 202 (acquisition means) estimates the rank in the moving direction for each of the person subjects detected in S401. The rank is a discrete value that increases by 1 in order, such as 1st, 2nd, 3rd... starting from the subject located at the head in the moving direction. The rank in the moving direction can be obtained, for example, by an inference process using a machine learning model such as a CNN that has been learned using, as a learning dataset, a combination of an image in which a plurality of subjects are captured and the rank in the moving direction of each subject area in the image. When using the detection result of the subject area in S401 during inference, the information of the subject area is also used during learning. For example, the detection result of the subject area can be used during learning and inference in any method, such as extracting the subject area from the original image and using it as the input image of the CNN.
[0042] FIG. 5(a) is an example of an image represented by the image data generated in S301. It is an image of a plurality of person subjects 401, 402, 403 moving in the direction indicated by the arrow (the arrow is not included in the image). Further, FIG. 5(b) is an example of the subject detection result in S401 for the image shown in FIG. 5(a) and the estimated rank for each subject area obtained in S402. Here, the human head is detected as the subject area, and the image coordinates corresponding to the center of gravity of the subject area are shown as the coordinates of the subject area. Also, the subject IDs 1 to 3 for specifying the subject area correspond to the person subjects 401 to 403, respectively.
[0043] In S403, the main subject operation unit 202 (determination means) determines the reliability of the rank estimated in S402. As an example, the main subject operation unit 202 determines the reliability of the rank estimated for the current frame based on the rank estimated in the past for the same subject area. Considering the shooting interval (generally 1 / 30 second in the case of video shooting), it is assumed that the variation from the previous estimated rank is ±1. Therefore, when the rank changes by ±2 or more, the main subject operation unit 202 determines that the reliability of the rank inferred by the machine learning model is low.
[0044] Specifically, when the following conditions are met, the main subject calculation unit 202 determines that the reliability of the ranking is low (reliability = 0), and when the conditions are not met, it determines that the reliability of the ranking is high (reliability = 1). Condition: The ranking of any subject has changed by more than a threshold value from the previous estimated ranking (In this embodiment, the threshold value = ±2, and when the ranking change is an absolute value, the threshold value = 2) When the current frame is the first frame, since there is no past ranking, the reliability is set to 0.
[0045] Here, the threshold value is set to 2, but the threshold value may be dynamically determined according to, for example, the number of subjects. For example, when there are a very large number of subjects, such as immediately after the start of a marathon or various races, the threshold value can be temporarily set to 3 or more. Also, the threshold value may be determined considering other conditions, such as making the threshold value larger as the shooting interval or the reciprocal of the frame rate, or the interval at which the ranking is estimated is larger. Also, not only the most recent ranking but also a plurality of past rankings may be considered.
[0046] In S404, the main subject calculation unit 202 (determination means) determines the main subject area from the subject areas detected in S401. If the determined main subject area does not exist (no main subject), when the main subject calculation unit 202 determines in S403 that the reliability of the ranking is high (reliability = 1), it determines the subject area ranked first in the current frame as the main subject area. On the other hand, when it is determined in S403 that the reliability of the ranking is low (reliability = 0), the main subject calculation unit 202 does not determine the main subject area.
[0047] If the determined main subject area exists (there is a main subject), the main subject calculation unit 202 determines whether to switch the main subject area based on the ranking obtained in S402 and the reliability determined in S403. Specifically, the main subject calculation unit 202 determines to switch the main subject area when both of the following Condition A and Condition B are met, and determines not to switch the main subject if one or more are not met. Condition A: The subject area ranked first in the estimated ranking has changed from the previous frame Condition B: The reliability of the ranking estimated in the current frame is high (reliability = 1).
[0048] When the main subject calculation unit 202 determines to switch the main subject area, it determines the subject area estimated to be ranked first in the current frame as the new main subject area. On the other hand, when the main subject calculation unit 202 determines not to switch the main subject area, it maintains the determined main subject area.
[0049] The switching operation of the main subject area in S404 will be further described with reference to FIG. 6. FIG. 6(a) shows an example of the ranking history for each subject area in the current frame (time t) and the past three frames. The ranking history is stored in, for example, the RAM 111 by the main subject calculation unit 202 and is sequentially updated after the execution of S402. Hereinafter, the subject areas of subject IDs 1 to 3 will be referred to as subject areas 1 to 3, respectively.
[0050] In this example, there is no change in the ranking until the previous frame (time t - 1), and the subject areas ranked first and second are swapped in the current frame (time t), so condition A is satisfied. Also, since the change in the ranking of each of the subject areas 1 to 3 between the previous frame and the current frame is within the range of ±1, the reliability of the ranking estimated in the current frame is high, and condition B is satisfied. Therefore, the main subject calculation unit 202 switches the main subject area from subject area 1 to subject area 2.
[0051] FIG. 6(b) shows another example of the ranking history. In this example, there is no change in the ranking until the previous frame (time t - 1), and the subject areas ranked first and third are swapped in the current frame (time t), so condition A is satisfied. On the other hand, the change in the ranking of subject areas 1 and 3 between the previous frame and the current frame is ±2 for both, which exceeds the threshold value. Therefore, the reliability of the ranking estimated in the current frame is low, and condition B is not satisfied. Therefore, the main subject calculation unit 202 does not switch the main subject area and maintains subject area 1, which was ranked first in the previous frame, as the main subject area.
[0052] FIG. 6(c) shows yet another example of the ranking history. In this example, there is no change in the ranking up to the previous frame (time t - 1), and the rankings of all subject regions change in the current frame (time t), so condition A is satisfied. On the other hand, the change in the rankings of subject regions 1 and 2 between the previous frame and the current frame is ±1, but the change in the ranking of subject region 3 is -2, exceeding the threshold. Therefore, the reliability of the ranking estimated in the current frame is low, and condition B is not satisfied. Accordingly, the main subject calculation unit 202 maintains the subject region 1 with the ranking of 1st place in the previous frame as the main subject region without switching the main subject region.
[0053] According to the present embodiment, since a machine learning model for estimating the positional relationship in the moving direction of a plurality of subjects based on the image of the current frame is used, it is not necessary to calculate a motion vector. Therefore, it is not necessary to hold the image of the previous frame. In addition, it is possible to suppress erroneous detection of the positional relationship when the positional relationship between the photographer and the subject or the angle of view changes.
[0054] In addition, by preventing the switching of the main subject region when it is determined that the reliability of the estimation result for the current frame by the machine learning model is low, it is possible to further suppress the switching of the incorrect subject region.
[0055] ●(Second Embodiment) Next, a second embodiment of the present invention will be described. In the first embodiment, the positional relationship (ranking) in the moving direction of each subject was estimated. In this embodiment, instead of the ranking, the leading likelihood proportional to the distance in the moving direction is estimated.
[0056] This embodiment is different from the first embodiment in the determination process of the main subject region performed in S308 in FIG. 3. The configuration of the image processing apparatus and other operations may be the same as those in the first embodiment. Therefore, hereinafter, the operation of the main subject calculation unit 202 in this embodiment will be mainly described.
[0057] FIG. 7 is a flowchart regarding the determination process of the main subject processing performed by the main subject calculation unit 202 in the present embodiment. In FIG. 7, the same reference numerals as those in FIG. 4 are given to the steps that execute the same processing as in the first embodiment, and the description thereof is omitted.
[0058] In S702, the main subject calculation unit 202 estimates the start likelihood for each of the subject regions detected in S401. Here, the start likelihood is a real number that can have a value from 0 to 1, and takes a larger value as it is located forward in the moving direction and a smaller value as it is located rearward. Also, the difference in the start likelihood between the subject regions is proportional to the distance between the subjects in the moving direction in the real space. That is, the magnitude of the difference in the start likelihood of a certain pair of subject regions represents the magnitude of the interval in the moving direction between the corresponding subjects in the real space.
[0059] Such a start likelihood can be obtained, for example, by an inference process using a machine learning model such as a CNN learned using a combination of an image in which a plurality of subjects are captured and the start likelihood of each subject region in the image as a learning dataset. When using the detection result of the subject region in S401 at the time of inference, the information of the subject region is also used at the time of learning. For example, the detection result of the subject region can be used at the time of learning and inference by any method such as extracting the subject region from the original image and using it as the input image of the CNN.
[0060] FIG. 8(a) is an example of an image similar to FIG. 5(a) except that the reference numerals of the human subject are different. FIG. 8(b) is an example of the subject detection result in S401 for the image shown in FIG. 8(a) and the start likelihood for each subject region obtained in S702. Here, the human head is also detected as a subject region, and the image coordinates corresponding to the center of gravity of the subject region are shown as the coordinates of the subject region. Also, the subject IDs 1 to 3 that identify the subject regions respectively correspond to the human subjects 801 to 803.
[0061] In S703, the main subject calculation unit 202 determines the reliability of the leading likelihood estimated in S702. As an example, for each subject area, the main subject calculation unit 202 determines the reliability of the leading likelihood estimated for the current frame based on the leading likelihood estimated in the past. Considering the shooting interval (generally 1 / 30 second in video shooting), it is unlikely that the relative distance to other subjects will change extremely from the frame in which the previous leading likelihood was estimated to the current frame. Therefore, for a subject area where the change in the difference in leading likelihoods from one or more other subject areas exceeds the threshold, the main subject calculation unit 202 determines that the reliability of the leading likelihood inferred by the machine learning model is low. Note that the difference in leading likelihoods is used instead of the leading likelihood for one subject area because the value of the leading likelihood can change significantly in a short time, for example, due to a panning operation of the camera.
[0062] For the subject area with subject ID i at time t (hereinafter referred to as subject area i), let the leading likelihood be h i (t), and the reliability of the leading likelihood be ω i (t). Also, let the difference in leading likelihoods between subject area i and subject j at time t be h i (t) - h j (t) be d i,j (t). In the example shown in FIG. 8, 1 ≤ i, j ≤ 3.
[0063] The main subject calculation unit 202 determines whether there exists a subject area j such that |d i,j (t) - d i,j (t - 1)| > θ for subject area i. If it is determined that such a subject area j exists, the main subject calculation unit 202 determines that the reliability of the leading likelihood estimated for subject area i in the current frame is low (ω i (t) = 0). Also, if such a subject area j does not exist, the main subject calculation unit 202 determines that the reliability of the leading likelihood estimated for subject area i in the current frame is high (reliability ω i (t) = 1). Here, θ is a predetermined threshold.
[0064] That is, for the subject region i in the current frame, when there is another subject region in which the change in the difference in leading likelihood exceeds the threshold from the previous frame, the main subject calculation unit 202 determines that the reliability of the leading likelihood estimated in the current frame is low. Note that when the current frame is the first frame, since there is no past leading likelihood, the reliability of the leading likelihood is low for all subject regions i (ω i (t)=0) is determined.
[0065] In S704, the main subject calculation unit 202 determines the main subject region from the subject regions detected in S401. If the determined main subject region does not exist (no main subject), when the reliability determined for the subject region with the highest leading likelihood in the current frame in S703 is high, the main subject calculation unit 202 determines the subject region with the highest leading likelihood as the main subject region. On the other hand, if the reliability determined for the subject region with the highest leading likelihood in the current frame in S703 is low, the main subject calculation unit 202 does not determine the main subject region.
[0066] If the determined main subject region exists (there is a main subject), the main subject calculation unit 202 determines whether to switch the main subject region based on the leading likelihood obtained in S702 and the reliability determined in S703. Specifically, the main subject calculation unit 202 determines to switch the main subject region when both of the following conditions A and B are satisfied, and determines not to switch the main subject if one or more are not satisfied. Condition A: The subject region with the highest estimated leading likelihood has changed from the previous frame Condition B: For the subject region with the highest leading likelihood estimated in the current frame, the reliability of the leading likelihood is high
[0067] When the main subject calculation unit 202 determines to switch the main subject region, it determines the subject region with the highest leading likelihood estimated for the current frame as the new main subject region. On the other hand, when the main subject calculation unit 202 determines not to switch the main subject region, it maintains the determined main subject region.
[0068] The switching operation of the main subject area in S704 will be further described with reference to FIG. 9. FIG. 9(a) shows an example of the history of the leading likelihood for each subject area in the current time t (current frame) and the past three frames. The history of the leading likelihood is stored, for example, in the RAM 111 by the main subject calculation unit 202 and is sequentially updated after the execution of S702. Hereinafter, the subject areas of subject IDs 1 to 3 will be referred to as subject areas 1 to 3, respectively. Also, let the threshold value θ = 0.1.
[0069] In this example, up to the previous frame (time t - 1), there is no change in the order of the magnitudes of the leading likelihoods of each subject area, and since the subject area with the maximum leading likelihood in the current frame (time t) has changed from subject area 1 to subject area 2, condition A is satisfied.
[0070] Next, when paying attention to the subject area 2 with the maximum leading likelihood and the other subject areas 1 and 3, d 2,1 (t)=h2(t)-h1(t)=0.63 - 0.61 = 0.02 d 2,1 (t - 1)=h2(t - 1)-h1(t - 1)=0.61 - 0.62 = -0.01 |d 2,1 (t)-d 2,1 (t - 1)|=|0.02 - (-0.01)| = 0.03 < θ d 2,3 (t)=h2(t)-h3(t)=0.63 - 0.46 = 0.17 d 2,3 (t - 1)=h2(t - 1)-h3(t - 1)=0.61 - 0.45 = 0.16 |d 2,3 (t)-d 2,3 (t - 1)|=|0.17 - 0.16| = 0.01 < θ Therefore, in S703, the main subject calculation unit 202 determines that the reliability of the leading likelihood estimated for the current frame is high and satisfies condition B. Therefore, in S704, the main subject calculation unit 202 switches the main subject area from subject area 1 to subject area 2.
[0071] FIG. 9(b) shows another example of the history of the leading likelihood. Since there is no change in the ranking of the magnitudes of the leading likelihoods of each subject region up to the previous frame (time t - 1), and the subject region with the maximum leading likelihood in the current frame (time t) has changed from subject region 1 to subject region 2, condition A is satisfied.
[0072] Next, when focusing on the subject region 2 with the maximum leading likelihood and the other subject regions 1 and 3, d 2,1 (t)=h2(t)-h1(t)=0.83 - 0.61 = 0.22 d 2,1 (t - 1)=h2(t - 1)-h1(t - 1)=0.53 - 0.62 = -0.09 |d 2,1 (t)-d 2,1 (t - 1)|=|0.22 - (-0.09)| = 0.31 > θ d 2,3 (t)=h2(t)-h3(t)=0.83 - 0.45 = 0.38 d 2,3 (t - 1)=h2(t - 1)-h3(t - 1)=0.53 - 0.44 = 0.09 |d 2,3 (t)-d 2,3 (t - 1)|=|0.38 - 0.09| = 0.29 > θ Therefore, in S703, the main subject calculation unit 202 determines that the reliability of the leading likelihood estimated for the current frame is low and does not satisfy condition B. Accordingly, in S704, the main subject calculation unit 202 maintains the subject region 1 with the maximum leading likelihood in the previous frame as the main subject region without switching the main subject region.
[0073] FIG. 9(c) shows yet another example of the history of the leading likelihood. Since there is no change in the ranking of the magnitudes of the leading likelihoods of each subject region up to the previous frame (time t - 1), and the subject region with the maximum leading likelihood in the current frame (time t) has changed from subject region 1 to subject region 2, condition A is satisfied.
[0074] Next, when focusing on the subject region 2 with the maximum leading likelihood and the other subject regions 1 and 3, d 2,1(t)=h2(t)-h1(t)=0.45 - 0.42 = 0.03 d 2,1 (t - 1)=h2(t - 1)-h1(t - 1)=0.60 - 0.62 = -0.02 |d 2,1 (t)-d 2,1 (t - 1)|=|0.03 - (-0.02)| = 0.05 < θ d 2,3 (t)=h2(t)-h3(t)=0.45 - 0.25 = 0.20 d 2,3 (t - 1)=h2(t - 1)-h3(t - 1)=0.60 - 0.44 = 0.16 |d 2,3 (t)-d 2,3 (t - 1)|=|0.20 - 0.16| = 0.05 < θ Therefore, in S703, the main subject calculation unit 202 determines that the reliability of the leading likelihood estimated for the current frame is high and satisfies condition B. Accordingly, in S704, the main subject calculation unit 202 switches the main subject area from subject area 1 to subject area 2.
[0075] In this embodiment, the reliability is obtained for all other subject areas in S703. However, other subject areas for calculating the change in the difference in the leading likelihood may be restricted. For example, the change in the difference in the leading likelihood may be calculated for each of the subject area with the maximum leading likelihood and some other subject areas with the top leading likelihoods. Conversely, the change in the difference in the leading likelihood may be calculated for all combinations of subject areas. In any case, if there is even one change in the difference in the leading likelihood that exceeds the threshold value, the main subject calculation unit 202 determines that the reliability of the leading likelihood is low.
[0076] Also, in this embodiment, as condition B in S704, it is conditioned that the reliability of the leading likelihood with respect to the subject area body with the maximum leading likelihood in the current frame is high. However, the condition may be changed, such as that the reliability of the leading likelihood is high for all of not only the subject area with the maximum leading likelihood but also the top N subject areas (the total number of subject areas ≥ N ≥ 2) of the leading likelihoods.
[0077] In this embodiment as well, the same effects as those of the first embodiment can be achieved.
[0078] ●(Third Embodiment) Next, a third embodiment of the present invention will be described. In this embodiment, the leading likelihood is used in the same manner as in the second embodiment, but the method for determining the reliability of the leading likelihood is different. In this embodiment, the determination process for determining the main subject area performed in S308 in FIG. 3 is different from that in the first or second embodiment, but the configuration of the image processing apparatus and other operations may be the same as those in the first or second embodiment. Therefore, hereinafter, the operation of the main subject calculation unit 202 in this embodiment will be mainly described.
[0079] FIG. 10 is a flowchart related to the determination process of the main subject process performed by the main subject calculation unit 202 in this embodiment. In FIG. 10, the same reference numerals as those in FIG. 4 are assigned to the steps for executing the same processes as those in the first embodiment, and the description thereof is omitted.
[0080] In S1002, the main subject calculation unit 202 acquires the leading likelihood and its reliability for each of the subject areas detected in S401. The leading likelihood is the same as that described in the second embodiment.
[0081] Also, the machine learning model used for estimating the leading likelihood in this embodiment is configured and learned to output the reliability of the inference together with the leading likelihood. The reliability is a parameter obtained from the output of, for example, an intermediate layer in the CNN, and is hereinafter referred to as the first reliability. Also, for the subject area i detected in the frame at time t, the leading likelihood estimated is h i (t), h i (t), and the first reliability thereof is λ i (t). λ i (t) is 0 (low reliability) or 1 (high reliability).
[0082] FIG. 11(a) is an example of an image similar to FIG. 5(a), except that the reference numerals of the human subject are different. Further, FIG. 11(b) shows an example of the subject detection result in S401 for the image shown in FIG. 11(a), and the leading likelihood and the first confidence level for each subject area acquired in S1002. Here too, the human head is detected as the subject area, and the image coordinates corresponding to the center of gravity of the subject area are shown as the coordinates of the subject area. Also, the subject IDs 1 to 3 that identify the subject areas correspond to the human subjects 1101 to 1103, respectively.
[0083] In S1003, the main subject calculation unit 202 determines the overall confidence level (referred to as the second confidence level) for the leading likelihood acquired in S1002. The main subject calculation unit 202 calculates the second confidence level based on the first confidence level λ i (t) acquired in S1002 for the current frame and the leading likelihood h i (t) acquired in the past.
[0084] Specifically, the main subject calculation unit 202 calculates the second confidence level ω2 i (t) of the leading likelihood for the subject area i at time t by the following formula.
Equation
[0085] The second confidence level ω2 i (t) becomes a low confidence level (=0) when either the confidence level output by the machine learning model or the confidence level based on the change amount of the difference in the leading likelihood is low (confidence level = 0).
[0086] In S1004, the main subject calculation unit 202 determines the main subject area from the subject areas detected in S401. If the determined main subject area does not exist (no main subject), and the second reliability determined in S1003 for the subject area with the maximum leading likelihood in the current frame is high, the main subject calculation unit 202 determines the subject area with the maximum leading likelihood as the main subject area. On the other hand, if the second reliability determined in S703 for the subject area with the maximum leading likelihood in the current frame is low, the main subject calculation unit 202 does not determine the main subject area.
[0087] If the determined main subject area exists (there is a main subject), the main subject calculation unit 202 determines whether to switch the main subject area based on the leading likelihood obtained in S702 and the second reliability determined in S703. Specifically, the main subject calculation unit 202 determines to switch the main subject area when both of the following condition A and condition B are satisfied, and determines not to switch the main subject if one or more are not satisfied. Condition A: The subject area with the maximum estimated leading likelihood has changed from the previous frame. Condition B: For the subject area with the maximum estimated leading likelihood in the current frame, the second reliability of the leading likelihood is high.
[0088] When the main subject calculation unit 202 determines to switch the main subject area, it determines the subject area with the maximum leading likelihood estimated for the current frame as the new main subject area. On the other hand, when the main subject calculation unit 202 determines not to switch the main subject area, it maintains the determined main subject area.
[0089] According to this embodiment, it is determined whether to switch the subject area in consideration of both the reliability obtained together with the leading likelihood and the reliability based on the change amount of the difference in the leading likelihood. Therefore, the subject area can be switched with higher accuracy than in the second embodiment.
[0090] (Other embodiments) The present invention can also be realized by supplying a program that implements one or more functions of the above-described embodiments to a system or apparatus via a network or a storage medium, and causing one or more processors in a computer of the system or apparatus to read and execute the program. It can also be realized by a circuit (for example, an ASIC) that implements one or more functions.
[0091] The disclosure of this embodiment includes the following image processing apparatus, image processing method, imaging apparatus, and program. (Item 1) An acquisition unit that uses a machine learning model to acquire information regarding the positional relationship in the direction for each of the plurality of subjects from an input image based on image data obtained by photographing a scene in which the plurality of subjects are moving in the same direction; Determination means for determining a main subject region among the regions of the plurality of subjects included in the image data based on the information. The determination means determines whether or not to determine the main subject region based on the reliability of the information. An image processing apparatus characterized by the above. (Item 2) The information is the ranking of the plurality of subjects in the direction, The determination means determines the region of the subject with a specific ranking among the plurality of subjects as the main subject region. The image processing apparatus according to Item 1, characterized by the above. (Item 3) The determination means determines that the reliability of the information is low and does not determine the main subject region when there is a subject among the plurality of subjects whose change in the ranking exceeds a predetermined threshold. The image processing apparatus according to Item 2, characterized by the above. (Item 4) The image data is one frame of a moving image, When there is a subject for which the change between the rank for the current frame and the rank for a past frame exceeds the threshold value, the determination means determines that the reliability of the information is low and does not determine the main subject area, in the image processing apparatus according to item 3. (Item 5) When there is no subject among the plurality of subjects for which the change in the rank exceeds a predetermined threshold value, the determination means determines that the reliability of the information is high and determines the area of the subject with the specific rank as the main subject area, in the image processing apparatus according to any one of items 2 to 4. (Item 6) The image processing apparatus according to any one of items 2 to 5, wherein the specific rank is the first rank. (Item 7) The information is the likelihood that the plurality of subjects are located at the head in the direction, The determination means determines the area of the subject with the maximum likelihood among the plurality of subjects as the main subject area, in the image processing apparatus according to item 1. (Item 8) When there is a subject among the plurality of subjects for which the change in the difference in likelihood from another subject exceeds a predetermined threshold value, the determination means determines that the reliability of the information is low and does not determine the main subject area, in the image processing apparatus according to item 7. (Item 9) The image data is one frame of a moving image, When there is a subject for which the change between the difference in likelihood for the current frame and the difference in likelihood for a past frame exceeds the threshold value, the determination means determines that the reliability of the information is low and does not determine the main subject area, in the image processing apparatus according to item 8. (Item 10) The determination means determines that the reliability of the information is low and does not determine the main subject area if, among the plurality of subjects, there is a subject for which the change in the difference in likelihood with respect to other subjects exceeds a predetermined threshold for the subject having the maximum likelihood in the current frame, in the image processing apparatus according to item 9. (Item 11) The image processing apparatus according to any one of items 8 to 10, wherein the threshold value has a value corresponding to the number of the plurality of subjects. (Item 12) The acquisition means acquires a first reliability that is the reliability of the likelihood together with the likelihood, The determination means determines whether to determine the main subject area based on the first reliability and a second reliability based on the change in the difference in likelihood between the plurality of subjects, in the image processing apparatus according to item 7. (Item 13) The image processing apparatus according to item 12, wherein the determination means does not determine the main subject area when it is determined that at least one of the first reliability and the second reliability is low. (Item 14) The image processing apparatus according to any one of items 7 to 13, wherein the magnitude of the difference in likelihood between the subjects represents the magnitude of the distance between the subjects in the real space. (Item 15) An image processing apparatus according to any one of items 1 to 14, An imaging apparatus comprising: an automatic focus detection means for performing focus detection so as to be in focus on the main subject area determined by the image processing apparatus. (Item 16) An image processing method executed by an image processing apparatus, Using a machine learning model, for each of the plurality of subjects, acquiring information regarding the positional relationship in the direction from an input image based on image data obtained by photographing a scene in which the plurality of subjects are moving in the same direction, Based on the said information, determining a main subject area among the areas of the plurality of subjects included in the image data; The said determining includes determining whether or not to determine the main subject area based on the reliability of the said information; An image processing method characterized by the above. (Item 17) A program for causing a computer to function as each means included in the image processing apparatus according to any one of Items 1 to 14.
[0092] The present invention is not limited to the content of the above-described embodiments, and various changes and modifications are possible without departing from the spirit and scope of the invention. Therefore, claims are attached to disclose the scope of the invention.
Explanation of Signs
[0093] 100... Camera, 101... Main body, 102... Arithmetic device, 120... Lens unit, 201... Control unit, 202... Main subject arithmetic unit
Claims
1. An acquisition means for acquiring information regarding the positional relationship in the direction for each of the plurality of subjects from an input image based on image data obtained by photographing a scene in which the plurality of subjects are moving in the same direction, using a machine learning model; Determination means for determining a main subject area among the areas of the plurality of subjects included in the image data based on the information; The determination means determines whether to determine the main subject area based on the reliability of the information. An image processing apparatus characterized by the above.
2. The information is the ranking of the plurality of subjects in the direction; The determination means determines the area of the subject with a specific ranking among the plurality of subjects as the main subject area. The image processing apparatus according to claim 1, characterized by the above.
3. When there is a subject among the plurality of subjects whose change in the ranking exceeds a predetermined threshold, the determination means determines that the reliability of the information is low and does not determine the main subject area. The image processing apparatus according to claim 2, characterized by the above.
4. The image data is one frame of a moving image; When there is a subject among the plurality of subjects whose change between the ranking for the current frame and the ranking for a past frame exceeds the threshold, the determination means determines that the reliability of the information is low and does not determine the main subject area. The image processing apparatus according to claim 3, characterized by the above.
5. When there is no subject among the plurality of subjects whose change in the ranking exceeds a predetermined threshold, the determination means determines that the reliability of the information is high and determines the area of the subject with the specific ranking as the main subject area. The image processing apparatus according to claim 2, characterized by the above.
6. The image processing apparatus according to claim 2, characterized in that the specific ranking is the first place.
7. The information is the likelihood that the plurality of subjects are located at the head in the direction; The determination means determines the area of the subject with the maximum likelihood among the plurality of subjects as the main subject area. The image processing apparatus according to claim 1, characterized by the above.
8. When there is a subject among the plurality of subjects whose change in the difference in likelihood from other subjects exceeds a predetermined threshold, the determination means determines that the reliability of the information is low and does not determine the main subject area. The image processing apparatus according to claim 7, characterized by the above.
9. The image data is one frame of a video, and when there is a subject for which a change in the difference in likelihoods for the current frame and the difference in likelihoods for a past frame exceeds the threshold, the determination means determines that the reliability of the information is low and does not determine the main subject area. The image processing apparatus according to claim 8, characterized in that.
10. Among the plurality of subjects, for the subject having the maximum likelihood in the current frame, when there is a subject for which a change in the difference in likelihoods from other subjects exceeds a predetermined threshold, the determination means determines that the reliability of the information is low and does not determine the main subject area. The image processing apparatus according to claim 9, characterized in that.
11. The image processing apparatus according to claim 8, characterized in that the threshold has a value corresponding to the number of the plurality of subjects.
12. The acquisition means acquires a first reliability, which is the reliability of the likelihood, together with the likelihood, and the determination means determines whether to determine the main subject area based on the first reliability and a second reliability based on a change in the difference in likelihoods between the plurality of subjects. The image processing apparatus according to claim 7, characterized in that.
13. The image processing apparatus according to claim 12, characterized in that the determination means does not determine the main subject area when it is determined that at least one of the first reliability and the second reliability is low.
14. The image processing apparatus according to claim 7, characterized in that the magnitude of the difference in likelihoods between subjects represents the magnitude of the distance between the subjects in real space.
15. An image processing apparatus according to any one of claims 1 to 14, and an autofocus detection means for performing focus detection so as to be in focus on the main subject area determined by the image processing apparatus. An imaging apparatus, characterized in that it has.
16. An image processing method executed by an image processing apparatus, using a machine learning model, from an input image based on image data obtained by photographing a scene in which a plurality of subjects are moving in the same direction, for each of the plurality of subjects, obtaining information regarding the positional relationship in the direction, and based on the information, determining a main subject area among the areas of the plurality of subjects included in the image data, wherein the determining includes determining whether to determine the main subject area based on the reliability of the information. An image processing method characterized by the following.
17. A program for causing a computer to function as each means included in the image processing apparatus according to any one of Claims 1 to 14.
Citation Information
Patent Citations
Subject tracking device and subject tracking method, and imaging apparatus
JP2022022767A