Electronic apparatus, image processing method, program, and storage medium
The electronic device uses subject and object detection, posture analysis, and depth information to accurately determine the main subject intended by the photographer, addressing the issue of unintended focus when faces are not detectable.
Patent Information
- Application Number
- JP2024027866
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2024-02-27
- Publication Date
- 2025-09-08
AI Technical Summary
Existing methods for determining a main subject in continuous shooting or video shooting fail to accurately identify the subject intended by the photographer when the subject's face is not detectable, leading to unintended subjects being focused on.
An electronic device with first and second detection means to identify subjects and objects, a reliability determination based on subject posture and depth information, and an adjustment mechanism to refine the likelihood of each subject being the main subject.
Enables accurate determination of the subject closest to the photographer's intention as the main subject, even when faces are not detectable, by using posture and depth information to adjust reliability.
Smart Images

Figure 2025130592000001_ABST
Abstract
Description
[Technical Field]
[0001] The present invention relates to an electronic device, an image processing method, a program, and a storage medium, and more particularly to a technique for determining a main subject in a captured image. [Background technology]
[0002] Conventionally, in continuous shooting or video shooting in which a plurality of shots are taken consecutively, when shooting a plurality of moving subjects detected by an imaging device, a main subject is determined from the plurality of subjects, and the determined main subject is kept in focus. Patent Document 1 discloses a method of determining the main subject by detecting a subject (object) and a face to be tracked, setting the main subject based on the position and movement of the tracked target and the position of the face, and performing focus adjustment on the set main subject. [Prior art documents] [Patent documents]
[0003] [Patent Document 1] Japanese Patent Application Publication No. 2018-66889 Summary of the Invention [Problem to be solved by the invention]
[0004] However, with the method disclosed in Patent Document 1, a subject whose face has not been detected cannot be set as the main subject, so if the subject intended by the photographer is facing away, for example, there is a risk that an unintended subject will be determined to be the main subject and continue to be focused on.
[0005] The present invention has been made in consideration of the above-mentioned problems, and has as its object to make it possible to determine, when a plurality of different subjects are detected, the subject that is closest to the photographer's intention as the main subject. [Means for solving the problem]
[0006] In order to achieve the above object, the electronic device of the present invention has a first detection means that detects one or more subjects from images obtained by repeatedly taking photographs, a second detection means that detects a predetermined object from the images, a determination means that determines a reliability indicating the likelihood that each subject is a main subject based on the posture of each subject detected by the first detection means, an adjustment means that adjusts the reliability based on the difference between information regarding the depth distance to each subject and information regarding the depth distance to each object, and a determination means that determines one of the subjects detected by the first detection means as the main subject based on the reliability adjusted by the adjustment means. [Effects of the Invention]
[0007] According to the present invention, when a plurality of different subjects are detected, it is possible to determine the subject that is closest to the photographer's intention as the main subject. [Brief explanation of the drawings]
[0008] [Figure 1] FIG. 1 is a block diagram showing the functional configuration of an imaging apparatus according to an embodiment of the present invention. [Figure 2] FIG. 2 is a block diagram showing the functional configuration of a main subject determination unit in the first embodiment. [Figure 3] 6 is a flowchart of a main subject determination process according to the first embodiment. [Figure 4] FIG. 3 is a conceptual diagram of information acquired by an orientation acquisition unit and an object detection unit in the first embodiment. [Figure 5] FIG. 2 is a diagram showing an example of the structure of a neural network according to the first embodiment. [Figure 6] 5A to 5C are diagrams for explaining an acquisition method in a depth information acquisition unit according to the first embodiment. [Figure 7] 4A and 4B are conceptual diagrams of information acquired by a subject-specific depth information acquisition unit and an object depth information acquisition unit in the first embodiment. [Figure 8] 5A and 5B are graphs showing an adjustment method of the main subject reliability adjustment unit in the first embodiment. [Figure 9]FIG. 10 is a block diagram showing the functional configuration of a main subject determination unit according to a second embodiment. [Figure 10] 10 is a flowchart of a main subject determination process according to the second embodiment. DETAILED DESCRIPTION OF THE INVENTION
[0009] Hereinafter, embodiments will be described in detail with reference to the accompanying drawings. Note that the following embodiments do not limit the scope of the invention claimed. Although multiple features are described in the embodiments, not all of these multiple features are necessarily essential to the invention, and multiple features may be combined arbitrarily. Furthermore, in the accompanying drawings, the same reference numerals are used to designate the same or similar components, and redundant explanations will be omitted.
[0010] In the following embodiments, the present invention will be described as being implemented in an imaging device such as a digital camera, but the present invention can be applied to any electronic device that can be equipped with a subject detection function and an imaging function. Such electronic devices include imaging devices as well as computer devices (personal computers, tablet computers, media players, PDAs, etc.), mobile phones, smartphones, game consoles, robots, drones, drive recorders, etc. These are merely examples, and the present invention can also be applied to other electronic devices.
[0011] First Embodiment Imaging device configuration 1 is a block diagram showing an example of the functional configuration of an image capturing device 100 according to this embodiment. The image capturing device 100 is capable of capturing and recording moving images and still images. The functional blocks within the image capturing device 100 are connected to each other via a bus 160 so that they can communicate with each other. The operation of the image capturing device 100 is realized by a main control unit 151, which is comprised of a central processing unit (CPU) or the like, executing a program to control the functional blocks.
[0012] The photographing lens 101 has a first fixed lens group 102, a zoom lens 111, an aperture 103, a second fixed lens group 121, a focus lens 131, a zoom motor 112, an aperture motor 104, and a focus motor 132. The first fixed lens group 102, the zoom lens 111, the aperture 103, the second fixed lens group 121, and the focus lens 131 constitute a photographing optical system. For convenience, the first fixed lens group 102, the zoom lens 111, the second fixed lens group 121, and the focus lens 131 are each represented by a single lens, but each may be composed of multiple lenses. The photographing lens 101 may also be configured as a detachable interchangeable lens.
[0013] An aperture control unit 105 controls the operation of an aperture motor 104 that drives the aperture 103, and changes the opening diameter of the aperture 103. A zoom control unit 113 controls the operation of a zoom motor 112 that drives the zoom lens 111, and changes the focal length (angle of view) of the photographing lens 101.
[0014] The focus control unit 133 calculates the defocus amount and defocus direction of the photographic lens 101 based on the phase difference between a pair of focus detection signals (image A and image B) obtained from the image sensor 141. The focus control unit 133 then converts the defocus amount and defocus direction into a drive amount and drive direction of the focus motor 132. Based on this drive amount and drive direction, the focus control unit 133 controls the operation of the focus motor 132 and drives the focus lens 131, thereby controlling the focus state of the photographic lens 101. In this way, the focus control unit 133 performs autofocus detection (AF) using a phase difference detection method. The AF method is not limited to the phase difference detection method. For example, the focus control unit 133 may perform contrast detection AF based on a contrast evaluation value obtained from an image signal obtained from the image sensor 141.
[0015] An object image formed on the imaging plane of the image sensor 141 by the photographing lens 101 is converted into an electrical signal (image signal) by a photoelectric conversion element included in each of a plurality of pixels arranged on the image sensor 141. In this embodiment, the image sensor 141 is assumed to have m horizontal rows and n vertical columns (where m and n are multiples) of pixels arranged in a matrix, with each pixel having a plurality of photoelectric conversion elements (photoelectric conversion regions) that share a single microlens. It is assumed here that two photoelectric conversion elements are provided, and a normal image can be obtained by adding the outputs (signal A and signal B) of these two photoelectric conversion elements. Furthermore, two images with parallax can be obtained by separately handling the outputs of the two photoelectric conversion elements. In the following description, a signal obtained by adding the outputs of the two photoelectric conversion elements is referred to as an image signal, and an image based on the image signal is simply referred to as an image. Furthermore, the outputs of the two photoelectric conversion elements are individually handled as a pair of parallax image signals, and the images based on the pair of parallax image signals are referred to as image A and image B, respectively, to distinguish them from one another.
[0016] The imaging control unit 143 controls the reading of signals from the imaging element 141 in accordance with instructions from the main control unit 151. Note that the A signal and the B signal may be read out separately from each pixel, or the A+B signal obtained by adding the A signal and the B signal, or either the A signal or the B signal may be read out, as long as the readings are performed so as to ultimately obtain an image signal and a pair of parallax image signals.
[0017] The signal read out from the imaging element 141 is supplied to a signal processing unit 142, which generates an image signal and a pair of parallax image signals. Then, the signal processing unit 142 performs signal processing such as noise reduction processing, A / D conversion processing, and automatic gain control processing, and outputs the signal to an imaging control unit 143. The imaging control unit 143 stores the image signal and the pair of parallax image signals received from the signal processing unit 142 in a RAM (random access memory) 154.
[0018] The image processing unit 152 applies predetermined image processing to the image signal stored in the RAM 154. The image processing applied by the image processing unit 152 includes, but is not limited to, so-called development processing such as white balance adjustment processing, color interpolation (demosaic) processing, and gamma correction processing, as well as signal format conversion processing and scaling processing. The image processing unit 152 can also generate information regarding subject brightness to be used for automatic exposure control (AE). Information regarding a specific subject region may be supplied from the main subject determination unit 162 and used for white balance adjustment processing, for example. When performing contrast detection AF, the image processing unit 152 may also generate an AF evaluation value. The image processing unit 152 stores the processed image data in the RAM 154.
[0019] When recording image data stored in RAM 154, main control unit 151 generates a data file according to the recording format by, for example, adding a predetermined header to the image data. At this time, main control unit 151 encodes the image data using compression / decompression unit 153 as necessary to compress the amount of information. Main control unit 151 records the generated data file on recording medium 157, such as a memory card.
[0020] Furthermore, when displaying image data stored in RAM 154, main control unit 151 uses image processing unit 152 to scale the image data so that it fits the display size of display unit 150, and then writes the scaled image data to an area of RAM 154 used as a video memory (VRAM area). Display unit 150 reads the image data for display from the VRAM area of RAM 154 and displays it on a display device such as an LCD or an organic EL display.
[0021] The imaging device 100 of this embodiment causes the display unit 150 to function as an electronic viewfinder (EVF) by instantly displaying on the display unit 150 each frame of image captured during standby for still image capture or during video recording. The video images (frame images) displayed when the display unit 150 functions as an EVF are called live view images or through images. Furthermore, when capturing a still image, the imaging device 100 displays the most recently captured still image on the display unit 150 for a certain period of time so that the user can check the capture results. These display operations are also realized under the control of the main control unit 151.
[0022] The operation unit 156 includes operation members such as switches, buttons, keys, a touch panel, and an eye-gaze input device that allow the user to input instructions to the imaging device 100. Inputs made through the operation unit 156 are detected by the main control unit 151 via the bus 160, and the main control unit 151 controls each unit to realize operations according to the inputs.
[0023] The main control unit 151 has one or more programmable processors such as a CPU or MPU, and controls each unit by, for example, loading a program stored in the storage unit 155 into the RAM 154 and executing it, thereby realizing the functions of the imaging device 100. The main control unit 151 also executes AE processing that automatically determines exposure conditions (shutter speed or accumulation time, aperture value, sensitivity) based on information about the brightness of the subject. Information about the brightness of the subject can be acquired from, for example, the image processing unit 152. The main control unit 151 can also determine exposure conditions based on the area of a specific subject, such as a person's face.
[0024] Of the determined exposure conditions, the main control unit 151 notifies the imaging control unit 143 of the accumulation time and sensitivity (gain), and notifies the aperture value to the aperture control unit 105. The imaging control unit 143 controls the operation of the image sensor 141 so that photography is performed according to the notified accumulation time, and sets the gain in the signal processing unit 142. Furthermore, the aperture control unit 105 drives the aperture motor 104 to control the opening diameter of the aperture 103 so that the notified aperture is achieved.
[0025] The battery 159 is managed by the power management unit 158 and supplies power to the entire imaging device 100 . The storage unit 155 stores the programs executed by the main control unit 151, setting values required for executing the programs, GUI data, user setting values, weights and bias values learned by machine learning (described later), etc. For example, when an instruction to transition from a power-off state to a power-on state is given by operating the operation unit 156, the programs stored in the storage unit 155 are read into a part of the RAM 154, and the main control unit 151 executes the programs.
[0026] The depth information acquisition unit 161 calculates depth information from the pair of parallax image signals stored in the RAM 154 and stores the calculated depth information in the RAM 154.
[0027] Main subject determination unit 162 determines the main subject from among the subjects detected in the image of each captured frame. The configuration and operation of main subject determination unit 162 will be described in detail later. The determination result by main subject determination unit 162 is notified to each processing block via bus 160.
[0028] The determination result by the main subject determination unit 162 can be used, for example, to automatically set the focus detection area. As a result, a tracking AF function for a specific subject area can be realized. AE processing can also be performed based on the luminance information of the focus detection area, and image processing (e.g., gamma correction processing, white balance adjustment processing, etc.) can also be performed based on the pixel values of the focus detection area. Furthermore, the main control unit 151 may superimpose on the displayed image an indicator (e.g., a rectangular frame surrounding the area) indicating the current main subject or an area indicating a subject or unique object detected by the main subject determination unit 162 through processing described below.
[0029] Main subject detection processing Next, the main subject determination process in this embodiment will be described with reference to FIGS. Fig. 2 is a block diagram showing the functional configuration of main subject determination unit 162, and Fig. 3 is a flowchart showing the main subject determination process. Note that, unless otherwise specified, each process in this flowchart is realized by each unit of image capture device 100 operating under the control of main control unit 151. Also, at the start of this flowchart, it is assumed that image capture device 100 is powered on and in live view imaging mode, and that a command to start capturing (recording) still images or videos can be given by operating operation unit 156. Note that, in the following explanation, a ball game involving multiple players is described as a shooting scene that is the target of the main subject determination process, but the shooting scene to which this embodiment can be applied is not limited to this.
[0030] First, in S301, the main control unit 151 controls the imaging control unit 143 to perform imaging with the imaging element 141, and stores in the RAM 154 the image signal A / D converted by the signal processing unit 142 and the pair of parallax image signals.
[0031] In S302, the subject detection unit 201 of the main subject determination unit 162 detects a subject (person) on the image of the image signal acquired in S301. Any known subject detection method may be used, and for example, the subject may be detected using a feature extraction process of a specific subject using CNN (Convolutional Neural Networks), or the subject to be detected may be registered in advance as a template and detected by template matching.
[0032] In S303, the posture acquisition unit 202 estimates the posture of each of the subjects detected by the subject detection unit 201 and acquires posture information. The content of the posture information to be acquired is determined according to the type of subject. In this case, since the subject is a person, the posture acquisition unit 202 acquires the positions of multiple joints of the person as posture information.
[0033] In S304, object detection unit 203 detects an object (hereinafter referred to as a "unique object") that is unique to the scene on the image of the image signal acquired in S301, and acquires the two-dimensional coordinates and size of the detected unique object on the image. The type of unique object to be detected is determined based on the shooting scene, and in this case, since the shooting scene is a ball game, object detection unit 203 detects a ball as the unique object.
[0034] Any known method can be used for object detection and pose estimation, such as the methods described in Reference 1 (Redmon, Joseph, et al., "You only look once: Unified, real-time object detection.", Proceedings of the IEEE conference on computer vision and pattern recognition, 2016) and Reference 2 (Cao, Zhe, et al., "Realtime multi-person 2D pose estimation using part affinity fields.", Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, 2017).
[0035] FIG. 4 is a conceptual diagram of information acquired by the posture acquisition unit 202 and the object detection unit 203. FIG. 4(a) shows an example of an image to be processed, which is a scene in which a subject 402 is about to kick a ball 403 and a subject 401 is about to interfere with the subject 402's action of kicking the ball 403. The important subject in this shooting scene is the subject 402 about to kick the ball. In this embodiment, posture information of the subject and information about the unique object are used to determine a main subject that is likely to be intended by the user as a target for imaging control. On the other hand, the subject 401 is a non-main subject. Here, a non-main subject refers to a subject other than the main subject.
[0036] FIG. 4(b) is a diagram showing an example of posture information of subject 401 and subject 402, as well as the position and size of ball 403. Joint 411 represents each joint of subject 401, and joint 412 represents each joint of subject 402. FIG. 4(b) shows an example in which the positions of the top of the head, neck, shoulders, elbows, wrists, hip joints, knees, and ankles are acquired as joints, but only some of these joints may be acquired, or positions other than joints may be acquired. Furthermore, posture information may be information on axes connecting joints, etc., in addition to joint positions, and any information representing the posture of the subject may be used. In the following explanation, a case in which joint positions are acquired as posture information will be described.
[0037] Orientation acquisition unit 202 acquires the two-dimensional coordinates (x, y) of joint 411 and joint 412 in the image. Here, the unit of (x, y) is pixels. Center of gravity position 413 represents the center of gravity position of ball 403, and arrow 414 represents the size of ball 403 in the image. Object detection unit 203 acquires the two-dimensional coordinates (x, y) of the center of gravity position of ball 403 in the image, and the number of pixels indicating the diameter, which is the size of ball 403 in the image.
[0038] In S305, the reliability calculation unit 204 calculates a reliability (probability) representing the likelihood of each subject being the main subject based on the coordinates of the joints acquired by the orientation acquisition unit 202 and at least one of the coordinates and size of the unique object acquired by the object detection unit 203. This reliability corresponds to the degree of possibility that the subject is the main subject of the image to be processed. A method for calculating the reliability will be described later. In this embodiment, a case will be described in which the probability that the subject is the main subject of the image to be processed is used as the reliability representing the likelihood of the subject being the main subject, but a value other than probability may also be used. For example, the reciprocal of the distance between the center of gravity of the subject and the center of gravity of the unique object can be used as the reliability.
[0039] Here, a description will be given of a method for calculating the probability that an object is likely to be the main subject based on the coordinates of each joint and the coordinates and size of each unique object, which is performed by the reliability calculation unit 204. In the following explanation, a case where a neural network, which is one of the machine learning techniques, is used will be described.
[0040] FIG. 5 is a diagram showing an example of the structure of a neural network. In FIG. 5, 501 denotes an input layer, 502 denotes a hidden layer, 503 denotes an output layer, 504 denotes neurons, and 505 denotes the connections between the neurons 504. Here, for convenience of illustration, only representative neurons and connecting lines are numbered. The number of neurons 504 in the input layer 501 is equal to the dimension of the input data, and the number of neurons in the output layer 503 is two. This corresponds to a two-class classification problem of determining whether or not an object is the main subject.
[0041] The line 505 connecting the i-th neuron 504 in the input layer 501 and the j-th neuron 504 in the hidden layer 502 has a weight w ji is given, and the value z j can be expressed by the following equations (1) and (2). TIFF2025130592000002.tif766h(z) = max(z,0) …(2) In formula (1), x i represents the value input to the i-th neuron 504 in the input layer 501. The sum is taken over all neurons 504 in the input layer 501 that are connected to the j-th neuron. j is called the bias, and is a parameter that controls the likelihood of firing the j-th neuron 504. The function h defined in equation (2) is an activation function called ReLU (Rectified Linear Unit). It is also possible to use other functions as the activation function, such as a sigmoid function.
[0042] Also, the value y kis given by the following equations (3) and (4). TIFF2025130592000003.tif770TIFF2025130592000004.tif961In equation (3), z j represents the value output by the j-th neuron 504 in the hidden layer 502, where i, k = 0, 1. 0 corresponds to a non-main subject, and 1 corresponds to the main subject. The sum is taken for all neurons in the hidden layer 502 that are connected to the k-th neuron. The function f defined by equation (4) is called a softmax function, and outputs a probability value that belongs to the k-th class. In this embodiment, f(y1) is used as the probability that the object is likely to be the main subject.
[0043] During learning, the coordinates of the subject's joints and the coordinates and sizes of the eigenobjects are input. Then, all weights and biases are optimized to minimize a loss function that uses the output probability and correct label. Here, the correct label takes two values: "1" for the main subject and "0" for a non-main subject. The loss function L can be the binary cross-entropy expressed in the following equation (5). TIFF2025130592000005.tif6110In equation (5), the subscript m represents the index of the object to be learned. m is the probability value output from the neuron 504 with k=1 in the output layer 503, and t m is the correct label. The loss function can be any function other than equation (5) that can measure the degree of agreement with the correct label, such as mean square error. By optimizing based on equation (5), it is possible to determine the weights and biases so that the correct label and the output probability value are closer.
[0044] The learned weights and bias values (learned data) are saved in advance in the storage unit 155 and stored in the RAM 154 as needed. A plurality of types of weights and bias values may be prepared depending on the scene. The reliability calculation unit 204 uses the learned weights and biases (results of machine learning performed in advance) to output the probability value f(y1) based on equations (1) to (4).
[0045] During learning, the state before moving on to an important action can be learned as the state of the main subject. For example, in the case of kicking a ball, the state in which the main subject swings his / her leg up to kick the ball can be learned as one of the states of the main subject. The reason for adopting this configuration is that when the main subject actually performs an important action, it is necessary to accurately control the image capture device 100. For example, if the reliability (probability) corresponding to the main subject exceeds a preset threshold, control to automatically record images or videos (recording control) can be initiated, allowing the user to capture important moments without missing them. In this case, information on the typical time from the state to be learned until the important action occurs may be used to control the image capture device 100.
[0046] Although a method for calculating reliability (probability) using a neural network has been described above, other machine learning methods such as support vector machines and decision trees may be used as long as they can classify an image as to whether it is the main subject or not. Furthermore, a function that outputs reliability or probability based on a certain model may be constructed, without being limited to machine learning.
[0047] In the scene shown in FIG. 4, both subject 401 and subject 402 are in a posture with their legs raised, and there is no significant difference in their coordinate distance from ball 403, nor is there any significant difference in their detected reliability.
[0048] In S306, the depth information acquisition unit 161 generates depth information using a pair of parallax image signals. Depth information is a type of information relating to the distance of each pixel in the depth direction relative to the position of the imaging device 100, and is also called a distance map, depth map, depth image, or distance image. Note that the depth information may be generated without using a pair of parallax image signals. For example, the depth information for each pixel may be generated by determining, for each pixel, the position of the focus lens 131 at which the contrast evaluation value is maximized.
[0049] Here, a method for calculating distance information to the subject at each pixel as an example of depth information will be described with reference to FIG. Assuming that image A 601 and image B 602 have been obtained, it can be seen that the light beam is refracted as shown by the solid line from the focal length of the photographing lens 101 and the distance information between the focus lens 131 and the image sensor 141. Therefore, it can be seen that the subject in focus is at position 605. Similarly, if image B 603 is obtained for image A 601, it can be seen that the subject in focus is at position 606, and if image B 604 is obtained, it can be seen that the subject in focus is at position 607. As described above, for each pixel, it is possible to calculate distance information to the subject at that pixel position from the relative position between image A containing that pixel and the corresponding image B.
[0050] For example, assume that an image A 601 and an image B 604 have been obtained in Fig. 6. In this case, a distance 608 from a pixel 609 at the midpoint, which corresponds to half the amount of shift between the images, to a position 607, or a defocus amount corresponding to the distance 608, is stored as distance information for the pixel 609. In this way, distance information can be obtained for each pixel, and depth information can be generated.
[0051] Alternatively, depth information may be generated by dividing an image into minute regions and calculating the defocus amount for each minute region. In this case, images A and B are generated from pixels included in the minute region, and their phase difference (image shift amount) is detected by correlation calculation and converted into a defocus amount. In this case, the generated depth information also indicates distance information for each pixel, but pixels included in the same minute region have the same distance information. The depth information acquisition unit 161 supplies the generated depth information to the subject depth information acquisition unit 205 and the object depth information acquisition unit 206.
[0052] In the above description, a case where a defocus amount, which is distance information of a subject or object, is acquired as depth information has been described in detail, but the present invention is not limited to this. As mentioned above, depth information includes a distance map, a depth map, a depth image, and a distance image. For example, a value obtained by normalizing the defocus amount by a focal depth (e.g., 1Fδ, where F is the aperture value and δ is the allowable circle of confusion diameter) may be acquired. Furthermore, for example, the defocus amount may be an actual distance indicating the actual distance from the image capture device 100 to the subject, or information indicating the image shift amount used to derive the defocus amount. Furthermore, the processes of S302 and S303, the process of S304, and the process of S306 may be performed in parallel. Next, in S307 , the subject depth information acquisition unit 205 acquires depth information of each subject detected by the subject detection unit 201 using the depth information acquired by the depth information acquisition unit 161 .
[0053] Furthermore, in S308 , the object depth information acquisition unit 206 acquires depth information of the unique object detected by the object detection unit 203 using the depth information acquired by the depth information acquisition unit 161 . The processing of S307 and the processing of S308 may be performed in parallel.
[0054] Fig. 7 is a conceptual diagram showing the results of applying processing by the subject depth information acquisition unit 205 and the object depth information acquisition unit 206 to the subjects 401 and 402 and the ball 403 in Fig. 4. The diagonal lines and shading in Fig. 7 each represent distance information, with the diagonal lines indicating closer to the image capture device 100. The subject 402 has the same depth as the ball 403, and the subject 401 is located further back (farther from the image capture device 100) than the subjects 402 and the ball 403.
[0055] In S309, the reliability adjustment unit 207 adjusts the reliability calculated by the reliability calculation unit 204, using the depth information acquired by the subject depth information acquisition unit 205 and the object depth information acquisition unit 206. For example, the reliability (probability) f(y1) output based on equations (1) to (4) is adjusted by multiplying it by a coefficient k corresponding to the absolute value of the difference between the subject distance information and the object distance information.
[0056] Fig. 8 is a graph showing the relationship between the coefficient (amount of change) and the absolute value of the difference between the distance information of the subject and the distance information of the specific object. As shown in Fig. 8, the coefficient takes a value between 0 and 1. When the difference is 0, the coefficient k is set to 1, and as the difference increases, the coefficient k approaches 0, thereby decreasing the reliability (probability).
[0057] The slope of the graph can also be changed according to the distance between the unique object and the subject on the image. d(n) is the absolute value of the difference in distance information between the subject and the unique object when the coefficient k is 0 when the distance between the subject and the unique object on the image is n, and the straight line 801 is the graph at that time. Similarly, d(m) is the absolute value of the difference in distance information between the subject and the unique object when the coefficient k is 0 when the distance between the subject and the unique object on the image is m, and the straight line 802 is the graph at that time. In this embodiment, as the distance between the subject and the object on the image decreases, the absolute value of the difference in distance information between the subject and the unique object when the coefficient k is 0 increases, thereby making the slope of the coefficient k gentler. In the example shown in FIG. 8, the distance is m <nとなる。
[0058] Furthermore, if the reliability f(y1) is equal to or greater than a predetermined threshold, the coefficient k may not be multiplied. The threshold may also be adjusted according to the absolute value of the difference in distance information between the subject and the object.
[0059] In this embodiment, the relationship between the coefficient and the absolute value of the difference in distance information between the subject and the object has been described as a linear function (straight line), but it may also be a logarithmic function or an exponential function. Also, in this embodiment, the minimum value of the coefficient k is set to 0, but the minimum value may also be set to a value greater than 0.
[0060] In the example shown in FIG. 7, the absolute value of the difference in distance information between subject 402 and ball 403 is smaller than that of subject 401, and therefore, as a result of the reliability adjustment, the rate of decrease in the reliability of the main subject of subject 402 is lower than the rate of decrease in the reliability of the main subject of subject 401.
[0061] In S310, the main subject determination unit 208 determines the main subject of the image acquired in S301 based on the reliability adjusted by the reliability adjustment unit 207, stores the main subject information of the determined main subject in RAM 154, and ends the processing.
[0062] 4, it is assumed that the subject with the highest reliability among the subjects whose reliability has been adjusted by reliability adjustment unit 207 is subject 402. At this time, if the current main subject stored in RAM 154 is subject 401, the main subject is changed from subject 401 to subject 402 with the highest reliability.
[0063] When changing the main subject, hysteresis may be provided, such as changing the main subject from subject 401 to subject 402 when subject 402 has the highest reliability a predetermined number of times in chronological order. In this case, the timing to change to subject 402 is determined using past main subject information stored in RAM 154. Also, the main subject may be changed when the distance on the coordinates between the current main subject and the subject with the highest reliability in the current frame image is constant. Also, the main subject may be changed using any chronological information.
[0064] As described above, according to the first embodiment, the image capturing device 100 acquires posture information, unique objects, and depth information for each of the multiple subjects detected from the processing target image. Then, for each of the multiple subjects, the image capturing device 100 calculates a reliability corresponding to the degree of likelihood that the subject is the main subject of the processing target image based on the posture information and unique objects. Then, by adjusting the calculated reliability using the depth information, the image capturing device 100 determines the main subject of the processing target image from among the multiple subjects based on the multiple reliabilities adjusted for the multiple subjects. This makes it possible to determine a main subject that is likely to meet the user's intention in an image containing multiple subjects.
[0065] In the above example, the main subject is determined for each frame image, but it may be determined every predetermined number of frame images.
[0066] <Second embodiment> Next, a second embodiment of the present invention will be described. Note that the same reference numerals will be used for the configurations and processes common to the first embodiment, and the description will be omitted.
[0067] The second embodiment will be described with reference to Fig. 9 and Fig. 10. Fig. 9 is a block diagram showing the functional configuration of main subject determination unit 162 in the second embodiment, and Fig. 10 is a flowchart of main subject determination processing in the second embodiment.
[0068] After the depth information is acquired by the depth information acquisition unit 161 in S306, in S1001 the reliability calculation unit 901 calculates the reliability (probability) of each subject, which indicates the likelihood of the subject being the main subject, based on the coordinates of the joints estimated by the posture acquisition unit 202, and at least one of the coordinates and size of the unique object acquired by the object detection unit 203, and the depth information acquired by the depth information acquisition unit 161. The method of calculating the probability is the same as in the first embodiment, and therefore description thereof will be omitted.
[0069] In the first embodiment, the coordinates of the joints of the subject and the coordinates and sizes of the eigenobjects are used during learning, and the loss function L is optimized based on equation (5). In contrast, in the second embodiment, depth information acquired by the depth information acquisition unit 161 is additionally input, and the loss function L is optimized based on equation (5).
[0070] In S1002, the main subject determination unit 902 determines the main subject of the image acquired in S301 based on the reliability calculated by the reliability calculation unit 901. Note that the main subject determination method is the same as in the first embodiment, and therefore description thereof will be omitted.
[0071] As described above, according to the second embodiment, by using depth information as a learning input in advance, the image capturing device 100 calculates a reliability corresponding to the degree of possibility that each of a plurality of subjects is a main subject based on the posture information, unique object, and depth information. Then, the calculated reliability is used to determine the main subject of the processing target image from among the plurality of subjects. This makes it possible to determine a main subject that is likely to meet the user's intention in an image containing multiple subjects.
[0072] <Modification> In the first and second embodiments described above, it has been explained that posture information of the detected subject is obtained without error and with an appropriate reliability. However, it is possible that, for example, if the movement of the main subject intended by the photographer is large and quick, acquisition of posture information may fail and an appropriate reliability may not be obtained. For example, in the scene shown in FIG. 4, if the subject 402 moves quickly and acquisition of posture information fails, the reliability of the subject 402 will be lower than that of the subject 401.
[0073] In this modified example, as a process to deal with such a case, when subject 402 is stored as the main subject in the main subject history stored in RAM 154 and subject 401 is selected as the main subject based on reliability, the following determination is made: That is, the absolute value of the difference in distance information between subject 401 and ball 403 is compared with a threshold. If the absolute value of the difference is equal to or greater than the threshold, subject 402 is maintained as the main subject regardless of reliability, and if it is smaller than the threshold, the main subject is changed to subject 401.
[0074] By controlling in this way, even if acquisition of posture information fails, frequent changes of the main subject can be prevented.
[0075] <Other embodiments> The present invention may be applied to a system made up of a plurality of devices, or to an apparatus made up of a single device.
[0076] The present invention can also be realized by supplying a program that realizes one or more functions of the above-described embodiments to a system or device via a network or a storage medium, and having one or more processors in the computer of the system or device read and execute the program.The present invention can also be realized by a circuit (e.g., ASIC) that realizes one or more functions.
[0077] <Summary> The disclosure of this embodiment includes the following configuration.
[0078] (Item 1) a first detection means for detecting one or more subjects from images obtained by repeatedly photographing; a second detection means for detecting a predetermined object from the image; a determination means for determining a reliability indicating a likelihood that each subject is a main subject based on the posture of each subject detected by the first detection means; an adjustment means for adjusting the reliability based on a difference between information about the distance to each subject in the depth direction and information about the distance to the object in the depth direction; a determining means for determining one of the subjects detected by the first detecting means as a main subject based on the reliability adjusted by the adjusting means; An electronic device comprising: (Item 2) The electronic device described in item 1, characterized in that the adjustment means reduces the decrease in reliability when the difference is a first value compared to when the difference is a second value greater than the first value. (Item 3) The electronic device described in item 2 is characterized in that the adjustment means reduces the rate of reduction according to the difference when the distance in the image between each subject and the object is a first distance more slowly than when the distance in the image is a second distance greater than the first distance. (Item 4) 4. The electronic device according to any one of items 1 to 3, wherein the adjustment unit does not adjust the reliability when the reliability is equal to or greater than a predetermined threshold value. (Item 5) the adjustment means adjusts the reliability by multiplying the reliability by a coefficient determined based on the difference; 5. The electronic device according to any one of items 1 to 4, wherein the coefficient is a value between 0 and 1. (Item 6) 6. The electronic device according to item 5, wherein the coefficient is defined by one of a linear function, a logarithmic function, and an exponential function of the difference. (Item 7) The electronic device described in any one of items 1 to 6, characterized in that the adjustment means obtains information about the distance in the depth direction for each pixel of the image, and obtains information about the distance to each of the subjects and the object based on the obtained information about the distance for each pixel. (Item 8) The electronic device described in any one of items 1 to 6, characterized in that the adjustment means obtains information about the distance in the depth direction for each of a plurality of regions obtained by dividing the image, and obtains information about the distance to each of the subjects and the object based on the obtained information about the distance for each region. (Item 9) 9. The electronic device described in any one of items 1 to 8, characterized in that the information about the distance includes a defocus amount, a value obtained by normalizing the defocus amount by a focal depth, a distance from the electronic device, and an image shift amount. (Item 10) 10. The electronic device according to any one of items 1 to 9, wherein the determining unit determines the reliability by machine learning. (Item 11) Item 11. The electronic device according to item 10, wherein the machine learning includes a neural network, a support vector machine, and a decision tree that have been trained on the subject. (Item 12) a first detection means for detecting one or more subjects from images obtained by repeatedly photographing; a second detection means for detecting a predetermined object from the image; an acquisition means for obtaining information on the distance in the depth direction to each subject detected by the first detection means and information on the distance in the depth direction to the object; a determination means for determining a reliability indicating a degree of possibility that each subject is a main subject based on the posture of the subject using machine learning; determining means for determining one of the subjects detected by the first detection means as a main subject based on the reliability; The electronic device is characterized in that the determination means uses, in the machine learning, learning data optimized using information regarding the distance of the main subject previously acquired by the acquisition means. (Item 13) Item 13. The electronic device according to item 12, further comprising a learning means for optimizing learning data used in the machine learning using information about the distance acquired by the acquisition means. (Item 14) The camera further includes a storage means for storing the main subject determined by the determination means, 13. The electronic device described in any one of items 1 to 12, wherein the determining unit selects one of the subjects detected by the first detecting unit as a main subject based on the reliability, and determines the main subject based on the selected main subject and the history of main subjects stored in the storage unit. (Item 15) Item 15. The electronic device described in item 14, characterized in that when the main subject stored in the storage means is detected by the first detection means and the stored main subject is different from the selected main subject, the determination means determines the main subject stored in the storage means to be the main subject if the difference in information regarding the depth direction distance between the selected main subject and the object is greater than the difference in depth direction distance between the main subject stored in the storage means and the object. (Item 16) Item 15. The electronic device described in item 14, characterized in that when the main subject stored in the storage means is detected by the first detection means and the stored main subject is different from the selected main subject, the determination means determines the selected subject as the main subject if the same subject is selected as the main subject a predetermined number of times in succession. (Item 17) 17. The electronic device according to any one of items 1 to 16, wherein the predetermined object is determined according to a scene in which the image is to be captured. (Item 18) 18. The electronic device according to any one of items 1 to 17, further comprising an imaging means for capturing the image. (Item 19) a first detection step of detecting one or more subjects from images obtained by repeatedly photographing; a second detection step of detecting a predetermined object from the image; a determination step of determining a reliability indicating a likelihood that each subject is a main subject based on the posture of each subject detected in the first detection step; an adjustment step of adjusting the reliability based on a difference between information about the distance to each subject in the depth direction and information about the distance to the object in the depth direction; a determination step of determining one of the subjects detected in the first detection step as a main subject based on the adjusted reliability; An image processing method comprising: (Item 20) a first detection step of detecting one or more subjects from images obtained by repeatedly photographing; a second detection step of detecting a predetermined object from the image; an acquisition step of obtaining information on the distance in the depth direction to each subject detected in the first detection step and information on the distance in the depth direction to the object; a determining step of determining a reliability indicating a high possibility that each subject is a main subject based on the posture of the subject using machine learning; a determination step of determining one of the subjects detected in the first detection step as a main subject based on the reliability, An image processing method characterized in that in the determination step, learning data optimized using information regarding the distance of the main subject previously acquired in the acquisition step is used in the machine learning. (Item 21) A program for causing a computer to function as each means of the electronic device described in any one of items 1 to 17. (Item 22) 22. A computer-readable storage medium storing the program according to item 21.
[0079] The invention is not limited to the above-described embodiments, and various changes and modifications can be made without departing from the spirit and scope of the invention. Accordingly, the following claims are appended to apprise the public of the scope of the invention. [Explanation of symbols]
[0080] 100: imaging device, 150: display unit, 151: main control unit, 152: image processing unit, 154: RAM, 155: storage unit, 161: depth information acquisition unit, 162: main subject determination unit, 201: subject detection unit, 202: posture acquisition unit, 203: object detection unit, 204, 901: reliability calculation unit, 205: subject depth information acquisition unit, 206: object depth information acquisition unit, 207: reliability adjustment unit, 208, 902: main subject determination unit
Claims
1. a first detection means for detecting one or more subjects from images obtained by repeatedly photographing; a second detecting means for detecting a predetermined object from the image; a determination means for determining a reliability indicating a likelihood that each subject is a main subject based on the posture of each subject detected by the first detection means; an adjustment means for adjusting the reliability based on a difference between information about the distance to each subject in the depth direction and information about the distance to the object in the depth direction; a determining means for determining one of the subjects detected by the first detecting means as a main subject based on the reliability adjusted by the adjusting means; An electronic device comprising:
2. 2. The electronic device according to claim 1, wherein the adjustment means reduces the decrease in the reliability when the difference is a first value compared to when the difference is a second value greater than the first value.
3. The electronic device according to claim 2, characterized in that the adjustment means makes the rate of decrease according to the difference gentler when the distance in the image between each subject and the object is a first distance than when the distance in the image is a second distance greater than the first distance.
4. 2. The electronic device according to claim 1, wherein the adjustment unit does not adjust the reliability when the reliability is equal to or greater than a predetermined threshold value.
5. the adjustment means adjusts the reliability by multiplying the reliability by a coefficient determined based on the difference; 2. The electronic device according to claim 1, wherein the coefficient is a value greater than or equal to 0 and less than or equal to 1.
6. 6. The electronic device according to claim 5, wherein the coefficient is defined by one of a linear function, a logarithmic function, and an exponential function of the difference.
7. The electronic device according to claim 1, characterized in that the adjustment means obtains information regarding the distance in the depth direction for each pixel of the image, and obtains information regarding the distance to each of the subjects and the object based on the obtained information regarding the distance for each pixel.
8. The electronic device according to claim 1, characterized in that the adjustment means calculates information regarding the depth distance for each of a plurality of regions into which the image is divided, and calculates information regarding the distance to each of the subjects and the object based on the calculated information regarding the distance for each region.
9. 2. The electronic device according to claim 1, wherein the information about the distance includes a defocus amount, a value obtained by normalizing the defocus amount by a focal depth, a distance from the electronic device, and an image shift amount.
10. The electronic device according to claim 1 , wherein the determining unit determines the reliability by machine learning.
11. The electronic device according to claim 10 , wherein the machine learning includes a neural network, a support vector machine, and a decision tree that have been trained on the subject.
12. The camera further includes a storage means for storing the main subject determined by the determination means, 2. The electronic device according to claim 1, wherein the determining unit selects one of the subjects detected by the first detecting unit as a main subject based on the reliability, and determines the main subject based on the selected main subject and the history of main subjects stored in the storage unit.
13. 13. The electronic device according to claim 12, wherein when the main subject stored in the storage means is detected by the first detection means and the stored main subject is different from the selected main subject, the determination means determines the main subject stored in the storage means to be the main subject if a difference in information regarding the distance in the depth direction between the selected main subject and the object is greater than a difference in the distance in the depth direction between the main subject stored in the storage means and the object.
14. 13. The electronic device according to claim 12, wherein when the main subject stored in the storage means is detected by the first detection means and the stored main subject is different from the selected main subject, the determination means determines the selected subject as the main subject when the same subject has been selected as the main subject a predetermined number of times in succession.
15. 2. The electronic device according to claim 1, wherein the predetermined object is determined according to a scene in which the image is to be captured.
16. a first detection means for detecting one or more subjects from images obtained by repeatedly photographing; a second detecting means for detecting a predetermined object from the image; an acquisition means for obtaining information on the distance in the depth direction to each subject detected by the first detection means and information on the distance in the depth direction to the object; a determination means for determining a reliability indicating a degree of possibility that each subject is a main subject based on the posture of the subject using machine learning; determining means for determining one of the subjects detected by the first detection means as a main subject based on the reliability; The electronic device is characterized in that the determination means uses, in the machine learning, learning data optimized using information regarding the distance of the main subject previously acquired by the acquisition means.
17. 17. The electronic device according to claim 16, further comprising: a learning unit that optimizes learning data used in the machine learning by using the information about the distance acquired by the acquisition unit.
18. The camera further includes a storage means for storing the main subject determined by the determination means, 17. The electronic device according to claim 16, wherein the determining unit selects one of the subjects detected by the first detecting unit as a main subject based on the reliability, and determines the main subject based on the selected main subject and the history of main subjects stored in the storage unit.
19. 19. The electronic device according to claim 18, wherein, when the main subject stored in the storage means is detected by the first detection means and the stored main subject is different from the selected main subject, the determination means determines the main subject stored in the storage means to be the main subject if a difference in information regarding the distance in the depth direction between the selected main subject and the object is greater than a difference in the distance in the depth direction between the main subject stored in the storage means and the object.
20. 19. The electronic device according to claim 18, wherein when the main subject stored in the storage means is detected by the first detection means and the stored main subject is different from the selected main subject, the determination means determines the selected subject as the main subject when the same subject has been selected as the main subject a predetermined number of times in succession.
21. 17. The electronic device according to claim 16, wherein the predetermined object is determined according to a scene in which the image is to be captured.
22. 22. The electronic device according to claim 1, further comprising an image capturing unit for capturing the image.
23. a first detection step of detecting one or more subjects from images obtained by repeatedly photographing; a second detection step of detecting a predetermined object from the image; a determination step of determining a reliability indicating a likelihood that each subject is a main subject based on the posture of each subject detected in the first detection step; an adjustment step of adjusting the reliability based on a difference between information about the distance to each subject in the depth direction and information about the distance to the object in the depth direction; a determination step of determining one of the subjects detected in the first detection step as a main subject based on the adjusted reliability; An image processing method comprising:
24. a first detection step of detecting one or more subjects from images obtained by repeatedly photographing; a second detection step of detecting a predetermined object from the image; an acquisition step of obtaining information on the distance in the depth direction to each subject detected in the first detection step and information on the distance in the depth direction to the object; a determining step of determining a reliability indicating a high possibility that each subject is a main subject based on the posture of the subject using machine learning; a determination step of determining one of the subjects detected in the first detection step as a main subject based on the reliability, An image processing method characterized in that in the determination step, learning data optimized using information regarding the distance of the main subject previously acquired in the acquisition step is used in the machine learning.
25. A program for causing a computer to function as each of the means of the electronic device according to any one of claims 1 to 21.
26. A computer-readable storage medium storing the program according to claim 25.
Citation Information
Patent Citations
Imaging apparatus
JP2018066889A