Focus adjustment device and method, image-capturing device, program and recording medium
Patent Information
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- CANON KK
- Filing Date
- 2023-06-13
- Publication Date
- 2026-06-16
AI Technical Summary
Conventional autofocus systems in imaging devices face inaccuracies in predicting focus position, leading to potential deterioration of image focus due to misprediction, especially with moving subjects.
The system detects predetermined parts of a subject, acquires and stores their focus states, predicts future focus states based on historical data, and adjusts focus only when the predicted difference meets a predetermined condition to minimize misprediction.
This approach effectively suppresses focus deterioration by accurately predicting and adjusting focus to maintain image clarity, even with moving subjects.
Smart Images

Figure 00000000_0000_ABST
Abstract
Description
[Technical field]
[0001] The present invention relates to a focus adjustment device and method, an imaging device, a program, and a storage medium, and more particularly to a technique for predicting a focal position of a subject. [Background technology]
[0002] In the autofocus (AF) control of conventional imaging devices, focus detection is generally performed in the area to be focused on within the imaging screen, and the focus lens is driven based on the result. In recent years, the pixels of imaging devices have become finer, making it possible to capture images with higher resolution, and therefore there is a demand for more accurate focus adjustment control.
[0003] On the other hand, Patent Document 1 discloses a method for predicting the future focal position of a moving subject by approximating the time-series change in focal position accompanying the movement of the subject using a pre-designed function. [Prior art documents] [Patent documents]
[0004] [Patent Document 1] JP 2001-21794 A Summary of the Invention [Problem to be solved by the invention]
[0005] However, depending on the movement of the subject, the prediction of the focus position may not be accurate, and when the prediction is incorrect, it may not be possible to obtain an image in a desired focus state.
[0006] The present invention has been made in consideration of the above problems, and aims to suppress deterioration of the focus state of an image due to an incorrect prediction of the focus position when performing focus adjustment by predicting a future focus position. [Means for solving the problem]
[0007] In order to achieve the above-mentioned object, the focus adjustment device of the present invention has a detection means for detecting a predetermined first portion and a second portion of a predetermined subject from an image obtained by photographing, an acquisition means for acquiring the focus states of the first portion and the second portion detected by the detection means, a storage means for storing the focus states of the first portion and the second portion acquired by the acquisition means, a prediction means for predicting the focus states of the first portion and the second portion at a third time after the first time from the focus states of the first portion and the second portion of an image obtained at a first time and the focus states of the first portion and the second portion of an image obtained at a second time before the first time stored in the storage means, and a focus adjustment means for performing focus adjustment processing, wherein the focus adjustment means performs the focus adjustment processing based on the predicted focus states of the first portion and the second portion when a difference between the focus state of the first portion and the focus state of the second portion predicted by the prediction means satisfies a predetermined condition. Effect of the Invention
[0008] According to the present invention, when focus adjustment is performed by predicting a future focus position, it is possible to suppress deterioration of the focus state of an image caused by an incorrect prediction of the focus position. [Brief description of the drawings]
[0009] [Figure 1] 1 is a block diagram showing a configuration of an imaging apparatus according to an embodiment of the present invention. [Diagram 2] 5 is a flowchart showing a predictive focus adjustment process in the first embodiment. [Diagram 3] 10 is a flowchart showing a predictive focus adjustment process in the second embodiment. [Figure 4] 5 is a diagram showing an example of the concept of an absolute difference value of defocus predicted values in the first embodiment. [Diagram 5] FIG. 11 is a diagram showing an example of the concept of a predicted range of a defocus range in the second embodiment. [Figure 6]FIG. 11 is a diagram showing the configuration of a defocus range inferring device according to the second embodiment. DETAILED DESCRIPTION OF THE PREFERRED EMBODIMENTS
[0010] Hereinafter, the embodiments will be described in detail with reference to the attached drawings. Note that the following embodiments do not limit the invention according to the claims. Although the embodiments describe a number of features, not all of these features are essential to the invention, and the features may be combined in any manner. Furthermore, in the attached drawings, the same reference numbers are used for the same or similar configurations, and duplicated descriptions are omitted.
[0011] <First embodiment> First, the configuration of an imaging device in this embodiment will be described with reference to Fig. 1. Fig. 1 is a block diagram showing the configuration of an imaging device 100. In this embodiment, a digital still camera capable of photographing a subject and recording video and still image data on various media such as tape, solid-state memory, optical disk, and magnetic disk will be described as the imaging device 100. However, the present invention is not limited to a digital still camera, and may be various electronic devices equipped with a camera function. For example, the present invention may be a video camera, a mobile communication terminal equipped with a camera function such as a mobile phone or a smartphone, a mobile computer equipped with a camera function, a mobile game machine equipped with a camera function, or the like.
[0012] Each component in the imaging device 100 is connected via a bus 160 and controlled by a CPU 151 (Central Processing Unit).
[0013] The lens unit 101 is configured to include a first fixed lens group 102, a zoom lens 111, an aperture 103, a second fixed lens group 121, and a focus lens 131, and may be configured integrally with the imaging device 100 or may be configured to be detachable. Aperture control unit 105 adjusts the amount of light during shooting by adjusting the aperture diameter of aperture 103 by driving aperture 103 via aperture motor 104 (AM) in accordance with instructions from CPU 151. At this time, CPU 151 determines the aperture diameter of aperture 103 using the luminance value of a specific subject area.
[0014] A zoom control unit 113 changes the focal length by driving the zoom lens 111 via a zoom motor 112 (ZM) in accordance with a command from the CPU 151 . The focus control unit 133 determines the drive amount for driving the focus motor 132 (FM) based on the focal shift amount (defocus amount) of the lens unit 101 with respect to a specific subject area. Then, the focus lens 131 is driven by the drive amount determined via the focus motor 132, thereby controlling the focus adjustment state. AF (autofocus) control is realized by the movement control of the focus lens 131 by the focus control unit 133 and the focus motor 132. Note that although the focus lens 131 is simply shown as a single lens in FIG. 1, it is usually composed of multiple lenses.
[0015] The image sensor 141 is a photoelectric conversion element that converts a subject image into an electric signal by photoelectric conversion, and converts an optical image of a subject (subject image) formed on the imaging surface of the image sensor 141 via the lens unit 101 into an electric signal. The image sensor 141 has light receiving elements arranged with m pixels in the horizontal direction and n pixels in the vertical direction. The electric signal (image signal) obtained by photoelectric conversion by the image sensor 141 is arranged as image data by the image signal processing unit 142 and output.
[0016] The image data output from the imaging signal processing unit 142 is sent to the imaging control unit 143 and temporarily stored in the RAM 154. The image data stored in the RAM 154 is compressed by the image compression / decompression unit 153 and then recorded on the image recording medium 157.
[0017] In parallel with this, the image data stored in RAM 154 is also sent to image processing unit 152, which performs processes such as reducing / enlarging the sent image data to a size appropriate for the purpose and calculating the similarity between image data. The image data reduced to a size for display is sent to monitor display 150 as appropriate. In addition, image processing unit 152 performs gamma correction, white balance processing, etc. on the sent image data based on the image signal of the subject area.
[0018] The monitor display 150 can display a preview image or a through image by displaying an image based on the image data reduced to a display size by the image processing unit 152. Furthermore, the result of subject detection by the subject detection unit 162 can be displayed superimposed on the image data using a rectangular frame or the like.
[0019] The position and orientation change acquisition unit 161 is configured with a position and orientation sensor such as a gyro, an acceleration sensor, or an electronic compass, and measures a change in the position and orientation of the imaging device 100 relative to the shooting scene. The acquired position and orientation change is stored in the RAM 154.
[0020] The subject detection unit 162 uses image data to detect a region in which a predetermined subject exists. This region may be output as rectangular information, or as a subject region map, which is an image in which pixel values indicate the "likelihood that a subject exists."
[0021] The RAM 154 can be used as a ring buffer to buffer image data of multiple images captured within a specified period of time, the detection results of the subject detection unit 162 corresponding to each image, and changes in position and orientation of the imaging device 100 acquired by the position and orientation change acquisition unit 161.
[0022] The operation unit 156 is an input interface including a touch panel, buttons, etc., and various operations can be performed by selecting and operating various functional icons displayed on the monitor display 150, etc.
[0023] CPU 151 determines the charge accumulation time of imaging element 141 and the setting value of gain when outputting from imaging element 141 to imaging signal processing unit 142, based on instructions from the operator inputted from operation unit 156 or the signal level of pixel signals of image data temporarily stored in RAM 154. Imaging control unit 143 receives instructions on the charge accumulation time and gain setting values from CPU 151 and controls imaging element 141.
[0024] The battery 159 is managed by a power management unit 158 and provides a stable power supply to the entire image capture device 100 . The flash memory 155 stores control programs necessary for the operation of the imaging device 100, parameters used for the operation of each unit, and the like. When the imaging device 100 is started up by a user's operation (when the power is switched from OFF to ON), the control programs and parameters stored in the flash memory 155 are loaded into a part of the RAM 154. The CPU 151 controls the operation of the imaging device 100 according to the control programs and constants loaded into the RAM 154.
[0025] The defocus calculation unit 163 calculates the defocus amount in an arbitrary region in the image. The defocus amount may be calculated for one point, or may be calculated at equal intervals over the entire image and output as a defocus map. The generated defocus information is stored in the RAM 154 and is referred to by the focus control unit 133. It should be noted that the above-described configuration is merely one example of the configuration of the imaging device 100.
[0026] Next, the flow of predictive focus adjustment processing by the image capturing apparatus 100 having the above-described configuration in this embodiment will be described with reference to FIG. In S200, an image (input image) captured by the imaging element 141 is acquired, and image data of the acquired input image is supplied from the imaging control unit 143 to each unit.
[0027] Next, in S201, the subject detection unit 162 performs subject detection processing on the input image, and detects a plurality of parts (portions) from the detected subject. In this embodiment, the subject to be detected is a person, and the head and torso of the person are detected. The subject detection unit 162 can perform subject detection using, for example, CNN (Convolutinal Neural Networks), but any method may be used as long as the subject can be detected. The subject detection unit 162 detects the head and torso of the detected person, and outputs information on a rectangular area indicating the detected head and torso. In addition, if either the head or the torso cannot be detected, information on a rectangular area indicating the detected head or torso may be output. In addition, if multiple heads or torsos are detected, which of them is the subject of interest is selected. The selection processing at this time may be any method, and is realized by selecting the one closest in distance from the detection position of the previous frame, for example.
[0028] In S202, CPU 151 determines whether the head and torso of a person were detected in S201, and if they were detected, that is, if there is a detection result of S201, the process proceeds to S203. On the other hand, if they were not detected, any processing is performed and then processing for the image of the current frame is terminated. For example, until the subject is detected again, the lens position of focus lens 131 is fixed without moving without predicting the defocus value of the next frame. Also, when either the head or the torso is detected, the defocus value of the detected head or torso may be calculated as information on the focus state, and may be stored in RAM 154 in association with the frame.
[0029] In S203, the CPU 151 calculates a defocus value for each of the head and torso regions of the detected subject, and stores the calculated defocus value in the RAM 154 in association with the frame.
[0030] In S204, the CPU 151 determines whether or not a defocus value corresponding to the head and torso regions of a past frame is stored in the RAM 154. In this embodiment, it is determined whether or not a defocus value of the immediately previous frame is stored, but it may also be determined whether or not a defocus value of an even earlier frame is stored. If a defocus value corresponding to the head and torso regions of a past frame is stored, the process proceeds to S205.
[0031] On the other hand, if it is determined in S204 that the defocus values corresponding to the head and torso regions of the past frame have not been saved, then any processing is performed and the processing for the image of that frame is terminated. For example, the processing for the image of the current frame is terminated after any processing is performed without predicting the defocus value of the next frame. As the any processing, for example, it is possible to perform focus adjustment by a conventional method using the defocus value detected in the current frame. Also, if a defocus value corresponding to the head or torso region has been obtained in the past frame, the defocus value of the next frame may be predicted using the history of the defocus value corresponding to the head or torso region, and the focus lens 131 may be driven based on the prediction result. As a prediction method, for example, it can be realized by obtaining a regression curve based on the least squares method using the defocus values corresponding to each region in the past frame and the current frame and the time of shooting.
[0032] In S205, the CPU 151 predicts the defocus value of the next frame using the defocus value of the past frame and the defocus value of the current frame calculated in S203. The prediction is performed for each region of the head and the torso. As a prediction method, for example, it can be realized by finding a regression curve based on the least squares method using the defocus values corresponding to each region in the past frame and the current frame and the time of shooting.
[0033] Then, in S206, the CPU 151 calculates the absolute difference between the defocus predicted values corresponding to the head and torso regions calculated in S205. It can be said that the smaller the absolute difference is, the closer the two regions are to each other in the depth direction at the predicted time.
[0034] Fig. 4 is a diagram showing an example of the concept of the absolute difference value of the defocus predicted value in this embodiment. Therefore, in order to make the explanation easier to understand, Fig. 4 shows the relative deviation of the defocus value of the head from the defocus value of the body, that is, the difference between the defocus values of the head and the body, with the defocus value of the body being set as the reference (0). In the following explanation, the defocus value of the body is called the reference defocus value, and the relative deviation of the defocus value of the head from the defocus value of the body is called the relative defocus amount.
[0035] That is, at time t (n-1) and time t n The head defocus amount is predicted from the head defocus amount relative to the reference defocus amount of the torso at a future time t (n+1) The absolute value of the relative defocus amount of the head in corresponds to the absolute difference value of the defocus predicted value calculated in S206.
[0036] In S207, the CPU 151 determines whether the absolute difference value calculated in S206 is less than a threshold value (a predetermined condition). Since the head and torso of the same subject should be located relatively close to each other, if the calculated absolute difference value is equal to or greater than the threshold value, it can be determined that the prediction of the head is highly likely to be incorrect.
[0037] For example, in the example in Figure 4, time t (n+1) This shows that the absolute value of the relative defocus amount of the head with respect to the reference defocus value of the body at time t, i.e., the absolute difference value of the defocus values, is large and exceeds the threshold value. In this way, when the calculated absolute difference value is equal to or greater than the threshold value, the predicted result of the head is not used, and the lens position is maintained or driven to a position equal to the threshold value. (n+1)By suppressing the out-of-focus state, it is possible to minimize the out-of-focus state of the head at time t (n+1) The head can be detected again, and a more likely prediction curve can be redrawn. After any other processing is performed, the processing for the image of the current frame is terminated. On the other hand, if the absolute difference value is less than the threshold value, the process proceeds to S208.
[0038] In S208, CPU 151 calculates the amount of drive of focus lens 131 for focusing on the head area based on the predicted result of the defocus value for the head area, and controls focus control unit 133 to drive focus lens 131.
[0039] In the above example, the subject is a person, and the defocus value corresponding to the head and torso area is calculated, but the subject is not limited to a person, and multiple parts of the subject may be detected, and the defocus value corresponding to the detected multiple parts may be calculated and predicted. In this case, by determining a main part among the multiple parts, it is possible to easily focus on a desired part.
[0040] As described above, according to the first embodiment, for a subject that is difficult to predict, it is possible to detect that the prediction has failed and to drive the focus lens based on the fact that the prediction has failed.
[0041] <Variation 1> In the first embodiment described above, in S207, it is determined whether the absolute difference between the predicted defocus values for the head and torso regions is less than a threshold value. This is processing that assumes that the defocus values of the head and torso of the same subject are somewhat close, but the degree of this depends on the distance between the camera and the subject. For example, it is considered that the defocus difference between the head and torso becomes smaller as the subject becomes farther from the camera. Therefore, an adjustment may be made so that the threshold value becomes smaller as the shooting distance becomes greater.
[0042] In the first embodiment described above, the head and torso regions of a person are the targets, but the method may be applied to other types of objects, such as animals and vehicles. In this case, the defocus difference threshold used in S207 may be adjusted according to the type of object. For example, the defocus difference may be larger for the head and torso of a horse than for a person, so the threshold may be set to a larger value.
[0043] <Second embodiment> Next, a second embodiment of the present invention will be described. Note that the configuration of the device in the second embodiment uses the imaging device 100 described in the first embodiment with reference to FIG. 1, and therefore the description will be omitted.
[0044] Fig. 3 is a flowchart showing the flow of processing in the second embodiment. Note that the same processes as those shown in Fig. 2 are denoted by the same S symbols, and descriptions thereof will be omitted as appropriate.
[0045] When it is determined that the head and torso of the current frame have been detected by the processing up to S202, in S303, the CPU 151 calculates a defocus range for each area of the detected head and / or torso as information on the focus state. For example, when the angle of view is narrow and the subject person is photographed large, the depth of the head or torso has a width, so the width of the defocus value in the depth direction is detected. The calculation of the defocus range may be performed by any method. In this embodiment, an inference device based on CNN shown in FIG. 6 is used. In addition to this, for example, it may be realized by calculating defocus values for multiple points of each part and selecting the minimum and maximum values among them.
[0046] 6 illustrates a configuration in the present embodiment in which the defocus range is calculated using CNN. The input unit 601 integrates the image output by the imaging signal processing unit 142, the defocus map output by the defocus calculation unit 163, and the object region map output by the object detection unit 162 as input information, as data having multiple channels, and inputs the data to a defocus range inference device 602. The input image, defocus map, and object region map are appropriately upsampled and downsampled to match the input size (resolution).
[0047] The defocus range inference unit 602 receives parameters generated by machine learning stored in the parameter storage unit 604 for the data input from the input unit 601, and infers the defocus range. The defocus range inference unit 602 outputs a defocus range corresponding to the part of the subject included in the image as an inference result. The output unit 603 outputs the defocus range of each part (head, torso, etc.) obtained from the defocus range inference unit 602 in association with meta information such as the ID of the image.
[0048] Although the case where the subject to be detected is a person has been described here, the information acquired from the parameter storage unit 604 may be switched depending on the type of subject to be detected. Although the parameters are costly to store, the parameters can be optimized depending on the type of subject, improving the accuracy of the output. Also, the subject region map may be generated for each part of the subject, or only the subject map of a specific part (for example, the torso) may be input as a representative part of the subject.
[0049] In this embodiment, the defocus range inference device 602 is configured by a machine-learned CNN, and infers the defocus range for each part of the subject. The defocus range inference device 602 may be realized by a circuit specialized for estimation processing by a GPU (graphics processing unit) or a CNN. The defocus range inference device 602 appropriately repeats convolution calculation in the convolution layer and pooling in the pooling layer on the data input from the input unit 601. Then, global average pooling processing (GAP) is performed to reduce the data. Next, the data subjected to GAP processing is input to a multi-layer perceptron. After processing of an arbitrary hidden layer, the value at one end of the defocus range of each part is output via the output layer.
[0050] When the defocus range is calculated in S303, in S304, the CPU 151 judges whether or not the defocus range corresponding to the head and torso area of the past frame is stored in the RAM 154. In this embodiment, it is judged whether or not the defocus range of the previous frame is stored, but it may be judged whether or not the defocus range of the frame before that is stored. If it is not stored, any processing is performed and then the processing for the image of the current frame is terminated. For example, the same processing as when NO is judged in S204 described above in the first embodiment may be performed using the defocus range. On the other hand, if it is stored, the process proceeds to S305.
[0051] In S305, the CPU 151 predicts the defocus range of the next frame using the defocus range of the past frame and the defocus range of the current frame calculated in S303. Any method of prediction may be used, but for example, prediction processing can be performed on the minimum and maximum values of the defocus range, and the results can be used as the minimum and maximum values of the defocus range of the next frame. Prediction is performed on each of the head and torso regions.
[0052] Next, in S306, the CPU 151 expands the predicted range corresponding to the region of the torso obtained in S305 by a factor of a. Hereinafter, the predicted range expanded by a factor of a is referred to as the "expanded predicted range." Note that a is a value of 1 or more, and may be a fixed value or may be changed according to circumstances. This expanded predicted range is regarded as the range in which the head may exist.
[0053] Fig. 5 is a diagram showing the concept of the predicted range of the defocus range in this embodiment. For ease of explanation, Fig. 4 shows the relative deviation of the defocus range of the head from the defocus range of the torso, with the defocus range of the torso as the reference. In the following explanation, the defocus range of the torso is called the reference defocus range, and the relative deviation of the defocus range of the head from the defocus range of the torso is called the relative defocus range.
[0054] time t (n-1) and t n Draw regression curves for the minimum and maximum values of the reference defocus range of the torso and the minimum and maximum values of the relative defocus range of the head, respectively, and calculate the regression curves for the future time t (n+1) The reference defocus range and the relative defocus range in each region are predicted. Note that the prediction may be performed using the median and width of the defocus range instead of the minimum and maximum values. Furthermore, since the head is nearly spherical, it is possible to predict only the median of the defocus range and fix the width, thereby applying prediction calculations according to the region.
[0055] And at time t (n+1) The predicted reference defocus range in is multiplied by a (multiple times) to obtain an expanded predicted range 501.
[0056] In S307, the CPU 151 determines whether or not the predicted defocus range of the head is included in the predicted enlarged range calculated in S306 (predetermined condition). If not included, it is determined that the prediction of the head is highly likely to be incorrect.
[0057] For example, in the example in Figure 5, time t (n+1) Since the predicted relative defocus range of the head in is not included within enlarged predicted range 501, it is determined that the prediction of the head is incorrect. In this case, as in the case of NO in S207, the head prediction result is not used and the lens position is maintained, etc., to prevent the head from being significantly out of focus. After any other processing is performed, the processing of the image of the current frame is terminated.
[0058] On the other hand, if the predicted defocus range of the head is included within the predicted enlargement range 501, the process proceeds to S308. In S308, CPU 151 calculates the driving amount of focus lens 131 for illuminating the head area based on the predicted defocus range corresponding to the head area, and controls focus control unit 133 to drive focus lens 131. The position to be focused may be the closest position within the predicted defocus range, or may be the center of the predicted defocus range. Furthermore, the aperture may be adjusted so that the entire range is in focus.
[0059] As described above, according to the second embodiment, for a subject that is difficult to predict, it is possible to detect that the prediction has failed and drive the focus lens based on the fact that the prediction has failed.
[0060] <Variation 2> In S306, the predicted defocus range of the torso is enlarged by a times, and in S307, it is determined whether the predicted defocus range of the head is included in the enlarged range. In the above-mentioned second embodiment, the head and torso regions of a person are targeted, but it may be performed for other types of objects, such as animals and vehicles, and in that case, the value of a (magnification) may be adjusted according to the type of subject. As an example, when comparing a person and a horse, the defocus range of the horse is likely to be different between the torso and face, so the value of a may be made larger than that of the person.
[0061] The value of a may also be adjusted according to the subject's posture. For example, if the horse is facing sideways as seen by the photographer, the defocus ranges of the body and face are likely to be close (i.e., a value of a close to 1 is appropriate). On the other hand, if the horse is facing forward as seen by the photographer, it is likely that there will be a difference in the defocus between the body and face (a value of a should be set large). The horse's orientation can be inferred to be facing sideways if the detection frame of the body is long horizontally. Alternatively, if tracking shows that the horse is moving sideways, it can be inferred to be facing sideways. Alternatively, if an eye detector is present and only one eye is detected, it may be determined that the horse is facing sideways. By inferring the posture using any of these methods and changing the value of a accordingly, it is possible to make the prediction success / failure judgment more reliable.
[0062] In the second embodiment described with reference to Fig. 5, the defocus range of the torso is enlarged equally in both the front and rear directions, but it may be enlarged in only one direction, or the enlargement ratio may be changed between the front and rear. For example, if the subject is a horse and faces the photographer (the face is closer to the photographer than the torso), it is possible to enlarge the defocus range of the torso only forward (toward the photographer).
[0063] <Modification 3> The threshold value used in S207 in the first embodiment and the value of the magnification ratio a used in S306 in the second embodiment can also be changed according to the reliability of the subject detection result.
[0064] For example, consider a case where the reliability of the defocus prediction for the head region is calculated based on the defocus prediction result for the torso region of a person. If the detection result of the torso is incorrect, the defocus prediction result for the torso performed using the defocus information of that region is also likely to be incorrect. Furthermore, if a deviation occurs from the defocus prediction result of the head as a result, the defocus prediction of the head may be determined to be incorrect even though it is correct. In such a case, it is possible to prevent the defocus prediction of the head from being determined to be incorrect by increasing the threshold value used in S207 or the value of the magnification rate a used in S306. In other words, it is possible to give these threshold values and the value of a larger value as the detection reliability of the torso becomes lower.
[0065] In addition, it is considered that the accuracy of defocus prediction decreases as the frame rate (frequency of executing subject detection and AF processing) decreases. Therefore, it is considered that the lower the frame rate, the larger the above-mentioned threshold value and the value of a are set.
[0066] <Modification 4> In the first and second embodiments described above, the reliability of the defocus prediction corresponding to the head area is calculated based on the defocus prediction result corresponding to the torso area of the person. This is based on the assumption that the movement of the torso is relatively slower than the movement of the head, and the prediction is less likely to be wrong, whereas the prediction of the head is more likely to be wrong. However, depending on the shooting scene, the opposite relationship may be true.
[0067] Therefore, the amount of movement may be calculated for each part, and the part with the smaller amount of movement may be used as the reference. Alternatively, if it is determined that the amount of movement of the head is smaller, any prediction process and AF process may be performed on the head, and the process may end without performing any prediction on the torso.
[0068] The amount of movement can be the distance traveled on the image plane within a predetermined time. It can be determined that the smaller the distance of movement, the smaller the amount of movement. This distance of movement can be determined by using the distance between the center coordinates of the subject detection result, or by using a tracking process such as template matching.
[0069] In addition, the change in appearance can be found by calculating the total amount of pixel change in the subject area. If this value is small, it can be estimated that the amount of movement of that part is small.
[0070] Alternatively, it is possible to calculate the amount of movement in the optical axis direction for each part by using the defocus value and the lens driving amount for each frame. A part with a small amount of movement in the optical axis direction may be selected as a reference part.
[0071] <Other embodiments> The present invention may be applied to a system made up of a plurality of devices, or to an apparatus made up of a single device.
[0072] The present invention can also be realized by supplying a program for implementing one or more of the functions of the above-described embodiments to a system or device via a network or a storage medium, and having one or more processors in the computer of the system or device read and execute the program. It can also be realized by a circuit (e.g., ASIC) that implements one or more of the functions.
[0073] <Summary> The disclosure of this embodiment includes the following configuration.
[0074] (Item 1) A detection means for detecting a first portion and a second portion of a predetermined subject from an image obtained by photographing the subject; an acquisition means for acquiring a focus state of each of the first portion and the second portion detected by the detection means; a storage means for storing the focus states of the first portion and the second portion acquired by the acquisition means; a prediction means for predicting a focus state of the first portion and the second portion at a third time point after the first time point, based on a focus state of the first portion and the second portion of an image obtained at a first time point and a focus state of the first portion and the second portion of an image obtained at a second time point before the first time point, the focus state being stored in the storage means; A focus adjustment unit that performs focus adjustment processing, The focus adjustment device is characterized in that the focus adjustment means performs the focus adjustment processing based on the predicted focus states of the first part and the second part when a difference between the focus state of the first part and the focus state of the second part predicted by the prediction means satisfies a predetermined condition. (Item 2) The focus adjustment device described in item 1 is characterized in that, when a difference between the focus state of the first portion and the focus state of the second portion predicted by the prediction means satisfies a predetermined condition, the focus adjustment means performs focus adjustment processing based on the predicted focus state of a predetermined one of the first portion and the second portion. (Item 3) The focus state is represented by a defocus value, and the predetermined condition is characterized in that an absolute value of a difference between the defocus value of the first portion and the defocus value of the second portion predicted by the prediction means is smaller than a predetermined threshold value. (Item 4) 4. The focus adjustment device according to item 3, wherein the threshold value is set smaller as the distance to the subject increases. (Item 5) 4. The focus adjustment device according to item 3, wherein the threshold value is changed depending on the type of the subject. (Item 6) 4. The focus adjustment device according to item 3, wherein the threshold value is increased as the reliability of the subject detection result decreases. (Item 7) 4. The focus adjustment device according to item 3, wherein the threshold value is increased as the frequency at which the detection means detects the first portion and the second portion decreases. (Item 8) The focus state is represented by a defocus range of the first portion and the second portion, and the predetermined condition is characterized in that the defocus prediction range of the first portion predicted by the prediction means is included in a range that is multiple times larger than the defocus prediction range of the second portion. The focus adjustment device described in item 1 or 2. (Item 9) 9. The focus adjustment device according to item 8, characterized in that a magnification for enlarging the defocus prediction range of the second portion is changed depending on the type of the subject. (Item 10) 9. The focus adjustment device according to item 8, characterized in that a magnification for enlarging the defocus prediction range of the second portion is changed depending on the orientation of the subject. (Item 11) 9. The focus adjustment device according to item 8, characterized in that a direction in which the defocus prediction range of the second portion is expanded is changed depending on the type of the subject. (Item 12) 9. The focus adjustment device according to item 8, characterized in that the lower the reliability of the detection result of the object, the larger the magnification for expanding the defocus prediction range of the second portion. (Item 13) The focus adjustment device according to item 8, characterized in that the magnification for expanding the defocus prediction range of the second portion is increased as the frequency of detection of the first portion and the second portion by the detection means decreases. (Item 14) further comprising a motion detection means for detecting an amount of motion of each of the first and second portions; The focus adjustment device according to any one of items 1 to 13, characterized in that the prediction means performs prediction based on the first portion or the second portion, whichever has a smaller amount of movement. (Item 15) A focus adjustment device according to any one of items 1 to 14, An imaging means for capturing the image; An imaging device comprising: (Item 16) Further comprising a lens unit including a focus lens, 16. The imaging device according to item 15, wherein the focus adjustment means determines a drive amount for driving the focus lens. (Item 17) The lens unit includes a focus lens, and is detachable from the lens unit. 16. The imaging device according to item 15, wherein the focus adjustment means determines a drive amount for driving the focus lens of the lens unit. (Item 18) a detection step in which a detection means detects a predetermined first portion and a second portion of a predetermined subject from an image obtained by photographing; an acquisition step in which an acquisition means acquires focus states of the first portion and the second portion detected in the detection step; a storage step in which a storage means stores the focus states of the first portion and the second portion acquired in the acquisition step; a prediction step in which a prediction means predicts a focus state of the first part and the second part at a third time point after the first time point from a focus state of the first part and the second part of an image obtained at a first time point and a focus state of the first part and the second part of an image obtained at a second time point before the first time point, the focus state being stored in the storage step; A focus adjustment step in which the focus adjustment means performs a focus adjustment process, A focus adjustment device characterized in that, in the focus adjustment process, when a difference between the focus state of the first portion predicted in the prediction process and the focus state of the second portion satisfies a predetermined condition, the focus adjustment process is performed based on the predicted focus states of the first portion and the second portion. (Item 19) A program for causing a computer to function as each of the means of the focus adjustment device according to any one of items 1 to 14. (Item 20) 20. A computer-readable storage medium storing the program according to item 19.
[0075] The invention is not limited to the above-described embodiments, and various modifications and variations are possible without departing from the spirit and scope of the invention. Accordingly, the following claims are appended to apprise the public of the scope of the invention. [Explanation of symbols]
[0076] 100: imaging device, 101: lens unit, 103: aperture, 111: zoom lens, 113: zoom control unit, 131: focus lens, 132: focus motor, 133: focus control unit, 141: imaging element, 142: imaging signal processing unit, 143: imaging control unit, 150: monitor display, 151: CPU, 152: image processing unit, 153: image compression / decompression unit, 154: RAM, 155: flash memory, 156: operation switch, 157: image recording medium, 162: subject detection unit, 163: defocus calculation unit
Claims
1. A detection means for detecting a predetermined first part and a predetermined second part of a predetermined subject from an image obtained by taking a photograph, An acquisition means for acquiring the focal state of the first part and the second part respectively detected by the detection means, A storage means for storing the focus states of the first and second parts acquired by the acquisition means, A prediction means predicts the focal state of the first and second parts of an image obtained at a first time, based on the focal state of the first and second parts of an image obtained at a second time, prior to the first time, as stored in the storage means, and the focal state of the first and second parts of an image obtained at a third time, prior to the first time. It has a focus adjustment means that performs focus adjustment processing, The focus adjustment device is characterized in that the focus adjustment means performs the focus adjustment process based on the predicted focus states of the first and second parts when the difference between the focus state of the first part and the focus state of the second part predicted by the prediction means satisfies predetermined conditions.
2. The focus adjustment device according to claim 1, characterized in that the focus adjustment means performs a focus adjustment process based on the predetermined predicted focus state of either the first part or the second part when the difference between the focus state of the first part and the focus state of the second part predicted by the prediction means satisfies the predetermined conditions.
3. The focus state is represented by a defocus value, and the predetermined condition is that the absolute value of the difference between the defocus value of the first part and the defocus value of the second part predicted by the prediction means is smaller than a predetermined threshold, as described in claim 1.
4. The focus adjustment device according to claim 3, characterized in that the threshold is reduced as the distance to the subject increases.
5. The focus adjustment device according to claim 3, characterized in that the threshold is changed according to the type of subject.
6. The focus adjustment device according to claim 3, characterized in that the threshold is increased as the reliability of the detection result of the subject decreases.
7. The focus adjustment device according to claim 3, characterized in that the threshold is increased as the frequency of detection of the first portion and the second portion by the detection means decreases.
8. The focus state is represented by the defocus ranges of the first and second parts, and the predetermined condition is that the predicted defocus range of the first part predicted by the prediction means is contained within a range that is multiple times larger than the predicted defocus range of the second part, as described in claim 1.
9. The focus adjustment device according to claim 8, characterized in that the magnification for expanding the defocus prediction range of the second portion is changed according to the type of subject.
10. The focus adjustment device according to claim 8, characterized in that the magnification for expanding the defocus prediction range of the second portion is changed according to the orientation of the subject.
11. The focus adjustment device according to claim 8, characterized in that the direction in which the defocus prediction range of the second portion is expanded is changed according to the type of subject.
12. The focus adjustment device according to claim 8, characterized in that the lower the reliability of the detection result of the subject, the greater the magnification for expanding the defocus prediction range of the second portion.
13. The focus adjustment device according to claim 8, characterized in that the less frequently the detection means detects the first portion and the second portion, the greater the magnification for expanding the defocus prediction range of the second portion.
14. The system further includes motion detection means for detecting the amount of movement of the first part and the second part, The focus adjustment device according to claim 1, characterized in that the prediction means makes a prediction based on the one of the first and second parts that has less movement.
15. A focus adjustment device according to any one of claims 1 to 14, An imaging means for capturing the aforementioned image and An imaging device characterized by having the following features.
16. It further comprises a lens unit including a focusing lens, The imaging apparatus according to claim 15, characterized in that the focus adjustment means determines the amount of drive for driving the focus lens.
17. The lens unit, including the focus lens, is detachable. The imaging apparatus according to claim 15, characterized in that the focus adjustment means determines the amount of drive for driving the focus lens of the lens unit.
18. The detection means includes a detection step of detecting a predetermined first part and a predetermined second part of a predetermined subject from an image obtained by taking a photograph, The acquisition means includes an acquisition step of acquiring the focal state of the first part and the second part respectively detected in the detection step, The storage means includes a storage step of storing the focus states of the first part and the second part acquired in the acquisition step, A prediction means includes a prediction step that predicts the focal state of the first and second parts of an image obtained at a first time and the focal state of the first and second parts of an image obtained at a second time prior to the first time, which is stored in the storage step, based on the focal state of the first and second parts of an image obtained at a third time after the first time. The focus adjustment means includes a focus adjustment step that performs a focus adjustment process, The focus adjustment method is characterized in that, in the focus adjustment step, when the difference between the focus state of the first part and the focus state of the second part predicted in the prediction step satisfies predetermined conditions, the focus adjustment process is performed based on the predicted focus states of the first part and the second part.
19. A program for causing a computer to function as one of the means of a focus adjustment device according to any one of claims 1 to 14.
20. A computer-readable storage medium storing the program described in claim 19.