Focus detection device and method, imaging apparatus, program, and storage medium

The focus detection device uses subject detection and historical data to predict and maintain focus on high-priority subjects despite temporary changes, ensuring effective focus adjustment.

JP2025105013APending Publication Date: 2025-07-10CANON KK
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
JP2023223261
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Filing Date
2023-12-28
Publication Date
2025-07-10

AI Technical Summary

Technical Problem

Existing focus detection technologies struggle to maintain focus on high-priority subjects when there are temporary state changes such as posture adjustments or changes in light conditions, leading to ineffective focus adjustment.

Method used

A focus detection device that includes subject detection, feature area setting, focus state detection, selection of priority areas, and prediction of focus state changes based on historical data to maintain focus on high-priority subjects.

Benefits of technology

Enables continuous focus on high-priority subjects despite temporary state changes by utilizing historical focus detection data for predictive focus adjustment.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2025105013000001_ABST
    Figure 2025105013000001_ABST
Patent Text Reader

Abstract

To continue to focus on a portion with higher priority even when there is a temporary change of states of a subject.SOLUTION: A focus detection device detects a subject and a feature area included in the subject from an image of image signals repeatedly output from an imaging apparatus, and sets a plurality of focus detection areas with a predetermined size to the detected subject and the feature area. The focus detection device detects a focus state in each of the plurality of set focus detection areas on the basis of the image signals, and from the plurality of focus detection areas, selects a first area satisfying a predetermined condition and a second area including the feature area. The focus detection device predicts a focus state of the second area in an arbitrary time, from a first history of the focus state of the first area, and a second history of a difference between the focus state of the first area and the focus state of the second area, based on the focus state, and determines an amount of drive of focusing means for focusing on the second area on the basis of the focus state of the second area.SELECTED DRAWING: Figure 13
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to a focus detection device and method, an imaging device, a program, and a storage medium, and particularly to a focus adjustment technique utilizing subject detection information.

Background Art

[0002] In recent years, cameras equipped with a focus adjustment function (hereinafter referred to as "AF") for automatically adjusting the focus position of a photographing lens have become widespread. As means for performing such focus adjustment, various AF methods such as an imaging surface phase difference AF method and a contrast AF method using an image sensor have been put into practical use.

[0003] Furthermore, in various AF methods, there is a technique for specifying the region of a main subject and focusing on it. Patent Document 1 discloses a control method for focusing on a site to be prioritized while avoiding a region where focus detection is difficult among the detected subjects.

[0004] Moreover, many of them have a servo shooting mode for driving a focus lens so that the focus is on not only a stationary subject but also a moving subject. The servo shooting mode has a function of predicting how the subject is moving from its past movement history. Patent Document 2 discloses a control method for holding a focus detection history for each site of the detected subject and switching the history used for prediction according to conditions.

Prior Art Documents

Patent Documents

[0005]

Patent Document 1

Patent Document 2

Summary of the Invention

Problems to be Solved by the Invention

[0006] However, in Patent Document 1 and Patent Document 2, when there are temporary state changes such as the posture of the subject and the way light hits, problems occur where the subject in the detection site with high priority cannot be detected, or focus detection cannot be performed, resulting in the problem that focus cannot be adjusted to the detection site with high priority.

[0007] The present invention has been made in view of the above problems, and an object thereof is to continue to focus on a site with high priority even when there is a temporary state change in the subject.

Means for Solving the Problems

[0008] To achieve the above object, the focus detection device of the present invention includes a subject detection means for detecting a subject and a feature area included in the subject from an image of an image signal repeatedly output from an imaging device; a setting means for setting a plurality of focus detection areas of a predetermined size in the subject and the feature area detected by the subject detection means; a detection means for detecting a focus state in each of the plurality of focus detection areas set by the setting means based on the image signal; a selection means for selecting a first area that satisfies a predetermined condition and a second area that includes the feature area from among the plurality of focus detection areas; a prediction means for predicting the focus state of the second area at an arbitrary time from a first history of the focus state of the first area and a second history of the difference between the focus state of the first area and the focus state of the second area based on the focus state detected by the detection means; and an acquisition means for obtaining a drive amount of a focus adjustment means for focusing on the second area based on the focus state of the second area predicted by the prediction means.

Effects of the Invention

[0009] According to the present invention, even when there is a temporary state change in the subject, it is possible to continue to focus on a site with high priority.

Brief Description of the Drawings

[0010]

Figure 1

Figure 2

Figure 3

Figure 4

Figure 5

Figure 6

Figure 7

Figure 8

Figure 9

Figure 10

Figure 11

Figure 12

Figure 13

Figure 14

Figure 15

Figure 16

Figure 17

Mode for Carrying Out the Invention

[0011] Hereinafter, embodiments will be described in detail with reference to the accompanying drawings. Note that the following embodiments do not limit the invention according to the claims. Although a plurality of features are described in the embodiments, not all of these plurality of features are essential to the invention, and the plurality of features may be arbitrarily combined. Further, in the accompanying drawings, the same or similar configurations are denoted by the same reference numerals, and redundant descriptions are omitted.

[0012] ● Configuration of Imaging Device FIG. 1 is a block diagram showing the configuration of an interchangeable-lens camera (hereinafter simply referred to as a “camera”) as an imaging device according to an embodiment of the present invention. The camera according to the present embodiment is an example of an imaging device equipped with a focus adjustment device to which the present invention is applied, and performs focus adjustment by an imaging plane phase difference detection method using an output signal from an image sensor that captures a subject image. Note that the present invention is applicable to, for example, a lens-integrated digital camera or a video camera. Further, the present invention can also be implemented in any electronic device equipped with a camera, such as a mobile phone, a personal computer (laptop, tablet, desktop type, etc.), a game machine, or the like.

[0013] As shown in FIG. 1, the camera is composed of a lens device (interchangeable lens) 100 and a camera body 200. When the lens device 100 is attached to the camera body 200 via a mount portion having an electrical contact unit 106, a lens controller 105 that comprehensively controls the operation of the lens device 100 and a system control unit 209 that comprehensively controls the operation of the entire camera can communicate with each other.

[0014] First, the configuration of the lens device 100 will be described. The lens device 100 includes a photographing lens 101 including a zoom mechanism, a diaphragm and shutter 102 for controlling the amount of light, a focus lens 103 for focusing on an image sensor 201 described later, a motor 104 for driving the focus lens, and a lens controller 105.

[0015] Next, the configuration of the camera body 200 will be described. The camera body 200 is configured to be able to acquire an image signal from the light beam that has passed through the imaging optical system of the lens device 100. In the camera body 200, the imaging element 201 is composed of a CCD or a CMOS sensor, receives the reflected light from the subject, converts it into signal charges corresponding to the incident light amount by a photodiode, and accumulates them. The signal charges accumulated in each photodiode are sequentially read out from the imaging element 201 as voltage signals (image signals) corresponding to the signal charges based on the drive pulses given from the timing generator 208 according to the commands of the system control unit 209.

[0016] The A / D conversion unit 202 performs A / D conversion on the voltage signal output from the imaging element 201. The A / D conversion unit 202 includes a CDS circuit that removes the output noise of the imaging element 201 and a non-linear amplification circuit that is performed before A / D conversion.

[0017] Furthermore, the camera body 200 includes an image processing unit 203, an AF signal processing unit 204, a format conversion unit 205, and a high-speed built-in memory (hereinafter referred to as "DRAM") 206 such as a random access memory. The DRAM 206 is used as a high-speed buffer for temporarily storing images or as a working memory for image compression and decompression, etc. The image recording unit 207 consists of a removable recording medium such as a memory card and its interface.

[0018] The system control unit 209 controls the entire camera, such as the timing generator 208 and the shooting sequence. Furthermore, the digital camera body 200 includes a lens communication unit 210 that communicates between the camera body 200 and the lens device 100, a subject detection unit 211, a subject movement prediction unit 219, and an image display memory 212 (hereinafter referred to as "VRAM"). The image display unit 213 displays the captured image, and also displays operation assistance, the camera state, the shooting screen during shooting, and the focus detection area.

[0019] In addition, the camera body 200 has various operation members for the user to operate the camera. The operation unit 214 includes, for example, a menu switch for performing various settings such as the shooting function of the camera and settings during image playback, an operation mode switching switch for switching between the shooting mode and the playback mode, and the like. The shooting mode switch 215 is a switch for selecting a shooting mode such as the macro mode and the sports mode, and the main switch 216 is a switch for turning on the power of the camera. Further, it is provided with a switch (hereinafter referred to as "SW1") 217 for performing a shooting standby operation such as AF and AE, and a shooting switch (hereinafter referred to as "SW2") 218 for performing shooting after the operation of SW1.

[0020] ● Configuration of the imaging device The imaging device 201 is composed of a CCD or a CMOS sensor. Each pixel of the imaging device 201 used in this embodiment is composed of two (a pair) of photodiodes A and B and one microlens provided for this pair of photodiodes A and B. In each pixel, a pair of optical images are formed on the pair of photodiodes A and B through the microlens for the incident light, and a pair of pixel signals (A signal and B signal) used for an AF signal described later are output from the pair of photodiodes A and B. In addition, an imaging signal (A + B signal) can be obtained by adding the outputs of the pair of photodiodes A and B.

[0021] By collecting the plurality of A signals and the plurality of B signals output from the plurality of pixels respectively, a pair of image signals (A image signal, B image signal) are obtained as an AF signal (focus detection signal) used for AF by the imaging plane phase difference detection method. The AF signal processing unit 204 performs a correlation operation on the A image signal and the B image signal, calculates the phase difference (hereinafter referred to as "image shift amount") which is the shift amount between the A image signal and the B image signal, and further obtains the defocus amount, the defocus direction, and the reliability (hereinafter collectively referred to as "focus detection information") of the photographing optical system from the calculated image shift amount. In this embodiment, the AF signal processing unit 204 obtains focus detection information for each area within the AF frame set as described later.

[0022] ● Operation of the imaging device Next, the shooting operation of the camera according to this embodiment will be described with reference to FIG. 2. FIG. 2 shows the flow of the shooting operation when performing still image shooting from the state of displaying a live view image. Note that each process in this flowchart is performed by the system control unit 209 executing a control program stored in a non-volatile memory (not shown).

[0023] First, in S201, the state of SW1 (217) is checked. If it is ON, the process proceeds to S202. In S202, the system control unit 209 performs an AF frame setting process described later, sets an AF frame for the AF signal processing unit 204, and proceeds to S203. In S203, an AF operation described later is performed for the areas of each AF frame set in S202, and the process proceeds to S204. In S204, the state of SW1 (217) is checked. If it is ON, the process proceeds to S205; otherwise, the process returns to S201. When SW1 (217) is ON, the process proceeds to S205. The state of SW2 (218) is checked. If it is ON, the process proceeds to S206; otherwise, the process returns to S201. In S206, shooting is performed, and after the shooting is completed, the process returns to S201.

[0024] ● AF frame setting process FIG. 3 is a flowchart for explaining the AF frame setting process performed in S202 of FIG. 2.

[0025] First, in S301, subject detection information is acquired from the subject detection unit 211. In the following description, the subject of this embodiment is a person, and the main areas in the person (within the subject) are detected. Here, the main areas are the areas of the pupils, face, and body in the person. However, the subject is not limited to a person. For example, it may be an animal, a vehicle, a train, etc., and the main areas may be determined based on the characteristics of the subject. As a method for detecting the subject and the main areas, a known learning method based on machine learning, a recognition process by image processing, or the like can be used.

[0026] For example, the types of machine learning include the following.

[0027] (1) Support Vector Machine (2) Covolutional Neural Network (3) Recurrent Neural Network

[0028] In addition, as an example of recognition processing, there is a method of extracting a skin color area from the gradation color of each pixel represented by image data and detecting a face based on the degree of matching with a face contour plate prepared in advance. Also, a method of performing face detection by extracting feature points of a face such as eyes, nose, mouth, etc. using well-known pattern recognition techniques is also well-known. Note that the detection methods for the main areas applicable to the present invention are not limited to these methods, and other methods may be used.

[0029] In S302, based on the subject detection information obtained from the subject detection unit 211, it is determined whether the subject has been detected, and if so, whether a plurality of main areas have been detected. If a plurality of main areas have been detected, the process proceeds to S303. If only one main area has been detected, the process proceeds to S304. If the subject has not been detected, the process proceeds to S308.

[0030] Here, the concept of the AF frame in the case where one main area is detected and the case where a plurality of main areas are detected will be described with reference to FIGS. 4 and 5. FIG. 4(a) shows a state where only the face area 401 has been detected, and FIG. 5(a) shows a state where the pupil area 501, the face area 502, and the body area 503 have been detected. Note that the subject detection unit 211 can acquire the type of the subject such as a person or an animal, the center coordinates in each detected main area, and the horizontal size and the vertical size, and outputs this information as subject detection information.

[0031] In S303, among the detected multiple main regions, the smallest main region is selected, and the smaller value of the horizontal size and the vertical size of the main region is set as MinA. That is, in the example of Fig. 5(a), the smaller value of the horizontal size and the vertical size of the pupil region 501 is set as MinA. Then, this value MinA is set as the length of one side (AF frame size) of one AF frame 504.

[0032] In S305, from the horizontal coordinates and the horizontal sizes of the detected multiple main regions, as shown in Fig. 5(b), the horizontal size H of the region including all the main regions is obtained, and the horizontal AF frame number is determined by dividing the obtained horizontal size H by the AF frame size MinA. Subsequently, in S307, from the vertical coordinates and the vertical sizes of the detected multiple main regions, as shown in Fig. 5(b), the vertical size V of the region including all the main regions is obtained, and the vertical AF frame number is determined by dividing the obtained vertical size V by the AF frame size MinA, and the AF frame setting process is terminated. Note that the order of the processes in S305 and S307 may be reversed or they may be performed in parallel.

[0033] In this embodiment, a square AF frame using the shorter side length of the smallest main region is used, but the present invention is not limited to this. For example, the horizontal size and the vertical size of the AF frame may be different, or the size of the AF frame may be set based on the size of the region including all the main regions and the number of AF frames that can be calculated by the system control unit 209.

[0034] On the other hand, when the number of detected main regions is not multiple, in S304, an AF frame with a predetermined size X is set for the detected face. As the predetermined size X, it may be the pupil size estimated from the face, or a size may be set that can ensure the S / N and have sufficient focusing performance considering low illuminance environments. In this embodiment, the predetermined size X is set as the estimated pupil size. Then, in S306, as shown in Fig. 4(b), the number of AF frames with the size X of the AF frame that includes the face region 401 and can cope with the movement of the face is set.

[0035] Also, when the subject is not detected, in S308, the size, setting position, and number of AF frames are determined by a known method. As a known method, for example, there are a method of setting one AF frame of a predetermined size in the center, a method of setting AF frames of a predetermined size at arbitrary multiple positions, and a method of dividing the screen into a plurality of regions and using each region as an AF frame. There are also various other methods such as setting by the user using the operation unit 214, and the present invention is not limited to these methods.

[0036] In addition, in the examples shown in FIGS. 4 and 5, a plurality of AF frames are set so as to include the main region, but the present invention is not limited to this. For example, a plurality of AF frames may be discretely set within the main region. When the AF frame is set by the above processing, the process proceeds to the process of S203 in FIG. 2.

[0037] Note that the above-described AF frame setting process does not have to be performed every time an image is input, and it may be performed once every time a plurality of images are input. In that case, when the AF frame setting process is not performed, the set AF frame is stored in the DRAM 206, and the AF frame stored at the timing closest in time may be read out and used.

[0038] ●AF Operation FIG. 6 is a flowchart for explaining the overall flow of the AF operation performed in S203. First, in S401, a focus detection process is performed to detect the defocus amount, and the process proceeds to S402. Note that the focus detection process will be described later. In S402, a main frame selection process described later is performed using the subject detection information obtained in S301, and the process proceeds to S403. In S403, the drive amount calculation of the focus lens described later is performed using the main frame selection result obtained in S402, and the process proceeds to S404. In S404, the focus lens drive amount obtained in S403 is transmitted to the lens communication unit 210 to drive the focus lens 103, and the AF operation is terminated.

[0039] ◆Focus Detection Process Next, the focus detection process performed in S401 will be described with reference to FIG. 7. First, in S501, from the image data output from the imaging device 201, an area of one unprocessed AF frame among the AF frames set in S202 is set as the focus detection area, and the process proceeds to S502. In S502, a pair of image signals (A image signal, B image signal), which are focus detection signals, are acquired from the pixels in the area of the imaging device 201 corresponding to the focus detection area set in S501, and the process proceeds to S503. In S503, after performing row addition averaging processing on the pair of image signals acquired in S502 in the vertical direction, the process proceeds to S504. By this process, the influence of noise in the image signals can be reduced.

[0040] In S504, filter processing is performed to extract signal components in a predetermined frequency band from the signal vertically row-added and averaged in S503, and the process proceeds to S505. In S505, a correlation amount is calculated from the signal filter-processed in S504, and the process proceeds to S506. In S506, a correlation change amount is calculated from the correlation amount calculated in S505, and the process proceeds to S507. In S507, an image shift amount is calculated from the correlation change amount calculated in S506, and the process proceeds to S508. In S508, a reliability representing how reliable the image shift amount calculated in S507 is is calculated, and the process proceeds to S509. In S509, the image shift amount is converted into a defocus amount.

[0041] In S510, it is determined whether there is an unprocessed AF frame among the AF frames set in S202. If there is an unprocessed AF frame, the process returns to S501 and the above process is repeated for the next unprocessed AF frame. If the processing for all the set AF frames has been completed, the focus detection process ends.

[0042] ◆ Main Frame Selection Process FIG. 8 is a flowchart for explaining the main frame selection process performed in S402. First, in S601, it is determined whether a subject is detected by the subject detection unit 211. If not detected, the process proceeds to S602, where the main frame is selected by a first selection method that does not use the subject detection information, and the main frame selection process ends. As the first selection method, for example, among the AF frames set in S308, a method of selecting a predetermined area within the screen (for example, an area including the center position) as the main frame can be considered, but it is not limited to this, and known techniques may also be used.

[0043] On the other hand, if the subject is detected, the process proceeds to S603, where the main frame is selected by a second selection method of selecting the main frame from among the AF frames for which focus detection has been performed, and the process proceeds to S604.

[0044] Here, the details of the main frame selection process by the second selection method performed in S603 will be described with reference to the flowcharts of FIGS. 9 and 12. First, in S701 of FIG. 9, it is determined whether the pupil of the subject is detected by the subject detection unit 211. If the pupil is detected, the process proceeds to S702; if not, the process proceeds to S703. In S702, the pupil area is added as the main frame selection area, and the process proceeds to S703.

[0045] In S703, it is determined whether the face of the subject is detected by the subject detection unit 211. If the face is detected, the process proceeds to S704; if not, the process proceeds to S705. In S704, the face area is added as the main frame selection area, and the process proceeds to S705.

[0046] In S705, it is determined whether the body of the subject is detected by the subject detection unit 211. If the body is detected, the process proceeds to S706; if not, the process proceeds to S707. In S706, the body area is added as the main frame selection area, and the process proceeds to S707.

[0047] In S707, it is determined whether the number of AF frames including the main frame selection area is equal to or greater than a predetermined first threshold. Here, the AF frame including the main frame selection area is an AF frame whose center is included in the main frame selection area, but it is not limited to this. For example, it may be an AF frame including at least a part of the main frame selection area, or it may be an AF frame in which a region equal to or greater than a predetermined ratio of the AF frame is applied to the main frame selection area. If the main frame selection area includes an AF frame equal to or greater than the first threshold, the process proceeds to S708. If not, the process proceeds to S711.

[0048] In S708, the defocus amount calculated for each AF frame including the main frame selection area is classified for each predetermined depth to create a histogram. Then, in S709, it is determined whether the peak value of the histogram created in S708 is equal to or greater than a predetermined second threshold. In this embodiment, the maximum number of AF frames in the histogram is normalized by the total number of AF frames and converted into a ratio, which is used as the peak value. If the peak value is equal to or greater than the second threshold, the process proceeds to S710. If it is less than the second threshold, the process proceeds to S711.

[0049] In S710, there is an AF frame equal to or greater than the first threshold within the main frame selection area, and within a certain depth (bin of the peak value), there is a state where the defocus amount of an AF frame equal to or greater than the overall second threshold exists (hereinafter referred to as the "first state"). Therefore, main frame selection processing corresponding to the first state is performed.

[0050] On the other hand, in S711, the AF frames included in the main frame selection area are less than the first threshold, or even if they are equal to or greater than the first threshold, the defocus amounts of the AF frames are scattered over a plurality of depths (hereinafter referred to as the "second state"). Therefore, main frame selection processing corresponding to the second state is performed.

[0051] In this embodiment, the main frame is selected by different methods in the different states described above.

[0052] FIG. 10 is a flowchart showing the main frame selection process in the first state performed in S710. First, in S801, a loop process for all AF frames including the main frame selection area is started to select a main frame from within the main frame selection area. In S802, it is determined whether the AF frame being processed is included in the bin indicating the peak value of the histogram. If so, the process proceeds to S803; otherwise, the loop process is repeated for other AF frames. In S803, it is determined whether the AF frame being processed is closer to the center of the main frame selection area than the currently selected main frame. If it is closer, the process proceeds to S804 to update the main frame; otherwise, the loop process is repeated for other AF frames. As the center of the main frame selection area, for example, it may be the center of the maximum width in the horizontal and vertical directions of the main frame selection area, or it may be the centroid of the main frame selection area.

[0053] When the loop in S801 ends, the main frame selection process in the first state ends. Through the above process, among the AF frames included in the bin indicating the peak value, the AF frame closest to the center of the main frame selection area is selected as the main frame.

[0054] FIG. 11 is a flowchart showing the main frame selection process in the second state performed in S711. In S901, it is determined whether the center of the main frame selection area is within a predetermined depth. If it is within the predetermined depth, the process proceeds to S902, and the AF frame including the center of the main frame selection area is set as the main frame. On the other hand, if the center of the main frame selection area is not within the predetermined depth, the process proceeds to S903.

[0055] In S903, a loop process for all frames is performed to select a main frame from among the AF frames including the set main frame selection area. The initial value of the main frame is set to information (such as the total number of frames + 1) that can determine that the main frame has not been selected, and the figure is omitted. In S904, it is determined whether the defocus amount of the AF frame being processed is within a predetermined depth and is closer to the subject side than the defocus amount of the currently selected main frame. If the conditions are met, the main frame is updated in S905.

[0056] In S906, it is determined whether the main frame can be selected by the loop of S903. If it cannot be selected, in S907, an AF frame including the center of the main frame selection area is set as the main frame. After finishing the above processing, the process returns to S603 in FIG. 8.

[0057] By the above processing, when the main frame is selected by the second selection method in S603, in S604, it is determined whether the pupil of the subject is detected by the subject detection unit 211. If the pupil is detected, the process proceeds to S605; if not, the main frame selection process ends. In S605, a process of selecting an AF frame representing a feature part (feature area) such as a pupil is performed.

[0058] FIG. 12 is a flowchart for explaining the feature part representative frame selection process performed in S605. First, in S1001, the feature part area is set to the pupil area, and the process proceeds to S1002. In this embodiment, since a person is assumed as the subject, the pupil is used as the feature part and the feature part area is set to the pupil area. However, the present invention is not limited to this, and any part can be set as the feature part.

[0059] In S1002, in order to select the feature part representative frame, a loop process is performed for all AF frames including the set feature part area. The initial value of the feature part representative frame is set to information (such as the total number of frames + 1) that can determine that the feature part representative frame has not been selected, and the figure is omitted. In S1003, it is determined whether the defocus amount of the AF frame being processed is within a predetermined depth and is closer to the subject than the selected feature part representative frame. If the condition is satisfied, in S1004, the feature part representative frame is updated. In S1005, it is determined whether the feature part representative frame can be selected by the loop of S1002. If it cannot be selected, in S1006, an AF frame including the center of the feature part area is set as the feature part representative frame. Thus, the main frame selection process ends, and the process returns to S402 in FIG. 6.

[0060] ◆Lens driving amount calculation process FIG. 13 is a flowchart for explaining the lens driving amount calculation process in S403.

[0061] First, in S1301, the focus detection result of the main frame selected in S402 is recorded, and the process proceeds to S1302. Since the focus detection result of the main frame is recorded every time focus detection is performed, the history of past focus detection results can be traced back.

[0062] In S1302, it is determined whether the feature part representative frame is selected. If it is selected, the process proceeds to S1303; if not, the process proceeds to S1304. In S1303, the relative position of the image plane position of the feature part obtained from the focus detection results of the main frame selected in S603 and the feature part representative frame selected in S605 is recorded. The relative position of the feature part is obtained by subtracting the image plane position of the main frame from the image plane position of the feature part representative frame, and indicates the relative position of the image plane position of the feature part representative frame with respect to the image plane position of the main frame. Since the relative position of the feature part is recorded every time focus detection is performed and the feature part representative frame is selected, the history of past relative positions can be traced back. When the recording of the relative position of the feature part is completed, the process proceeds to S1304. In S1304, the subject movement prediction unit 219 performs a prediction process of predicting the image plane position using the history of the focus detection results, and the process proceeds to S1305. The details of the prediction process will be described later.

[0063] In S1305, the driving target of the focus lens is set to the predicted image plane position, and the lens driving calculation process ends.

[0064] FIG. 14 is a flowchart for explaining the prediction process performed in S1304. First, in S1401, a prediction process for the entire subject is performed using the focus detection result of the main frame recorded in S1301.

[0065] In the prediction process of the entire subject, the predicted position of the entire subject is obtained by drawing a prediction curve from the past focus detection history. 1501 in FIG. 15 shows an example of the prediction curve, where the vertical axis represents the image plane position and the horizontal axis represents time. The larger the image plane position on the vertical axis, the farther the distance, and the history shown in FIG. 15 shows the state of following a subject approaching the photographer (camera). As described above, the AF operation in S203 is periodically performed in the servo shooting mode, and T1 to T5 each represent the execution time of the AF operation performed in S203.

[0066] The predicted image plane position is obtained, for example, by using the past image plane positions and their respective focus detection times to derive a prediction curve by the batch least squares method, and calculating the image plane position at the predicted time based on this curve. The predicted time indicates the time at the time of shooting if SW2 is ON, or the current focus detection time otherwise. Also, the prediction method is not limited to the batch least squares method. For example, when using the sequential least squares method, it is not necessary to store a plurality of past focus detection histories, and the parameters calculated during the previous AF operation can be recorded to substitute for the past focus detection history. When the prediction process of the entire subject is completed, the process proceeds to S1402.

[0067] In S1402, the relative position of the feature part is predicted using the history of the relative position of the feature part recorded in S1303.

[0068] In the prediction of the relative position of the feature part, the predicted change amount is obtained by drawing a prediction curve from the history of the past relative positions. 1601 in FIG. 16 shows an example of the prediction curve, where the vertical axis represents the relative position and the horizontal axis represents time. The larger the relative position on the vertical axis, the farther the feature part is from the entire subject. The history shown in FIG. 16 shows the state of following a subject with a feature part that has jumped out towards the photographer (camera) side, such as a flying bird, as it approaches the photographer. As described above, the AF operation in S203 is periodically performed in the servo shooting mode, and T1 to T5 each represent the execution time of the AF operation performed in S203.

[0069] The predicted relative position of the target can be obtained by, for example, deriving a prediction curve using the past image plane positions and the respective focus detection times through the batch least squares method, and calculating the image plane position at the predicted future time based on this curve. The predicted future time indicates the time at the time of shooting if SW2 is ON, or the current focus detection time otherwise. Also, the prediction method is not limited to the batch least squares method. For example, when using the sequential least squares method, it is not necessary to store a plurality of past focus detection histories, and the parameters calculated during the previous AF operation can be recorded to substitute for the past focus detection histories.

[0070] Also, the depth width of the subject may be estimated in advance at the time of focus detection, and the relative position of the feature part may be expressed using the depth width of the subject. For example, the relative position of the feature part may be predicted by predicting the ratio of the depth width of the subject to the relative position of the feature part by the batch least squares method or the sequential least squares method. Or, it may be approximated and expressed as a single vibration model with the depth width of the subject as the amplitude. As the estimation of the depth width of the subject, for example, the depth width obtained by summing up the depth widths of all bins including the AF frame including the main frame selection area using the histogram created at the time of selecting the main frame of the entire subject may be used as the depth width of the subject. Also, the depth width of the subject may be estimated in advance by learning to be able to use the detected subject type, shooting distance, and image information, and the depth width of the subject estimated thereby may be used. Note that when there is no record of the feature part being detected in the past, the relative position of the feature part is set to 0. When the prediction process of the relative position of the feature part is completed, the process proceeds to S1403.

[0071] In S1403, the prediction process of the feature part is performed using the prediction result of the entire subject obtained in S1401 and the prediction result of the relative position of the feature part obtained in S1402.

[0072] Assuming that the prediction formula for the entire subject is a function L(t) of time t and the prediction formula for the relative position of the feature part is l(t), the prediction formula L′(t) for the feature part can be expressed by the following formula (1). L'(t) = L(t) + l(t) …(1)

[0073] FIG. 17 shows the predicted curve 1702 of the feature part when the predicted curve of the entire subject is FIG. 15 and the predicted curve 1701 of the relative position of the feature part is FIG. 16. When the prediction process of the feature part is completed, the process proceeds to S1404. In S1404, the prediction result of the feature part obtained in S1403 is set as the predicted image plane position, and the prediction process is terminated. Note that the prediction of the relative position of the feature part shown in S1403 is calculated even when the feature part has not been detected, as long as there is a history of the feature part being detected at least once in the past. If the model of the predicted curve is made stable in advance by measures such as using long time-series data and reducing the curvature, even when part of the information of the feature part is missing, the relative position of the feature part can be estimated using the past history, and the position where the feature part is presumed to exist can be continuously tracked.

[0074] Also, when there is no record of the feature part being detected in the past, by setting the relative position of the feature part to 0, the prediction formula L′(t) of the feature part is the following formula (2) L'(t) = L(t) …(2) and the focus detection result of the main frame becomes the prediction result of the feature part.

[0075] As described above, according to the present embodiment, by using the past focus detection history of a plurality of focus detection regions included in the same subject, even when there is a temporary state change of the subject, it is possible to continuously focus on the detection site with high priority.

[0076] In the above description, the processing is described as being performed using the defocus amount. However, even if it is other than the defocus amount, as long as it is information indicating the focus state, the above-described processing can be performed using the information.

[0077] Also, in the above description, the prediction process is described as being performed based on the image plane position. However, the present invention is not limited to this. For example, it is also possible to perform prediction using information indicating the focus state such as the defocus amount and the focal length. Note that the image plane position is also included in the focus state in a broad sense.

[0078] <Other embodiments> Note that the present invention may be applied to a system composed of a plurality of devices or to an apparatus composed of a single device.

[0079] In addition, the present invention can also be realized by supplying a program that realizes one or more functions of the above-described embodiments to a system or apparatus via a network or a storage medium, and causing one or more processors in a computer of the system or apparatus to read and execute the program. It can also be realized by a circuit (for example, ASIC) that realizes one or more functions.

[0080] <Summary> The disclosure of the present embodiment includes the following configurations.

[0081] (Item 1) Subject detection means for detecting a subject and a feature region included in the subject from an image of an image signal repeatedly output from an imaging device; Setting means for setting a plurality of focus detection regions of a predetermined size on the subject and the feature region detected by the subject detection means; Detection means for detecting a focus state in each of the plurality of focus detection regions set by the setting means based on the image signal; Selection means for selecting a first region that satisfies a predetermined condition and a second region that includes the feature region from among the plurality of focus detection regions; Prediction means for predicting a focus state of the second region at an arbitrary time from a first history of the focus state of the first region and a second history of a difference between the focus state of the first region and the focus state of the second region based on the focus state detected by the detection means; Obtaining means for obtaining a driving amount of a focus adjustment means for focusing on the second region based on the focus state of the second region predicted by the prediction means A focus detection device, characterized by comprising: (Item 2) The prediction means predicts the in-focus state of the first area at the arbitrary time based on the first history, and predicts the difference at the arbitrary time based on the second history. The focus detection device according to item 1, characterized in that. (Item 3) The prediction means predicts the in-focus state of the first area and the difference at the arbitrary time by the batch least squares method or the sequential least squares method based on the first history and the second history. The focus detection device according to item 2, characterized in that. (Item 4) It further has an estimation means for estimating the depth width of the subject, The prediction means predicts the difference at the arbitrary time by predicting the ratio of the depth width and the difference by the batch least squares method or the sequential least squares method. The focus detection device according to item 2, characterized in that. (Item 5) It further has an estimation means for estimating the depth width of the subject, The prediction means predicts the difference at the arbitrary time by approximating the depth width of the subject to a single vibration model with the amplitude. The focus detection device according to item 2, characterized in that. (Item 6) The prediction means predicts the in-focus state of the second area at the arbitrary time by adding the predicted in-focus state of the first area and the difference at the arbitrary time. The focus detection device according to any one of items 2 to 5, characterized in that. (Item 7) The estimation means classifies the plurality of focus detection areas into a plurality of depths with a predetermined depth width based on the in-focus states of the plurality of focus detection areas, and estimates the depth width obtained by summarizing the depths including the plurality of focus detection areas as the depth width of the subject. The focus detection device according to item 4 or 5, characterized in that. (Item 8) The estimation means performs learning for estimating the depth width of the subject based on the type of the subject, the shooting distance, and the image information, and estimates the depth width of the subject. The focus detection device according to item 4 or 5, characterized in that. (Item 9) When there is no such second history, the prediction means predicts the relative position of the feature region at the arbitrary time as 0, and the focus detection device according to any one of items 1 to 8. (Item 10) When the plurality of focus detection regions are equal to or greater than a predetermined first threshold value, the selection means classifies the plurality of focus detection regions into a plurality of depths having a predetermined depth width based on the focus states of the plurality of focus detection regions, and selects the first region based on the classification result. The focus detection device according to any one of items 1 to 9. (Item 11) When the ratio of the focus detection regions at the depth where the focus detection regions are most classified among the plurality of depths is equal to or greater than a predetermined second threshold value, among the focus detection regions classified at the depth, the focus detection region closest to the center of the plurality of focus detection regions is selected as the first region. The focus detection device according to item 10. (Item 12) When the plurality of focus detection regions are less than a predetermined first threshold value, and when the ratio of the focus detection regions at the depth where the focus detection regions are most classified among the plurality of depths is less than the second threshold value, if the focus detection region closest to the center of the plurality of focus detection regions is included in the depth width of a predetermined depth, the focus detection device according to item 11, wherein the focus detection region is selected as the first region. (Item 13) When the plurality of focus detection regions are less than a predetermined first threshold value, and when the ratio of the focus detection regions at the depth where the focus detection regions are most classified among the plurality of depths is less than the second threshold value, if the focus detection region closest to the center of the plurality of focus detection regions is not included in the depth width of a predetermined depth, the focus detection device according to item 11 or 12, wherein the focus detection region showing the closest focus state in the depth width of the predetermined depth is selected as the first region. (Item 14) The focus detection device according to any one of Items 1 to 13, wherein when the subject is a person, the feature region is the pupil of the person. (Item 15) The focus detection device according to any one of Items 1 to 14, wherein the focus state includes the image plane position. (Item 16) The focus detection device according to any one of Items 1 to 14, wherein the focus state includes the defocus amount. (Item 17) Imaging means, The focus detection device according to any one of Items 1 to 16, An imaging device, characterized by comprising the same. (Item 18) The focus adjustment means, Drive means for driving the focus adjustment means by the drive amount obtained by the acquisition means at any given time, The imaging device according to Item 17, further comprising the same. (Item 19) A subject detection step of detecting a subject and a feature region included in the subject from an image of an image signal repeatedly output from an imaging device; A setting step of setting a plurality of focus detection regions of a predetermined size for the subject and the feature region detected in the subject detection step; A detection step of detecting a focus state in each of the plurality of focus detection regions set in the setting step based on the image signal; A selection step of selecting a first region that satisfies a predetermined condition and a second region that includes the feature region from among the plurality of focus detection regions; A prediction step of predicting a focus state of the second region at an arbitrary time from a first history of the focus state of the first region and a second history of a difference between the focus state of the first region and the focus state of the second region based on the focus state detected in the detection step; An acquisition step of obtaining a drive amount of focus adjustment means for focusing on the second region based on the focus state of the second region predicted in the prediction step A focus detection method characterized by having (Item 20) A program for causing a computer to function as each means of the focus detection apparatus according to any one of Items 1 to 16. (Item 21) A computer-readable storage medium storing the program according to Item 20.

[0082] The invention is not limited to the above embodiments, and various changes and modifications are possible without departing from the spirit and scope of the invention. Therefore, claims are appended to disclose the scope of the invention.

Description of Reference Numerals

[0083] 100: Lens device, 103: Focus lens, 105: Lens controller, 200: Camera body, 201: Image sensor, 204: AF signal processing unit, 209: System control unit, 210: Lens communication unit, 211: Subject detection unit, 219: Subject movement prediction unit

Claims

1. Subject detection means for detecting a subject and a feature area included in the subject from an image of an image signal repeatedly output from an imaging device; Setting means for setting a plurality of focus detection areas of a predetermined size for the subject and the feature area detected by the subject detection means; Detection means for detecting a focus state in each of the plurality of focus detection areas set by the setting means based on the image signal; Selection means for selecting a first area that satisfies a predetermined condition and a second area that includes the feature area from among the plurality of focus detection areas; Prediction means for predicting the focus state of the second area at an arbitrary time from a first history of the focus state of the first area and a second history of the difference between the focus state of the first area and the focus state of the second area based on the focus state detected by the detection means; Acquisition means for obtaining a driving amount of a focus adjustment means for focusing on the second area based on the focus state of the second area predicted by the prediction means A focus detection device characterized by comprising the same.

2. The prediction means predicts the focus state of the first area at the arbitrary time based on the first history, and predicts the difference at the arbitrary time based on the second history. The focus detection device according to claim 1.

3. The prediction means predicts the focus state of the first area and the difference at the arbitrary time by the batch least squares method or the sequential least squares method based on the first history and the second history. The focus detection device according to claim 2.

4. Further comprising estimation means for estimating the depth width of the subject, The prediction means predicts the difference at the arbitrary time by predicting a ratio of the depth width and the difference by the batch least squares method or the sequential least squares method. The focus detection device according to claim 2.

5. Further comprising estimation means for estimating the depth width of the subject, The prediction means predicts the difference at the arbitrary time by approximating a simple vibration model having the depth width of the subject as an amplitude. The focus detection device according to claim 2.

6. The focus detection device according to claim 2, wherein the prediction means predicts the focus state of the second region at an arbitrary time by adding the predicted focus state of the first region and the difference at the arbitrary time.

7. The estimation means according to claim 4, wherein based on the focus states of the plurality of focus detection regions, the plurality of focus detection regions are classified into a plurality of depths with a predetermined depth width, and a depth width obtained by summing up the depths including the plurality of focus detection regions is estimated as the depth width of the subject.

8. The focus detection device according to claim 4, wherein the estimation means performs learning for estimating the depth width of a subject based on the type of the subject, the shooting distance, and the image information, and estimates the depth width of the subject.

9. The focus detection device according to claim 5, wherein the estimation means classifies the plurality of focus detection regions into a plurality of depths with a predetermined depth width based on the focus states of the plurality of focus detection regions, and estimates a depth width obtained by summing up the depths including the plurality of focus detection regions as the depth width of the subject.

10. The focus detection device according to claim 5, wherein the estimation means performs learning for estimating the depth width of a subject based on the type of the subject, the shooting distance, and the image information, and estimates the depth width of the subject.

11. The focus detection device according to claim 1, wherein when there is no second history, the prediction means predicts the relative position of the feature region at the arbitrary time as 0.

12. The focus detection device according to claim 1, wherein when the plurality of focus detection regions are equal to or greater than a predetermined first threshold, the selection means classifies the plurality of focus detection regions into a plurality of depths with a predetermined depth width based on the focus states of the plurality of focus detection regions, and selects the first region based on the classification result.

13. The focus detection device according to claim 12, wherein when the ratio of the focus detection regions in the depth where the focus detection regions are most classified among the plurality of depths is equal to or greater than a predetermined second threshold, among the focus detection regions classified into the depth, the focus detection region closest to the center of the plurality of focus detection regions is selected as the first region.

14. The selection means selects, as the first region, a focus detection region included in a depth range of a predetermined depth when the number of the plurality of focus detection regions is less than a predetermined first threshold value, and when a ratio of the focus detection regions at a depth at which the focus detection regions are most classified among the plurality of depths is smaller than the second threshold value, and when a focus detection region closest to the center of the plurality of focus detection regions is included in the depth range of the predetermined depth. The focus detection device according to claim 13, characterized in that

15. The selection means selects, as the first region, a focus detection region indicating the closest focus state within a depth range of a predetermined depth when the number of the plurality of focus detection regions is less than a predetermined first threshold value, and when a ratio of the focus detection regions at a depth at which the focus detection regions are most classified among the plurality of depths is smaller than the second threshold value, and when a focus detection region closest to the center of the plurality of focus detection regions is not included in the depth range of the predetermined depth. The focus detection device according to claim 13, characterized in that

16. The focus detection device according to claim 1, characterized in that when the subject is a person, the feature region is the pupil of the person.

17. The focus detection device according to claim 1, characterized in that the focus state includes an image plane position.

18. The focus detection device according to claim 1, characterized in that the focus state includes a defocus amount.

19. Imaging means, A focus detection device according to any one of claims 1 to 18 An imaging device, characterized by comprising

20. The focus adjustment means, At any time, driving means for driving the focus adjustment means by the driving amount obtained by the acquisition means The imaging device according to claim 19, further comprising

21. A subject detection step of detecting a subject and a feature region included in the subject from an image of an image signal repeatedly output from an imaging device; A setting step of setting a plurality of focus detection regions having a predetermined size for the subject and the feature region detected in the subject detection step; A detection step of detecting a focus state in each of the plurality of focus detection regions set in the setting step based on the image signal; A selection step of selecting a first region satisfying a predetermined condition and a second region including the feature region from among the plurality of focus detection regions; A prediction step of predicting a focus state of the second region at an arbitrary time from a first history of the focus state of the first region and a second history of a difference between the focus state of the first region and the focus state of the second region, based on the focus state detected in the detection step; An acquisition step of obtaining a drive amount of a focus adjustment means for focusing on the second region, based on the focus state of the second region predicted in the prediction step; A focus detection method characterized by comprising: **Claim 22** A program for causing a computer to function as each means of the focus detection apparatus according to any one of claims 1 to 18. **Claim 23** A computer-readable storage medium storing the program according to claim 22.

Citation Information

Patent Citations

  • Imaging apparatus and method for controlling the same, program, and storage medium

    JP2021173803A

  • Focus adjustment device, control method thereof, program, storage medium, and imaging device

    JP7066388B2