Focus detection device, focus detection method, and program
The focus detection device uses a machine learning model to analyze brightness and phase difference data from dual pixel groups to enhance autofocus accuracy and speed in various challenging scenes.
Patent Information
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- SAMSUNG ELECTRONICS CO LTD
- Filing Date
- 2024-10-29
- Publication Date
- 2026-05-15
Smart Images

Figure 2026078622000001_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to a focus detection device, a focus detection method, and a program.
Background Art
[0002] In cameras of smartphones and the like, an image sensor capable of detecting a phase difference is used. Patent Document 1 discloses a technique for detecting an image shift amount using a neural network that outputs a defocus function or the like when phase difference pixel data is input.
Prior Art Documents
Patent Documents
[0003]
Patent Document 1
Summary of the Invention
Problems to be Solved by the Invention
[0004] Conventional phase difference detection methods have a problem that the detection accuracy is low in various scenes such as a scene where the defocus amount is too large, a low-contrast or low-illuminance scene, and a scene where a horizontal stripe pattern appears.
[0005] The present invention has been made to solve such problems, and an object thereof is to provide a focus detection device, a focus detection method, and a program that achieve fast and accurate autofocus in more scenes.
Means for Solving the Problems
[0006] A focus detection device according to an embodiment includes: an image sensor including first and second pixel groups for phase difference detection; A processing unit that uses a machine learning model to detect the main subject and the amount of defocus from the output of the brightness of the first and second pixel groups, the result of a phase calculation based on the brightness output, and the RGB information based on the output of the image sensor. It is equipped with.
[0007] One embodiment of the focus detection method is: A machine learning model is used to detect the main subject and the amount of defocus from the output of the brightness of the first and second pixel groups for phase difference detection, the result of a phase calculation based on the brightness output, and RGB information based on the output of the image sensor including the first and second pixel groups. Includes.
[0008] The program of one embodiment is: On the computer, A machine learning model is used to detect the main subject and the amount of defocus from the output of the brightness of the first and second pixel groups for phase difference detection, the result of a phase calculation based on the brightness output, and RGB information based on the output of the image sensor including the first and second pixel groups. Make it run. [Effects of the Invention]
[0009] The present invention provides a focus detection device, a focus detection method, and a program that enable high-speed and accurate autofocus in a wider range of scenes. [Brief explanation of the drawing]
[0010] [Figure 1] This is a block diagram illustrating the configuration of the imaging device according to Embodiment 1. [Figure 2] This is a schematic diagram illustrating training data. [Figure 3] This is a block diagram illustrating the configuration of the machine learning model according to Embodiment 1. [Figure 4] This is a flowchart illustrating the operation of the main subject / defocus amount detection unit according to Embodiment 1. [Figure 5] This is a schematic diagram showing the verification results. [Figure 6] This is a schematic diagram showing the verification results. [Modes for carrying out the invention]
[0011] For clarity, the following descriptions and drawings have been omitted and simplified as appropriate. Furthermore, the same elements are denoted by the same reference numerals in each drawing, and redundant explanations have been omitted where necessary.
[0012] Embodiment 1 Figure 1 is a block diagram showing the configuration of an imaging device 10 according to Embodiment 1. The imaging device 10 includes a lens 11, an image sensor 12, an image sensor control unit 13, an image signal processing unit 14, an image storage unit 15, an image display unit 16, a phase difference calculation processing unit 17, a main subject / defocus amount detection unit 18, an autofocus main control unit 19, and a lens control unit 110.
[0013] The imaging device 10 may include a processor and memory, although these are not shown in the diagram. The memory may include storage located separately from the processor. The memory may store machine learning models, such as neural networks.
[0014] The functions of each component of the imaging device 10 may be realized by a processor executing a program loaded into memory. Alternatively, each component of the imaging device 10 may be realized by dedicated hardware such as circuits or semiconductor chips. The functions of each component may also be realized by a combination of hardware and software.
[0015] The light passing through the lens 11 reaches the light-receiving surface of the imaging device 12 and forms an image of the subject. The imaging device 12 may be a CCD (Charge Coupled Device) image sensor or a CMOS (Complementary Metal Oxide Semiconductor) image sensor that converts an optical signal into an electrical signal. A plurality of RGB pixels are arranged on the light-receiving surface of the imaging device 12.
[0016] The imaging device 12 includes a first pixel group and a second pixel group for phase difference detection. A pair of pixels for phase difference detection is also referred to as a phase difference pixel. The first pixel group corresponds to one of the pair of phase difference pixels, and the second pixel group corresponds to the other of the pair of phase difference pixels. The pair of phase difference pixels may be arranged side by side horizontally or vertically.
[0017] Specifically, the imaging device 12 may be a 2PD (Photodiode) sensor. In this case, each RGB pixel includes a pair of photodiodes, a color filter, and a microlens. With the 2PD sensor, accurate and fast autofocus can be achieved without degrading the image quality. Note that the imaging device 12 is not limited to the 2PD sensor, and a plurality of pairs of phase difference pixels may be arranged among the RGB pixels.
[0018] The imaging device control unit 13 controls the imaging by the imaging device 12. The imaging device control unit 13 may control the sensitivity of the imaging device 12 according to a control signal generated automatically or a control signal input manually.
[0019] The image signal processing unit 14 performs analog-to-digital conversion processing on the analog signal output from the imaging device 12 to obtain outputs (also referred to as Raw data) regarding the luminance of the first and second pixel groups. The image signal processing unit 14 outputs the Raw data to the phase difference calculation processing unit 17.
[0020] Furthermore, the image signal processing unit 14 generates an captured image. Specifically, the image signal processing unit 14 generates a single captured image by combining the raw data output from the first and second pixel groups. Alternatively, the image signal processing unit 14 may generate an captured image according to the output of RGB pixels prepared separately from the pixels for phase difference detection. The image signal processing unit 14 performs processing on the captured image, such as noise reduction, demosaicing, and auto white balance. The image signal processing unit 14 outputs the captured image to the image storage unit 15 and the image display unit 16.
[0021] The auto white balance described above will now be explained. The image signal processing unit 14 divides the captured image into multiple blocks and calculates statistical information (e.g., average value, integrated value) of the RGB pixel values for each block as RGB information. The image signal processing unit 14 then performs auto white balance using a correction coefficient based on the RGB information. The image signal processing unit 14 outputs the RGB information to the main subject / defocus amount detection unit 18. The image signal processing unit 14 may also be equipped with hardware for calculating the RGB information. Alternatively, the image signal processing unit 14 may output information indicating the RGB values for each pixel of the captured image as RGB information to the main subject / defocus amount detection unit 18.
[0022] The image storage unit 15 records the captured image. The image storage unit 15 may be composed of a storage device such as memory or an SSD (Solid State Drive).
[0023] The image display unit 16 displays the captured image. The image display unit 16 may be composed of a display device such as an organic EL (Electro-Luminescence) display or a liquid crystal display.
[0024] The phase difference calculation processing unit 17 performs phase calculations based on raw data. Note that if the image sensor 12 is a 2PD sensor, it is not necessary to perform phase difference calculations using all phase difference pixels. The phase difference calculation processing unit 17 outputs the result of the phase calculation to the main subject / defocus amount detection unit 18. Note that the function for performing phase difference calculations may be provided in the main subject / defocus amount detection unit 18, and the phase difference calculation processing unit 17 may be omitted.
[0025] The phase difference calculation processing unit 17, for example, fixes the raw data of an arbitrary line in the first pixel group and calculates the difference or correlation between the two raw data sets while shifting the raw data of the corresponding line in the second pixel group left or right. The difference or correlation is an example of the result of the phase calculation. The result of the phase calculation may also be a phase difference. The result of the phase calculation may include calculation results for each of multiple lines.
[0026] The main subject / defocus amount detection unit 18 uses a machine learning model to detect one or more main subjects and the amount of defocus relative to the main subjects from the raw data, the results of the phase calculation, and the RGB information. The main subject / defocus amount detection unit 18 outputs the detected defocus amount to the autofocus main control unit 19. The main subject / defocus amount detection unit 18 corresponds to the calculation processing unit described above.
[0027] Figure 2 is a schematic diagram illustrating training data for a machine learning model. Each training data set 20 provides, as ground truth values, a primary subject (1 or more) and a defocus amount for the primary subject, for the raw data output from the first pixel group (e.g., the left phase-difference pixel group) and the second pixel group (e.g., the right phase-difference pixel group), the result of phase calculations based on the raw data, and RGB information. A primary subject is a subject that is expected to be the target of autofocus. The ground truth values may also include the position and size of the primary subject.
[0028] When collecting training data, one possible method is to fix the position of the imaging device 10 and the position of the main subject, and gradually move the position of the lens 11. Since the position of the imaging device 10 and the position of the main subject are fixed, the correct value of the defocus amount can be determined from the drive amount of the lens 11. Note that the method of collecting training data is arbitrary, and training data may be collected by means of simulation, for example.
[0029] A machine learning model can be trained to take raw data, the results of phase calculations, and RGB information as input data, and use the main subject and the amount of defocus as the ground truth values, so that it takes the input data as input and outputs the ground truth values. The machine learning model can be implemented by, for example, a neural network, but is not limited to this. The training of the machine learning model is achieved by executing a predetermined program on a processor. The training of the machine learning model may be performed by the imaging device 10, but it may also be performed by a computer different from the imaging device 10.
[0030] Next, an example of processing of raw data performed during inference using a machine learning model will be described. Note that this processing may also be performed during training to generate the machine learning model. The main subject / defocus amount detection unit 18 performs a process of taking the difference between the brightness value of each pixel in the first pixel group and the second pixel group and the brightness value of a pixel located a predetermined number of pixels (e.g., 4 pixels, 8 pixels) away from that pixel in a predetermined direction, and generates new raw data based on the difference. For example, the main subject / defocus amount detection unit 18 may subtract the brightness value of a pixel located a predetermined number of pixels to the right from the brightness value of each pixel in the first pixel group on the left, and set the subtraction result as the new brightness value. For example, the main subject / defocus amount detection unit 18 may subtract the brightness value of a pixel located a predetermined number of pixels to the left from the brightness value of each pixel in the second pixel group on the right, and set the subtraction result as the new brightness value. This difference processing makes it possible to emphasize the contrast of components of a specific frequency, making it easier to detect the main subject and the amount of defocus. Furthermore, due to shading of lens 11, i.e., brightness unevenness, the output of the left phase-difference pixel has a downward offset to the right, and the output of the right phase-difference pixel has an upward offset to the right. However, this offset can be removed by differential processing. Differential processing can reduce the time required for correction compared to the normal process of multiplying the output of the image sensor 12 by a correction coefficient.
[0031] Next, we will provide a supplementary explanation regarding RGB information. As already explained, RGB information may be statistical information (e.g., average value) of the RGB values for each block calculated for auto white balance. Alternatively, RGB information may be information indicating the RGB values for each pixel of the captured image. However, using the RGB values for each block allows for faster detection of the amount of defocus.
[0032] Figure 3 is a block diagram illustrating the configuration of the machine learning model 30. The encoding layer 31 of the machine learning model 30 receives, for example, 2PD raw data, the result of phase calculation, and RGB information as input data. The encoding layer 31 extracts a feature map 32 from the input data. The feature map 32 is matrix-like data that represents the features of the input data. The decoding layer 33 outputs a probability map that shows the probability that each pixel is classified as the main subject based on the feature map 32. The detection layer 34 outputs a defocus amount based on the feature map 32 and the probability map.
[0033] The main subject / defocus amount detection unit 18 inputs raw data, phase calculation results, and RGB information to the machine learning model 30. As a result, the machine learning model 30 outputs the defocus amount.
[0034] Referring again to Figure 1, the imaging device 10 may be further equipped with a ToF (Time of Flight) sensor and a LIDAR (Light Detection and Ranging) sensor, which are not shown. The input data for the machine learning model 30 may further include at least one of the sensing results from the ToF sensor and the sensing results from the LIDAR.
[0035] The main subject / defocus amount detection unit 18 may pre-resize the data (e.g., raw data) input to the machine learning model 30. This allows the main subject / defocus amount detection unit 18 to support image sensors 12 with various pixel counts and to detect the amount of defocus more quickly.
[0036] Figure 4 is a flowchart illustrating the operation of the main subject / defocus amount detection unit 18. First, the main subject / defocus amount detection unit 18 acquires raw data, the result of phase calculation, and RGB information (step S101). Next, the main subject / defocus amount detection unit 18 performs a process to take the difference between the brightness value of each pixel in the raw data and the brightness value of a pixel far from that pixel (step S102). Next, the main subject / defocus amount detection unit 18 resizes the data (e.g., raw data) to be input to the machine learning model (step S103). Next, the main subject / defocus amount detection unit 18 inputs the resized data to the machine learning model 30 (step S104), and the machine learning model 30 outputs the main subject area and the defocus amount (step S105). Note that the order of steps S102 and S103 may be reversed. It is also possible to resize the raw data before performing the difference processing.
[0037] Referring again to Figure 1, the autofocus main control unit 19 outputs a control signal to the lens control unit 110 based on the defocus amount provided by the main subject / defocus amount detection unit 18.
[0038] The lens control unit 110 adjusts the position of the lens 11 based on control signals provided by the autofocus main control unit 19. The lens control unit 110 may also consist of a motor that drives the lens 11 according to the control signals.
[0039] Referring to Figures 5 and 6, the effects of Embodiment 1 will be explained. The graph on the left shows the results of detecting the amount of defocus using the conventional technology, and the graph on the right shows the results of detecting the amount of defocus using Embodiment 1. Since the conventional technology does not detect the main subject, the position of the main subject was manually set within the frame. In each graph, the horizontal axis represents the position of the lens 11, and the vertical axis represents the amount of defocus. In each graph, the straight line represents the correct value, also called ground truth, and the curve represents the detection result of Embodiment 1.
[0040] Figure 5 shows the verification results for scenes with moiré patterns and low-light scenes with a BV of -6.47. Figure 6 shows the verification results for low-contrast scenes and scenes with horizontal stripes. From Figures 5 and 6, it can be seen that Embodiment 1 can improve the accuracy of defocus detection in various scenes that conventional technology struggles with.
[0041] Embodiment 1 is configured to detect the main subject and the amount of defocus based on raw data, phase calculation results, and RGB information, thereby improving the accuracy of defocus detection. Furthermore, Embodiment 1 can handle scenes where various subjects are located at different positions. Since Embodiment 1 detects the main subject and the amount of defocus simultaneously, processing time can be reduced.
[0042] In the examples described above, the program includes a set of instructions (or software code) that, when loaded into a computer, cause the computer to perform one or more of the functions described in the embodiments. The program may be stored on a non-temporary computer-readable medium or a physical storage medium. Examples, but not limited to, include RAM (random-access memory), ROM (read-only memory), flash memory, SSD (solid-state drive), or other memory technologies, CD-ROM, DVD (digital versatile disc), Blu-ray® disc, or other optical disc storage, magnetic cassette, magnetic tape, magnetic disk storage, or other magnetic storage devices. The program may be transmitted over a temporary computer-readable medium or a communication medium. Examples, but not limited to, include, a temporary computer-readable medium or a communication medium that includes an electrical, optical, acoustic, or other form of propagating signal.
[0043] The present invention is not limited to the embodiments described above, and can be modified as appropriate without departing from the spirit of the invention. [Explanation of Symbols]
[0044] 10 Imaging device 11 lenses 12 Image sensor 13 Image sensor control unit 14 Image signal processing unit 15 Image storage section 16 Image display section 17 Phase difference calculation processing unit 18 Main subject / defocus amount detection unit 19 Autofocus Main Control Unit 110 Lens control unit 20 Training Data 30 Machine Learning Models 31 Encoding Layers 32 Feature Map 33 Decode Layer 34 detection layer
Claims
1. An image sensor including first and second pixel groups for phase difference detection, A processing unit that uses a machine learning model to detect the main subject and the amount of defocus from the output of the brightness of the first and second pixel groups, the result of a phase calculation based on the brightness output, and the RGB information based on the output of the image sensor. A focus detection device equipped with the following features.
2. Before detecting the main subject and the amount of defocus, the processing unit performs a process to take the difference between the brightness value of each pixel in the first and second pixel groups and the brightness value of a pixel located a predetermined number of pixels away from the pixel in a predetermined direction. The focus detection device according to claim 1.
3. The RGB information represents statistical information of the RGB pixel values of each of the multiple blocks into which the captured image is divided, and is used in the auto white balance calculation. The focus detection device according to claim 1 or 2.
4. The processing unit resizes the data input to the machine learning model before detecting the main subject and the amount of defocus. The focus detection device according to claim 1 or 2.
5. The training data for the machine learning model includes, as ground truth values, one or more main subject regions and the amount of defocus for the main subject regions. The focus detection device according to claim 1 or 2.
6. A machine learning model is used to detect the main subject and the amount of defocus from the output of the brightness of the first and second pixel groups for phase difference detection, the result of a phase calculation based on the brightness output, and RGB information based on the output of the image sensor including the first and second pixel groups. A focus detection method including the following.
7. On the computer, A machine learning model is used to detect the main subject and the amount of defocus from the output of the brightness of the first and second pixel groups for phase difference detection, the result of a phase calculation based on the brightness output, and RGB information based on the output of the image sensor including the first and second pixel groups. A program that executes the command.