System and method for predicting focus convergence direction using machine learning model

By using a machine learning model to predict the direction of lens movement and combining it with PDAF technology, the problem of image blurring caused by inaccurate lens positioning was solved, achieving a more efficient and accurate autofocus effect.

CN122070700APending Publication Date: 2026-05-19QUALCOMM INC
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
QUALCOMM INC
Filing Date
2023-10-27
Publication Date
2026-05-19

AI Technical Summary

Technical Problem

Existing imaging systems often fail to autofocus effectively when lens positioning is inaccurate, especially phase detection autofocus (PDAF), which traditionally cannot work accurately, resulting in blurred or out-of-focus images.

Method used

A trained machine learning model is used to process the focusing data to predict the direction of lens movement. The lens position is adjusted by an actuator to improve image focusing. The focusing accuracy is improved by combining phase detection autofocus (PDAF) and machine learning model.

Benefits of technology

When the lens positioning is inaccurate, it achieves more efficient and accurate autofocus, which is superior to traditional PDAF and improves image clarity and focusing effect.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122070700A_ABST
    Figure CN122070700A_ABST
Patent Text Reader

Abstract

Imaging systems and techniques are described. In some examples, an imaging system processes at least focus data (e.g., and in some cases, image data) using a trained machine learning model to identify predicted movement directions (e.g., phase shift symbols) for moving a lens to improve image focusing (e.g., positioning toward a focus lens). The focus data is associated with at least one focus pixel of image data captured using the image sensor. An imaging system moves a lens in a predicted direction of movement (e.g., using an actuator) to improve image focusing. In some examples, an imaging system captures an image using an image sensor after moving a lens in a predicted direction of movement.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to imaging. More specifically, this application relates to systems and methods for automatically predicting the focusing convergence direction of a camera using a machine learning model. Background Technology

[0002] Many devices include one or more cameras. For example, a smartphone or tablet includes a front-facing camera for capturing selfie images and a rear-facing camera for capturing scene images (such as landscapes or other scenes of interest to the device user). A camera can use its image sensor to capture images, which may include an array of photodetectors. Some devices can analyze the image data captured by the image sensor to detect objects within that image data. Sometimes, a camera can be used to capture images of a scene that includes one or more people. Summary of the Invention

[0003] An imaging system and techniques are described. In some examples, the imaging system uses a trained machine learning model to process at least the focus data (e.g., and in some cases, image data) to identify a predicted movement direction (e.g., phase shift sign) for moving the lens to improve image focus (e.g., positioning towards the focusing lens). The focus data is associated with at least one focused pixel of image data captured using an image sensor. The imaging system moves the lens in the predicted movement direction (e.g., using an actuator) to improve image focus. In some examples, the imaging system uses an image sensor to capture an image after moving the lens in the predicted movement direction.

[0004] In another example, an apparatus for imaging is provided, the apparatus including at least one memory and at least one processor coupled to the at least one memory. The at least one processor is configured to: process at least focus data using a trained machine learning model to identify a predicted movement direction for moving a lens to improve image focus, wherein the focus data is associated with at least one focused pixel of image data captured using an image sensor; and move the lens in the predicted movement direction to improve image focus.

[0005] According to at least one example, a method for imaging is provided. The method includes: processing at least focus data using a trained machine learning model to identify a predicted movement direction for moving a lens to improve image focus, wherein the focus data is associated with at least one focus pixel of image data captured using an image sensor; and moving the lens in the predicted movement direction to improve image focus.

[0006] In another example, a non-transitory computer-readable medium is provided having instructions stored thereon, which, when executed by one or more processors, cause the one or more processors to: use a trained machine learning model to process at least focus data to identify a predicted direction of movement for moving a lens to improve image focus, wherein the focus data is associated with at least one focused pixel of image data captured using an image sensor; and to move the lens in the predicted direction of movement to improve image focus.

[0007] In another example, an apparatus for imaging is provided. The apparatus includes: components for processing at least focus data using a trained machine learning model to identify a predicted movement direction for moving a lens to improve image focus, wherein the focus data is associated with at least one focused pixel of image data captured using an image sensor; and components for moving the lens in the predicted movement direction to improve image focus.

[0008] In some aspects, the device is part of and / or includes the following: wearable devices, extended reality devices (e.g., virtual reality (VR) devices, augmented reality (AR) devices, or mixed reality (MR) devices), head-mounted display (HMD) devices, wireless communication devices, mobile devices (e.g., mobile phones and / or mobile cell phones and / or so-called "smartphones" or other mobile devices), cameras, personal computers, laptop computers, server computers, vehicles or computing devices or components of vehicles, another device, or combinations thereof. In some aspects, the device includes one or more cameras for capturing one or more images. In some aspects, the device also includes a display for displaying one or more images, notifications, and / or other displayable data. In some aspects, the device described above may include one or more sensors (e.g., one or more inertial measurement units (IMUs), such as one or more gyroscopes, one or more gyroscope testers, one or more accelerometers, any combination thereof, and / or other sensors).

[0009] This summary is not intended to identify key or essential features of the claimed subject matter, nor is it intended to be used in isolation to define the scope of the claimed subject matter. This subject matter should be understood with reference to the appropriate portions of the entire specification, any or all drawings, and each claim.

[0010] The foregoing and other features and aspects will become more apparent from the following description, claims and accompanying drawings. Attached Figure Description

[0011] The exemplary aspects of this application are described in detail below with reference to the following figures:

[0012] Figure 1 This is a block diagram illustrating an example architecture of an image capture and processing system based on some examples;

[0013] Figure 2A This is a cross-sectional view illustrating a phase detection autofocus (PDAF) camera system that is in phase and therefore in "focus" state according to some examples;

[0014] Figure 2B This is an example based on some examples. Figure 2A A cross-sectional view of a PDAF camera system in "front-focus" mode with different phases;

[0015] Figure 2C This is an example based on some examples. Figure 2A A cross-sectional view of a PDAF camera system in "back-focus" mode with different phases;

[0016] Figure 3A This is a top view illustration of a pixel array of an image sensor having a mask that partially covers the photodiodes of the focused pixels, according to some examples;

[0017] Figure 3B It is based on some example identifiers Figure 3A Legend of the components in the diagram;

[0018] Figure 3C This is a top view illustration of a pixel array of an image sensor, according to some examples, having two side-by-side focusing pixels covered by a 2-pixel by 1-pixel microlens;

[0019] Figure 3D This is a top view illustration of a pixel array of an image sensor, according to some examples, having four adjacent focusing pixels covered by a 2-pixel by 2-pixel microlens;

[0020] Figure 3E This is a top view illustration of a pixel array of an image sensor according to some examples, wherein at least one focused pixel has two photodiodes;

[0021] Figure 3F This is a top view illustration of a pixel array of an image sensor according to some examples, wherein at least one focused pixel has four photodiodes;

[0022] Figure 4A This is a cross-sectional side view of a photodiode of an image sensor according to some examples, which is partially covered by a mask to represent a focused pixel;

[0023] Figure 4BThis is a cross-sectional side view of two photodiodes in a pixel array of an image sensor according to some examples. These two photodiodes are covered by 2-pixel by 1-pixel microlenses, representing the focused pixels.

[0024] Figure 5 This is a graph illustrating the phase shift detected by phase detection autofocus (PDAF), the corresponding PDAF confidence, the predicted direction determined using a trained machine learning model, and the corresponding prediction direction confidence, based on some examples.

[0025] Figure 6 This is a block diagram illustrating a machine learning system that uses, uses (e.g., infers), and updates (e.g., further trains) an ML model associated with a machine learning (ML) prediction engine to determine the predicted direction of moving a lens to improve focusing, based on some examples.

[0026] Figure 7 This is a block diagram illustrating an imaging system including an ML prediction engine, a PDAF engine, and a focusing convergence engine, based on some examples.

[0027] Figure 8 This is a block diagram illustrating, based on some examples, the process for determining lens movement for autofocus by a focusing convergence engine based on an ML prediction engine and a PDAF engine;

[0028] Figure 9 This is a block diagram illustrating an imaging system based on some examples, including an ML prediction engine, a PDAF engine, a CDAF and a focus value (FV) statistics engine, and a focus convergence engine.

[0029] Figure 10 This is a block diagram illustrating, based on some examples, the process for determining lens movement for autofocus by a focusing convergence engine based on an ML prediction engine, a PDAF engine, and a CDAF and FV statistical engine.

[0030] Figure 11 This is a block diagram illustrating various aspects of the process for determining lens movement for autofocus by a focusing and convergence engine, based on some examples.

[0031] Figure 12 This is a block diagram illustrating examples of neural networks that can be used for imaging operations, based on some examples;

[0032] Figure 13A This is a perspective view illustrating a head-mounted display (HMD) used as part of an imaging system, based on some examples;

[0033] Figure 13B This is an example based on some examples. Figure 13AThe head-mounted display (HMD) is viewed through the perspective of the user wearing it;

[0034] Figure 14A This is a perspective view illustrating the front surface of a mobile phone, including a front-facing camera and used as part of an imaging system, based on some examples.

[0035] Figure 14B This is a perspective view of the rear surface of a mobile phone, exemplified by some examples, including a rear camera and which can be used as part of an imaging system;

[0036] Figure 15 This is a flowchart illustrating a process for imaging based on some examples; and

[0037] Figure 16 This is a diagram illustrating an example of a computing system used to implement some of the aspects described in this article. Detailed Implementation

[0038] Certain aspects of this disclosure are provided below. Some of these aspects may be applied independently, and some may be applied in combination, as will be apparent to those skilled in the art. Specific details are set forth in the following description for purposes of explanation in order to provide a thorough understanding of the various aspects of this application. However, it will be apparent that various aspects may be practiced without these specific details. The accompanying drawings and descriptions are not intended to be limiting.

[0039] The following description provides only exemplary aspects and is not intended to limit the scope, applicability, or configuration of this disclosure. Rather, the following description of the exemplary aspects will provide those skilled in the art with a description that can be used to implement the exemplary aspects. It should be understood that various changes may be made to the function and arrangement of the elements without departing from the spirit and scope of this application as set forth in the appended claims.

[0040] A camera is a device that uses an image sensor to receive light and capture image frames (such as still images or video frames). The terms "image," "image frame," and "frame" are used interchangeably herein. Cameras can be configured with various image capture and image processing settings. Different settings produce images with different appearances. Camera settings, such as ISO, exposure time, aperture size, aperture value, shutter speed, focus, and gain, are determined and applied before or during the capture of one or more image frames. For example, settings or parameters can be applied to the image sensor used to capture one or more image frames. Other camera settings can configure post-processing of one or more image frames, such as changes to contrast, brightness, saturation, sharpness, level, curves, or color. For example, settings or parameters can be applied to a processor (e.g., an image signal processor or ISP) used to process one or more image frames captured by the image sensor.

[0041] Devices including cameras can analyze image data captured by image sensors to detect, identify, classify, and / or track objects within that image data. For example, by detecting and / or identifying objects in multiple video frames, the device can track the movement of those objects over time.

[0042] An imaging system and techniques are described. In some examples, the imaging system uses a trained machine learning model to process at least the focus data (e.g., and in some cases, image data) to identify a predicted movement direction (e.g., phase shift sign) for moving the lens to improve image focus (e.g., positioning towards the focusing lens). The focus data is associated with at least one focused pixel of image data captured using an image sensor. The imaging system moves the lens in the predicted movement direction (e.g., using an actuator) to improve image focus. In some examples, the imaging system uses an image sensor to capture an image after moving the lens in the predicted movement direction.

[0043] The imaging system and techniques described in this paper offer numerous technological improvements over existing imaging systems, such as more efficient autofocus (PDAF) even when the lens is in a lens position where phase detection autofocus (PDAF) traditionally cannot operate accurately. The imaging system and techniques described in this paper provide a solution to the problems encountered when using PDAF at certain lens positions, and solve these problems more efficiently and accurately than CDAF, active autofocus, or using an additional (auxiliary) camera.

[0044] Various aspects of this application will be described with reference to the accompanying drawings. Figure 1 This is a block diagram illustrating the architecture of an image capture and processing system 100. The image capture and processing system 100 includes various components for capturing and processing images of one or more scenes (e.g., an image of scene 110). The image capture and processing system 100 may capture individual images (or photographs) and / or capture video comprising multiple images (or video frames) in a specific sequence. A lens 115 of the system 100 faces scene 110 and receives light from scene 110. The lens 115 bends the light toward an image sensor 130. The light received by the lens 115 passes through an aperture controlled by one or more control mechanisms 120 and is received by the image sensor 130. In some examples, scene 110 is a scene in the environment. In some examples, scene 110 is a scene of at least a portion of a user's face. For example, scene 110 may be a scene of one or both eyes of a user and / or at least a portion of a user's face.

[0045] One or more control mechanisms 120 may control exposure, focus, and / or zoom based on information from image sensor 130 and / or information from image processor 150. One or more control mechanisms 120 may include multiple mechanisms and components; for example, control mechanism 120 may include one or more exposure control mechanisms 125A, one or more focus control mechanisms 125B, and / or one or more zoom control mechanisms 125C. One or more control mechanisms 120 may also include additional control mechanisms besides those illustrated, such as controls for analog gain, flash, HDR, depth of field, and / or other image capture properties.

[0046] The focus control mechanism 125B of the control mechanism 120 can obtain focus settings. In some examples, the focus control mechanism 125B stores the focus settings in a memory register. Based on the focus settings, the focus control mechanism 125B can adjust the positioning of the lens 115 relative to the positioning of the image sensor 130. For example, based on the focus settings, the focus control mechanism 125B can move the lens 115 closer to or further away from the image sensor 130 by actuating a motor or servo, thereby adjusting the focus. In some cases, additional lenses, such as one or more microlenses above each photodiode of the image sensor 130, may be included in the system 100, each of which bends light received from the lens 115 toward the corresponding photodiode before it reaches that photodiode. The focus settings can be determined by contrast detection autofocus (CDAF), phase detection autofocus (PDAF), or some combination thereof. The focus settings can be determined using the control mechanism 120, the image sensor 130, and / or the image processor 150. The focus settings may be referred to as image capture settings and / or image processing settings.

[0047] The exposure control mechanism 125A of the control mechanism 120 can obtain the exposure settings. In some cases, the exposure control mechanism 125A stores the exposure settings in a memory register. Based on the exposure settings, the exposure control mechanism 125A can control the aperture size (e.g., aperture size or aperture coefficient), the duration of the aperture opening (e.g., exposure time or shutter speed), the sensitivity of the image sensor 130 (e.g., ISO speed or film speed), the analog gain applied by the image sensor 130, or any combination thereof. The exposure settings may be referred to as image capture settings and / or image processing settings.

[0048] The zoom control mechanism 125C of the control mechanism 120 can obtain zoom settings. In some examples, the zoom control mechanism 125C stores the zoom settings in a memory register. Based on the zoom settings, the zoom control mechanism 125C can control the focal length of an assembly (lens assembly) of lens elements including lens 115 and one or more additional lenses. For example, the zoom control mechanism 125C can control the focal length of the lens assembly by actuating one or more motors or servos to move one or more lenses in the lens relative to each other. The zoom settings may be referred to as image capture settings and / or image processing settings. In some examples, the lens assembly may include a parfocal zoom lens or a variable focal length zoom lens. In some examples, the lens assembly may include a focusing lens (in some cases, this focusing lens may be lens 115) that first receives light from scene 110, where the light then passes through a focusless zoom system between the focusing lens (e.g., lens 115) and image sensor 130 before reaching image sensor 130. In some cases, a focusless zoom system may include two positive (e.g., converging, convex) lenses with equal or similar focal lengths (e.g., within a threshold difference), with a negative (e.g., diverging, concave) lens between them. In some cases, the zoom control mechanism 125C moves one or more lenses in the focusless zoom system, such as one or both positive lenses and the negative lens.

[0049] Image sensor 130 includes one or more arrays of photodiodes or other photosensitive elements. Each photodiode measures the amount of light that ultimately corresponds to a specific pixel in the image generated by image sensor 130. In some cases, different photodiodes may be covered by different color filters, and thus light matching the color of the color filter covering the photodiode can be measured. For example, Bayer color filters include red, blue, and green color filters, where each pixel of the image is generated based on red light data from at least one photodiode covered by the red color filter, blue light data from at least one photodiode covered by the blue color filter, and green light data from at least one photodiode covered by the green color filter. Other types of color filters may use yellow, magenta, and / or cyan (also known as "emerald green") color filters as alternatives to or complements to red, blue, and / or green color filters. Some image sensors may have no color filters at all and may alternatively use different photodiodes (in some cases stacked vertically) throughout the pixel array. Different photodiodes in the pixel array can have different spectral sensitivity profiles, thus responding to light of different wavelengths. Monochrome image sensors may also lack color filters and therefore lack color depth.

[0050] In some cases, image sensor 130 may alternatively or additionally include an opaque mask and / or a reflective mask that blocks light from reaching certain photodiodes or portions of certain photodiodes at certain times and / or from certain angles, which can be used for phase detection autofocus (PDAF). Image sensor 130 may also include an analog gain amplifier for amplifying the analog signal output from the photodiodes and / or an analog-to-digital converter (ADC) for converting the analog signal output from the photodiodes (and / or amplified by the analog gain amplifier) ​​into a digital signal. In some cases, certain components or functions discussed with respect to one or more control mechanisms in control mechanism 120 may alternatively or additionally be included in image sensor 130. Image sensor 130 may be a charge-coupled device (CCD) sensor, an electron multiplication CCD (EMCCD) sensor, an active pixel sensor (APS), a complementary metal-oxide-semiconductor (CMOS), an N-type metal-oxide-semiconductor (NMOS), a hybrid CCD / CMOS sensor (e.g., sCMOS), or some other combination thereof.

[0051] The image processor 150 may include one or more processors, such as one or more image signal processors (ISPs) (including ISP 154), one or more host processors (including host processor 152), and / or one or more of any other type of processor 1610 discussed with respect to the computing system 1600. The host processor 152 may be a digital signal processor (DSP) and / or other types of processor. In some specific implementations, the image processor 150 is a single integrated circuit or chip (e.g., referred to as a system-on-a-chip or SoC) including the host processor 152 and the ISP 154. In some cases, the chip may also include one or more input / output ports (e.g., input / output (I / O) port 156), a central processing unit (CPU), a graphics processing unit (GPU), a broadband modem (e.g., 3G, 4G, or LTE, 5G, etc.), memory, and connectivity components (e.g., Bluetooth). ™This includes components such as the Global Positioning System (GPS), any combination thereof, and / or other components. I / O port 156 may include any suitable input / output port or interface according to one or more protocols or specifications, such as Inter-Integrated Circuit 2 (I2C) interface, Inter-Integrated Circuit 3 (I3C) interface, Serial Peripheral Interface (SPI) interface, Serial General Purpose Input / Output (GPIO) interface, Mobile Industrial Processor Interface (MIPI) (such as MIPI CSI-2 physical (PHY) layer ports or interfaces), Advanced High Performance Bus (AHB) bus, any combination thereof, and / or other input / output ports. In an exemplary example, host processor 152 may communicate with image sensor 130 using an I2C port, and ISP 154 may communicate with image sensor 130 using a MIPI port.

[0052] Image processor 150 can perform multiple tasks, such as demosaicing, color space conversion, image frame downsampling, pixel interpolation, automatic exposure (AE) control, automatic gain control (AGC), CDAF, PDAF, automatic white balance, merging image frames to form an HDR image, image recognition, object recognition, feature recognition, receiving input, managing output, managing memory, or some combination thereof. Image processor 150 can store image frames and / or processed images in random access memory (RAM) 140 and / or 1620, read-only memory (ROM) 145 and / or 1625, cache, memory units, another storage device, or some combination thereof.

[0053] Various input / output (I / O) devices 160 may be connected to the image processor 150. I / O devices 160 may include a display screen, keyboard, keypad, touchscreen, touchpad, touch-sensitive surface, printer, any other output device 1635, any other input device 1645, or some combination thereof. In some cases, text may be input into the image processing device 105B via the physical keyboard or keypad of the I / O device 160, or via a virtual keyboard or keypad on the touchscreen of the I / O device 160. I / O devices 160 may include one or more ports, jacks, or other connectors that enable wired connections between the system 100 and one or more peripheral devices, through which the system 100 receives data from and / or sends data to one or more peripheral devices. I / O devices 160 may include one or more wireless transceivers that enable wireless connections between the system 100 and one or more peripheral devices, through which the system 100 receives data from and / or sends data to one or more peripheral devices. Peripheral devices may include any type of I / O device 160 discussed earlier, and they can be considered I / O devices 160 in themselves once they are coupled to ports, jacks, wireless transceivers or other wired and / or wireless connectors.

[0054] In some cases, the image capture and processing system 100 may be a single device. In other cases, the image capture and processing system 100 may be two or more independent devices, including an image capture device 105A (e.g., a camera) and an image processing device 105B (e.g., a computing device coupled to the camera). In some embodiments, the image capture device 105A and the image processing device 105B may be coupled together, for example, via one or more wires, cables, or other electrical connectors, and / or wirelessly coupled together via one or more wireless transceivers. In some embodiments, the image capture device 105A and the image processing device 105B may be disconnected from each other.

[0055] like Figure 1 As shown, the vertical dashed line will Figure 1 The image capture and processing system 100 is divided into two parts, namely image capture device 105A and image processing device 105B. Image capture device 105A includes a lens 115, a control mechanism 120, and an image sensor 130. Image processing device 105B includes an image processor 150 (including an ISP 154 and a host processor 152), RAM 140, ROM 145, and I / O devices 160. In some cases, certain components illustrated in image capture device 105A (such as ISP 154 and / or host processor 152) may be included in image capture device 105A.

[0056] Image capture and processing system 100 may include electronic devices such as mobile or landline phones (e.g., smartphones, cellular phones, etc.), desktop computers, laptop or notebook computers, tablet computers, set-top boxes, televisions, cameras, display devices, digital media players, video game consoles, video streaming devices, Internet Protocol (IP) cameras, or any other suitable electronic devices. In some examples, image capture and processing system 100 may include one or more wireless transceivers for wireless communication (such as cellular network communication, 1602.11 Wi-Fi communication, wireless local area network (WLAN) communication, or some combination thereof). In some specific implementations, image capture device 105A and image processing device 105B may be different devices. For example, image capture device 105A may include a camera device, and image processing device 105B may include a computing device, such as a mobile phone, desktop computer, or other computing device.

[0057] Although the image capture and processing system 100 is shown to include certain components, those skilled in the art will understand that the image capture and processing system 100 may include more than [other components]. Figure 1 The components shown herein are additional components. Components of the image capture and processing system 100 may include software, hardware, or one or more combinations of software and hardware. For example, in some embodiments, components of the image capture and processing system 100 may include electronic circuitry or other electronic hardware, and / or may be implemented using electronic circuitry or other electronic hardware, which may include one or more programmable electronic circuits (e.g., microprocessors, GPUs, DSPs, CPUs, and / or other suitable electronic circuits); and / or may include computer software, firmware, or any combination thereof, and / or may be implemented using computer software, firmware, or any combination thereof to perform the various operations described herein. Software and / or firmware may include one or more instructions stored on a computer-readable storage medium and executable by one or more processors of an electronic device implementing the image capture and processing system 100.

[0058] Figure 2A This is a cross-sectional view illustrating a phase-detection autofocus (PDAF) camera system 200 in phase and therefore in focus state 250. The PDAF camera system 200 can be... Figure 1An example of an image capture and processing system 100. Light 275 may propagate from a subject 205 (e.g., an apple) through a lens 210, which focuses the scene having subject 205 onto an image sensor 285 (not fully depicted). Subject 205 may represent an object in the scene (e.g., scene 110). Subject 205 may represent a region of interest (ROI) in scene 110. In some examples, the ROI may include one or more objects. In some examples, the ROI may include one or more portions of an object. In an exemplary example, subject 205 (and / or ROI) may include the face of a specific person in the scene. Image sensor 285 may be... Figure 1 An example of an image sensor 130 in an image capture and processing system 100. Image sensor 285 includes focusing photodiodes 225A and 225B corresponding to focused pixels. Focusing photodiodes 225A and 225B may be associated with one or two focused pixels of the pixel array of image sensor 285 (e.g., focusing photodiodes 225A and 225B may be two photodiodes sharing a single focused pixel 220, or focusing photodiode 225A may be associated with a first focused pixel and focusing photodiode 225B may be associated with a second focused pixel, with both focused pixels sharing a single microlens 220). In some cases, light 275 may propagate through at least one microlens 220 before reaching focusing photodiodes 225A and 225B. When camera system 200 is in... Figure 2A In focus state 250, light rays 275 can ultimately converge on the plane corresponding to the positioning of focusing photodiodes 225A and 225B. When the camera system 200 is in focus state 250... Figure 2A In the focusing state 250, the light 275 can also converge at the focal plane 215 (also known as the image plane) after passing through the lens 210 but before reaching the microlens 220 and / or the focusing photodiodes 225A and 225B.

[0059] because Figure 2A The PDAF camera system 200 is in focus state 250, so the data from the focusing photodiodes 225A and 225B are aligned, as shown here by image 270A, which shows a clear and sharp representation of the subject 205 due to this alignment, compared with the data from the focusing photodiodes 225A and 225B respectively. Figure 2B and Figure 2CThe misalignment of the subject 205 caused by out-of-phase states (e.g., front focus state 240 and rear focus state 245) represents the opposite. Focus state 250 can also be referred to as an "in-phase" state because the data from focusing photodiodes 225A and 225B have no phase shift, or have a very small phase shift (e.g., the phase shift drops below a predetermined phase shift threshold). In some cases, the phase shift may be referred to as phase difference or phase disparity.

[0060] Figure 2B This is an example Figure 2A A cross-sectional view of the PDAF camera system 200 in front-focusing state 240. Figure 2B PDAF camera system 200 and Figure 2A The PDAF camera system 200 is the same, but the lens 210 is moved closer to the body 205 and further away from the focusing photodiodes 225A and 225B (e.g., using an actuator 280, such as a voice coil motor, linear actuator, other motor, or other actuator), and is therefore in a forward-focusing state 240. The lens positioning in the focusing state 250 remains the same. Figure 2B The center is drawn as a dashed outline for reference, with double arrows indicating the movement of the lens between the "Front Focus" 240 lens position and the "Focus" 250 lens position.

[0061] When the camera system 200 is in Figure 2B In the pre-focusing state 240, the light ray 275 can ultimately converge at the plane (indicated by the dashed line) before the positioning of the focusing photodiodes 225A and 225B (i.e., between the microlens 220 and the focusing photodiodes 225A and 225B). The light ray 275 can also converge at the positioning (indicated by another dashed line) before the focal plane 215 after passing through the lens 210 but before reaching the microlens 220 and / or the focusing photodiodes 225A and 225B. Because... Figure 2B In the PDAF camera system 200, the light 275 is out of phase in the front focusing state 240, so the data from the focusing photodiodes 225A and 225B are misaligned, which is represented here by the image 270B, which illustrates the misalignment of the subject 205 by diagonal and vertical shadows. The direction of the misalignment in the image 270B is related to the front focusing state 240, and the distance of the misalignment in the image 270B is related to the distance of the lens 210 from its positioning in the focusing state 250.

[0062] Figure 2C This is an example Figure 2A A cross-sectional view of the PDAF camera system 200 in the back-focusing state 245. Figure 2C PDAF camera system 200 and Figure 2AThe PDAF camera system 200 is identical, but the lens 210 is moved further away from the body 205 and closer to the focusing photodiodes 225A and 225B (e.g., using actuator 280, such as a voice coil motor, linear actuator, other motor, or other actuator), and is therefore in a back-focus state 245 (also referred to as the "back-focus" state). The lens positioning in focus state 250 is still drawn as a dashed outline for reference, with double-sided arrows indicating the movement of the lens between the lens positioning in back-focus state 245 and the lens positioning in focus state 250.

[0063] When the camera system 200 is in Figure 2C In the post-focusing state 245, the light ray 275 can ultimately converge at a plane (indicated by the dashed line) beyond the positioning of the focusing photodiodes 225A and 225B. The light ray 275 can also converge at a position beyond the focal plane 215 (indicated by another dashed line) after passing through the lens 210 but before reaching the microlens 220 and / or the focusing photodiodes 225A and 225B. Because... Figure 2C In the PDAF camera system 200, the light 275 is out of phase in the back-focusing state 245, so the data from the focusing photodiodes 225A and 225B is misaligned. This is represented here by image 270C, which illustrates the misalignment of the subject 205 through diagonal and vertical shading. The direction of the misalignment in image 270C is related to the back-focusing state 245, and the distance of the misalignment in image 270C is related to the distance of the lens 210 from its positioning in the focusing state 250. The misalignment is illustrated differently in image 270C compared to image 270B. For example, the vertical shading of the subject 205 represents the left side represented by the diagonal shading of the subject 205 illustrated in image 270B, while the vertical shading of the subject 205 represents the right side represented by the diagonal shading of the subject 205 illustrated in image 270C. In some examples, the image with the forefocus state 240 and the image with the backfocus state 245 may appear blurry, distorted, or include ghosting artifacts similar to those illustrated in images 270B and 270C (e.g., represented by two different representations of the subject 205 in each image).

[0064] When light 275 converges in front of the planes of focusing photodiodes 225A and 225B, as in the forward focusing state 240, or converges beyond the planes of focusing photodiodes 225A and 225B, as in the backward focusing state 245, the resulting image generated by the image sensor 285 may be out of focus or blurred. In the case of an out-of-focus image, if lens 210 is in the backward focusing state 245, lens 210 can be moved forward (towards the body 205 and away from photodiodes 225A and 225B), or if lens is in the forward focusing state 240, lens can be moved backward (away from the body 205 and towards photodiodes 225A and 225B). Lens 210 can be moved forward or backward within a positioning range (e.g., using actuator 280), in some cases, the positioning range having a predetermined length R representing the possible range of movement of the lens within the camera system 200. The camera system 200 or its computing system can determine the distance and direction for positioning the lens 210 to focus the image based on one or more phase shift values, which are calculated as the difference between data from two focusing photodiodes (such as focusing photodiodes 225A and 225B) receiving light from different directions. The direction of movement of the lens 210 may correspond to the direction in which the data from focusing photodiodes 225A and 225B are determined to be out of phase, or whether the phase shift is positive or negative. The distance of movement of the lens 210 may correspond to the degree or amount of out of phase, or the absolute value of the phase shift, as determined from the data from focusing photodiodes 225A and 225B.

[0065] In some examples, the direction in which the light 275 is out of phase (e.g., between the front focus state 240 and the rear focus state 245) can be referred to as the sign of the phase shift (e.g., positive or negative). The sign of the phase shift can indicate the direction in which the lens 210 should be moved (e.g., by actuator 280) to correct the phase shift and bring the image in phase and focused (e.g., towards the subject 205 or towards the image sensor 285). In some examples, the degree of the light 275 being out of phase (e.g., how much the light is out of phase, or how far the light is from being in phase or focused) can be referred to as the magnitude of the phase shift. The magnitude of the phase shift can indicate the distance at which the lens 210 should be moved (e.g., by actuator 280) to correct the phase shift and bring the image in phase and focused. In focus state 250, the phase shift can be zero or close to zero.

[0066] In some examples, Figure 2A , Figure 2B and Figure 2CThe PDAF camera system 200 may include various additional components, such as lenses, mirrors, partial reflection (PR) mirrors, prisms, photodiodes, image sensors, other components of the image capture and processing system 100, other components associated with the camera, other components associated with optical equipment, or combinations thereof. In some cases, the focusing photodiodes 225A and 225B may be referred to as PDAF photodiodes, PDAF photodetectors, PDAF diodes, PDAF pixels, phase detection (PD) photodiodes, PD diodes, PDAF pixel photodiodes, PDAF pixel diodes, PDAF pixel photodetectors, PD pixel photodiodes, PD pixel diodes, PD pixel photodetectors, PD pixels, PD photodetectors, focused pixel photodiodes, focused pixel diodes, focused pixels, focused photodetectors, pixel photodiodes, pixel diodes, photodiodes, diodes, pixels, photodetectors, or combinations thereof.

[0067] Figure 3A This is a top view illustrating a pixel array 300 of an image sensor having a mask (e.g., mask 302A-302B) that partially covers the photodiodes of the focused pixels. Image sensors (e.g., image sensor 130, image sensor 285) of camera systems (e.g., image capture and processing system 100, PDAF camera system 200) may include pixel arrays, such as... Figure 3A The pixel array 300 may include a photodiode array. Figure 3A In this context, the photodiode is filtered by a color filter (e.g., a Bayer color filter or other types of color filters discussed below) and such as Figure 3B The microlens 318, as indicated in Figure 310, covers the area. The photodiode of the focusing pixel is also partially covered by... Figure 3A The mask 320 covers the pixel array 300.

[0068] Figure 3B It is a sign Figure 3A The figure 310 shows the components. The figure 310 identifies the microlens 318, which is a single pixel, as a circle, and the mask 320 as a dark shaded rectangle. Figure 3BFigure 310 also identifies three squares with different patterns, each representing a color filter 312, 314, and 316, with each filter used for one of three different colors: red, green, or blue. That is, the square with the first pattern (e.g., a mottled pattern) represents the color filter 312 used for the first color, which could be, for example, green; the square with the second pattern (e.g., a pattern of diagonal stripes between the lower left and upper right directions) represents the color filter 314 used for the second color, which could be, for example, blue; and the square with the third pattern (e.g., a pattern of diagonal stripes between the lower right and upper left directions) represents the color filter 316 used for the third color, which could be, for example, red.

[0069] These color filters are arranged in... Figure 3A , Figure 3C and Figure 3D In the color filter array (CFA) above the photodiode array in pixel arrays 300, 330, and 340. Figure 3B The colors (and the number of colors) indicated in Legend 310 and in Figure 3A , Figure 3C and Figure 3D The arrangement of color filters illustrated in pixel arrays 300, 330, and 340 should be understood as exemplary and not construed as limiting. In some examples, red, green, and blue color filters are used in image sensors. In some examples, a CFA using red, green, and blue color filters may be referred to as a Bayer color filter, a Bayer color filter array, a Bayer CFA, or a Bayer color filter CFA. A single color filter in a Bayer CFA may be referred to as a Bayer color filter. In some examples, a Bayer CFA includes more green Bayer color filters than red or blue Bayer color filters, for example, in a ratio of 50% green, 25% red, and 25% blue, to simulate the sensitivity of the human eye's physiological structure to green light. Bayer CFAs with these ratios of color filters are sometimes referred to as BGGR, RGBG, GRGB, or RGGB, and in the presence of color filter 312, at a ratio greater than... Figure 3A , Figure 3C and Figure 3D The high proportions of color filters 314 and 316 in the pixel arrays 300, 330, and 340 are reflected. In some examples, in such Bayer CFAs, green is treated as two colors, which can be referred to as "Gr" and "Gb" respectively.

[0070] CFAs can use a variety of color schemes and may include more colors, fewer colors, any of the colors discussed above (e.g., red, green, and / or blue), other colors not discussed above, or combinations thereof. For example, some CFAs use cyan, yellow, and magenta filters instead of the red, green, and blue Bayer filter scheme. In an arrangement called Cyan-Yellow-Yellow-Magenta (CYYM), 50% of the filters are yellow, 25% are cyan, and 25% are magenta. Some filters also add a fourth green filter to three cyan, yellow, and magenta filters, collectively called the Cyan-Yellow-Green-Magenta (CYGM) filter. Some CFAs use red, green, blue, and “emerald green” or cyan, called the RGBE color scheme. In some cases, a mixture or combination of the Bayer, CYYM, CYGM, or RGBE color schemes may be used. In some cases, color filters for one or more colors in the Bayer, CYYM, CYGM, or RGBE color schemes can be omitted, leaving only two colors or even one color in some situations. Although Figure 3B Figure 310 precisely lists three color filters 312, 314, and 316, and provides green, red, and blue as examples to conform to the Bayer color filter color scheme. However, it should be understood that more than or fewer colors can be used in a CFA, and the colors can vary, including, for example, red, green, blue, cyan, magenta, yellow, emerald green, white (transparent), or some combination thereof. Some image sensors may lack color filters and may choose to use different photodiodes throughout the pixel array (e.g., photodiodes stacked in a direction perpendicular to the plane surface of the image sensor). Different photodiodes have different spectral sensitivity profiles and therefore respond to light of different wavelengths. Monochrome or grayscale image sensors may also lack color filters and therefore lack color information.

[0071] Figure 3A The pixel array 300 is illustrated as having two pixels used for phase detection autofocus (PDAF), which are referred to herein as focusing pixels, but may also be referred to as PDAF pixels or phase detection (PD) pixels. Other pixels not used for PDAF may simply be referred to as imaging pixels 304. Figure 3A In the pixel array 300, any pixel without mask 320 is an imaging pixel 304, even if only two imaging pixels 304 are specifically marked. Although in Figure 3ATwo focused pixels are illustrated in pixel array 300, both in the same column but separated by three rows of imaging pixels. However, different pixel arrays (not depicted) may have any number of focused pixels (i.e., one or more focused pixels), which may be arranged in any possible pattern or arrangement. In some cases, the pattern of focused pixels may be repeated across the pixel array, for example, as “patches” of size 8 pixels by 8 pixels or 16 pixels by 16 pixels.

[0072] Figure 3A The two focused pixels illustrated are partially covered by masks 320, which are labeled masks 302A and 302B, respectively. Each of the masks 320 may be a mask or shield made of an opaque and / or reflective material (such as metal). Each mask 320 limits the amount and direction of light striking the photodiode of the focused pixel partially covered by the mask. Masks 302A and 302B each limit how much light arrives from a specific direction and strikes the underlying focused pixel photodiode, and are positioned above two different focused pixel diodes in opposite directions to produce a pair of left and right images. For example, mask 302A is positioned on the left side of the first focused pixel, thus leaving the right side of the first focused pixel to receive light entering from the right (right image). Mask 302B is positioned on the right side of the second focused pixel, thus leaving the left side of the second focused pixel to receive light entering from the left (left image). Because both focused pixels are illustrated as being partially covered by mask 320, their focusing photodiodes effectively receive 50% of the light that the imaging photodiode (which will not be covered by the mask) at the same location on the pixel array would receive.

[0073] Any number of focused pixels can be included in the pixel array of an image sensor. Left and right focused pixel pairs can be adjacent to each other or spaced apart by one or more imaging pixels 304. The two pixels from a left and right focused pixel pair can both be in the same row and / or the same column of the pixel array, or in different rows and / or different columns, or some combination thereof. Although masks 302A and 302B are shown within pixel array 300 as masking the left and right portions of the focused pixel photodiode, this is for illustrative purposes only. Focused pixel mask 320 may alternatively mask the top or bottom portion of the focused pixel photodiode, thereby generating top and bottom images (or “up” and “down” images) based on focused pixel data received by the focused pixels. Similar to the left and right focused pixel pairs, top and bottom focused pixel pairs can both be in the same row and / or the same column of the pixel array, or in different rows and / or different columns, or some combination thereof. The pixel array of an image sensor may have a focused pixel having a mask 320 on the left side of a focused pixel, a mask 320 on the right side of a second focused pixel, a mask 320 on the top side of a third focused pixel, a mask 320 on the bottom side of a fourth focused pixel, and more focused pixels having any of these types of masks 320 in some examples. Using focused pixels with masks 320 along multiple axes (e.g., left-right focused pixel pairs and top-bottom focused pixel pairs) can improve autofocus quality. One reason why autofocus quality can be improved by using a focus pixel with a mask 320 along multiple axes is that using the mask 320 only along the left and right sides of the focus pixel photodiode for PDAF may result in poor focus on a scene or subject with many horizontal edges (i.e., lines appearing along the left-right axis relative to the orientation of the focus pixel and the mask 320), and using the mask 320 only along the top and bottom sides of the focus pixel photodiode for PDAF may result in poor focus on a scene or subject with many vertical edges (i.e., lines appearing along the up-down axis relative to the orientation of the focus pixel and the mask 320).

[0074] Some PDAF camera systems are not like Figure 3A Instead of using a mask 320 on the focused pixel as in the example, multiple pixels are covered under a single microlens, which can be called an on-chip lens (OCL). Figure 3C A top view of a pixel array configuration is shown, in which two side-by-side focused pixels are covered by 2-pixel by 1-pixel microlenses. Figure 3D A top view of a pixel array configuration is shown, in which four adjacent focused pixels are covered by 2-pixel by 2-pixel microlenses. Figure 3C and Figure 3D Pixel arrays 330 and 340 can also be based on Figure 3B Let's use Figure 310 to explain.

[0075] Figure 3C This is a top view illustration of a pixel array 330 of an image sensor having two side-by-side focusing pixels covered by a 2-pixel by 1-pixel microlens 332. Figure 3D This is a top view illustrating a pixel array 340 of an image sensor having four adjacent focusing pixels covered by 2-pixel by 2-pixel microlenses 342. (See reference) Figure 3C and Figure 3D , Figure 3C 2-pixel by 1-pixel microlens 332 and Figure 3D Both the 2-pixel by 2-pixel microlens 342 span multiple adjacent focal pixels (i.e., the microlens covers multiple adjacent focal pixel photodiodes), and both can limit the amount and / or direction of light striking the focal pixel photodiodes of those focal pixels. Figure 3C Microlens 332 covers two horizontally adjacent focused pixels of pixel array 330, enabling the generation of focused pixel data from two focused photodiodes. Focused pixel data from the left focused pixel (marked "L") represents light approaching from the left side of pixel array 330, and focused pixel data from the right focused pixel (marked "R") represents light approaching from the right side of pixel array 330. While microlens 332 is shown within pixel array 330 spanning left and right adjacent pixels / diodes (e.g., in the horizontal direction), this is for illustrative purposes only. Alternatively, a 2-pixel by 1-pixel microlens 332 may span the top and bottom adjacent pixels / diodes (e.g., in the vertical direction), thereby generating top-bottom (or top-and-bottom) focused photodiode pairs and corresponding pixel data.

[0076] Similarly, Figure 3D The microlens 342 covers a 2-pixel square of four adjacent focused pixels of the pixel array 340, enabling the generation of focused pixel data from all four photodiodes within the square. The focused pixel data from the four adjacent focused pixels therefore includes data from the top-left pixel (in...). Figure 3D The pixel data marked "UL" represents the focused pixel data of light approaching from the upper left of pixel array 340, and the pixel data from the upper right (in the pixel array 340). Figure 3D The pixel data marked "UR" represents the focused pixel data of light approaching from the upper right of pixel array 340, and the pixel data from the lower left (in the pixel). Figure 3D The pixel data marked "BL" represents the focused pixel data of light approaching from the lower left of pixel array 340, and the data from the lower right pixel (in... Figure 3DThe data marked "BR" represents focused pixel data of light approaching from the lower right of pixel array 340. Any number of focused pixels may be included in the pixel array, and may include one or more horizontally oriented (left-right) 2-pixel by 1-pixel microlenses 332, one or more vertically oriented (up-down) 2-pixel by 1-pixel microlenses 332, one or more 2-pixel by 2-pixel microlenses 342, or combinations thereof.

[0077] Refer again Figure 3C and Figure 3D Once the pixel array captures a frame, thus capturing the focus pixel data for each focused pixel, the focus pixel data from paired focus pixels can be compared to each other. For example, focus pixel data from the left focus pixel photodiode can be compared to focus pixel data from the right focus pixel photodiode, and focus pixel data from the top focus pixel photodiode can be compared to focus pixel data from the bottom focus pixel photodiode. If the compared focus pixel data values ​​differ, the difference is called phase shift, also known as phase difference, defocus value, or separation error. Figure 3D The focused pixels under the 2-pixel by 2-pixel microlens 342 essentially have two vertically adjacent horizontally oriented focused pixel pairs and / or two horizontally adjacent vertically oriented focused pixel pairs. Therefore, focused pixel data from the UL focused pixel can be compared with focused pixel data from the BL focused pixel (as a top / bottom pair), focused pixel data from the UR focused pixel can be compared with focused pixel data from the BR focused pixel (as a top / bottom pair), focused pixel data from the UL focused pixel can be compared with focused pixel data from the UR focused pixel (as a left / right pair), focused pixel data from the BL focused pixel can be compared with focused pixel data from the BR focused pixel (as a left / right pair), or some combination thereof. In some cases, focused pixel data can be compared between pixels that are diagonally opposite each other (along two axes). For example, focused pixel data from the UL focused pixel can be compared with focused pixel data from the BR focused pixel, and / or focused pixel data from the BL focused pixel can be compared with focused pixel data from the UR focused pixel.

[0078] Although Figure 3C The focused pixels under the 2-pixel by 1-pixel microlens 332 and Figure 3D The focused pixels under the 2-pixel by 2-pixel microlens 342 are all exemplified as color filters 312 with the first color, but this is not necessary. In some cases, the normal pattern of the CFA of the pixel array can continue under the 2-pixel by 1-pixel microlens 332 and / or under the 2-pixel by 2-pixel microlens 342.

[0079] Figure 3EThis is a top view illustrating a pixel array 350 of an image sensor, wherein at least one focused pixel has two photodiodes. Specifically, in Figure 3E The image illustrates a four-pixel by four-pixel pixel array 350 with four focused pixels. Each of the four focused pixels in the pixel array 350 includes two photodiodes, wherein the left and right photodiodes of the photodiode pair in each focused pixel are labeled "L" and "R," respectively. A focused pixel with two photodiodes (such as...) Figure 3E The focused pixel (often referred to as a dual photodiode (2PD) focused pixel) is sometimes called a dual photodiode (2PD) focused pixel.

[0080] Figure 3E One of the 2PD focusing pixels is designated as 2PD focusing pixel 352. The left photodiode (L) of 2PD focusing pixel 352 is designated as "left photodiode 354L", and the right photodiode (R) of 2PD focusing pixel 352 is designated as "right photodiode 354R". For each captured frame, the left photodiode 354L and the right photodiode 354R can capture light received by 2PD focusing pixel 352 from different angles. For a given frame, the data captured by the left photodiode 354L can be referred to as the left image or left image data, while the data captured by the right photodiode 354R can be referred to as the right image or right image data. The left image data and the right image data can be compared to determine the phase shift.

[0081] Figure 3E The pixel array 350 illustrated is a "sparse" 2PD pixel array, where only some pixels in the pixel array 350 include two photodiodes (i.e., focusing pixels). The remaining pixels are imaging pixels and include only a single photodiode. However, in some cases, a "dense" 2PD pixel array may be used alternatively, where each pixel in the pixel array (or a higher percentage of pixels in the pixel array) includes two photodiodes (e.g., 2PD), four photodiodes (e.g., 4PD), or more than one other number of photodiodes, and in some cases may simultaneously function as both a focusing pixel and an imaging pixel, or may switch between functioning as a focusing pixel in one frame and an imaging pixel in another frame. Although Figure 3EAll 2PD focused pixels are shown as “horizontal” 2PD focused pixels with a left photodiode and a right photodiode, but this arrangement is exemplary. In some examples, the pixel array with 2PD focused pixels may additionally or alternatively include “vertical” focused pixels with a top (“upper”) photodiode and a bottom (“lower”) photodiode and / or photodiodes arranged diagonally opposite each other. Since using only horizontal focused pixels can sometimes limit the recognition of horizontal edges in an image, and using only vertical focused pixels can sometimes limit the recognition of vertical edges in an image, using both horizontal and vertical focused pixels can improve focus quality by performing well even in images with many horizontal and / or vertical edges.

[0082] Figure 3F This is a top view illustration of a pixel array 360 of an image sensor, wherein at least one focused pixel has four photodiodes. Figure 3F The pixel array 360 illustrated includes focused pixels, each of which comprises four diodes, commonly referred to as a 4PD focused pixel or quadrature phase detection (QPD) focused pixel. For example, 4PD focused pixel 362 in... Figure 3F The 4PD focusing pixel 362 is labeled and includes the top left photodiode labeled "UL", the top right photodiode labeled "UR", the bottom left photodiode labeled "BL", and the bottom right photodiode labeled "BR". Data from each photodiode of the 4PD focusing pixel 362 can be compared with data from adjacent photodiodes of the 4PD focusing pixel 362 to determine the phase difference. For example, photodiode data from the UL photodiode can be compared with photodiode data from the BL photodiode (as a top / bottom pair), photodiode data from the UR photodiode can be compared with photodiode data from the BR photodiode (as a top / bottom pair), photodiode data from the UL photodiode can be compared with photodiode data from the UR photodiode (as a left / right pair), photodiode data from the BL photodiode can be compared with photodiode data from the BR photodiode (as a left / right pair), or some combination thereof. In some examples, photodiode data from the 4PD focusing pixel 362 can be alternatively or additionally compared between photodiodes that are diagonally (along two axes) opposite each other. For example, photodiode data from the UL photodiode of the 4PD focused pixel 362 can be compared with photodiode data from the BR photodiode of the 4PD focused pixel 362, and / or photodiode data from the BL photodiode of the 4PD focused pixel 362 can be compared with photodiode data from the UR photodiode of the 4PD focused pixel 362.

[0083] Figure 3F The pixel array 360 illustrated is a "sparse" 4PD pixel array, where only some pixels in the pixel array 360 include four photodiodes (i.e., focus pixels). The remaining pixels are imaging pixels and include only a single photodiode. However, in some cases, a "dense" 4PD pixel array may be used alternatively, where each pixel in the pixel array (or a higher percentage of pixels in the pixel array) includes two photodiodes (e.g., 2PD), four photodiodes (e.g., 4PD), or more than one other number of photodiodes, and in some cases may simultaneously function as both focus pixels and imaging pixels, or may switch between functioning as focus pixels in one frame and imaging pixels in another frame. Although Figure 3F All 4PD focused pixels are shown relative to Figure 3F The illustrated orientation has a “horizontal” 4PD focusing pixel with a left photodiode and a right photodiode, but this arrangement is exemplary. A pixel array with 4PD focusing pixels may additionally or alternatively include “vertical” focusing pixels with a top (“upper”) photodiode and a bottom (“lower”) photodiode (e.g., perpendicular to the illustrated “left” and “right” orientations) and / or photodiodes arranged diagonally relative to each other. Since using only horizontal focusing pixels can sometimes limit the recognition of horizontal edges in an image, and using only vertical focusing pixels can sometimes limit the recognition of vertical edges in an image, using a combination of horizontal and vertical focusing pixels can improve focus quality well, even in images with many horizontal and / or vertical edges.

[0084] In some cases, pixel arrays can use one or more pairs of focused pixels with a mask 320 (e.g. Figure 3A (as illustrated in the example), a pair or more pairs of focused pixels covered by a 2-pixel by 1-pixel microlens 332 (e.g.) Figure 3C (as illustrated in the example), a group or more groups of focused pixels covered by a 2-pixel by 2-pixel microlens 342 (e.g.) Figure 3D (Example shown), one or more 2PD focused pixels 352 (e.g.) Figure 3E (as shown in the example) and / or one or more 4PD focused pixels 362 (e.g. Figure 3F (Example) A certain combination. In some cases, in Figure 3A The focused pixels in any of the configurations illustrated in Figure 2F and discussed therein can be arranged in a vertical and / or horizontal tiled pattern, such as... Figure 3E and Figure 3F Tiled patterns of 2PD and 4PD focused pixels.

[0085] Figure 4AThis is a cross-sectional side view of a photodiode 420A of an image sensor, partially covered by a mask 420, representing a focused pixel 400. The side view of the focused pixel 400 illustrates a single-pixel microlens 318 above a color filter 410A above the mask 320, which covers the left side of the photodiode 420A. Light 450B entering from the right side of the microlens 318 passes through the color filter 410A and reaches the photodiode 420A, while light 450A entering from the left side of the microlens 318 is reflected by the mask 320. Although a similar pixel with a mask 320 above the right side of the photodiode 420A is not illustrated, it should be understood that this can be achieved by horizontal flipping. Figure 4A This can be achieved through examples. In some examples, the mask 320 may be positioned above the color filter 410A and / or the microlens 318 (e.g., closer to the scene from which the light 450A originates).

[0086] Figure 4B This is a cross-sectional side view of two photodiodes 420B-420C of an image sensor pixel array, which are covered by a 2-pixel by 1-pixel microlens 432, representing the focused pixel 440. Figure 4B The side view of the focused pixel 440 illustrates a 2-pixel by 1-pixel microlens 332 above a color filter 410B on the left and another adjacent color filter 410C on the right, wherein the left color filter 410B is above the left photodiode 420B, and the right color filter 410C is above the right photodiode 420C. Two light rays 450C and 450D entering from the left side of the microlens 332 pass through the left color filter 410B and reach the left photodiode 420B, while two light rays 450E and 450F entering from the right side of the microlens 332 pass through the right color filter 410C and reach the right photodiode 420C.

[0087] Figure 4A and Figure 4B Each of the color filters 410A, 410B, and 410C can be a color filter for any color previously described relative to color filters 312, 314, and 316. That is, although Figure 4A and Figure 4B Red, green, and blue are listed as example colors to conform to the Bayer color scheme, but each of the color filters 410A, 410B, and 410C can represent another color, such as cyan, yellow, magenta, emerald green, or white (transparent). Although in Figure 4A and Figure 4B The color filters 410A, 410B, and 410C are all illustrated with the same pattern, which is consistent with... Figures 3A to 3DThe pattern of color filter 312 is matched, but the three color filters 410A, 410B, and 410C do not need to all represent color filters that have the same color as each other, and do not need to represent color filters that have the same color as each other. Figures 3A to 3D The color filter 312 represents the same color. In some examples, all three color filters 410A, 410B, and 410C represent different colors. In some examples, any two (or all three) of color filters 410A, 410B, and / or 410C may share a color. In some examples, color filters may be omitted (e.g., color filters 410A, 410B, and / or 410C).

[0088] Figure 5 This is a graph 500 illustrating the phase shift 520 detected by phase detection autofocus (PDAF), the corresponding PDAF confidence 530, the predicted direction 525 determined using a trained machine learning model, and the corresponding predicted direction confidence 535. Graph 500 includes a horizontal axis representing lens positioning 515, extending from positioning zero, as close as possible to the image sensor (and as far as possible from the subject), to positioning 870, as far away from the image sensor (and as close as possible to the subject). Graph 500 includes two vertical axes, including a phase shift axis 505 indicating the phase shift of the out-of-phase image (and thus indicating the direction in which the lens should be moved to make the image in-phase and therefore focused). The phase shift 520 detected by phase detection autofocus (PDAF) and the predicted direction 525 determined using a trained machine learning model are plotted relative to the phase shift axis 505. Positive values ​​along the phase shift axis 505 indicate that the lens should be moved further away from the image sensor and closer to the subject. The negative values ​​along phase shift axis 505 indicate that the lens should be moved further away from the subject and closer to the image sensor. The two vertical axes in graph 500 also include a confidence axis 510 indicating confidence values ​​from zero (indicating low confidence) to 1200 (indicating high confidence). PDAF confidence 530 and prediction direction confidence 535 are plotted relative to confidence axis 510.

[0089] PDAF works well in many scenarios, generally allowing the camera to focus faster than CDAF. However, in some examples, PDAF can be unreliable for certain lens positions and may provide incorrect phase shift signs (e.g., positive or negative), thus suggesting incorrect lens movement directions to the camera (e.g., towards the subject or towards the image sensor). For example, in graph 500, the phase shift 520 determined by PDAF fluctuates widely between large positive and large negative numbers between zero and 500 for the lens position (e.g., when the lens is close to the image sensor and far from the subject). The phase shift 520 detected by PDAF is illustrated as a solid black line within graph 500. These negative numbers represent an error 540 in the sign of the phase shift 520 determined by PDAF. The closer the phase shift 520 determined by PDAF is to the focus 545 lens position (e.g., approximately 690 in graph 500), the more reliable it becomes. In some examples, PDAF may additionally or alternatively be unreliable and may provide an incorrect phase shift sign (e.g., positive or negative), and thus suggest an incorrect direction of lens movement to the camera (e.g., towards the subject or towards the image sensor) at lens positions that are farther from the image sensor and closer to the subject. PDAF confidence 530 indicates a low confidence level (e.g., close to zero) of phase shift 520 determined using PDAF across lens positions where error 540 is exemplified, a higher confidence level (e.g., a peak value exceeding 100) near and at the focusing lens position 545, and then again indicates a lower confidence level (e.g., close to zero) at lens positions that are farther from the image sensor and closer to the subject than the focusing lens position 545. PDAF confidence 530 is illustrated as a black dashed line within graph 500.

[0090] Relying on a phase shift 520, as determined by PDAF, for autofocusing can be useful, reliable, and efficient for certain ranges of lens positioning (e.g., the range surrounding and including the focus 545 lens positioning). However, relying on a phase shift 520, as determined by PDAF, for other ranges of lens positioning (such as lens positioning illustrated in graph 500 as including error 540 in the notation of phase shift 520 determined by PDAF) can be counterproductive. Because error 540 includes an incorrect notation, relying on a phase shift 520 determined by PDAF for such lens positioning may cause the camera to move the lens in the wrong direction (e.g., away from the focus positioning instead of towards it), potentially increasing the time spent focusing the camera on the subject (e.g., the region of interest).

[0091] In some examples, the camera may use CDAF to determine unreliable lens positioning ranges via phase shifts determined by PDAF, for example, for lens positioning ranges that include an error 540 in the notation of phase shift 520 as determined by PDAF in graph 500. However, CDAF is generally slower and less efficient than PDAF, for example, introducing more time delay during focusing, using more power (e.g., during repeated actuation of actuator 280), and introducing more wear (e.g., during repeated actuation of actuator 280).

[0092] In some examples, the camera system may use PDAF-based focus determination from an auxiliary camera for a lens positioning range where phase shift determination via PDAF is unreliable in the main camera, such as for a lens positioning range that includes an error 540 in the sign of a phase shift 520 determined by PDAF as shown in graph 500. For example, an ultra-wide (UW) camera may have a shorter focal length than a non-UW camera and can therefore provide a reliable phase shift (e.g., phase shift sign and / or phase shift magnitude) determination for a lens positioning range for which PDAF provides an unreliable phase shift sign determination for the main camera. However, using two cameras each time focusing requires significantly more power than using one camera and may introduce errors based on small differences between the two cameras (e.g., manufacturing differences, lens differences, aperture size differences, exposure time differences, sensor differences, calibration differences, color filter differences, image processing differences, rolling shutter versus global shutter, etc.).

[0093] In some examples, the camera system may use one or more trained machine learning (ML) models to determine a predicted direction 525 in which the lens should be moved to approach the focus position 545. In some examples, the predicted direction 525 may be a prediction of the sign (e.g., positive or negative) of the phase shift 520. The predicted direction 525 is illustrated as a solid white line within the graph 500, outlined by a solid black line around the white line. For visualization purposes, a value of +10 for the predicted direction 525 is used to indicate a positive sign of the phase shift, indicating that the lens should move away from the image sensor and toward the subject to approach the focus lens position 545. Similarly, a value of -10 for the predicted direction 525 is used to indicate a negative sign of the phase shift, indicating that the lens should move away from the subject and toward the image sensor to approach the focus lens position 545. The prediction direction confidence 535 associated with the predicted direction 525 determined by the trained ML model indicates that the trained ML model can predict the predicted direction 525 with relatively high confidence relative to the PDAF confidence 530 across all lens positions. The predicted orientation confidence score 535 is illustrated as a black dashed line within graph 500. Similar to the PDAF confidence score 530, the predicted orientation confidence score 535 peaks at focus 545 positioning (e.g., a confidence value of 1023). However, the confidence value of the predicted orientation confidence score 535 ranges from approximately 530 to approximately 950 across graph 500, while the confidence value of the PDAF confidence score 530 ranges from approximately 0 to approximately 120 across graph 500. Therefore, the trained ML model is able to predict the predicted orientation 525 with higher confidence compared to determining the phase shift 520 via PDAF across all lens positioning. The trained ML model can be lightweight because it only generates values ​​indicating whether the phase shift is positive or negative, indicating whether the lens is to move closer or further away from the image sensor to approach focus 545 lens positioning. In some cases, the trained ML model can also predict what lens positioning is for a 545-lens focus, where the phase shift is zero. Therefore, in some examples, the trained ML model can output an indication of zero phase shift, indicating that the lens is in a 545-lens focus position and should not be moved.

[0094] Figure 6This is a block diagram illustrating a machine learning system 600 for training, using (e.g., inferring), and updating (e.g., further training) an ML model 625 associated with a machine learning (ML) prediction engine 620 to determine a predicted direction 635 for moving a lens to improve focusing. The machine learning system 600 includes a machine learning (ML) engine 620 that generates, trains, uses (e.g., infers), and / or updates (e.g., further trains) one or more ML models 625. ML model 625 may include, for example, one or more neural networks (NN) (e.g., neural network 1200), convolutional neural networks (CNN), time-delayed neural networks (TDNN), deep networks (DN), autoencoders (AE), variational autoencoders (VAE), deep belief networks (DBN), recurrent neural networks (RNN), residual neural networks (e.g., RestNet), U-Net-style neural networks, generative adversarial networks (GAN), conditional generative adversarial networks (cGAN), feedforward networks, networks with fully connected layers, trained support vector machines (SVM), trained random forests (RF), computer vision (CV) systems, autoregressive (AR) models, sequence-to-sequence (Seq2Seq) models, large language models (LLM), deep learning systems, classifiers, transformer models, attention models, multi-head attention models, or combinations thereof. In the example of ML model 625 including LLM, LLM may include, for example, a generative pre-trained transformer (GPT) (e.g., GPT-2, GPT-3, GPT-3.5, GPT-4, etc.), DaVinci or variants thereof, or a transformer using MIT. ® Long-chain LLM, Path Language Model (PaLM), Large-scale Language Model Meta ® AI (LLaMA), Language Model for Dialogue Applications (LaMDA), Bidirectional Encoder Representation from Transformer (BERT), Falcon (e.g., 40B, 7B, 1B), Orca, Phi-1, StableLM, any variant of the previously listed LLMs, or combinations thereof.

[0095] exist Figure 6The graphic representation of ML model 625 illustrates a set of circles connected to another set of circles. Each of these circles may represent a node, neuron, perceptron, layer, a portion thereof, or a combination thereof. These circles are arranged in columns. The white circles in the leftmost column represent the input layer. The white circles in the rightmost column represent the output layer. The two columns of shaded circles between the white circles in the leftmost and rightmost columns each represent hidden layers. An ML model may include more or fewer hidden layers than the two illustrated, but may include at least one hidden layer. In some examples, layers and / or nodes represent interconnected filters, and information associated with the filters is shared between different layers, where each layer retains information while processing it. Lines between nodes may represent node-to-node interconnections along which information is shared. Lines between nodes may also represent weights (e.g., numerical weights) between nodes, which may be tuned, updated, added, and / or removed during training and / or updating of ML model 625. In some cases, certain nodes (e.g., nodes in hidden layers) can transform that information by applying activation functions (e.g., filters) to the information of each input node, such as applying convolution functions, shrinking, scaling, data transformations, and / or any other suitable functions.

[0096] In some examples, the ML model 625 may include a feedforward network, in which case there are no feedback connections where the network's output is fed back into itself. In some cases, the ML model 625 may include a recurrent neural network, which may have loops that allow information to be carried across nodes as input is read. In some cases, the network may include a convolutional neural network, which may not link each node in a layer to every other node in the next layer.

[0097] One or more inputs 605 may be provided to the ML model 625. The ML model 625 may be trained by the ML engine 620 (e.g., based on training data 660) to generate one or more outputs 630. In some examples, the input 605 includes a phase detection buffer 610, which includes phase detection data and / or focus data. In some examples, for example, the phase detection buffer 610 may include 2PD focus data or focus data corresponding to masked focus pixels, such as left and right focus data, upper and lower focus data, or combinations thereof. In some examples, the phase detection buffer 610 may include 4PD focus data (e.g., corresponding to at least one 4PD focus pixel, such as...). Figure 3F 4PD focused pixels (362), such as top left focus data, bottom left focus data, top right focus data, bottom right focus data, or combinations thereof. Phase shift detected by PDAF (such as...) Figure 5 The phase shift 520 may be included in the phase detection buffer 610.

[0098] In some examples, input 605 includes image data 615, such as raw image data (e.g., Bayer image data) from an image sensor (e.g., image sensor 130, image sensor 285). The image data may include focus data from focused pixels, image data from image pixels, or a combination thereof. In some examples, image data 615 includes processing such as demosaicing, tone mapping, denoising, pixel correction, automatic white balance, processing operations discussed in connection with ISP 154, or a combination thereof.

[0099] The output 630 generated by the ML model 625 in response to inputs in input 605 (e.g., in response to PD buffer 610 and / or image data 615) includes a predicted direction 635 in which the lens will move (e.g., toward the image sensor and away from the subject, or toward the subject and away from the image sensor) to move the lens from its current position (e.g., the position of capture input 605) toward a focusing lens position (e.g., focus state 250, focus 545 lens position). The predicted direction 635 may be referred to as and / or may be computed as a prediction of the sign (e.g., positive, negative, or zero) of a phase shift (e.g., phase shift 520). An example of the predicted direction 635 generated by the ML model 625 is illustrated as follows. Figure 5 The predicted direction 525 of the curve 500. In some examples, the output 630 generated by the ML model 625 may also include a confidence level 640 associated with the generation of the predicted direction 635, as in the predicted direction confidence level 535. In some examples, the output 630 generated by the ML model 625 may also include a prediction of the magnitude of the phase shift, and / or a prediction of the distance by which the lens is moved to approach and / or reach the focusing lens position (e.g., focusing state 250, focusing lens position 545).

[0100] In some examples, the ML system includes one or more feedback engines 645 that generate and / or provide feedback 650 about the output 630. In some examples, feedback 650 indicates the degree to which the output 630 is aligned with the corresponding expected output, the degree to which the output 630 serves its intended purpose, or a combination thereof. In some examples, feedback engine 645 includes a loss function, a reward model (e.g., another ML model used to score the output 630), a discriminator, an error function (e.g., in backpropagation), user interface feedback received from the user via the user interface, or a combination thereof. In some examples, feedback 650 may include one or more alignment scores that score the level of alignment between the output 630 and the expected output and / or the intended purpose.

[0101] For example, in an exemplary example, feedback engine 645 may include a loss function that compares output 630 (e.g., predicted direction 635) with a ground truth (e.g., the direction in which the lens actually needs to be moved to reach focus lens positioning) to generate feedback 650 (e.g., the result of the loss function). In some examples, ML system 600 may determine the ground truth via PDAF with pixel binning, via CDAF, via active autofocus (e.g., using a depth sensor such as a laser rangefinder or time-of-flight sensor), or a combination thereof. In some examples, feedback engine 645 may receive feedback 650 from the camera after the lens has been moved according to predicted direction 635. For example, if the camera moves the lens according to predicted direction 635 and the focus value (FV) increases to indicate that the movement brings the lens closer to focus lens positioning, then feedback 650 is positive, indicating that predicted direction 635 was successfully predicted. On the other hand, if the camera moves the lens according to predicted direction 635 and the focus value (FV) decreases to indicate that the movement moves the lens further away from focus lens positioning, then feedback 650 is negative, indicating that predicted direction 635 was not successfully predicted.

[0102] The ML engine 620 of the ML system can update (further train) the ML model 625 based on feedback 650 to perform updates 655 (e.g., further training) of the ML model 625 based on feedback 650. In some examples, feedback 650 includes positive feedback, such as indicating that output 630 is closely aligned with the expected output and / or serves its intended purpose (e.g., movement of the lens in the predicted direction 635 actually brings the lens closer to the focusing lens positioning). In some examples, feedback 650 includes negative feedback, such as indicating a mismatch between output 630 and the expected output, and / or that output 630 does not serve its intended purpose (e.g., movement of the lens in the predicted direction 635 actually moves the lens further away from the focusing lens positioning). For example, a large amount of loss and / or error (e.g., loss function exceeding a threshold) can be interpreted as negative feedback, while a small amount of loss and / or error (e.g., loss function less than a threshold) can be interpreted as positive feedback. Similarly, a high amount of alignment (e.g., exceeding a threshold) can be interpreted as positive feedback, while a low amount of alignment (e.g., less than a threshold) can be interpreted as negative feedback. In response to positive feedback in feedback 650, ML engine 620 may perform update 655 to update ML model 625 to strengthen and / or enhance the weights associated with the generation of output 630 (e.g., predicted direction 635, confidence 640) to encourage ML engine 620 to generate similar output 630 given similar input 605 (e.g., PD buffer 610 and / or image data 615). In response to negative feedback in feedback 650, ML engine 620 may perform update 655 to update ML model 625 to weaken and / or remove the weights associated with the generation of output 630 (e.g., predicted direction 635, confidence 640) to prevent ML engine 620 from generating similar output 630 given similar input 605 (e.g., PD buffer 610 and / or image data 615).

[0103] In some examples, before using ML model 625 to generate output 630 based on input 605, ML engine 620 performs initial training of ML model 625. During initial training, ML engine 620 may train ML model 625 based on training data 660. In some examples, training data 660 includes examples of inputs (of any input type discussed with respect to input 605), outputs (of any output type discussed with respect to output 630), and / or feedback (of any feedback type discussed with respect to feedback 650). In exemplary examples, training data 660 may include a PD buffer (as in PD buffer 610), image data (as in image data 615), a prediction direction corresponding to the input PD buffer and / or image data (as in prediction direction 635), and feedback indicating whether the prediction direction is a good or bad prediction given the input PD buffer and / or image data (as in feedback 650).

[0104] Figure 7 This is a block diagram illustrating an imaging system 700 including an ML prediction engine 710, a PDAF engine 705, and a focusing and converging engine 710. Within the imaging system 700, as per... Figure 6 As discussed, input 605 (e.g., PD buffer 610 and / or image data 615) is fed into ML prediction engine 620 to generate output 630 (e.g., prediction direction 635 and / or confidence 640 associated with the generation of prediction direction 635). Within imaging system 700, input 605 (e.g., PD buffer 610 and / or image data 615) is also fed into PDAF engine 705 to determine phase shift (e.g., sign and / or magnitude), for example, as per [reference to image data]. Figure 1 , Figures 2A to 2C , Figures 3A to 3F , Figures 4A to 4B and Figure 5 The PDAF process discussed herein. In some examples, PDAF engine 705 also generates a PDAF confidence level associated with the determined phase shift (e.g., PDAF confidence level 530). In an exemplary example, PDAF engine 705 specifically determines the phase shift and / or PDAF confidence level based on a PD buffer (e.g., PD buffer 610) in input 605.

[0105] The output 630 of the ML prediction engine 620 (e.g., prediction direction 635, confidence level 640) and the output of the PDAF engine 705 (e.g., phase shift and / or corresponding PDAF confidence level) are input to the focus convergence engine 710. The focus convergence engine 710 can determine the direction and / or distance of the lens moving the camera based on the outputs of the ML prediction engine 620 and the PDAF engine 705 provided to the focus convergence engine 710. Examples of focus convergence engines 710 include focus convergence engines 805, 910, 1005, 1105, or combinations thereof.

[0106] Figure 8This is a block diagram illustrating a process 800 for determining lens movement for autofocus by a focusing and convergence engine 805 based on an ML prediction engine 620 and a PDAF engine 705. The focusing and convergence engine 805 can determine at decision 810 whether a predicted direction (e.g., predicted direction 635) is reliable based on a predicted direction (e.g., predicted direction 635) from the ML prediction engine 620 and / or an associated confidence level (e.g., confidence level 640). For example, in some examples, the focusing and convergence engine 805 can make decision 810 by comparing the confidence level associated with the predicted direction (e.g., confidence level 640) with a confidence threshold. If the confidence level meets and / or exceeds the confidence threshold, the focusing and convergence engine 805 can make decision 810 by determining that the predicted direction (e.g., predicted direction 635) is reliable. If the confidence level is less than (or, in some examples, equal to) the confidence threshold, the focusing convergence engine 805 can make a decision 810 by determining that the predicted direction (e.g., predicted direction 635) is unreliable.

[0107] The convergence engine 805 can determine at decision 815 whether the sign and / or magnitude of the phase shift (e.g., phase shift 520) determined by the PDAF engine 705 is reliable based on the phase shift (e.g., phase shift 520) and / or the associated PDAF confidence level (e.g., PDAF confidence level 530). For example, in some examples, the convergence engine 805 can make decision 815 by comparing the confidence level associated with the phase shift (e.g., phase shift 520) with a confidence level threshold. If the confidence level meets and / or exceeds the confidence level threshold, the convergence engine 805 can make decision 815 by determining that the sign and / or magnitude of the phase shift (e.g., phase shift 520) determined via PDAF is reliable. If the confidence level is less than (or, in some examples, equal to) the confidence threshold, the focusing convergence engine 805 can make a decision 815 by determining that the sign and / or magnitude of the phase shift (e.g., phase shift 520) is unreliable. In some examples, the confidence threshold for decision 815 matches the confidence threshold for decision 810. In some examples, the confidence threshold for decision 815 is different from (e.g., greater than or less than) the confidence threshold for decision 810.

[0108] The focusing convergence engine 805 can determine at decision 820 whether the symbol matches between the prediction direction (e.g., prediction direction 635) and the phase shift (e.g., phase shift 520) based on the prediction direction from the ML prediction engine 620 and the phase shift from the PDAF engine 705. In other words, the focusing convergence engine 805 can determine at decision 820 whether the prediction direction (e.g., prediction direction 635) and the phase shift (e.g., phase shift 520) indicate the same direction of movement of the lens toward the focusing lens positioning (e.g., focusing state 250, focusing lens positioning 545).

[0109] In some examples, if the focusing and converging engine 805 determines yes for decision 810, yes for decision 815, and yes for decision 820, then at operation 825, the focusing and converging engine 805 can move the lens by a distance based on the magnitude of the phase shift in the direction indicated by both the predicted direction and the phase shift (e.g., by actuating actuator 280).

[0110] In some examples, if the focusing and converging engine 805 determines no to decision 810 or decision 815, then at operation 835, the focusing and converging engine 805 may move the lens in a direction indicated by the predicted direction (and / or phase shift) by a first specified distance (e.g., by actuating actuator 280). The first specified distance may be selected from a set of predetermined distances and may, for example, represent a smaller distance from that set of predetermined distances. Because operation 835 is reached when the focusing and converging engine 805 has low confidence in the reliability of the predicted direction (e.g., predicted direction 635) or the phase shift (e.g., phase shift 520) or both, moving the lens a small distance can offset any potential risks associated with the possibility of an incorrect predicted direction or a phase shift with an incorrect sign (e.g., as in error 540).

[0111] In some examples, if the focusing and converging engine 805 determines yes for decision 810, yes for decision 815, and no for decision 820, then at operation 830, the focusing and converging engine 805 can move the lens in the direction indicated by the predicted direction (and / or phase shift) by a second specified distance (e.g., by actuating actuator 280). The second specified distance can be selected from a set of predetermined distances and can, for example, represent a larger distance in that set of predetermined distances (e.g., a distance greater than the first specified distance of operation 835). Because operation 830 is reached when the focusing and converging engine 805 has a high confidence level in the reliability of the predicted direction (e.g., predicted direction 635), moving the lens a large distance can potentially improve efficiency by requiring less movement to reach the focusing lens position. Because operation 830 is reached when the focusing and converging engine 805 has low confidence in the reliability of the phase shift (e.g., phase shift 520), moving the lens a predetermined distance (opposite to the distance indicated by the magnitude of the phase shift as in operation 825) can compensate for any error in the magnitude of the phase shift (e.g., as in error 540).

[0112] In some examples, if the focus convergence engine 805 determines yes to decision 810 or decision 815 (but not both), then the focus convergence engine 805 may perform operation 830 instead of operation 835.

[0113] Figure 9 This is a block diagram illustrating an imaging system 900 including an ML prediction engine 620, a PDAF engine 705, a CDAF and a focus value (FV) statistics engine 905, and a focus convergence engine 910. Within the imaging system 900, as per... Figure 6 As discussed, input 605 (e.g., PD buffer 610 and / or image data 615) is fed into ML prediction engine 620 to generate output 630 (e.g., prediction direction 635 and / or confidence 640 associated with the generation of prediction direction 635). Within imaging system 900, input 605 (e.g., PD buffer 610 and / or image data 615) is also fed into PDAF engine 705 to determine phase shift (e.g., sign and / or magnitude), for example, as per [reference to image data]. Figure 1 , Figures 2A to 2C , Figures 3A to 3F , Figures 4A to 4B , Figure 5 , Figure 7 and Figure 8The PDAF process discussed herein. In some examples, PDAF engine 705 also generates a PDAF confidence level (e.g., PDAF confidence level 530) associated with the determined phase shift. In an exemplary example, PDAF engine 705 specifically determines the phase shift and / or PDAF confidence level based on a PD buffer (e.g., PD buffer 610) in input 605. Within imaging system 900, input 605 (e.g., PD buffer 610 and / or image data 615) is also fed into CDAF and FV statistics engine 905 to determine a focus value (FV) associated with the image data in input 605 (e.g., based on detected contrast in the image data). In some examples, CDAF and FV statistics engine 905 also generates a CDAF confidence level associated with the determined FV. In an exemplary example, CDAF and FV statistics engine 905 specifically determines the FV and / or CDAF confidence level based on image data (e.g., image data 615) in input 605. In some cases, the CDAF and FV statistics engine 905 may be referred to as the CDAF and FV engine, CDAF engine, CDAF statistics engine, FV engine, FV statistics engine, or a combination thereof.

[0114] The outputs 630 of the ML prediction engine 620 (e.g., prediction direction 635, confidence level 640), the outputs of the PDAF engine 705 (e.g., phase shift and / or corresponding PDAF confidence level), and / or the outputs of the CDAF and FV statistics engine 905 (e.g., FV and / or corresponding CDAF confidence level) are input into the focus convergence engine 910. The focus convergence engine 910 can determine the direction and / or distance of the lens moving the camera based on the outputs of the ML prediction engine 620, the PDAF engine 705, and / or the CDAF and FV statistics engine 905 provided to the focus convergence engine 910. Examples of focus convergence engines 910 include focus convergence engines 710, 805, 1005, 1105, or combinations thereof.

[0115] Figure 10This is a block diagram illustrating a process 1000 for determining lens movement for autofocus using a focusing and convergence engine 1005 based on an ML prediction engine 620, a PDAF engine 705, and a CDAF and FV statistics engine 905. Similar to the focusing and convergence engine 805, the focusing and convergence engine 1005 receives a predicted direction (e.g., predicted direction 635) and / or an associated confidence level (e.g., confidence level 640) from the ML prediction engine 620, and makes a determination 810 based on this predicted direction and / or the associated confidence level. Similarly, similar to the focusing and convergence engine 805, the focusing and convergence engine 1005 receives a phase shift (e.g., phase shift 520) and / or an associated PDAF confidence level (e.g., PDAF confidence level 530) from the PDAF engine 705, and makes a determination 815 based on this phase shift and / or the associated PDAF confidence level. Similar to the focusing convergence engine 805, the focusing convergence engine 1005 makes a decision 820 based on whether the symbol matches between the prediction direction (e.g., prediction direction 635) and the phase shift (e.g., phase shift 520).

[0116] If the focusing and convergence engine 1005 determines no at decision 815, indicating that the sign and / or magnitude of the phase shift (e.g., phase shift 520) determined by the PDAF engine 705 is unreliable, then the focusing and convergence engine 1005 may make decision 1010. At decision 1010, the focusing and convergence engine 1005 may determine whether the number of lens positions for which the CDAF and FV statistics engine 905 has generated focus values ​​(FVs) exceeds the FV quantity threshold N. The focusing and convergence engine 1005 may obtain from the CDAF and FV statistics engine 905 the number of lens positions for which the CDAF and FV statistics engine 905 have generated FVs. In some examples, the FV quantity threshold N may be 2, 3, 4, 5, 6, 7, 8, 9, 10, or another number greater than one.

[0117] In some examples, if the focus and convergence engine 1005 determines that decision 1010 is negative, then the focus and convergence engine 1005 can continue to operation 830. In some examples, if the focus and convergence engine 1005 determines that decision 1010 is positive, then the focus and convergence engine 1005 can continue to decision 1015.

[0118] The focus and convergence engine 1005 can determine at decision 1015 whether the FV determined by the CDAF and FV statistics engine 905 is reliable based on the FV generated for the current image data (e.g., input 605) and / or the associated CDAF confidence level. For example, in some examples, the focus and convergence engine 1005 can make decision 1015 by comparing the confidence level associated with the FV to a confidence level threshold. In some examples, the confidence level is at least partially based on the number of lens localizations for which the CDAF and FV statistics engine 905 has generated the FV (as discussed with respect to decision 1010), such that a higher number of FVs are associated with a higher confidence level, and a lower number of FVs are associated with a lower confidence level. If the confidence level meets and / or exceeds the confidence level threshold, the focus and convergence engine 1005 can make decision 1015 by determining that the FV determined by the CDAF and FV statistics engine 905 is reliable. If the confidence level is less than (or, in some examples, equal to) the confidence threshold, the focusing convergence engine 1005 can make a decision 1015 by determining that the FV is unreliable, as determined by the CDAF and FV statistics engine 905. In some examples, the confidence threshold for decision 1015 matches the confidence threshold for decision 810 and / or decision 815. In some examples, the confidence threshold for decision 1015 is different from (e.g., greater than or less than) the confidence threshold for decision 810 and / or decision 815.

[0119] In some examples, if the focus and convergence engine 1005 determines that decision 1015 is negative, then the focus and convergence engine 1005 can continue with operation 830. In some examples, if the focus and convergence engine 1005 determines that decision 1015 is positive, then the focus and convergence engine 1005 can continue with operation 830.

[0120] Figure 11 This is a block diagram illustrating various aspects of a process 1100 for determining lens movement for autofocus, performed by a focusing and converging engine 1105. The focusing and converging engine 1105 may be, for example, part of a focusing and converging engine 710, a focusing and converging engine 805, a focusing and converging engine 910, or a focusing and converging engine 1005. Figure 11 Procedure 1100 provides additional context for operations 825, 830, and 835.

[0121] In some examples, at operation 825, the focusing convergence engine 1105 moves the lens based on the PDAF defocus value. In some examples, the PDAF defocus value may be calculated from or equal to the phase shift value. In some examples, the PDAF defocus value may be a measure of how far the current lens position is from the focusing lens position (e.g., focus state 250, focus state 545 lens position).

[0122] In some examples, at operation 830, the focusing convergence engine 1105 moves the lens in the direction predicted by the prediction direction determined by the ML prediction engine 620 (e.g., prediction direction 635).

[0123] In some examples, at operation 830, the focus convergence engine 1105 makes a decision 1110 based on the current lens position of the lens. If the focus convergence engine 1105 determines at decision 1110 that the current lens position is close to the near end (e.g., within a threshold distance of the near end), then at operation 1120, the focus convergence engine 1105 may move the lens from the current lens position to a lens position associated with a minimum working distance associated with the camera. In some examples, the term "near end" may refer to a lens that is close to the image sensor and far from the subject. In an illustrative example, in Figure 5 In the context of graph 500, the lens positioning associated with the minimum working distance of the camera can be, for example, at lens positioning 150. In some examples, the threshold distance is equivalent to the difference between the lens positioning associated with the minimum working distance (e.g., 150) and the minimum lens positioning (e.g., 0).

[0124] If the focus convergence engine 1105 determines at determination 1110 that the current lens positioning is not close to the near end and / or near the far end (e.g., within a threshold distance of the far end), then at operation 1125, the focus convergence engine 1105 may move the lens from the current lens positioning to a lens positioning associated with a hyperfocal distance (e.g., hyperfocal distance) related to the camera. In some examples, the term "far end" may refer to a lens that is close to the subject and far from the image sensor. In an exemplary example, in Figure 5 In the context of graph 500, the lens positioning associated with the minimum working distance of the camera can be, for example, at lens positioning 800. In some examples, the threshold distance is equivalent to the difference between the maximum lens positioning (e.g., 950) and the lens positioning associated with the hyperfocal distance (e.g., 800).

[0125] Figure 12This is a block diagram illustrating an example of a neural network (NN) 1200 that can be used for imaging operations. The neural network 1200 can include any type of deep network, such as a convolutional neural network (CNN), an autoencoder, a deep belief network (DBN), a recurrent neural network (RNN), a residual neural network (e.g., ResNet), a generative adversarial network (GAN), and / or other types of neural networks. The neural network 1200 can be an example of an ML model 625 of an ML prediction engine 620, a trained ML model operating 1505, a neural network or other machine learning model running on a computing system 1600, or a combination thereof. The neural network 1200 can be used by various systems discussed herein, such as image capture and processing system 100, PDAF camera system 200, ML system 600, imaging system 700, focusing and converging engine 710, focusing and converging engine 805, imaging system 900, focusing and converging engine 910, focusing and converging engine 1005, focusing and converging engine 1105, head-mounted display (HMD) 1310, mobile phone 1410, imaging system of execution process 1500, computing system 1600, or combinations thereof.

[0126] The input layer 1210 of the neural network 1200 includes input data. The input data of the input layer 1210 may include data representing pixels of one or more input image frames. In some examples, the input data of the input layer 1210 includes data representing pixels of images captured using the image capture and processing system 100, images captured using the PDAF camera system 200, and data from... Figures 3A to 3F The focus data of any focus pixel in the focus pixel, from Figures 3A to 3F Image data of any imaging pixel in the imaging pixel, from Figures 4A to 4B The focus data of any of the photodiodes 420A-420C, phase shift 520 data (e.g., associated with a PDAF engine), PD buffer 610, image data 615 (e.g., raw image data, such as Bayer image data), and information about... Figure 6Other inputs discussed include 605, phase-shift data (e.g., associated with PDAF engine 705), CDAF FV data (e.g., associated with CDAF and FV statistics engine 905), images captured by one of cameras 1330A-1330D, images captured by one of cameras 1430A-1430D, focus data of operation 1505, image data corresponding to the focus data of operation 1505, image data captured using input device 1645, image data captured using any other image sensor described herein, any other image data described herein, or combinations thereof. Images may include image data from image sensors, including raw pixel data (including single color per pixel based on, for example, a Bayer color filter) or processed pixel values ​​(e.g., RGB pixels of an RGB image).

[0127] The neural network 1200 includes multiple hidden layers 1212, 1212B through 1212N. Hidden layers 1212, 1212B through 1212N comprise "N" hidden layers, where "N" is an integer greater than or equal to one. The multiple hidden layers can include as many layers as needed for a given application. The neural network 1200 also includes an output layer 1214, which provides the output resulting from the processing performed by the hidden layers 1212, 1212B through 1212N.

[0128] Output layer 1214 provides output data for operations performed using NN 1200. In some examples, output layer 1214 may provide a predicted direction (e.g., predicted direction 635), a predicted phase shift sign, a predicted phase shift magnitude, a confidence level (e.g., confidence level 640) associated with any one or more of the previously listed predictions, or a combination thereof, for moving the lens to achieve a focusing lens position (e.g., focusing state 250 and / or focusing lens position 545). In some examples, one of input layer 1210, output layer 1214, or hidden layers 1212A-1212N may generate intermediate data for generating further outputs.

[0129] Neural network 1200 is a multi-layer neural network of interconnected filters. Each filter can be trained to learn features representing the input data. Information associated with these filters is shared between different layers, and each layer retains information while processing it. In some cases, neural network 1200 may include a feedforward network, in which case there are no feedback connections where the network's output is fed back into itself. In some cases, network 1200 may include a recurrent neural network, which may have loops that allow information to be carried across nodes as input is read in.

[0130] In some cases, information can be exchanged between layers via node-to-node (neuron-to-neuron) interconnections (synapses). In some cases, the network may include a convolutional neural network, which may not link every node in one layer to every other node in the next layer. In a network in which information is exchanged between layers, nodes in input layer 1210 can activate a set of nodes in the first hidden layer 1212A. For example, as shown, each input node in input layer 1210 can be connected to each node in the first hidden layer 1212A. Nodes in the hidden layer can transform information by applying an activation function (e.g., a filter) to the information of each input node. The information derived from this transformation can then be passed to nodes in the next hidden layer 1212B and can activate these nodes, which can perform their own specified functions. Example functions include convolution, shrinking, magnification, data transformation, and / or any other suitable function. The output of hidden layer 1212B can then activate nodes in the next hidden layer, and so on. The output of the last hidden layer 1212N can activate one or more nodes of the output layer 1214, which provides the processed output image. In some cases, although nodes in the neural network 1200 (e.g., nodes 1216, nodes 1218) are shown as having multiple output lines, the nodes have a single output and are shown as all lines output from the node representing the same output value.

[0131] In some cases, each node or interconnection between nodes may have weights, which are a set of parameters derived from the training of the neural network 1200. For example, an interconnection between nodes may represent a piece of information about what the interconnected nodes have learned. Interconnections may have tunable numerical weights that can be tuned (e.g., based on the training dataset), allowing the neural network 1200 to adapt to the input and learn as more data is processed. For example, example weight 1220 is illustrated along the interconnection between nodes 1216 and 1218. Other interconnections between other nodes of the neural network 1200 may have different corresponding weights. In some examples, nodes of the neural network 1200 (e.g., nodes 1216, 1218) have corresponding biases or bias offsets that can also be tuned in the neural network 1200. In some examples, interconnections between nodes of the neural network 1200 (such as the interconnection corresponding to example weight 1220) have corresponding biases or bias offsets that can also be tuned in the neural network 1200, for example, during training.

[0132] The neural network 1200 is pre-trained to process features from the data in the input layer 1210 using different hidden layers 1212, 1212B to 1212N in order to provide output through the output layer 1214.

[0133] Figure 13AThis is a perspective view 1300 illustrating a head-mounted display (HMD) 1310 used as part of a sensor data processing system. The HMD 1310 may be, for example, an augmented reality (AR) head-mounted device, a virtual reality (VR) head-mounted device, a mixed reality (MR) head-mounted device, an extended reality (XR) head-mounted device, or some combination thereof. The HMD 1310 includes a first camera 1330A and a second camera 1330B along the front of the HMD 1310. The HMD 1310 includes a third camera 1330C and a fourth camera 1330D, which face the user's eyes when the user's eyes are facing the display 1340. In some examples, the HMD 1310 may have only a single camera with a single image sensor. In some examples, in addition to the first camera 1330A, the second camera 1330B, the third camera 1330C, and the fourth camera 1330D, the HMD 1310 may also include one or more additional cameras. In some examples, in addition to the first camera 1330A, the second camera 1330B, the third camera 1330C, and the fourth camera 1330D, the HMD 1310 may also include one or more additional sensors. In some examples, the first camera 1330A, the second camera 1330B, the third camera 1330C, and / or the fourth camera 1330D may be examples of the image capture and processing system 100, the image capture device 105A, the image processing device 105B, the image sensor 285, or combinations thereof.

[0134] HMD 1310 may include one or more displays 1340 visible to user 1320 (who wears HMD 1310 on their head). In some examples, HMD 1310 may include one display 1340 and two viewfinders. The two viewfinders may include a left viewfinder for user 1320's left eye and a right viewfinder for user 1320's right eye. The left viewfinder may be oriented such that user 1320's left eye sees the left side of the display. The right viewfinder may be oriented such that user 1320's right eye sees the right side of the display. In some examples, HMD 1310 may include two displays 1340, including a left display that displays content to user 1320's left eye and a right display that displays content to user 1320's right eye. The one or more displays 1340 of HMD 1310 may be a digital "passthrough" display or an optical "passthrough" display.

[0135] The HMD 1310 may include one or more earpieces 1335, which can be used as speakers and / or headphones to output audio to one or both ears of the user of the HMD 1310. Figure 13A and Figure 13BAn earpiece 1335 is illustrated, but it should be understood that the HMD 1310 may include two earpieces, one for each of the user's ears (left and right). In some examples, the HMD 1310 may also include one or more microphones (not shown). In some examples, the audio output by the HMD 1310 to the user through one or more earpieces 1335 may include or be based on audio recorded using one or more microphones.

[0136] Figure 13B This is an example Figure 13A A head-mounted display (HMD) is being viewed from perspective 1350 by user 1320. User 1320 wears HMD 1310 on their head, above their eyes. HMD 1310 can capture images using a first camera 1330A and a second camera 1330B. In some examples, HMD 1310 displays one or more output images towards user 1320's eyes using display 1340. In some examples, the output images may include processed image data (e.g., images autofocused using ML prediction engine 620). The output images may be based on images captured by the first camera 1330A and the second camera 1330B (e.g., image sensor 130, image sensor 285), for example, overlaid with processed image data (e.g., images autofocused using ML prediction engine 620). The output images can provide a stereoscopic view of the environment, in some cases overlaid with processed content and / or with other modifications. For example, HMD 1310 may display a first display image to the right eye of user 1320, the first display image being based on an image captured by a first camera 1330A. HMD 1310 may display a second display image to the left eye of user 1320, the second display image being based on an image captured by a second camera 1330B. For example, HMD 1310 may provide overlaid processed content in the display image, the overlaid processed content being superimposed on the images captured by the first camera 1330A and the second camera 1330B. A third camera 1330C and a fourth camera 1330D may capture images of the eyes before, during, and / or after the user views the display image displayed by display 1340. Thus, sensor data from the third camera 1330C and / or the fourth camera 1330D may capture the user's eye (and / or other parts of the user) response to the processed content. The earpiece 1335 of HMD 1310 is illustrated as being in the ear of user 1320. The HMD 1310 can output audio to the user 1320 via earpiece 1335 and / or via another earpiece (not shown) in the other ear (not shown) of the user 1320.

[0137] Figure 14AThis is a perspective view 1400 illustrating the front surface of a mobile phone 1410, which includes a front-facing camera and can be used as part of a sensor data processing system. The mobile phone 1410 can be, for example, a cellular phone, satellite phone, portable game console, music player, health tracking device, wearable device, wireless communication device, laptop computer, mobile device, any other type of computing device or computing system discussed herein, or a combination thereof.

[0138] The front surface 1420 of the mobile phone 1410 includes a display 1440. The front surface 1420 of the mobile phone 1410 includes a first camera 1430A and a second camera 1430B. The first camera 1430A and the second camera 1430B can be oriented towards the user, including the user's eyes, while displaying processed image data (e.g., an autofocused image using the ML prediction engine 620) on the display 1440.

[0139] A first camera 1430A and a second camera 1430B are illustrated in the bezel surrounding the display 1440 on the front surface 1420 of the mobile phone 1410. In some examples, the first camera 1430A and the second camera 1430B may be positioned in a notch or cutout cut out of the display 1440 on the front surface 1420 of the mobile phone 1410. In some examples, the first camera 1430A and the second camera 1430B may be under-display cameras positioned between the display 1440 and the rest of the mobile phone 1410, such that light passes through a portion of the display 1440 before reaching the first camera 1430A and the second camera 1430B. In perspective view 1400, the first camera 1430A and the second camera 1430B are front-facing cameras. The first camera 1430A and the second camera 1430B face a direction perpendicular to the planar surface of the front surface 1420 of the mobile phone 1410. The first camera 1430A and the second camera 1430B can be two of one or more cameras in the mobile phone 1410. In some examples, the front surface 1420 of the mobile phone 1410 may have only a single camera.

[0140] In some examples, the display 1440 of the mobile phone 1410 displays one or more output images to a user using the mobile phone 1410. In some examples, the output images may include processed image data (e.g., an image autofocused using the ML prediction engine 620). The output images may be based on images captured by the first camera 1430A, the second camera 1430B, the third camera 1430C, and / or the fourth camera 1430D (e.g., captured by image sensor 130 or image sensor 285), for example, overlaid with processed image data (e.g., an image autofocused using the ML prediction engine 620).

[0141] In some examples, in addition to the first camera 1430A and the second camera 1430B, the front surface 1420 of the mobile phone 1410 may also include one or more additional cameras. In some examples, in addition to the first camera 1430A and the second camera 1430B, the front surface 1420 of the mobile phone 1410 may also include one or more additional sensors. In some cases, the front surface 1420 of the mobile phone 1410 includes more than one display 1440. For example, one or more displays 1440 may include one or more touchscreen displays.

[0142] Mobile phone 1410 may include one or more speakers 1435A and / or other audio output devices (e.g., headphones or headsets or their connectors) that can output audio to one or both ears of the user of mobile phone 1410. Figure 14A A speaker 1435A is illustrated, but it should be understood that the mobile phone 1410 may include more than one speaker and / or other audio devices. In some examples, the mobile phone 1410 may also include one or more microphones (not shown). In some examples, the audio output by the mobile phone 1410 to the user through one or more speakers 1435A and / or other audio output devices may include or be based on audio recorded using one or more microphones.

[0143] Figure 14B This is a perspective view 1450 illustrating the rear surface 1460 of a mobile phone that includes a rear-facing camera and can be used as part of a sensor data processing system. The mobile phone 1410 includes a third camera 1430C and a fourth camera 1430D on the rear surface 1460 of the mobile phone 1410. The third camera 1430C and the fourth camera 1430D in perspective view 1450 are rear-facing. The third camera 1430C and the fourth camera 1430D face a direction perpendicular to the planar surface of the rear surface 1460 of the mobile phone 1410.

[0144] The third camera 1430C and the fourth camera 1430D may be two of one or more cameras in the mobile phone 1410. In some examples, the rear surface 1460 of the mobile phone 1410 may have only a single camera. In some examples, the rear surface 1460 of the mobile phone 1410 may also include one or more additional cameras in addition to the third camera 1430C and the fourth camera 1430D. In some examples, the rear surface 1460 of the mobile phone 1410 may also include one or more additional sensors in addition to the third camera 1430C and the fourth camera 1430D. In some examples, the first camera 1430A, the second camera 1430B, the third camera 1430C, and / or the fourth camera 1430D may be examples of an image capture and processing system 100, an image capture device 105A, an image processing device 105B, an image sensor 285, or a combination thereof.

[0145] Mobile phone 1410 may include one or more speakers 1435B and / or other audio output devices (e.g., headphones or headsets or their connectors) that can output audio to one or both ears of the user of mobile phone 1410. Figure 14B A speaker 1435B is illustrated, but it should be understood that the mobile phone 1410 may include more than one speaker and / or other audio devices. In some examples, the mobile phone 1410 may also include one or more microphones (not shown). In some examples, the mobile phone 1410 may include one or more microphones along and / or adjacent to the rear surface 1460 of the mobile phone 1410. In some examples, the audio output by the mobile phone 1410 to the user through one or more speakers 1435B and / or other audio output devices may include or be based on audio recorded using one or more microphones.

[0146] Mobile phone 1410 may use display 1440 on front surface 1420 as a transparent display. For example, display 1440 may display output images, such as processed image data (e.g., an image autofocused using ML prediction engine 620). The output image may be based on images captured by third camera 1430C and / or fourth camera 1430D (e.g., from image sensor 130, image sensor 285), such as overlaid with processed image data (e.g., an image autofocused using ML prediction engine 620). First camera 1430A and / or second camera 1430B may capture images of the user's eyes (and / or other parts of the user) before, during, and / or after the output image with processed content is displayed on display 1440. In this way, sensor data from first camera 1430A and / or second camera 1430B may capture the user's eyes (and / or other parts of the user) reaction to the processed content.

[0147] Figure 15 This is a flowchart illustrating a process 1500 for imaging. Process 1500 can be performed by an imaging system. In some examples, the imaging system may include, for example, an image capture and processing system 100, an image capture device 105A, an image processing device 105B, an image processor 150, and an ISP. 154. Host processor; 152. PDAF camera system; 200. Lens; 210. Focusing photodiodes 225A-225B; Actuator; 280. Image sensor; 285. Pixel array; 300. Pixel array; 330. Pixel array; 340. Pixel array; 350. Pixel array; 360. Focusing pixel; 400. Focusing pixel; 440. ML system; 600. ML prediction engine; 620. ML model; 625. Imaging system; 700. PDAF engine; 705. Focusing convergence engine; 710. Focusing convergence engine; 805. Imaging system; 900. CDAF and FV statistical engine; 905. Focusing convergence engine; 910. Focusing convergence engine; 1005. Focusing convergence engine; 1105. Neural network; 1200. Head-mounted display (HMD); 1310. Mobile phone; 1410. Computing system; 1600. Processor; 1610. Device, system, non-transitory computer-readable medium coupled to a processor, or combinations thereof.

[0148] At operation 1505, the imaging system (or at least one subsystem thereof) is configured and can use a trained machine learning model to process at least the focus data to identify a predicted direction of movement for moving the lens to improve image focus. The focus data is associated with at least one focused pixel of image data captured using an image sensor.

[0149] In some examples, in order to process at least the focused data using a trained machine learning model (as in operation 1505), the imaging system (or at least one subsystem thereof) is configured and can process at least the focused data and image data using a trained machine learning model. The image data includes at least one image pixel.

[0150] In some examples, the imaging system (or at least one subsystem thereof) is configured to receive (operation 1505) focus data, for example, from an image sensor that captures focus data. Examples of focus data (operation 1505) include focus data from focus pixels of an image captured using the image capture and processing system 100, focus data from other focus pixels of the focusing photodiodes 225A-225B and / or the image sensor 285, and focus data from... Figures 3A to 3F The focus data of any focus pixel in the focus pixel, from Figures 4A to 4B Focus data of any of the photodiodes 420A-420C, phase shift 520 data (e.g., associated with the PDAF engine), PD buffer 610, and information about... Figure 6 Other inputs discussed include 605, phase-shift data (e.g., associated with PDAF engine 705), focus data provided to input layer 1210 of NN 1200, focus data of focus pixels from any of cameras 1330A-1330D, focus data of focus pixels from any of cameras 1430A-1430D, focus data of focus pixels from input device 1645, focus data of focus pixels from any other image sensor described herein, any other focus data described herein, or combinations thereof.

[0151] In some examples, the imaging system (or at least one subsystem thereof) is configured to receive image data (operation 1505) from, for example, an image sensor that captures image data. Examples of image data (operation 1505) include data representing pixels of an image captured using the image capture and processing system 100, an image captured using the PDAF camera system 200, and data from... Figures 3A to 3F The focus data of any focus pixel in the focus pixel, from Figures 3A to 3F Image data of any imaging pixel in the imaging pixel, from Figures 4A to 4B The focus data of any of the photodiodes 420A-420C, phase shift 520 data (e.g., associated with a PDAF engine), PD buffer 610, image data 615 (e.g., raw image data, such as Bayer image data), and information about... Figure 6Other inputs discussed include 605, phase-shifted data (e.g., associated with PDAF engine 705), CDAF FV data (e.g., associated with CDAF and FV statistics engine 905), image data provided to input layer 1210 of NN 1200, images captured by one of cameras 1330A-1330D, images captured by one of cameras 1430A-1430D, image data captured using input device 1645, image data captured using any other image sensor described herein, any other image data described herein, or combinations thereof. Examples of image sensors include image sensor 130, image sensor 285, pixel array 300, pixel array 330, pixel array 340, pixel array 350, pixel array 360, image sensor with focused pixels 400, image sensor with focused pixels 440, image sensor capturing focus data and / or image data 615 in PD buffer 610, first camera 1330A, second camera 1330B, third camera 1330C, fourth camera 1330D, first camera 1430A, second camera 1430B, third camera 1430C, fourth camera 1430D, input device 1645, image sensor capturing any of the focus data and / or image data previously listed as examples of images, another image sensor described herein, another camera described herein, another sensor described herein, or combinations thereof.

[0152] In some examples, processing at least focused data using a trained machine learning model (as in operation 1505) includes processing both at least focused data and image data using a trained machine learning model. For example, see reference... Figure 6 Both focused data (e.g., from PD buffer 610) and image data (e.g., image data 615) can be fed as input (e.g., input 605) into a trained machine learning model (e.g., ML model 625). In some examples, the image data includes at least one image pixel (e.g., such as imaging pixel 304, as opposed to the focused pixel).

[0153] At operation 1510, the imaging system (or at least one subsystem thereof) is configured to allow the lens to be moved in the predicted direction of movement to improve image focusing.

[0154] In some examples, moving the lens in the predicted direction of movement (as in operation 1510) includes moving the lens a specified distance in the predicted direction of movement. Examples of the specified distance may include a first specified distance in operation 835, a second specified distance in operation 830, a distance based on a magnitude of phase shift as in operation 825, another distance discussed herein, or a combination thereof.

[0155] In some examples, the imaging system (or at least one subsystem thereof) is configured to and can select a specified distance from a plurality of predetermined distances (e.g., associated with operation 825, operation 830 and / or operation 835) based on a confidence level associated with the predicted direction of movement (e.g., as discussed with regard to decision 810 of focusing convergence engine 805 and / or focusing convergence engine 1005).

[0156] In some examples, the imaging system (or at least one subsystem thereof) is configured to determine a second predicted direction of movement based on the sign of the phase shift associated with the focused data. The imaging system can select a specified distance from a plurality of predetermined distances (e.g., associated with operation 825, operation 830, and / or operation 835) based on whether the predicted direction of movement matches the second predicted direction of movement (e.g., as discussed in decision 820 regarding focusing convergence engines 805 and / or focusing convergence engines 1005). In some examples, the imaging system (or at least one subsystem thereof) is configured to also select a specified distance from a plurality of predetermined distances based on a confidence level associated with the second predicted direction of movement (e.g., as discussed in decision 815 regarding focusing convergence engines 805 and / or focusing convergence engines 1005).

[0157] In some examples, the imaging system (or at least one subsystem thereof) is configured to determine a second predicted direction of movement based on corresponding focus values ​​associated with multiple lens positioning of the lens. The imaging system may select a specified distance from a plurality of predetermined distances based on whether the predicted direction of movement matches the second predicted direction of movement (e.g., as discussed with respect to decision 1015 of the focusing convergence engine 1005, and / or using a decision similar to decision 820, but using the CDAF and FV statistical engine 905). In some examples, the imaging system (or at least one subsystem thereof) is configured to also select a specified distance from a plurality of predetermined distances based on a confidence level associated with the second predicted direction of movement (e.g., as discussed with respect to decision 1015 of the focusing convergence engine 1005).

[0158] In some examples, moving the lens in the predicted movement direction (as in operation 1510) includes moving the lens from a first lens position to a second lens position in the predicted movement direction. In some examples, multiple lens positions include at least a first lens position.

[0159] In some examples, the imaging system (or at least one subsystem thereof) is configured to determine the predicted movement distance based on the magnitude of the phase shift associated with the focus data. In some examples, moving the lens in the predicted movement direction (as in operation 1510) includes moving the lens in the predicted movement direction by a predicted movement distance. For example, operation 825 could be an example of moving the lens in the predicted movement direction by a predicted movement distance.

[0160] In some examples, the imaging system (or at least one subsystem thereof) is configured to determine the predicted movement distance based on corresponding focus values ​​associated with multiple lens positioning of the lens. In some examples, moving the lens in the predicted movement direction (as in operation 1510) includes moving the lens in the predicted movement direction by the predicted movement distance. For example, in some examples, operations 830 and 835 (in...) Figure 10 The "yes" or "no" decision in 1015 can be an example of the lens moving in the predicted direction of movement to predict the distance traveled. In some examples, the predicted distance traveled can be predicted based on the trend of FV (e.g., similar to graph 500, a line or curve drawn using FV along lens positioning on one axis and another axis), based on the expected maximum FV value (e.g., based on other images captured using the same image sensor and / or similar image sensors), based on the characteristics of the image sensor or camera, or a combination thereof.

[0161] In some examples, the imaging system (or at least one subsystem thereof) is configured to determine a first focus value corresponding to image data and a second focus value corresponding to secondary image data captured by an image sensor after the lens moves in a predicted direction of movement. In some examples, the imaging system (or at least one subsystem thereof) is configured to update (e.g., as in update 655) a trained machine learning model based on the predicted direction of movement and the difference between the first and second focus values. For example, refer to... Figure 6The predicted direction of movement and the difference between the first and second focus values ​​can be examples of feedback 650 for updating the ML model 625. For example, if the difference between the first and second focus values ​​indicates an increase in focus value after moving the lens in the predicted direction of movement, this can be interpreted as positive feedback. Based on positive feedback, updates to the trained ML model can strengthen and / or reinforce weights associated with the generation of the predicted direction of movement to encourage the trained ML model to generate similar predicted directions of movement given similar inputs (e.g., similar focus data and / or image data). On the other hand, if the difference between the first and second focus values ​​indicates a decrease in focus value after moving the lens in the predicted direction of movement, this can be interpreted as negative feedback. Based on negative feedback, updates to the trained ML model can weaken and / or remove weights associated with the generation of the predicted direction of movement to prevent the trained ML model from generating similar predicted directions of movement given similar inputs (e.g., similar focus data and / or image data).

[0162] In some examples, moving the lens in the predicted direction of movement includes actuating a motor and / or actuator (e.g., actuator 280). In some examples, the motor and / or actuator may refer to a voice coil motor (VCM), another type of motor, a linear actuator, another type of actuator, or a combination thereof.

[0163] In some examples, the imaging system (or at least one subsystem thereof) is configured to and can capture an image using an image sensor after the lens has moved in the predicted direction of movement (as in operation 1510). In some aspects, the imaging system (or at least one subsystem thereof) is configured to and can output an image (e.g., using output device 1635 and / or communication interface 1640). In some aspects, the imaging system (or at least one subsystem thereof) is configured to and can display the image using a display (e.g., which is part of output device 1635) (e.g., causing an image to be displayed). In some aspects, the imaging system (or at least one subsystem thereof) is configured to and can cause the image to be transmitted (e.g., causing an image to be sent) to a receiving device using a communication interface and / or a communication transceiver (e.g., which is part of output device 1635 and / or communication interface 1640).

[0164] In some examples, the process described in this article (e.g., Figure 1 Figure 2, Figure 3, Figure 4 Figure 5 , Figure 6 , Figure 7 , Figure 8 , Figure 9 , Figure 10 , Figure 11 , Figure 12 The corresponding processes in Figures 13 and 14 Figure 15 The process 1500 and / or other processes described herein may be performed by a computing device or apparatus. In some examples, the processes described herein may be performed by an image capture and processing system 100, an image capture device 105A, an image processing device 105B, an image processor 150, an ISP 154, a host processor 152, a PDAF camera system 200, a lens 210, focusing photodiodes 225A-225B, an actuator 280, an image sensor 285, a pixel array 300, a pixel array 330, a pixel array 340, a pixel array 350, a pixel array 360, a focusing pixel 400, a focusing pixel 440, an ML system 600, an ML prediction engine 620, an ML model 625, an imaging system 700, a PDAF engine 705, and a focusing convergence system. Engine 710, focusing and converging engine 805, imaging system 900, CDAF and FV statistical engine 905, focusing and converging engine 910, focusing and converging engine 1005, focusing and converging engine 1105, neural network 1200, head-mounted display (HMD) 1310, mobile phone 1410, imaging system with execution process 1500, computing system 1600, processor 1610, device, system, non-transitory computer-readable medium coupled to processor or a combination thereof.

[0165] Computing devices may include any suitable device, such as mobile devices (e.g., mobile phones), desktop computing devices, tablet computing devices, wearable devices (e.g., VR headsets, AR headsets, AR glasses, network-connected watches or smartwatches or other wearable devices), server computers, autonomous vehicles or computing devices of autonomous vehicles, robotic devices, televisions, and / or any other computing device with the resource capability to perform the processes described herein. In some cases, a computing device or apparatus may include various components, such as one or more input devices, one or more output devices, one or more processors, one or more microprocessors, one or more microcomputers, one or more cameras, one or more sensors, and / or other components configured to perform the steps of the processes described herein. In some examples, a computing device may include a display, a network interface configured to communicate and / or receive data, any combination thereof, and / or other components. The network interface may be configured to communicate and / or receive Internet Protocol (IP)-based data or other types of data.

[0166] A component capable of implementing a computing device in a circuit. For example, the component may include electronic circuitry or other electronic hardware, and / or may be implemented using electronic circuitry or other electronic hardware, which may include one or more programmable electronic circuits (e.g., a microprocessor, graphics processing unit (GPU), digital signal processor (DSP), central processing unit (CPU), and / or other suitable electronic circuitry), and / or may include computer software, firmware, or any combination thereof for performing the various operations described herein, and / or may be implemented using computer software, firmware, or any combination thereof for performing the various operations described herein.

[0167] The processes described herein are illustrated as logic flow diagrams, block diagrams, or conceptual diagrams, whose operations represent sequences of operations that can be implemented by hardware, computer instructions, or combinations thereof. In the context of computer instructions, each operation represents a computer-executable instruction stored on one or more computer-readable storage media that, when executed by one or more processors, performs the described operation. Generally, computer-executable instructions include routines, programs, objects, components, data structures, etc., that perform a particular function or implement a particular data type. The order in which the operations are described is not intended to be construed as limiting, and any number of the described operations can be combined in any order and / or in parallel to implement the process.

[0168] Additionally, the processes described herein can be executed under the control of one or more computer systems configured with executable instructions, and can be implemented as code (e.g., executable instructions, one or more computer programs, or one or more applications) that executes jointly on one or more processors, implemented in hardware, or a combination of the foregoing. As noted above, the code can be stored on a computer-readable or machine-readable storage medium, for example, in the form of a computer program comprising multiple instructions executable by one or more processors. The computer-readable or machine-readable storage medium can be non-transitory.

[0169] Figure 16 This is a diagram illustrating an example of a system used to implement certain aspects of this technology. Specifically, Figure 16 An example of a computing system 1600 is illustrated. This computing system can be any computing device, such as an internal computing system, a remote computing system, a camera, or any component thereof, wherein the components of the system communicate with each other using connection 1605. Connection 1605 can be a physical connection using a bus, or a direct connection to processor 1610, such as in a chipset architecture. Connection 1605 can also be a virtual connection, a networking connection, or a logical connection.

[0170] In some aspects, computing system 1600 is a distributed system in which the functions described herein can be distributed across a data center, multiple data centers, a peer-to-peer network, etc. In some aspects, one or more of the described system components represent a plurality of such components, each of which performs some or all of the functions described for that component. In some aspects, the components can be physical or virtual devices.

[0171] Example system 1600 includes at least one processing unit (CPU or processor) 1610 and a connection 1605 that couples various system components, including system memories 1615 such as read-only memory (ROM) 1620 and random access memory (RAM) 1625, to processor 1610. Computing system 1600 may include a cache 1612 of high-speed memory that is directly connected to, closely proximate to, or integrated into processor 1610.

[0172] Processor 1610 may include any general-purpose processor and hardware or software services (such as services 1632, 1634, and 1636 stored in storage device 1630 and configured to control processor 1610), as well as dedicated processors in which software instructions are incorporated into the actual processor design. Processor 1610 may be a substantially completely independent computing system containing multiple cores or processors, buses, memory controllers, caches, etc. Multi-core processors may be symmetric or asymmetric.

[0173] To enable user interaction, the computing system 1600 includes an input device 1645 that can represent any number of input mechanisms, such as a microphone for voice, a touch-sensitive screen for gesture or graphical input, a keyboard, a mouse, motion input, voice input, etc. The computing system 1600 may also include an output device 1635 that can be one or more of a plurality of output mechanisms. In some instances, a multi-mode system allows the user to provide multiple types of input / output to communicate with the computing system 1600. The computing system 1600 may include a communication interface 1640, which typically governs and manages user input and system output. The communication interface can perform or facilitate the receipt and / or transmission of wired or wireless communications using wired and / or wireless transceivers, including utilizing audio jacks / plugs, microphone jacks / plugs, Universal Serial Bus (USB) ports / plugs, Apple... ® Lightning ® Ports / plugs, Ethernet ports / plugs, fiber optic ports / plugs, dedicated wired ports / plugs, Bluetooth ® Wireless signal transmission, Bluetooth ® Low-power (BLE) wireless signal transmission, IBEACON ®The communication interface 1640 may include one or more Global Navigation Satellite System (GNSS) receivers or transceivers for determining the location of the computing system 1600 based on one or more signals received from one or more satellites associated with one or more GNSS systems. This includes wireless signal transmission, radio frequency identification (RFID) wireless signal transmission, near field communication (NFC) wireless signal transmission, dedicated short range communication (DSRC) wireless signal transmission, 1602.11 Wi-Fi wireless signal transmission, wireless local area network (WLAN) signal transmission, visible light communication (VLC) wireless signal transmission, Global Microwave Access Interoperability (WiMAX) wireless signal transmission, infrared (IR) wireless signal transmission, public switched telephone network (PSTN) signal transmission, integrated services digital network (ISDN) signal transmission, 3G / 4G / 5G / LTE cellular data network wireless signal transmission, ad hoc network signal transmission, radio wave signal transmission, microwave signal transmission, infrared signal transmission, visible light signal transmission, ultraviolet light signal transmission, wireless signal transmission along the electromagnetic spectrum, or some combination thereof. GNSS systems include, but are not limited to, the U.S. Global Positioning System (GPS), Russia's Global Navigation Satellite System (GLONASS), China's BeiDou Navigation Satellite System (BDS), and Europe's Galileo GNSS. There are no limitations on operation on any particular hardware configuration, and therefore the underlying features here can be easily replaced to obtain improved hardware or firmware configurations as they are developed.

[0174] Storage device 1630 may be a non-volatile and / or non-transitory and / or computer-readable storage device, and may be a hard disk or other type of computer-readable medium capable of storing data accessible by a computer, such as magnetic tape, flash memory cards, solid-state storage devices, digital multifunction disks, cartridges, floppy disks, hard disks, magnetic tapes, magnetic stripes, any other magnetic storage media, flash memory, memristor memory, any other solid-state storage, CD-ROM discs, rewritable CD discs, DVD discs, Blu-ray discs (BDD discs), holographic discs, another optical medium, secure digital (SD) cards, microSD cards, Memory Sticks. ®Cards, smart card chips, EMV chips, Subscriber Identity Module (SIM) cards, mini / micro / nano / micro SIM cards, another integrated circuit (IC) chip / card, random access memory (RAM), static RAM (SRAM), dynamic RAM (DRAM), read-only memory (ROM), programmable read-only memory (PROM), erasable programmable read-only memory (EPROM), electrically erasable programmable read-only memory (EEPROM), flash EPROM, cache memory (L1 / L2 / L3 / L4 / L5 / L#), resistive random access memory (RRAM / ReRAM), phase change memory (PCM), spin-transfer torque RAM (STT-RAM), another memory chip or cassette and / or combinations thereof.

[0175] Storage device 1630 may include software services, servers, etc., which enable the system to perform functions when the code defining such software is executed by processor 1610. In some aspects, hardware services that perform specific functions may include software components for performing functions stored in computer-readable media connected to necessary hardware components such as processor 1610, connection 1605, output device 1635, etc.

[0176] As used herein, the term "computer-readable medium" includes, but is not limited to, portable or non-portable storage devices, optical storage devices, and various other media capable of storing, containing, or carrying instructions and / or data. Computer-readable media may include non-transitory media in which data can be stored and which do not include carrier waves and / or transient electronic signals propagating wirelessly or over a wired connection. Examples of non-transitory media include, but are not limited to, magnetic disks or magnetic tapes, optical storage media (such as compact optical discs (CDs) or digital versatile discs (DVDs)), flash memory, memory, or memory devices. Computer-readable media may store code and / or machine-executable instructions thereon, which may represent procedures, functions, subroutines, programs, routines, subroutines, modules, software packages, classes, or any combination of instructions, data structures, or program statements. A code segment may be coupled to another code segment or hardware circuitry by passing and / or receiving information, data, arguments, parameters, or memory contents. Information, arguments, parameters, data, etc., may be passed, forwarded, or transmitted using any suitable means, including memory sharing, messaging, token passing, network transmission, etc.

[0177] In some respects, computer-readable storage devices, media, and memories may include cables or wireless signals containing bit streams, etc. However, when referred to, non-transitory computer-readable storage media explicitly exclude media such as energy, carrier signals, electromagnetic waves, and the signals themselves.

[0178] Specific details are provided in the foregoing description to provide a thorough understanding of the aspects and examples presented herein. However, those skilled in the art will understand that these aspects can be practiced without these specific details. For clarity, in some instances, the technology may be presented as comprising individual functional blocks, including functional blocks containing devices, device components, steps or routines in methods embodied in software or a combination of hardware and software. Additional components may be used in addition to those shown in the figures and / or described herein. For example, circuits, systems, networks, processes and other components may be shown as components in block diagram form to avoid obscuring these aspects in unnecessary detail. In other instances, well-known circuits, processes, algorithms, structures and techniques may be shown without unnecessary detail to avoid obscuring the aspects.

[0179] Various aspects described above can be presented as processes or methods, depicted as flowcharts, diagrams, data flow graphs, structure diagrams, or block diagrams. Although a flowchart can describe operations as a sequential process, many operations within an operation can be executed in parallel or concurrently. Furthermore, the order of operations can be rearranged. A process terminates when its operations are completed, but a process may have additional steps not included in the accompanying diagrams. A process can correspond to a method, function, procedure, subroutine, subroutine, etc. When a process corresponds to a function, the termination of the process may correspond to the function returning to the calling function or the main function.

[0180] The processes and methods described in the examples above can be implemented using stored computer-executable instructions or computer-executable instructions otherwise obtainable from a computer-readable medium. Such instructions may include, for example, instructions and data that configure, or otherwise configure, a general-purpose computer, special-purpose computer, or processing device to perform a function or group of functions. The portion may be accessible via a network of the computer resources used. The computer-executable instructions may be, for example, binary, intermediate format instructions, such as assembly language, firmware, source code, etc. Examples of computer-readable media that can be used to store instructions, information used, and / or information created during the methods according to the described examples include disks or optical discs, flash memory, USB devices with non-volatile memory, networked storage devices, etc.

[0181] Devices implementing the processes and methods according to these disclosures may include hardware, software, firmware, middleware, microcode, hardware description languages, or any combination thereof, and may take any of a variety of form factors. When implemented as software, firmware, middleware, or microcode, program code or code segments (e.g., computer program products) for performing necessary tasks may be stored in a computer-readable or machine-readable medium. A processor may perform the necessary tasks. Typical examples of form factors include laptop computers, smartphones, mobile phones, tablet devices, or other small form factor personal computers, personal digital assistants, rack-mounted devices, standalone devices, etc. The functionality described herein may also be embodied in peripheral devices or interlocking cards. By further example, such functionality may also be implemented on circuit boards of different chips or different processes executed on a single device.

[0182] Instructions, media for delivering such instructions, computing resources for executing them, and other structures for supporting such computing resources are example components for providing the functionality described in this disclosure.

[0183] In the foregoing description, aspects of this application have been described with reference to their specific aspects, but those skilled in the art will recognize that this application is not limited thereto. Therefore, although illustrative aspects of this application have been described in detail herein, it is to be understood that the inventive concepts may be embodied and employed in various other ways, and the appended claims are not intended to be construed as including these variations unless limited by prior art. The various features and aspects of the applications described above may be used individually or in combination. Furthermore, aspects may be utilized in any number of environments and applications beyond those described herein without departing from the broader spirit and scope of this specification. Therefore, the specification and drawings should be considered illustrative rather than restrictive. For illustrative purposes, the methods are described in a particular order. It should be understood that, in alternative aspects, the methods may be performed in a different order than described.

[0184] Those skilled in the art will understand that, without departing from the scope of this description, the less than (“<”) and greater than (“>”) symbols or terms used herein may be replaced with less than or equal to (“>”) respectively. ") and greater than or equal to (" The symbol ) is used instead.

[0185] When a component is described as being “configured” to perform certain operations, such configuration can be achieved, for example, by designing electronic circuits or other hardware to perform the operations, by programming programmable electronic circuits (e.g., microprocessors or other suitable electronic circuits) to perform the operations, or any combination thereof.

[0186] The phrase “coupled to” means any component that is physically connected directly or indirectly to another component, and / or any component that communicates directly or indirectly with another component (e.g., connected to another component via a wired or wireless connection and / or other suitable communication interface).

[0187] The claim language or other language that states "at least one of" and / or "one or more of" in a set indicates that one member of the set or multiple members of the set (in any combination) satisfy the claim. For example, the claim language stating "at least one of A and B" means A, B, or A and B. In another example, the claim language stating "at least one of A, B, and C" means A, B, C, or A and B, or A and C, or B and C, or A and B and C. The language "at least one of" and / or "one or more of" in a set does not limit the set to items listed in the set. For example, the claim language stating "at least one of A and B" may mean A, B, or A and B, and may additionally include items not listed in the set of A and B.

[0188] The various exemplary logic blocks, modules, circuits, and algorithm steps described in conjunction with the aspects disclosed herein can be implemented as electronic hardware, computer software, firmware, or combinations thereof. To clearly illustrate this interchangeability between hardware and software, various exemplary components, blocks, modules, circuits, and steps have been broadly described above in terms of their functionality. Whether such functionality is implemented as hardware or software depends on the specific application and the design constraints imposed on the overall system. Those skilled in the art may implement the described functionality in different ways for each specific application, but such implementation decisions should not be construed as departing from the scope of this application.

[0189] The techniques described herein can also be implemented in electronic hardware, computer software, firmware, or any combination thereof. Such techniques can be implemented in any of a variety of devices, such as general-purpose computers, wireless communication devices (mobile phones), or integrated circuit devices with multiple uses, including applications in wireless communication devices (mobile phones) and other devices. Any feature described as a module or component can be implemented together in an integrated logic device or separately as discrete but interoperable logic devices. If implemented in software, these techniques can be implemented at least in part by a computer-readable data storage medium comprising program code including instructions that, when executed, perform one or more of the methods described above. The computer-readable data storage medium can form part of a computer program product, which may include packaging materials. The computer-readable medium may include memory or data storage media, such as random access memory (RAM) (such as synchronous dynamic random access memory (SDRAM)), read-only memory (ROM), non-volatile random access memory (NVRAM), electrically erasable programmable read-only memory (EEPROM), flash memory, magnetic or optical data storage media, etc. Additionally or alternatively, the technology may be implemented at least in part by a computer-readable communication medium that carries or conveys program code in the form of instructions or data structures that can be accessed, read and / or executed by a computer, such as propagated signals or waves.

[0190] The program code can be executed by a processor, which may include one or more processors, such as one or more digital signal processors (DSPs), general-purpose microprocessors, application-specific integrated circuits (ASICs), field-programmable arrays (FPGAs), or other equivalent integrated or discrete logic circuits. Such processors can be configured to perform any of the techniques described in this disclosure. A general-purpose processor may be a microprocessor; however, in alternatives, the processor may be any conventional processor, controller, microcontroller, or state machine. The processor may also be implemented as a combination of computing devices, such as a combination of a DSP and a microprocessor, multiple microprocessors, one or more microprocessors combined with a DSP core, or any other such configuration. Therefore, as used herein, the term "processor" may refer to any of the foregoing structures, any combination of the foregoing structures, or any other structure or means suitable for implementing the techniques described herein. Furthermore, in some aspects, the functionality described herein may be provided within dedicated software or hardware modules configured for encoding and decoding, or incorporated into a combined video encoder-decoder (CODEC).

[0191] The exemplary aspects of this disclosure include:

[0192] Aspect 1. An apparatus for imaging, the apparatus comprising: at least one memory; and at least one processor coupled to the at least one memory and configured to: process at least focus data using a trained machine learning model to identify a predicted movement direction for moving a lens to improve image focus, wherein the focus data is associated with at least one focused pixel of image data captured using an image sensor; and move the lens in the predicted movement direction to improve image focus.

[0193] Aspect 2. The apparatus according to aspect 1, wherein the at least one processor is configured to process at least the focus data and the image data using the trained machine learning model, wherein the image data includes at least one image pixel.

[0194] Aspect 3. The apparatus according to any one of Aspects 1 to 2, wherein the at least one processor is configured to move the lens a specified distance in the predicted movement direction to move the lens in the predicted movement direction.

[0195] Aspect 4. The apparatus according to aspect 3, wherein the at least one processor is configured to select the specified distance from a plurality of predetermined distances based on a confidence level associated with the predicted direction of movement.

[0196] Aspect 5. The apparatus according to any one of Aspects 3 to 4, wherein the at least one processor is configured to: determine a second predicted movement direction based on the sign of the phase shift associated with the focus data; and select the specified distance from a plurality of predetermined distances based on whether the predicted movement direction matches the second predicted movement direction.

[0197] Aspect 6. The apparatus according to aspect 5, wherein the at least one processor is configured to: further select the specified distance from the plurality of predetermined distances based on a confidence level associated with the second predicted direction of movement.

[0198] Aspect 7. The apparatus according to any one of Aspects 3 to 6, wherein the at least one processor is configured to: determine a second predicted movement direction based on corresponding focus values ​​associated with a plurality of lens positions of the lens; and select the specified distance from a plurality of predetermined distances based on whether the predicted movement direction matches the second predicted movement direction.

[0199] Aspect 8. The apparatus according to aspect 7, wherein the at least one processor is configured to: further select the specified distance from the plurality of predetermined distances based on a confidence level associated with the second predicted direction of movement.

[0200] Aspect 9. The apparatus according to any one of Aspects 7 to 8, wherein the at least one processor is configured to: move the lens from a first lens position to a second lens position in the predicted movement direction to move the lens in the predicted movement direction, wherein the plurality of lens positions includes at least the first lens position.

[0201] Aspect 10. The apparatus according to any one of Aspects 1 to 9, wherein the at least one processor is configured to: determine a predicted movement distance based on an amount of phase shift associated with the focusing data; and move the lens by the predicted movement distance in the predicted movement direction to move the lens in the predicted movement direction.

[0202] Aspect 11. The apparatus according to any one of Aspects 1 to 10, wherein the at least one processor is configured to: determine a predicted movement distance based on corresponding focus values ​​associated with a plurality of lens positions of the lens; and move the lens by the predicted movement distance in the predicted movement direction to move the lens in the predicted movement direction.

[0203] Aspect 12. The apparatus according to any one of Aspects 1 to 11, wherein the at least one processor is configured to: determine a first focus value corresponding to the image data; determine a second focus value corresponding to secondary image data captured by the image sensor after the lens moves in the predicted movement direction; and update the trained machine learning model based on the predicted movement direction and the difference between the first focus value and the second focus value.

[0204] Aspect 13. The apparatus according to any one of Aspects 1 to 12, wherein the at least one processor is configured to actuate a motor to move the lens in the predicted movement direction.

[0205] Aspect 14. The apparatus according to any one of Aspects 1 to 13, wherein the at least one processor is configured to: after the lens has moved in the predicted movement direction, use the image sensor to capture an image.

[0206] Aspect 15. The apparatus according to aspect 14, wherein the at least one processor is configured to output the image.

[0207] Aspect 16. The apparatus according to any one of Aspects 14 to 15, wherein the at least one processor is configured to display the image using a display.

[0208] Aspect 17. The apparatus according to any one of Aspects 14 to 16, wherein the at least one processor is configured to: transmit the image to a receiving device using a communication interface.

[0209] Aspect 18. The apparatus according to any one of Aspects 1 to 17, wherein the apparatus comprises at least one of a head-mounted display (HMD), a mobile phone, or a wireless communication device.

[0210] Aspect 19. A method for imaging, the method comprising: processing at least focus data using a trained machine learning model to identify a predicted movement direction for moving a lens to improve image focus, wherein the focus data is associated with at least one focus pixel of image data captured using an image sensor; and moving the lens in the predicted movement direction to improve image focus.

[0211] Aspect 20. The method according to aspect 19, wherein using the trained machine learning model to process at least the focus data includes using the trained machine learning model to process at least the focus data and the image data, wherein the image data includes at least one image pixel.

[0212] Aspect 21. The method according to any one of Aspects 19 to 20, wherein moving the lens in the predicted movement direction includes moving the lens in the predicted movement direction by a specified distance.

[0213] Aspect 22. The method according to aspect 21, the method further comprising: selecting the specified distance from a plurality of predetermined distances based on a confidence level associated with the predicted direction of movement.

[0214] Aspect 23. The method according to any one of Aspects 21 to 22, the method further comprising: determining a second predicted movement direction based on the sign of a phase shift associated with the focused data; and selecting the specified distance from a plurality of predetermined distances based on whether the predicted movement direction matches the second predicted movement direction.

[0215] Aspect 24. The method according to aspect 23, the method further comprising: selecting the specified distance from the plurality of predetermined distances based on a confidence level associated with the second predicted direction of movement.

[0216] Aspect 25. The method according to any one of Aspects 21 to 24, the method further comprising: determining a second predicted movement direction based on corresponding focus values ​​associated with a plurality of lens positions of the lens; and selecting the specified distance from a plurality of predetermined distances based on whether the predicted movement direction matches the second predicted movement direction.

[0217] Aspect 26. The method according to aspect 25, the method further comprising: selecting the specified distance from the plurality of predetermined distances based on a confidence level associated with the second predicted direction of movement.

[0218] Aspect 27. The method according to any one of Aspects 25 to 26, wherein moving the lens in the predicted movement direction includes moving the lens in the predicted movement direction from a first lens position to a second lens position, wherein the plurality of lens positions includes at least the first lens position.

[0219] Aspect 28. The method according to any one of Aspects 19 to 27, the method further comprising: determining a predicted movement distance based on a magnitude of phase shift associated with the focusing data, wherein moving the lens in the predicted movement direction includes moving the lens in the predicted movement direction by the predicted movement distance.

[0220] Aspect 29. The method according to any one of Aspects 19 to 28, the method further comprising: determining a predicted movement distance based on corresponding focus values ​​associated with a plurality of lens positions of the lens, wherein moving the lens in the predicted movement direction includes moving the lens in the predicted movement direction by the predicted movement distance.

[0221] Aspect 30. The method according to any one of Aspects 19 to 29, the method further comprising: determining a first focus value corresponding to the image data; determining a second focus value corresponding to secondary image data captured by the image sensor after the lens has moved in the predicted movement direction; and updating the trained machine learning model based on the predicted movement direction and the difference between the first focus value and the second focus value.

[0222] Aspect 31. The method according to any one of Aspects 19 to 30, wherein moving the lens in the predicted movement direction includes an actuating motor.

[0223] Aspect 32. The method according to any one of aspects 19 to 31, the method further comprising: after the lens is moved in the predicted movement direction, using the image sensor to capture an image.

[0224] Aspect 33. The method according to aspect 32, the method further comprising: outputting the image.

[0225] Aspect 34. The method according to any one of aspects 32 to 33, the method further comprising: using a display to display the image.

[0226] Aspect 35. The method according to any one of Aspects 32 to 34, the method further comprising: using a communication interface to transmit the image to a receiving device.

[0227] Aspect 36. A non-transitory computer-readable medium having instructions stored thereon, which, when executed by one or more processors, cause the one or more processors to perform any one of aspects 1 to 35.

[0228] Aspect 37. An apparatus for sensor data processing, the apparatus comprising one or more components for performing operations according to any one of aspects 1 to 35.

Claims

1. An apparatus for imaging, the apparatus comprising: At least one memory; and At least one processor, the at least one processor being coupled to the at least one memory and being configured to: A trained machine learning model is used to process at least the focus data to identify a predicted direction of movement for moving the lens to improve image focus, wherein the focus data is associated with at least one focused pixel of image data captured using an image sensor. as well as The lens is moved in the predicted direction of movement to improve the image focus.

2. The apparatus of claim 1, wherein the at least one processor is configured to: The trained machine learning model is used to process at least the focus data and the image data, wherein the image data includes at least one image pixel.

3. The apparatus of claim 1, wherein the at least one processor is configured to: The lens is moved a specified distance in the predicted direction of movement to move the lens in the predicted direction of movement.

4. The apparatus of claim 3, wherein the at least one processor is configured to: The specified distance is selected from a plurality of predetermined distances based on the confidence level associated with the predicted direction of movement.

5. The apparatus of claim 3, wherein the at least one processor is configured to: The second predicted movement direction is determined based on the sign of the phase shift associated with the focused data; and The specified distance is selected from a plurality of predetermined distances based on whether the predicted movement direction matches the second predicted movement direction.

6. The apparatus of claim 5, wherein the at least one processor is configured to: The specified distance is also selected from the plurality of predetermined distances based on the confidence level associated with the second predicted direction of movement.

7. The apparatus of claim 3, wherein the at least one processor is configured to: A second predicted direction of movement is determined based on corresponding focus values ​​associated with the multiple lens positions of the lens; and The specified distance is selected from a plurality of predetermined distances based on whether the predicted movement direction matches the second predicted movement direction.

8. The apparatus of claim 7, wherein the at least one processor is configured to: The specified distance is also selected from the plurality of predetermined distances based on the confidence level associated with the second predicted direction of movement.

9. The apparatus of claim 7, wherein the at least one processor is configured to: The lens is moved from a first lens position to a second lens position in the predicted movement direction to move the lens in the predicted movement direction, wherein the plurality of lens positions includes at least the first lens position.

10. The apparatus of claim 1, wherein the at least one processor is configured to: The predicted movement distance is determined based on the magnitude of the phase shift associated with the focused data; and The lens is moved by the predicted movement distance in the predicted movement direction to move the lens in the predicted movement direction.

11. The apparatus of claim 1, wherein the at least one processor is configured to: The predicted movement distance is determined based on the corresponding focus values ​​associated with the multiple lens positions of the lens; and The lens is moved by the predicted movement distance in the predicted movement direction to move the lens in the predicted movement direction.

12. The apparatus of claim 1, wherein the at least one processor is configured to: Determine a first focus value corresponding to the image data; Determine a second focus value corresponding to the secondary image data captured by the image sensor after the lens has moved in the predicted movement direction; and The trained machine learning model is updated based on the predicted direction of movement and the difference between the first focus value and the second focus value.

13. The apparatus of claim 1, wherein the at least one processor is configured to: An actuated motor is used to move the lens in the predicted direction of movement.

14. The apparatus of claim 1, wherein the at least one processor is configured to: After the lens moves in the predicted direction of movement, the image sensor is used to capture an image.

15. The apparatus of claim 14, wherein the at least one processor is configured to: Output the image.

16. The apparatus of claim 14, wherein the at least one processor is configured to: Use a monitor to display the image.

17. The apparatus of claim 14, wherein the at least one processor is configured to: The image is sent to the receiving device using a communication interface.

18. The apparatus of claim 1, wherein the apparatus comprises at least one of a head-mounted display (HMD), a mobile phone, or a wireless communication device.

19. A method for imaging, the method comprising: A trained machine learning model is used to process at least the focus data to identify a predicted direction of movement for moving the lens to improve image focus, wherein the focus data is associated with at least one focused pixel of image data captured using an image sensor. as well as The lens is moved in the predicted direction of movement to improve the image focus.

20. The method of claim 19, wherein using the trained machine learning model to process at least the focus data comprises using the trained machine learning model to process at least the focus data and the image data, wherein the image data comprises at least one image pixel.

21. The method of claim 19, wherein moving the lens in the predicted movement direction comprises moving the lens a specified distance in the predicted movement direction.

22. The method according to claim 21, further comprising: The specified distance is selected from a plurality of predetermined distances based on the confidence level associated with the predicted direction of movement.

23. The method according to claim 21, further comprising: The second predicted movement direction is determined based on the sign of the phase shift associated with the focused data; as well as The specified distance is selected from a plurality of predetermined distances based on whether the predicted movement direction matches the second predicted movement direction.

24. The method according to claim 23, further comprising: The specified distance is also selected from the plurality of predetermined distances based on the confidence level associated with the second predicted direction of movement.

25. The method according to claim 21, further comprising: The second predicted direction of movement is determined based on the corresponding focus values ​​associated with the multiple lens positions of the lens; as well as The specified distance is selected from a plurality of predetermined distances based on whether the predicted movement direction matches the second predicted movement direction.

26. The method according to claim 25, further comprising: The specified distance is also selected from the plurality of predetermined distances based on the confidence level associated with the second predicted direction of movement.

27. The method of claim 25, wherein moving the lens in the predicted movement direction comprises moving the lens in the predicted movement direction from a first lens position to a second lens position, wherein the plurality of lens positions includes at least the first lens position.

28. The method according to claim 19, further comprising: The predicted movement distance is determined based on the magnitude of the phase shift associated with the focusing data, wherein moving the lens in the predicted movement direction includes moving the lens in the predicted movement direction by the predicted movement distance.

29. The method according to claim 19, further comprising: The predicted movement distance is determined based on corresponding focus values ​​associated with multiple lens positions of the lens, wherein moving the lens in the predicted movement direction includes moving the lens in the predicted movement direction by the predicted movement distance.

30. The method according to claim 19, further comprising: Determine a first focus value corresponding to the image data; Determine a second focus value corresponding to the secondary image data captured by the image sensor after the lens moves in the predicted movement direction; as well as The trained machine learning model is updated based on the predicted direction of movement and the difference between the first focus value and the second focus value.