Methods, apparatus, media and products for image processing

By introducing machine learning algorithms into the autofocus algorithm and using phase detection information to determine the focal length, the problem of inaccurate focusing in low light and complex scenes is solved, and better image capture results are achieved.

CN119137968BActive Publication Date: 2025-10-28QUALCOMM INC
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202380040109.5
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Priority Date
2022-05-19
Filing Date
2023-03-29
Publication Date
2025-10-28
Estimated Expiration
2043-03-29

AI Technical Summary

Technical Problem

Existing autofocus algorithms struggle to accurately determine phase shift in low-light conditions and complex scenarios, leading to inaccurate focusing by image capture devices, which affects image quality and user experience.

Method used

Machine learning algorithms, particularly phase detection-based machine learning (MLPD), are employed to determine the focal length of a scene using phase shift information from an image sensor, and the movement of the lens is controlled by machine learning algorithms to improve focusing.

Benefits of technology

It improves focusing accuracy and stability in low-light and complex scenes, reduces image blur, and enhances the user experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119137968B_ABST
    Figure CN119137968B_ABST
Patent Text Reader

Abstract

The present disclosure provides systems, methods, and devices for capturing images using an autofocus (AF) algorithm. In a first aspect, a method for autofocus includes: receiving first image data of a first scene from a first image sensor of a first camera, the first image data including phase shift information; determining a first focal length for the first scene based on a machine learning algorithm by inputting the phase shift information into the machine learning algorithm; and controlling a focus position of the first camera based on the first focal length. Other aspects and features are also claimed and described.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This disclosure relates generally to image processing, and more specifically to an autofocus system for an image capture device. Some features enable and provide improved image processing, including the use of machine learning in determining the focal length of a scene. Background Technology

[0002] An image capture device is a device capable of capturing one or more digital images (whether still images for photographs or sequences of images for video). Capture devices can be integrated into a variety of devices. For example, an image capture device may include a standalone digital camera or digital video camera, a wireless communication device with a camera (such as a mobile phone, cellular, or satellite radio phone), a personal digital assistant (PDA), a panel or tablet device, a gaming device, a computer device (such as a webcam or video surveillance camera), or other devices with digital imaging or video capabilities.

[0003] Autofocus (AF) algorithms improve the user experience by allowing an image capture device to keep its focus on the item of interest. When the AF algorithm fails to properly focus the image capture device, the image captured by the device appears blurry. Users viewing a blurry image perceive the blur as reduced image quality. Photos or videos captured by an image capture device when it is not properly focused result in a degraded user experience, especially when the scene is changing rapidly and the image capture device cannot be reconfigured for a second attempt to capture the scene. Summary of the Invention

[0004] The following outlines some aspects of this disclosure to provide a basic understanding of the techniques discussed. This outline is not an exhaustive summary of all intended features of this disclosure, and is neither intended to identify key or essential elements of all aspects of this disclosure, nor to depict the scope of any or all aspects of this disclosure. The sole purpose of this outline is to present some concepts of one or more aspects of this disclosure in a generalized form as a prelude to the more detailed description that follows.

[0005] One example autofocus (AF) algorithm is phase detection autofocus (PDAF), which operates using phase detection information obtained from the image sensor of the image capture device. With PDAF, the phase shift between the left and right phase images generated by the left and right phase detectors represents the focus level. When the AF algorithm controls the lens coupled to the image sensor used to acquire image data, the sign of the phase shift determines the direction of lens movement, and the magnitude of the phase shift determines the amount of lens movement. If the lens is in focus, a small or zero phase shift can be determined when the left and right phase images are the same. If the lens is in defocus, a large phase shift can be determined.

[0006] The PDAF operation used to determine phase shift in an AF algorithm can be based on machine learning to provide better performance in specific scene conditions where the AF algorithm struggles to determine the phase shift. For example, in low-light conditions, machine learning for phase detection (MLPD) can be used to determine the phase shift, and the machine learning-based phase shift value can be used to control the lens of an image capture device. In other examples, repeating patterns or moiré patterns can be detected, and MLPD can be used to determine the phase shift used to control the lens. MLPD can be applied to other or all scene conditions as part of an AF algorithm executed on an image capture device. In some implementations, MLPD can be selectively used as part of the AF algorithm by examining criteria to determine if there are specific scene conditions that benefit from machine learning. In some implementations, MLPD can be configured with different training weights based on the scene condition to provide improved phase shift determination using machine learning algorithms such as neural networks.

[0007] In one aspect of the invention, a method for image processing (such as in an image capture device) includes: receiving first image data of a first scene from a first image sensor of a first camera, the first image data including phase shift information; determining a first focal length for the first scene by inputting the phase shift information into the machine learning algorithm based on a machine learning algorithm; and controlling the focus position of the first camera based on the first focal length.

[0008] In an additional aspect of this disclosure, an apparatus includes at least one processor and a memory coupled to the at least one processor. The at least one processor is configured to perform operations including: receiving first image data of a first scene from a first image sensor of a first camera, the first image data including phase shift information; determining a first focal length for the first scene by inputting the phase shift information into the machine learning algorithm based on a machine learning algorithm; and controlling the focus position of the first camera based on the first focal length.

[0009] In a further aspect of the invention, an apparatus includes: means for receiving first image data of a first scene from a first image sensor of a first camera, the first image data including phase shift information; means for determining a first focal length for the first scene by inputting the phase shift information into a machine learning algorithm based on the machine learning algorithm; and means for controlling the focus position of the first camera based on the first focal length.

[0010] In an additional aspect of this disclosure, a non-transitory computer-readable medium stores instructions that, when executed by a processor, cause the processor to perform operations. The operations include: receiving first image data of a first scene from a first image sensor of a first camera, the first image data including phase shift information; determining a first focal length for the first scene by inputting the phase shift information into the machine learning algorithm based on a machine learning algorithm; and controlling the focus position of the first camera based on the first focal length.

[0011] Image capture devices (devices capable of capturing one or more digital images, whether still images, photographs, or video sequences) can be integrated into a variety of devices. For example, image capture devices may include standalone digital cameras or digital video cameras, wireless communication devices equipped with cameras (such as mobile phones, cellular, or satellite radio phones), personal digital assistants (PDAs), panel or tablet devices, gaming devices, computer devices (such as webcams, video surveillance cameras), or other devices with digital imaging or video capabilities.

[0012] Generally, this disclosure relates to image processing techniques for digital cameras having an image sensor and an image signal processor (ISP). The ISP can be configured to control the capture of image frames from one or more image sensors and process one or more image frames from said one or more image sensors to generate a view of a scene in a corrected image frame. The corrected image frame may be part of a sequence of image frames forming a video sequence. The video sequence may include other image frames received from the image sensor or other image sensors and / or other corrected image frames based on input from the image sensor or another image sensor. In some embodiments, processing of one or more image frames may be performed within the image sensor, such as in a binning module. The image processing techniques described in the embodiments disclosed herein may be performed by circuitry, such as a binning module, in the image sensor, in the image signal processor (ISP), in the application processor (AP), or in a combination of two or more of these components.

[0013] In one example, an image signal processor (ISP) may receive instructions to capture a sequence of image frames in response to the loading of software (such as a camera application) to generate a preview display from an image capture device. The ISP may be configured to generate a single output frame stream based on image frames received from one or more image sensors. The single output frame stream may include raw image data from the image sensors, merged image data from the image sensors, or corrected image frames processed by one or more algorithms within the ISP, such as in a merging module. For example, image frames obtained from image sensors (which may have undergone some processing before being output to the ISP) may be processed within the ISP by an image post-processing engine (IPE) and / or other image processing circuitry for performing one or more of tone mapping, portrait lighting, contrast enhancement, gamma correction, etc.

[0014] After the image signal processor (ISPM) determines an output frame representing the scene using image correction (such as merging as described in the various embodiments herein), the output frame can be displayed on a device display as a single still image and / or as part of a video sequence, saved to a storage device as a picture or video sequence, transmitted over a network, and / or printed to an output medium. For example, the ISPM can be configured to acquire input frames of image data (e.g., pixel values) from different image sensors and subsequently generate corresponding output frames of image data (e.g., preview display frames, still image captures, frames for video, frames for object tracking, etc.). In other examples, the ISPM can output frames of image data to various output devices and / or camera modules for further processing, such as for 3A parameter synchronization (e.g., autofocus (AF), auto white balance (AWB), and auto exposure control (AEC)), to generate video files via the output frames, to configure frames for display, to configure frames for storage, to transmit frames via a network connection, etc. In other words, an image signal processor can obtain incoming frames from one or more image sensors, each coupled to one or more camera lenses, and can then generate an output frame stream and output the output frame stream to various output destinations.

[0015] In some aspects, corrected image frames can be generated by combining various aspects of the image correction disclosed herein with other computational photography techniques such as high dynamic range (HDR) photography or multi-frame noise reduction (MFNR). In the case of HDR photography, the first and second image frames are captured using different exposure times, different apertures, different lenses, and / or other characteristics that can result in improved dynamic range of the fused image when combining the two image frames. In some aspects, the method can be performed for MFNR photography, wherein the first and second image frames are captured using the same or different exposure times, and the first and second image frames are fused to generate a corrected first image frame that has reduced noise compared to the captured first image frame.

[0016] In some aspects, the device may include an image signal processor or processor (e.g., an application processor) that includes specific functionalities for camera control and / or processing, such as enabling or disabling the merging module or otherwise controlling aspects of image correction. The methods and techniques described herein may be performed entirely by the image signal processor or processor, or various operations may be separated between the image signal processor and the processor, and in some aspects across additional processors.

[0017] The device may include one, two, or more image sensors, such as a first image sensor. When multiple image sensors are present, the first image sensor may have a larger field of view (FOV) than the second image sensor, or the first image sensor may have a different sensitivity or a different dynamic range than the second image sensor. In one example, the first image sensor may be a wide-angle image sensor, and the second image sensor may be a long-range image sensor. In another example, the first sensor is configured to acquire an image through a first lens having a first optical axis, and the second sensor is configured to acquire an image through a second lens having a second optical axis different from the first optical axis. Additionally or alternatively, the first lens may have a first magnification, and the second lens may have a second magnification different from the first magnification. This configuration may occur in a lens cluster on a mobile device, such as where multiple image sensors and associated lenses are located at offset positions on the front or rear of the mobile device. Additional image sensors with larger, smaller, or the same field of view may be included. The image correction techniques described herein can be applied to image frames captured from any of the image sensors in a multi-sensor device.

[0018] In an additional aspect of this disclosure, an apparatus configured for image processing and / or image capture is disclosed. The apparatus includes components for capturing image frames. The apparatus also includes one or more components for capturing data representing a scene, such as image sensors (including charge-coupled device (CCD), Bayer filter sensors, infrared (IR) detectors, ultraviolet (UV) detectors, complementary metal-oxide-semiconductor (CMOS) sensors), and time-of-flight detectors. The apparatus may further include components for focusing and / or directing light onto one or more image sensors (including simple lenses, compound lenses, spherical lenses, and aspherical lenses). These components can be controlled to capture a first image frame and / or a second image frame input to the image processing techniques described herein.

[0019] Other aspects, features, and specific embodiments will become apparent to those skilled in the art when they review the following description of particular exemplary aspects in conjunction with the accompanying drawings. Although features may be discussed hereinafter with reference to certain aspects and drawings, each aspect may include one or more of the advantageous features discussed herein. In other words, while one or more aspects may be discussed having certain advantageous features, one or more such features may also be used depending on the aspect. Similarly, although exemplary aspects may be discussed below as aspects of an apparatus, system, or method, exemplary aspects can be implemented in various apparatuses, systems, and methods.

[0020] This method can be embedded as computer program code in a computer-readable medium, the computer program code including instructions that cause a processor to perform the steps of the method. In some embodiments, the processor may be part of a mobile device including: a first network adapter configured to transmit data, such as images or videos as recorded data or streaming data, via a first network connection among a plurality of network connections; and a processor coupled to the first network adapter and memory. The processor enables the corrected image frames described herein to be transmitted via a wireless communication network, such as a 5G NR communication network.

[0021] The features and technical advantages of the examples according to this disclosure have been summarized quite extensively above in order to provide a better understanding of the specific embodiments described below. Additional features and advantages will be described below. The disclosed concepts and specific examples can be readily used as the basis for modifying or designing other structures for achieving the same purpose as this disclosure. Such equivalent constructions do not depart from the scope of protection of the appended claims. The characteristics of the concepts disclosed herein, in both their organization and manner of operation, and the associated advantages, will be better understood by considering the following description in conjunction with the accompanying drawings. Each drawing is provided for illustrative and descriptive purposes and not as a limitation of the claims.

[0022] While aspects and implementations are described herein by way of example, those skilled in the art will understand that additional implementations and use cases may arise in many other arrangements and scenarios. The innovations described herein can be implemented across many different platform types, devices, systems, shapes, sizes, and package arrangements. For example, aspects and / or uses may arise via integrated chip implementations and other devices based on non-modular components (e.g., end-user devices, vehicles, communication devices, computing devices, industrial equipment, retail / purchasing devices, medical devices, AI-enabled devices, etc.). While some examples may or may not specifically relate to use cases or applications, the described innovations can have a wide range of applicability. The scope of implementations can range from chip-level or modular components to non-modular, non-chip-level implementations, and further to aggregated, distributed, or original equipment manufacturer (OEM) devices or systems incorporating one or more aspects of the described innovations. In some practical contexts, devices incorporating the described aspects and features may also necessarily include additional components and features for implementing and practicing the claimed and described aspects. For example, the transmission and reception of wireless signals necessarily involve multiple components for analog and digital purposes (e.g., hardware components including antennas, radio frequency (RF) chains, power amplifiers, modulators, buffers, processors, interleavers, adders / summers, etc.). The innovations described herein are intended to be implemented in a variety of devices, chip-level components, systems, distributed arrangements, end-user equipment, etc., with different sizes, shapes, and constructions. Attached Figure Description

[0023] A further understanding of the nature and advantages of this disclosure can be achieved by referring to the following figures. In the figures, similar components or features may have the same reference numerals. Furthermore, various components of the same type can be distinguished by adding a dash after the reference numerals and a second reference numeral for differentiation between similar components. If only the first reference numeral is used in the specification, the description applies to any one of the similar components having the same first reference numeral, regardless of the second reference numeral.

[0024] Figure 1 A block diagram of an example device 100 for performing image capture from one or more image sensors is shown.

[0025] Figure 2 This is a block diagram illustrating an autofocus (AF) system using machine learning according to some embodiments of the present disclosure.

[0026] Figure 3 This is a block diagram illustrating the use of an autofocus system with multiple available machine learning configurations according to some embodiments of the present disclosure.

[0027] Figure 4 This is a block diagram illustrating a PDAF decision module for determining the application of machine learning in an autofocus system according to some embodiments of the present disclosure.

[0028] Figure 5 This is a flowchart illustrating a method for controlling the focal length of a camera using machine learning according to some embodiments of this disclosure.

[0029] Figure 6 This is a flowchart illustrating a method for determining a machine learning configuration for determining the focal length of a camera according to some embodiments of the present disclosure.

[0030] Figure 7 These are examples of methods for determining weights in a machine learning algorithm according to some embodiments of this disclosure.

[0031] Figure 8 This is a block diagram illustrating a feedback system for updating machine learning used in an autofocus system according to some embodiments of the present disclosure.

[0032] The same reference numerals and names in different figures denote the same elements. Detailed Implementation

[0033] The specific embodiments described below with reference to the accompanying drawings are intended as a description of various configurations and are not intended to limit the scope of this disclosure. Rather, the specific embodiments include specific details for providing a thorough understanding of the subject matter of the invention. It will be apparent to those skilled in the art that these specific details are not necessary in every case, and in some instances, well-known structures and components are shown in block diagram form for clarity of presentation.

[0034] This disclosure provides systems, apparatus, methods, and computer-readable media that support image capture operations. Autofocus (AF) performed by an image capture device can use machine learning algorithms to determine information for assisting in controlling the image capture device to improve the focus of photographs or videos captured by the image capture device. For example, the machine learning algorithm can use at least a portion of image frames captured by the image capture device to determine a representation of the focal length of the scene (such as a phase shift value). Lenses can be controlled based on the focal length representation to change the focus of the image capture device in order to improve the focus of additional captured image frames.

[0035] Specific embodiments of the subject matter described in this disclosure can be implemented to achieve one or more of the following potential advantages or benefits. In some aspects, this disclosure provides techniques for improving autofocus (AF) by providing better focus information for obtaining a better-focused image frame from an image sensor. In some aspects, this disclosure provides techniques for improving autofocus (AF) by providing faster focusing for obtaining an image frame from an image sensor to reduce the latency experienced by the user of an image capture device. In low-light scenes, machine learning-based AF can provide more reliable and stable results than those obtainable without machine learning. Furthermore, reliable results can be obtained in challenging scene conditions such as moiré patterns and other repeating patterns, as well as in low-light conditions.

[0036] Example devices for capturing image frames using one or more image sensors, such as smartphones, may include a configuration of two, three, four, or more cameras on the rear (e.g., the side opposite the user's display) or front (e.g., the same side as the user's display) of the device. Devices with multiple image sensors include one or more image signal processors (ISPs), computer vision processors (CVPs) (e.g., AI engines), or other suitable circuitry for processing the images captured by the image sensors. One or more image signal processors may provide the processed image frames to a memory and / or processor (such as an application processor, image front-end (IFE), image processing engine (IPE), or other suitable processing circuitry) for further processing, such as encoding, storage, transmission, or other manipulation.

[0037] As used herein, an image sensor can refer to the image sensor itself and any specific other components coupled to the image sensor for generating image frames for processing by an image signal processor or other logic circuitry, or for storage in memory (whether short-term buffers or long-term non-volatile memory). For example, an image sensor can include other components of a camera, including a shutter, buffers, or other readout circuitry for accessing the individual pixels of the image sensor. An image sensor can also refer to an analog front-end or other circuitry for converting analog signals into a digital representation of an image frame, which is provided to digital circuitry coupled to the image sensor.

[0038] In the following description, numerous specific details (e.g., examples of specific components, circuits, and processes) are set forth to provide a thorough understanding of this disclosure. As used herein, the term "coupled" means a direct connection or a connection via one or more intermediate components or circuits. Additionally, specific terminology is set forth in the following description and for illustrative purposes to provide a thorough understanding of this disclosure. However, it will be understood by those skilled in the art that implementing the teachings disclosed herein may not require these specific details. In other instances, known circuits and devices are illustrated in block diagram form to avoid obscuring the teachings of this disclosure.

[0039] Certain portions of the following detailed description are presented using other symbolic representations of procedures, logic blocks, processes, and data bit operations within computer memory. In this disclosure, programs, logic blocks, procedures, etc., are conceived as self-consistent sequences of steps or instructions that lead to desired results. These steps are steps that require physical operations on physical quantities. Although not strictly necessary, these physical quantities typically take the form of electrical or magnetic signals that can be stored, transferred, combined, compared, and otherwise manipulated within a computer system.

[0040] In the accompanying drawings, a single block can be described as performing one or more functions. The one or more functions performed by this block can be performed in a single component or across multiple components, and / or can be performed using hardware, software, or a combination of hardware and software. To clearly illustrate this interchangeability between hardware and software, various exemplary components, blocks, modules, circuits, and steps are described below in general terms of their functionality. Whether such functionality is implemented as hardware or software depends on the specific application and the design constraints imposed on the system as a whole. Those skilled in the art can implement the described functionality in different ways for each specific application, but such specific implementation decisions should not be construed as departing from the scope of this disclosure. Moreover, the example device may include components other than those shown, including well-known components such as processors, memory, etc.

[0041] The aspects of this disclosure are applicable to any electronic device that includes or is coupled to two or more image sensors capable of capturing image frames (or “frames”). Furthermore, the aspects of this disclosure can be implemented in image sensors or devices coupled to image sensors having the same or different capabilities and characteristics (e.g., resolution, shutter speed, sensor type, etc.). Additionally, the aspects of this disclosure can be implemented in devices for processing image frames, whether or not the device includes or is coupled to image sensors, such as processing devices capable of retrieving stored images for processing, including processing devices existing in cloud computing systems.

[0042] Unless otherwise specifically stated, it will be apparent from the following discussion that, throughout this application, the use of terms such as “access,” “receive,” “transmit,” “use,” “select,” “determine,” “normalize,” “multiply,” “average,” “monitor,” “compare,” “apply,” “update,” “measure,” “derive,” “set,” “generate,” etc., refers to the actions and processes of a computer system or similar electronic computing device that manipulate and transform data represented as physical (electronic) quantities in the registers and memories of the computer system into other data similarly represented as physical quantities in the registers, memories, or other such information storage, transmission, or display devices of the computer system.

[0043] The terms "device" and "apparatus" are not limited to one or a specific number of physical objects (e.g., a smartphone, a camera controller, a processing system, etc.). As used herein, a device can be any electronic device having one or more portions that implement at least some parts of this disclosure. Although the term "device" is used in the following description and examples to describe various aspects of this disclosure, the term "device" is not limited to a particular configuration, type, or number of objects. As used herein, an apparatus can include a device or part of a device for performing the described operations.

[0044] Figure 1A block diagram of an example device 100 for performing image capture from one or more image sensors is shown. Device 100 may include or be otherwise coupled to an image signal processor 112 for processing image frames from one or more image sensors, such as a first image sensor 101, a second image sensor 102, and a depth sensor 140. In some specific embodiments, device 100 also includes or is coupled to a processor 104 and a memory 106 for storing instructions 108. Device 100 may also include or be coupled to a display 114 and an input / output (I / O) component 116. I / O component 116 may be used for user interaction, such as a touchscreen interface and / or physical buttons. I / O component 116 may also include a network interface for communicating with other devices, including a wide area network (WAN) adapter 152, a local area network (LAN) adapter 153, and / or a personal area network (PAN) adapter 154. An example WAN adapter is a 4G LTE or 5G NR wireless network adapter. An example LAN adapter 153 is an IEEE 802.11 WiFi wireless network adapter. Example PAN adapter 154 is a Bluetooth wireless network adapter. Each of adapters 152, 153, and / or 154 may be coupled to an antenna including multiple antennas configured for primary and diversity reception and / or configured to receive a specific frequency band. Device 100 may also include or be coupled to a power supply 118 for device 100, such as a battery or a component coupling device 100 to an energy source. Device 100 may also include or be coupled to... Figure 1 Additional features or components not shown. In one example, a wireless interface that may include multiple transceivers and a baseband processor may be coupled to or included in the WAN adapter 152 for use in a wireless communication device. In yet another example, an analog front end (AFE) for converting analog image frame data into digital image frame data may be coupled between image sensors 101 and 102 and image signal processor 112.

[0045] The device may include or be coupled to a sensor hub 150, which interfaces with a sensor to receive data about the movement of the device 100, data about the environment surrounding the device 100, and / or other non-camera sensor data. One example non-camera sensor is a gyroscope, a device configured to measure rotation, orientation, and / or angular velocity to generate motion data. Another example non-camera sensor is an accelerometer, a device configured to measure acceleration, which can also be used to determine velocity and distance traveled by appropriately integrating the measured acceleration, and one or more of acceleration, velocity, and / or distance may be included in the generated motion data. In some aspects, the gyroscope in an electronic image stabilization system (EIS) may be coupled to the sensor hub or directly to the image signal processor 112. In another example, the non-camera sensor may be a Global Positioning System (GPS) receiver.

[0046] Image signal processor 112 can receive image data, such as data used to form image frames. In one embodiment, a local bus connection couples image signal processor 112 to image sensors 101 and 102 of a first camera and a second camera, respectively. In another embodiment, a wired interface couples image signal processor 112 to an external image sensor. In yet another embodiment, a wireless interface couples image signal processor 112 to image sensors 101 and 102.

[0047] The first camera may include a first image sensor 101 and a corresponding first lens 131. The second camera may include a second image sensor 102 and a corresponding second lens 132. Each of lenses 131 and 132 may be controlled by an associated autofocus (AF) algorithm 133 executed in the ISP 112, which adjusts lenses 131 and 132 to focus on a specific focal plane at a certain scene depth from image sensors 101 and 102. The AF algorithm 133 may be assisted by a depth sensor 140.

[0048] First image sensor 101 and second image sensor 102 are configured to capture one or more image frames. Lenses 131 and 132 focus light onto image sensors 101 and 102 respectively via one or more apertures for receiving light, one or more shutters for blocking light outside the exposure window, one or more color filter arrays (CFAs) for filtering light outside a specific frequency range, one or more analog front ends for converting analog measurements into digital information, and / or other suitable components for imaging. First lens 131 and second lens 132 may have different fields of view to capture different representations of the scene. For example, first lens 131 may be an ultra-wide (UW) lens, and second lens 132 may be a wide (W) lens. Multiple image sensors may include combinations of ultra-wide (high field of view (FOV)) sensors, wide sensors, long-range sensors, and ultra-long-range (low FOV) sensors. That is, each image sensor can be configured by hardware configuration and / or software settings to obtain different but overlapping fields of view. In one configuration, the image sensors are configured with different lenses with different magnifications, resulting in different fields of view. Sensors can be configured such that the UW sensor has a larger FOV than the W sensor, the W sensor has a larger FOV than the T sensor, and the T sensor has a larger FOV than the UT sensor. For example, a sensor configured for a wide FOV can capture a field of view in the range of 64 to 84 degrees, a sensor configured for an ultra-side FOV can capture a field of view in the range of 100 to 140 degrees, a sensor configured for a long-range FOV can capture a field of view in the range of 10 to 30 degrees, and a sensor configured for an ultra-long-range FOV can capture a field of view in the range of 1 to 8 degrees.

[0049] Image signal processor 112 processes image frames captured by image sensors 101 and 102. Although Figure 1 The illustrated device 100 includes two image sensors 101 and 102 coupled to the image signal processor 112, but any number (e.g., one, two, three, four, five, six, etc.) of image sensors can be coupled to the image signal processor 112. In some aspects, a depth sensor such as depth sensor 140 can be coupled to the image signal processor 112 and the output from the depth sensor can be processed in a similar manner to that of image sensors 101 and 102. Furthermore, any number of additional image sensors or image signal processors can be present for the device 100.

[0050] In some embodiments, the image signal processor 112 may execute instructions from memory, such as instructions 108 from memory 106, instructions stored in a separate memory coupled to or included in the image signal processor 112, or instructions provided by processor 104. Additionally or alternatively, the image signal processor 112 may include specific hardware (such as one or more integrated circuits (ICs)) configured to perform one or more operations described in this disclosure. For example, the image signal processor 112 may include one or more image front-ends (IFEs) 135, one or more image post-processing engines (IPEs) 136, and / or one or more automatic exposure compensation (AEC) engines 134. AF 133, AEC 134, IFE 135, and IPE 136 may each include dedicated circuitry embodied as software code executed by the ISP 112 and / or a combination of hardware within the ISP 112 and software code executed on the ISP 112.

[0051] In some embodiments, memory 106 may include a non-transitory or non-transitory computer-readable medium storing computer-executable instructions 108 to perform all or part of one or more of the operations described in this disclosure. In some embodiments, instructions 108 include a camera application (or other suitable application) to be executed by device 100 for generating images or videos. Instructions 108 may also include other applications or programs executed by device 100, such as an operating system and specific applications other than those for image or video generation. A camera application, such as one executed by processor 104, may cause device 100 to generate images using image sensors 101 and 102 and image signal processor 112. Memory 106 may also be accessed by image signal processor 112 to store processed frames, or may be accessed by processor 104 to obtain processed frames. In some embodiments, device 100 does not include memory 106. For example, device 100 may be circuitry including image signal processor 112, and the memory may be external to device 100. Device 100 may be coupled to external memory and configured to access that memory to write output frames for display or long-term storage. In some implementations, device 100 is a system-on-a-chip (SoC) that integrates an image signal processor 112, a processor 104, a sensor hub 150, a memory 106, and an input / output component 116 into a single package.

[0052] In some embodiments, at least one of the image signal processor 112 or processor 104 executes instructions to perform various operations described herein, including noise reduction operations. For example, execution of instructions may instruct the image signal processor 112 to begin or end capturing image frames or sequences of image frames, wherein the capture includes noise reduction as described in the embodiments herein. In some embodiments, processor 104 may include one or more general-purpose processor cores 104A capable of executing scripts or instructions (such as instructions 108 stored in memory 106) of one or more software programs. For example, processor 104 may include one or more application processors configured to execute a camera application (or other suitable application for generating images or videos) stored in memory 106.

[0053] When executing a camera application, processor 104 can be configured to command image signal processor 112 to perform one or more operations with reference to image sensor 101 or 102. For example, the camera application may receive a command to start a video preview display, and upon receiving such a command, capture and process video comprising a sequence of image frames from one or more image sensors 101 or 102. Image correction, such as using cascaded IPE, can be applied to one or more image frames in the sequence. Executing instructions 108 by processor 104 outside the camera application can also cause device 100 to perform any number of functions or operations. In some embodiments, in addition to the ability to execute software to cause device 100 to perform multiple functions or operations (such as those described herein), processor 104 may also include an IC or other hardware (e.g., an artificial intelligence (AI) engine 124). In some other embodiments, device 100 does not include processor 104, such as when all the described functionalities are configured in image signal processor 112.

[0054] In some embodiments, display 114 may include one or more suitable displays or screens that allow the user to interact and / or present items (such as previews of image frames captured by image sensors 101 and 102) to the user. In some embodiments, display 114 is a touch-sensitive display. I / O component 116 may be or include any suitable mechanism, interface, or device to receive input (such as commands) from the user and provide output to the user via display 114. For example, I / O component 116 may include (but is not limited to) a graphical user interface (GUI), keyboard, mouse, microphone, speaker, squeezable bezel, one or more buttons (e.g., power button), slider, switch, etc.

[0055] Although shown coupled to each other via processor 104, components such as processor 104, memory 106, image signal processor 112, display 114, and I / O components 116 may be coupled to each other in various other arrangements, such as via one or more local buses, which are not shown for simplicity. While image signal processor 112 is illustrated as separate from processor 104, image signal processor 112 may be the core of processor 104, which is an application processor unit (APU) included in a system-on-a-chip (SoC), or otherwise included in processor 104. Although device 100 is referenced in the examples herein to perform aspects of this disclosure, some device components may not be included. Figure 1 The details are shown to prevent obscuring aspects of this disclosure. Additionally, other components, the number of components, or combinations of components may be included in suitable equipment for performing aspects of this disclosure. Therefore, this disclosure is not limited to the configuration of a particular device or component, including device 100.

[0056] The AF algorithm 133 of the image capture device 100 can be based on a reference. Figures 2 to 7 The described implementation scheme is configured in some or all of the following ways. Figure 2This is a block diagram illustrating an autofocus (AF) system using machine learning according to some embodiments of the present disclosure. A first image sensor 101 of camera 103 outputs image data 202, which may include phase shift information and / or a representation of the scene in the field of view of image sensor 101. The phase shift information may include sensor data from one or more phase shift sensors embedded in the first image sensor 101. The scene representation may include Bayer-formatted image data, although other image formats may be used to represent the scene. An ISP 112 is coupled to the first camera 103 to receive image data 202, such as by means of a bus that couples image sensor 101 to the ISP 112. Although functionality may be described as part of the ISP 112, other components may be configured to perform similar functionality in controlling image capture devices and in providing autofocus capability in the first camera 103. For example, a processor (such as processor 104) or other general-purpose or fixed-function logic circuitry may be configured to provide the autofocus functionality described herein. In some embodiments, the autofocus functionality may be performed by the ISP 112 in conjunction with other processing circuitry, such as an AI engine 124. For example, AI engine 124 may be embedded in processor 104 and coupled to ISP 112 via a bus. As another example, AI engine 124 may be a separate component in device 100 and coupled to ISP 112 via a bus. As another example, AI engine 124 may be embedded in ISP 112. In some embodiments, one or more of ISP 112, processor 104, and AI engine 124 may be co-located as a single integrated circuit (IC) on a shared substrate, sharing the shared substrate with other functionalities including sensor hub 150, memory 106, and / or I / O components 116.

[0057] ISP 112 outputs focus information to first camera 103 for controlling the focal length of first camera 103. Camera 103 can use the focus information to adjust the position of first lens 131. ISP 112 can be coupled via the same bus that transmits image data 202, via a different bus, or via other control signals. In some embodiments, the focus information may include the estimated focal length of the scene in the field of view of camera 103, determined by ISP 112. Camera 103 can control lens 131 based on the received focal length by directly setting the lens position to the focal length. In some embodiments, the focus information may include different data used by camera 103 to determine the position of lens 131 and adjust the position of lens 131 based on the determined position. For example, the focus information may be a focus shift value indicating the distance and / or direction for moving the position of lens 131. Camera 103 can determine a new absolute position value of lens 131 based on the current position value and the focus shift value, and adjust lens 131 based on the new absolute position value.

[0058] The autofocus (AF) algorithm 133 executed by ISP 112 may include several components, including a conventional phase detection autofocus (PDAF) module 212A, a PDAF decision module 212B, and a machine learning phase detection (MLPD) module 212C. The conventional PDAF module 212A may include one or more non-machine learning AF algorithms, such as a phase detection AF algorithm that estimates the amount by which one image arriving at the PD sensor is shifted from another image arriving at the PD sensor. In some embodiments, the conventional PDAF module 212A may include cross-correlation determination between images. The PDAF decision module 212B may determine how to determine the focus information sent to camera 103. For example, the PDAF decision module 212B may determine whether to use one or both of the conventional PDAF module 212A and the MLPD module 212C to determine the focus information. As another example, the PDAF decision module 212B may determine how to use the MLPD module 212C to adjust the focus information generated by the conventional PDAF module 212A and / or how to use the conventional PDAF module 212A to adjust the focus information generated by the MLPD module 212C. In some implementations, the AF algorithm 133 may include one or more of modules 212A, 212B, and 212C, and may include other modules for performing the autofocus function. For example, the AF algorithm 133 may perform autofocus using only machine learning in the MLPD module 212C for determining focus information 204, such that the AF algorithm 133 does not include modules 212A and 212B. As another example, the AF algorithm 133 may include tracking algorithms such as object tracking or (specifically) face tracking. Object tracking may cooperate with one or both of the conventional PDAF module 212A and MLPD module 212C for determining focus information.

[0059] Figure 3 The diagram shows a configuration of AF algorithm 133 with modules 212A, 212B and 212C. Figure 3This is a block diagram illustrating an autofocus system using machine learning with multiple available machine learning configurations according to some embodiments of the present disclosure. AF algorithm 133 may receive phase shift information corresponding to left and right images corresponding to a first scene from a phase detector (PD) buffer, and may additionally receive raw image data corresponding to a representation of the first scene. Conventional PDAF module 212A may determine first focus information, which is provided to PDAF decision module 212B. PDAF decision module 212B may determine whether to provide the first focus information as the output of the AF algorithm based on one or more criteria. For example, a criterion may specify that AF algorithm 133 outputs the first focus information when the confidence level associated with the first focus information exceeds a threshold confidence level. As another example, a single or additional criterion may specify that AF algorithm 133 outputs the first focus information when a specific scene condition is detected (e.g., when an object is detected, when a specific object is detected, or when a specific light level is detected).

[0060] MLPD module 212C can determine second focus information to replace or adjust the first focus information. PDAF decision module 212B can determine whether to provide second focus information as the output of the AF algorithm based on one or more criteria. For example, a criterion may specify that the output of AF algorithm 133 is the second focus information when a specific scene condition is detected (such as when a light level below a threshold level is detected or when the confidence level of PDAF module 212A is below a threshold level). In some implementations, MLPD module 212C uses filters from convolutional layers to determine the second focus information for noise reduction, which achieves better focus information, especially in noisy images captured in low-light scenes, although the second focus information may be available in all scenes.

[0061] The MLPD module 212C can be activated and deactivated by the PDAF to reduce power consumption during the execution of the AF algorithm 133. The MLPD module 212C can consume more power than the conventional PDAF module 212A, and therefore deactivating the MLPD module 212C reduces power consumption when the conventional PDAF module 212A provides sufficient focusing capability. This reduced power consumption is particularly advantageous when the image capture device executing the AF algorithm 133 operates based on battery power (e.g., via a mobile device). Furthermore, deactivating the MLPD module 212C reduces the processing resources consumed, making those resources available for other algorithms to the extent that processing resources are shared between the AF algorithm 133 and other algorithms. Additionally, the MLPD module 212C can add a delay to the determination of focus information, which can be eliminated when scene conditions allow the conventional PDAF module 212A to provide sufficient focus information.

[0062] MLPD module 212C may include MLPD preprocessing module 312 and MLPD decision module 314. MLPD preprocessing module 312 processes image data representing a scene to obtain input for a machine learning algorithm. For example, the machine learning algorithm may be configured to receive a fixed-size image as input for predictions using focus information. In some embodiments, the machine learning algorithm is a neural network trained with a specific image size, and therefore the image data may be preprocessed to obtain samples of the same image size used to train the neural network. Preprocessing may include cropping, resizing, padding, or other image processing of the image data to determine the input for the machine learning algorithm. In some embodiments, preprocessing may include determining a region of interest (ROI) in the scene corresponding to the trained size and cropping the image data to the ROI for input to the machine learning algorithm. For example, the ROI may be determined through object recognition and / or user input. For example, an image capture device may allow a user to specify the region of interest by touching a location and / or defining an area on a display corresponding to the desired focus. As another example, preprocessing may include determining the ROI in the scene, cropping the image data to the ROI, and then resizing and / or padding the cropped image data to match the trained size for input into a machine learning algorithm.

[0063] In some implementations, a single ROI can be applied to the MLPD module 212C to output multiple focusing results. For example, the region of interest, corresponding to a window surrounding a portion of the image data, can be divided into multiple N sub-windows, such as a 4×3 array of sub-windows. A weighted average can be determined from the N sub-windows and used to determine the focusing information of the ROI based on the N sets of statistics determined from the N sub-windows.

[0064] The MLPD decision module 314 can determine the configuration of the machine learning algorithm when determining focus information. This can improve focus results in specific scene conditions that are difficult for phase detection algorithms. For example, focusing on moiré patterns or repeating patterns in a scene can be improved using a machine learning algorithm specifically trained for image scenes that include such patterns.

[0065] Figure 4 An example implementation of the standard PDAF decision module 212B that applies the AF algorithm is shown in the figure. Figure 4This is a block diagram illustrating a PDAF decision module for determining the application of machine learning in an autofocus system according to some embodiments of the present disclosure. At block 402, the PDAF decision module 212B receives conventional PDAF focus information and determines whether the focus information has a confidence level higher than a threshold level. If so, at block 406, the conventional PDAF focal length is sent to the camera as a first focal length. For example, at block 406, the ISP 112 of the PDAF decision module 212B may send a command to the camera to focus at the focal length corresponding to the output of the conventional PDAF module 212A. If the confidence level is lower than the confidence threshold, at block 404, image data is provided to the MLPD module 212C and a second focal length is determined based on machine learning.

[0066] exist Figure 5 The document illustrates a method for performing autofocus on an image capture device according to an embodiment described herein. Figure 5 This is a flowchart illustrating a method for controlling the focal length of a camera using machine learning according to some embodiments of the present disclosure. Method 500 includes receiving first image data including phase shift information at block 502. The first image data can be received from a first image sensor of a first camera of an image capturing device. The first image sensor may include multiple phase detectors at different locations to record corresponding phase information pairs of left and right images. The phase detectors may be embedded in the image sensor configured to obtain image data including a representation of a scene. At block 504, a first focal length for a first scene is determined based on a machine learning algorithm. For example, an MLPD module 212C may be used to determine the first focal length as focus information. The determination at block 504 may include, for example, a PDAF determination module 212B determining whether to use the MLPD module 212C. At block 506, the focal length of the first camera is controlled using machine learning based on the first focal length determined at block 504.

[0067] Determining the first focal length for the first scene at frame 504 may include, for example: Figure 6 The machine learning configuration shown is based on the scenario conditions. Figure 6FIG. 0 is a flowchart illustrating a method of determining a machine learning configuration for determining a focal length of a camera according to some embodiments of the present disclosure. Method 600 may include determining at block 602 whether the scene includes a repeating pattern. If so, then at block 604, the machine learning is configured by loading weights trained from the repeating pattern. If not, then at block 606, method 600 includes determining whether the scene condition reflects normal usage. If so, then at block 608, the machine learning is configured by loading weights trained from normal usage scenes. If not, then at block 610, method 600 includes determining whether the scene condition reflects a multi-depth scene. If so, then at block 612, the machine learning is configured by loading weights trained from multi-depth low light scenes. In some embodiments, the weights of block 612 may be determined by a weighted average, such as described in reference to Figure 7 After loading the weights according to one of blocks 604, 608, 612, and 614, method 600 includes determining a focus shift and a confidence level using the machine learning in MLPD module 212C at block 616.

[0068] Figure 7 FIG. 6 is an illustration of determining weights for a machine learning algorithm according to some embodiments of the present disclosure. Image portion 702 may extend from coordinates (x0, y0) in one corner of example image portion 702 to coordinates (x4, y3) in the opposite corner, and the example image portion is divided into 12 sub-windows arranged in a 4×3 grid, each sub-window having its own determined focal length. Image portion 702 may represent a fixed-size image that the machine learning algorithm is configured to receive by training with training images of the same fixed-size image as image portion 702. Region of interest (ROI) 704 for focusing may be a smaller portion of image portion 702, such as a face portion of a larger image portion. ROI 704 may be a floating window that moves around within fixed-size image portion 702 corresponding to the movement of an object of interest within image portion 702. ROI 704 may be defined by coordinates (fx0, fy0) in one corner and coordinates (fx1, fy1) in the opposite corner, where x0 < fx0 < x4 and y0 < fy0 < y3 and x0 < fx1 < x4 and y0 < fy1 < y3.

[0069] The phase difference (PD) of ROI 704 may be determined by a weighted average of the depth of focus within each of the sub-windows within image portion 702 based on the amount of overlap of ROI 704 with each of the sub-windows. As an example of a portion 706, the overlap may be determined based on the following equation:

[0070]

[0071] The weights for other parts of ROI 704 can be determined similarly. Weights can be applied to each sub-window of image part 702 based on the following equation: i,j The MLPD output of each sub-window of the MLPD module 212C in the image section 702 i,j To determine the phase difference of ROI 704:

[0072]

[0073] The weighted approach allows the overlapping area of ​​ROI 704 and image portion 702 to have a significant impact on the determined phase difference, and allows non-overlapping areas to have zero weight, so that the final MLPD output can be controlled by ROI 704.

[0074] A trainable machine learning algorithm can be used to determine the focus information to identify the weight set through offline training. For example, a sample dataset including images and associated underlying facts can be used to determine the weights that can be configured at boxes 604, 608, 612, and 614. As another example, weights can be retrieved from a configuration file, and the configuration file can be retrieved from a remote server.

[0075] Feedback can be used during the operation of the image capture device to update the machine learning algorithm used for autofocus, such as Figure 8 As shown in the image. Figure 8 This is a block diagram illustrating a feedback system for updating machine learning used in an autofocus system according to some embodiments of the present disclosure. Method 800 includes operating MLPD module 212C to generate prediction 804. Prediction 804 can be determined during the capture of image data by an image capture device, such that the machine learning algorithm is updated in real time during the use of the image capture device. Alternatively, aspects of method 800 can be performed offline, such as when the image capture device is inserted. At block 802, phase detection autofocus (PDAF) is performed based on grid or other conditions to provide base facts. At block 806, the base facts of block 802 and the prediction of block 804 are compared, such as by calculating a loss for focus shift and associated confidence level, wherein the calculated loss indicates the difference between the base facts and the prediction. The calculation in block 806 is used to update the stored weights in block 808, wherein MLPD module 212C uses the updated weights to determine future focus information.

[0076] In one or more aspects, techniques for supporting image processing (such as in an image capture device) may include additional aspects, such as any single aspect or any combination of aspects described below or in conjunction with one or more other processes or devices described elsewhere herein. In a first aspect, supporting image processing may include means configured to perform operations including: receiving first image data of a first scene from a first image sensor of a first camera, the first image data including phase shift information; determining a first focal length for the first scene based on a machine learning algorithm by inputting the phase shift information into the machine learning algorithm; and controlling the focus position of the first camera based on the first focal length. Additionally, the means may perform or operate according to one or more aspects described below.

[0077] In some embodiments, the apparatus includes a wireless device, such as a UE, configured to communicate via a wireless network and equipped with one or more cameras. In some embodiments, the apparatus may include at least one processor and memory coupled to the processor. The processor may be configured to perform the operations described herein with reference to the apparatus. In some other embodiments, the apparatus may include a non-transitory computer-readable medium having program code recorded thereon, and the program code may be executable by a computer to cause the computer to perform the operations described herein with reference to the apparatus. In some embodiments, the apparatus may include one or more components configured to perform the operations described herein. In some embodiments, a method of wireless communication may include one or more of the operations described herein with reference to the apparatus.

[0078] In a second aspect, in conjunction with the first aspect, the first image data further includes a representation of the first scene, and the apparatus is configured to perform operations including: determining characteristics of the first scene based on the first image data; and configuring the machine learning algorithm based on the characteristics of the first scene.

[0079] In a third aspect, in conjunction with one or more of the first or second aspects, the machine learning algorithm is configured to load multiple weights based on the characteristics of the first scene, wherein the characteristics of the first scene include at least one of a repeating pattern scene, a multi-depth scene, or a low-light scene.

[0080] In a fourth aspect, in conjunction with one or more of the first to third aspects, the apparatus is further configured to perform operations including: determining a second focal length based on the phase shift information in the absence of the machine learning algorithm; and determining a confidence level associated with the second focal length.

[0081] In a fifth aspect, in conjunction with one or more of the first to fourth aspects, the focus position of the first camera is controlled based on the confidence level, such that when a first criterion of the confidence level is met, the focus position of the first camera is controlled based on the first focal length, and when a second criterion of the confidence level is met, the focus position of the first camera is controlled based on the second focal length.

[0082] In the sixth aspect, the first focal length is determined based on the first criterion that satisfies the confidence level, in conjunction with one or more of the first to fifth aspects, such that the first focal length is not determined when the first criterion is not met.

[0083] In a seventh aspect, in conjunction with one or more of the first to sixth aspects, receiving the first image data further includes receiving a representation of the first scene, wherein determining the first focal length based on the machine learning algorithm further includes inputting at least a portion of the representation of the first scene into the machine learning algorithm.

[0084] In an eighth aspect, in conjunction with one or more of the first to seventh aspects, the apparatus is further configured to perform operations including: processing the first image data to determine a window of the first image data based on a region of interest, wherein determining the first focal length based on the machine learning algorithm includes inputting the window of the first image data into the machine learning algorithm, wherein the size of the window corresponds to the input size of the machine learning algorithm.

[0085] In a ninth aspect, in conjunction with one or more of the first to eighth aspects, the apparatus is further configured to perform operations including: determining a plurality of sub-windows of the region of interest, wherein determining the first focal length for the first scene includes determining a plurality of first focal lengths based on the machine learning algorithm by inputting the plurality of sub-windows, and wherein controlling the focus position of the first camera is based on the plurality of first focal lengths.

[0086] In the tenth aspect, in combination with one or more of the first to ninth aspects, the focus position of the first camera is controlled based on a weighted average of the plurality of first focal lengths.

[0087] In an eleventh aspect, in conjunction with one or more of the first to tenth aspects, the apparatus is further configured to perform operations including: determining a second focal length based on the phase shift information in the absence of the machine learning algorithm; and updating the machine learning algorithm based on the first focal length and the second focal length.

[0088] In a twelfth aspect, in conjunction with one or more of the first to eleventh aspects, the device includes a first camera having the first image sensor and a lens, wherein the device is configured with an autofocus (AF) algorithm to control the position of the lens based on a determined focal length (such as the focal length determined by the machine learning algorithm).

[0089] Those skilled in the art will understand that information and signals can be represented using any of a variety of different techniques and skills. For example, data, instructions, commands, information, signals, bits, symbols, and chips that may be mentioned throughout the above description can be represented by voltage, current, electromagnetic waves, magnetic fields or magnetic particles, light fields or optical particles, or any combination thereof.

[0090] In this article, relative to Figures 1 to 8 The components, functional blocks, and modules described include processors, electronic devices, hardware devices, electronic components, logic circuits, memory, software code, firmware code, etc., or any combination thereof. Software should be interpreted broadly as instructions, instruction sets, code, code segments, program code, programs, subroutines, software modules, application programs, software applications, software packages, routines, subroutines, objects, executables, threads of execution, procedures, and / or functions, regardless of whether it is referred to as software, firmware, middleware, microcode, hardware description languages, or other terms. Furthermore, the features discussed herein may be implemented via dedicated processor circuitry, via executable instructions, or a combination thereof.

[0091] Those skilled in the art will further understand that the various exemplary logic blocks, modules, circuits, and algorithm steps described in conjunction with this disclosure can be implemented as electronic hardware, computer software, or a combination of both. To clearly illustrate this interchangeability between hardware and software, various exemplary components, blocks, modules, circuits, and steps have been described in general terms of their functionality. Whether such functionality is implemented as hardware or software depends on the specific application and the design constraints imposed on the system as a whole. Those skilled in the art can implement the described functionality in different ways for each specific application, but such specific implementation decisions should not be construed as departing from the scope of this disclosure. Those skilled in the art will also readily recognize that the order or combination of components, methods, or interactions described herein are merely examples, and that components, methods, or interactions of various aspects of this disclosure can be combined or performed in ways other than those illustrated and described herein.

[0092] The various exemplary logic components, logic blocks, modules, circuits, and algorithmic processes described in conjunction with the specific implementations disclosed herein can be implemented as electronic hardware, computer software, or a combination of both. The interchangeability of hardware and software has been broadly described in terms of functionality and illustrated in the aforementioned exemplary components, blocks, modules, circuits, and processes. Whether such functionality is implemented in hardware or software depends on the specific application and the design constraints imposed on the overall system.

[0093] Hardware and data processing means for implementing the various exemplary logic components, logic blocks, modules, and circuits described in connection with the aspects disclosed herein can be implemented or executed using general-purpose single-chip or multi-chip processors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs) or other programmable logic devices, discrete gate or transistor logic components, discrete hardware components, or any combination thereof designed to perform the functions described herein. A general-purpose processor can be a microprocessor, or any conventional processor, controller, microcontroller, or state machine. In some embodiments, the processor can be implemented as a combination of computing devices, such as a combination of a DSP and a microprocessor, multiple microprocessors, one or more microprocessors combined with a DSP core, or any other such configuration. In some embodiments, specific processes and methods can be performed by circuitry specific to a given function.

[0094] In one or more aspects, the described functionality may be implemented in hardware, digital electronic circuits, computer software, firmware, including the structures disclosed in this specification and their structural equivalents, or in any combination thereof. Specific implementations of the subject matter described in this specification may also be implemented as one or more computer programs, i.e., one or more modules of computer program instructions encoded on a computer storage medium for execution by a data processing apparatus or for controlling the operation of a data processing apparatus.

[0095] If implemented in software, the functionality can be stored as one or more instructions or code on or transmitted through a computer-readable medium. The processes of the methods or algorithms disclosed herein can be implemented in a processor-executable software module that can reside on a computer-readable medium. Computer-readable media include both computer storage media and communication media, including any medium that can be implemented to transfer a computer program from one location to another. Storage media can be any available medium that a computer can access. By way of example and not limitation, such computer-readable media may include random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), CD-ROM or other optical disc storage devices, magnetic disk storage devices or other magnetic storage devices, or any other medium capable of storing desired program code in the form of instructions or data structures and accessible to a computer. Furthermore, any connection may be appropriately referred to as a computer-readable medium. As used herein, disks and optical discs include compact optical discs (CDs), laser discs, optical discs, digital versatile optical discs (DVDs), floppy disks, and Blu-ray discs, wherein disks typically magnetically reproduce data, while optical discs optically reproduce data using lasers. Combinations of the above should also be included within the scope of computer-readable media. Additionally, the operation of a method or algorithm may be as a set of code and instructions or any combination of code and instructions, located on a machine-readable medium and a computer-readable medium that may be incorporated into a computer program product.

[0096] Various modifications to the specific embodiments described herein will be apparent to those skilled in the art, and the general principles defined herein may be applied to other specific embodiments without departing from the spirit or scope of this disclosure. Therefore, the claims are not intended to be limited to the specific embodiments shown herein, but are to be accorded the broadest scope consistent with this disclosure, the principles disclosed herein, and the novel features.

[0097] Additionally, it will be readily understood by those skilled in the art that the terms “upper” and “lower” are sometimes used to describe the figures and to indicate relative positions on a correctly oriented page corresponding to the orientation of the figures, and may not reflect the correct orientation of any device as implemented.

[0098] Certain features described in this specification in the context of a single embodiment may also be implemented in combination in a single embodiment. Conversely, various features described in the context of a single embodiment may also be implemented individually or in any suitable sub-combination in multiple embodiments. Furthermore, although features may be described above as operating in certain combinations and even originally claimed in this way, one or more features from the claimed combination may be removed from that combination in some cases, and the claimed combination may be for sub-combinations or variations thereof.

[0099] Similarly, although operations are depicted in a specific order in the accompanying drawings, this should not be construed as requiring such operations to be performed in the indicated specific order or sequential order, or to perform all illustrated operations to achieve the desired result. Furthermore, the drawings may schematically depict one or more example processes in the form of flowcharts. However, other operations not depicted may be incorporated into the schematically illustrated example processes. For example, one or more additional operations may be performed before, after, simultaneously with, or between any illustrated operation. In some cases, multitasking and parallel processing are advantageous. Moreover, the separation of the various system components in the embodiments described above should not be construed as requiring such separation in all embodiments, and it should be understood that the described program components and systems can generally be integrated together in a single software product or packaged into multiple software products. Additionally, some other embodiments also fall within the scope of the appended claims. In some cases, the actions recited in the claims may be performed in a different order and still achieve the desired result.

[0100] As used herein (including in the claims), the term “or” in a list of two or more items means that any one of the listed items may be used alone, or any combination of two or more listed items may be used. For example, if a composition is described as containing component A, B, or C, the composition may contain A alone; B alone; C alone; a combination of A and B; a combination of A and C; a combination of B and C; or a combination of A, B, and C. Furthermore, as used herein (including in the claims), “or” in a list of items beginning with “at least one” indicates a separate list, such that a list such as “at least one of A, B, or C” refers to A or B or C or AB or AC or BC or ABC (i.e., A and B and C) or any combination of any of these items. The term “substantially” is defined as largely but not necessarily exactly what is specified (and includes what is specified; for example, substantially 90 degrees includes 90 degrees, and substantially parallel includes parallel), as understood by one of ordinary skill in the art. In any disclosed specific implementation, the term “substantially” may be used in place of “[percentage]” for the specified content, where the percentage includes 0.1%, 1%, 5%, or 10%.

[0101] The prior description of this disclosure is provided to enable any person skilled in the art to make or use this disclosure. Various modifications to this disclosure will be apparent to those skilled in the art, and the general principles defined herein may be applied to other variations without departing from the spirit or scope of this disclosure. Therefore, this disclosure is not intended to be limited to the examples and designs described herein, but is to be accorded the widest scope consistent with the principles and novel features disclosed herein.

Claims

1. A method for image processing, the method comprising: First image data of a first scene is received from the first image sensor of the first camera, the first image data including phase shift information and a representation of the first scene; The characteristics of the first scene are determined based on the first image data, wherein the characteristics of the first scene include at least one of a repeating pattern scene, a multi-depth scene, or a low-light scene. Configure a machine learning algorithm based on the characteristics of the first scenario, wherein the configuration includes loading multiple weights based on the characteristics of the first scenario; The first focal length for the first scene is determined by inputting the phase shift information into the machine learning algorithm. as well as The focus position of the first camera is controlled based on the first focal length.

2. The method according to claim 1, further comprising: The second focal length is determined based on the phase shift information without the machine learning algorithm described above. as well as Determine the confidence level associated with the second focal length. The focus position of the first camera is controlled based on the confidence level, such that when a first criterion of the confidence level is met, the focus position of the first camera is controlled based on the first focal length, and when a second criterion of the confidence level is met, the focus position of the first camera is controlled based on the second focal length.

3. The method of claim 2, wherein the first focal length is determined based on the first criterion that satisfies the confidence level, such that the first focal length is not determined when the first criterion is not satisfied.

4. The method of claim 1, wherein receiving the first image data further includes receiving a representation of the first scene, wherein determining the first focal length based on the machine learning algorithm further includes inputting at least a portion of the representation of the first scene into the machine learning algorithm.

5. The method according to claim 4, further comprising: Processing the first image data to determine a window of the first image data based on a region of interest, wherein determining the first focal length based on the machine learning algorithm includes inputting the window of the first image data into the machine learning algorithm. The size of the window corresponds to the input size of the machine learning algorithm.

6. The method according to claim 5, further comprising: Multiple sub-windows are defined for the region of interest. Determining the first focal length for the first scenario includes determining multiple first focal lengths based on the machine learning algorithm by inputting the multiple sub-windows into the machine learning algorithm, and The focus position of the first camera is controlled based on the plurality of first focal lengths.

7. The method of claim 6, wherein controlling the focus position of the first camera is based on a weighted average of the plurality of first focal lengths.

8. The method according to claim 1, further comprising: The second focal length is determined based on the phase shift information without the machine learning algorithm described above. as well as The machine learning algorithm is updated based on the first focal length and the second focal length.

9. An apparatus for image processing, the apparatus comprising: Memory, the memory storing processor-readable code; and At least one processor coupled to the memory, the at least one processor being configured to execute processor-readable code to cause the at least one processor to perform operations including: First image data of a first scene is received from the first image sensor of the first camera, the first image data including phase shift information and a representation of the first scene; The characteristics of the first scene are determined based on the first image data, wherein the characteristics of the first scene include at least one of a repeating pattern scene, a multi-depth scene, or a low-light scene. Configure a machine learning algorithm based on the characteristics of the first scenario, wherein the configuration includes loading multiple weights based on the characteristics of the first scenario; The first focal length for the first scene is determined by inputting the phase shift information into the machine learning algorithm. as well as The focus position of the first camera is controlled based on the first focal length.

10. The apparatus of claim 9, wherein the at least one processor is configured to execute processor-readable code to cause the at least one processor to perform further operations including: Determining the second focal length based on the phase shift information without the aforementioned machine learning algorithm; and Determine the confidence level associated with the second focal length. The focus position of the first camera is controlled based on the confidence level, such that when a first criterion of the confidence level is met, the focus position of the first camera is controlled based on the first focal length, and when a second criterion of the confidence level is met, the focus position of the first camera is controlled based on the second focal length.

11. The apparatus of claim 10, wherein the first focal length is determined based on the first criterion satisfying the confidence level, such that the first focal length is not determined when the first criterion is not satisfied.

12. The apparatus of claim 9, wherein receiving the first image data further comprises receiving a representation of the first scene, wherein determining the first focal length based on the machine learning algorithm further comprises inputting at least a portion of the representation of the first scene into the machine learning algorithm.

13. The apparatus of claim 12, wherein the at least one processor is configured to execute processor-readable code to cause the at least one processor to perform further operations including: Processing the first image data to determine a window of the first image data based on a region of interest, wherein determining the first focal length based on the machine learning algorithm includes inputting the window of the first image data into the machine learning algorithm, wherein the size of the window corresponds to the input size of the machine learning algorithm.

14. The apparatus of claim 13, wherein the at least one processor is configured to execute processor-readable code to cause the at least one processor to perform further operations including: Multiple sub-windows are defined for the region of interest. Determining the first focal length for the first scenario includes determining multiple first focal lengths based on the machine learning algorithm by inputting the multiple sub-windows, and The focus position of the first camera is controlled based on the plurality of first focal lengths.

15. The apparatus of claim 14, wherein controlling the focus position of the first camera is based on a weighted average of the plurality of first focal lengths.

16. The apparatus of claim 9, wherein the at least one processor is configured to execute processor-readable code to cause the at least one processor to perform further operations including: Determining the second focal length based on the phase shift information without the aforementioned machine learning algorithm; and The machine learning algorithm is updated based on the first focal length and the second focal length.

17. A non-transitory computer-readable medium storing instructions, which, when executed by a processor, cause the processor to perform operations including: First image data of a first scene is received from the first image sensor of the first camera, the first image data including phase shift information and a representation of the first scene; The characteristics of the first scene are determined based on the first image data, wherein the characteristics of the first scene include at least one of a repeating pattern scene, a multi-depth scene, or a low-light scene. Configure a machine learning algorithm based on the characteristics of the first scenario, wherein the configuration includes loading multiple weights based on the characteristics of the first scenario; The first focal length for the first scene is determined by inputting the phase shift information into the machine learning algorithm. as well as The focus position of the first camera is controlled based on the first focal length.

18. The non-transitory computer-readable medium of claim 17, wherein the operation further comprises: The second focal length is determined based on the phase shift information without the machine learning algorithm described above. as well as Determine the confidence level associated with the second focal length. The focus position of the first camera is controlled based on the confidence level, such that when a first criterion of the confidence level is met, the focus position of the first camera is controlled based on the first focal length, and when a second criterion of the confidence level is met, the focus position of the first camera is controlled based on the second focal length.

19. The non-transitory computer-readable medium of claim 18, wherein the first focal length is determined based on the first criterion satisfying the confidence level, such that the first focal length is not determined when the first criterion is not satisfied.

20. A computer program product comprising computer-readable instructions, which, when executed by a processor, cause the processor to perform the method according to any one of claims 1-8.

Citation Information

Patent Citations

  • System and method for improved camera flash

    CN111164508A

  • Novel three-dimensional reconstruction method based on monocular unmanned aerial vehicle multi-frame RGB image

    CN114092640A