Multi-frame time processing

By analyzing statistical data of image frames through machine learning models, identifying and anchoring image frames and replacing pixels, the problem of poor subject poses or expressions in multi-frame images is solved, improving image quality, especially in group portraits and action portrait scenes.

CN122207262APending Publication Date: 2026-06-12QUALCOMM INC
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
QUALCOMM INC
Filing Date
2023-11-17
Publication Date
2026-06-12

AI Technical Summary

Technical Problem

Existing technologies struggle to effectively address issues such as poor poses or expressions of the subject across different image frames when processing multi-frame images, leading to a decline in image quality, especially in group portraits and action portraits.

Method used

By using machine learning models to analyze statistical data of image frames, anchor image frames are identified, and replacement pixels in other image frames are selected based on this. The anchor image frames are then modified to improve the subject's pose and expression. Multiple image frames are then merged to generate a better image representation.

Benefits of technology

It improves image noise, tone, and overall quality, especially in group portraits and action portrait scenes, ensuring that the subjects' facial expressions are more natural and clear.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122207262A_ABST
    Figure CN122207262A_ABST
Patent Text Reader

Abstract

The present disclosure provides systems, methods, and apparatuses for image signal processing that support multi-frame image processing, such as for group portraits. In one aspect, a method of image processing includes determining an anchor image frame from a plurality of image frames, replacing first pixels of a region of the anchor image frame with second pixels of a first replacement image frame of the plurality of image frames, and determining an output image frame by merging the anchor image frame with at least one additional image frame. Other aspects and features are also claimed and described.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This disclosure relates generally to image processing, and more specifically to multi-frame image processing. Several features enable and provide improved image processing, including improved portrait photographs involving multiple subjects. Background Technology

[0002] An image capture device is a device capable of capturing one or more digital images (whether still images for photographs or sequences of images for video). Capture devices can be integrated into a variety of devices. For example, an image capture device may include a standalone digital camera or digital video camera, a wireless communication device with a camera (such as a mobile phone, cellular, or satellite radio phone), a personal digital assistant (PDA), a panel or tablet device, a gaming device, a computing device (such as a webcam, video surveillance camera), or other devices with digital imaging or video capabilities. Summary of the Invention

[0003] The following summary outlines some aspects of this disclosure to provide a basic understanding of the techniques discussed. This summary is not an exhaustive overview of all the intended features of this disclosure, nor is it intended to identify key or essential elements of all aspects of this disclosure, nor to define the scope of any or all aspects of this disclosure. The sole purpose of this summary is to present, in a general form, some concepts of one or more aspects of this disclosure as a prelude to the more detailed description given later.

[0004] In some aspects, multi-frame image processing may involve processing multiple image frames using an anchor image frame as a base reference. The anchor image frame may have regions that are better represented in other image frames. For example, a subject may have its eyes closed in the anchor image frame but not in subsequent image frames. Artificial intelligence (AI) statistics about the image frames can be used to identify the anchor image frame, as well as other image frames that better represent the subject. The anchor image frame can be modified by replacing pixels in the anchor image frame with pixels from replacement image frames. The modified anchor image frame can then be used in multi-frame image processing by merging multiple image frames to obtain a better representation of the scene.

[0005] Some examples of anchored image frames that can benefit from this replacement of statistical data determined by machine learning (ML) models are group portraits and action portraits of humans and pets, which often produce artifacts such as one or more subjects having poor posture (e.g., closed eyes or an unhappy expression). In other examples, some subjects may be blurred due to subject movement (e.g., moving children and pets). These problems can be a result of no single image frame having the best facial appearance for each person in the portrait. Embodiments of this disclosure include processing multiple image frames and analyzing the image frames using artificial intelligence (AI) in an ML model to determine information such as statistical data about at least one feature of the subjects (e.g., open eyes, happy expression, non-neutral expression). Based on the AI ​​data, one frame is identified as the anchored image frame from the available frames. Based on the anchored image frame and the statistical data determined by the ML model, other image frames including better expressions for some subjects in the anchored image frame can be identified.

[0006] For individual subjects, better frames can be masked and merged with anchor image frames, so that a better representation of the subject is placed in the anchor image frame (e.g., by replacing pixels in the anchor image frame with pixels from a replacement image frame). This can be repeated for multiple subjects based on the problem subjects identified in the anchor image frames. After stitching the improved anchor image frames together with other image frames, the anchor image frames are used in temporal merging with other temporally close image frames (e.g., image frames captured sequentially in time, within several image frames, or within a predetermined time period, such as less than 1 second or less than 5 times the exposure duration) to improve the noise, tone, and other overall characteristics of the resulting image frames. The resulting image frames can be displayed directly on the display of the image capture device as part of a preview operation of a real-time representation of the scene. The preview operation can bypass encoding tools to improve the responsiveness of the preview operation. The resulting image frames can optionally be encoded as picture files and stored in memory.

[0007] Example methods embodying some of the techniques of this disclosure include the following steps, although the invention should not be considered limited to this particular sequence of steps, require all specific steps in this sequence, and / or disallow other steps to be performed as part of the method. First, anchor image frame selection is performed based on statistics determined by an ML model from a preview stream of image frames. This reduces the latency of anchor frame selection because the statistics are already available. Second, AI metrics are used to identify subjects with poor poses in the anchor image frames. Third, for the identified subjects, other frames are searched to better render the subjects. If the anchor image frames already have the best rendering for all subjects, further processing can be skipped to produce the final output. In some embodiments, additional steps may include providing a user interface (UI) pop-up (e.g., as part of a camera application) that indicates the anchor image frames along with other image frames having better poses for the subjects. In some embodiments, replacement may be performed automatically (e.g., based on predetermined criteria rather than user input). Whether by user or automatically, image processing may perform the replacement of the subjects selected in the anchor image frames from replacement image frames. Then, by performing temporal and / or spatial processing using anchored image frames, output image frames are generated from the newly composed anchored image frames to produce portrait photographs.

[0008] In one aspect of this disclosure, a method for image processing includes receiving, by at least one processor, a plurality of image frames captured at different times; determining, by at least one processor, an anchor image frame from the plurality of image frames; replacing, by at least one processor, a first pixel of a region of the anchor image frame with a second pixel of a first replacement image frame from the plurality of image frames; and determining, by at least one processor, an output image frame based on pixel values ​​of the anchor image frame and at least one additional image frame by merging the anchor image frame with at least one additional image frame. When at least one processor includes more than one processor, each of these processors may be configured to perform all operations, or these processors may perform individual operations of the configured operations to perform the operations collaboratively.

[0009] In an additional aspect of this disclosure, an apparatus includes at least one processor and a memory coupled to the at least one processor. The at least one processor is configured to perform operations including: receiving a plurality of image frames captured at different times; determining an anchor image frame from the plurality of image frames; replacing a first pixel of a region of the anchor image frame with a second pixel of a first replacement image frame from the plurality of image frames; and determining an output image frame based on the anchor image frame and the at least one additional image frame by merging the first plurality of pixels of the anchor image frame with a second plurality of pixels of at least one additional image frame.

[0010] In an additional aspect of this disclosure, an apparatus includes components for receiving, by at least one processor, a plurality of image frames captured at different times; components for determining, by at least one processor, an anchor image frame from the plurality of image frames; components for replacing, by at least one processor, a first pixel of a region of the anchor image frame with a second pixel of a first replacement image frame from the plurality of image frames; and components for determining, by at least one processor, an output image frame based on the anchor image frame and the at least one additional image frame by merging a first plurality of pixels of the anchor image frame with a second plurality of pixels of at least one additional image frame.

[0011] In an additional aspect of this disclosure, a non-transitory computer-readable medium stores instructions that, when executed by a processor, cause the processor to perform operations. The operations include receiving, by at least one processor, a plurality of image frames captured at different times; means for determining, by at least one processor, an anchor image frame from the plurality of image frames; means for at least one processor to replace a first pixel of a region of the anchor image frame with a second pixel of a first replacement image frame from the plurality of image frames; and means for at least one processor to determine an output image frame based on the anchor image frame and the at least one additional image frame by merging a first plurality of pixels of the anchor image frame with a second plurality of pixels of at least one additional image frame.

[0012] The image processing methods described herein can be performed by an image capture device and / or on image data captured by one or more image capture devices. An image capture device (a device capable of capturing one or more digital images, whether still photographs or video sequences) can be incorporated into a variety of devices. By way of example, an image capture device may include a standalone digital camera or digital video camera, a wireless communication device equipped with a camera (such as a mobile phone, cellular, or satellite radio phone), a personal digital assistant (PDA), a panel or tablet device, a gaming device, a computing device (such as a webcam, video surveillance camera), or other devices with digital imaging or video capabilities.

[0013] The image processing techniques described herein can relate to a digital camera having an image sensor and processing circuitry (e.g., an application-specific integrated circuit (ASIC), a digital signal processor (DSP), a graphics processing unit (GPU), or a central processing unit (CPU)). An image signal processor (ISP) may include one or more of these processing circuits and is configured to perform operations to acquire image data for processing according to the image processing techniques described herein and / or those involved in the image processing techniques described herein. An ISP may be configured to control the capture of image frames from one or more image sensors and to determine one or more image frames from said one or more image sensors to generate a view of a scene in an output image frame. The output image frame may be part of a sequence of image frames forming a video sequence. The video sequence may include additional image frames received from an image sensor or other image sensors.

[0014] In an example application, an image signal processor (ISP) may receive instructions for capturing a sequence of image frames in response to the loading of software, such as a camera application, to generate a preview display from an image capture device. The ISP may be configured to generate a single output image frame stream based on image frames received from one or more image sensors. The single output image frame stream may include raw image data from the image sensors, merged image data from the image sensors, or corrected image data processed by one or more algorithms within the ISP. For example, image frames may be processed by an image post-processing engine (IPE) and / or other image processing circuitry to process the image frames obtained from the image sensors (which may have undergone some processing before being output to the ISP), thereby performing one or more of tone mapping, portrait lighting, contrast enhancement, gamma correction, etc. The output image frames from the ISP may be stored in memory and retrieved by an application processor executing the camera application, which may perform further processing on the output image frames to adjust their appearance and reproduce them on a display for user viewing.

[0015] After an image signal processor and / or application processor (such as the image processing techniques described in the various embodiments herein) determines an output image frame representing a scene, the output image frame may be displayed on a device display as a single still image and / or as part of a video sequence, saved to a storage device as a picture or video sequence, transmitted over a network, and / or printed to an output medium. For example, an image signal processor (ISP) may be configured to acquire input frames of image data (e.g., pixel values) from one or more image sensors and subsequently generate corresponding output image frames (e.g., preview display frames, still image captures, frames for video, frames for object tracking, etc.). In other examples, the image signal processor may output image frames to various output devices and / or camera modules for further processing, such as for 3A parameter synchronization (e.g., autofocus (AF), auto white balance (AWB), and auto exposure control (AEC)), to generate video files via the output frames, to configure frames for display, to configure frames for storage, to transmit frames via a network connection, etc. Generally, an image signal processor (ISP) can obtain incoming frames from one or more image sensors, generate an output frame stream, and output the output frame stream to various output destinations.

[0016] In some aspects, output image frames can be generated by combining various aspects of the image correction disclosed herein with other computational photographic techniques such as high dynamic range (HDR) photography or multi-frame noise reduction (MFNR). In the case of HDR photography, the first and second image frames are captured using different exposure times, different apertures, different lenses, and / or other characteristics that can result in improved dynamic range of the fused image when combining the two image frames. In some aspects, the method can be performed for MFNR photography, wherein the first and second image frames are captured using the same or different exposure times, and the first and second image frames are fused to generate a corrected first image frame that has reduced noise compared to the captured first image frame.

[0017] In some aspects, the device may include an image signal processor or processor (e.g., an application processor) that includes specific functionalities for camera control and / or processing, such as enabling or disabling the merging module or otherwise controlling aspects of image correction. The methods and techniques described herein may be performed entirely by the image signal processor or processor, or the various operations may be separated between the image signal processor and the processor, and in some aspects across additional processors.

[0018] The device may include one, two, or more image sensors, such as a first image sensor. When multiple image sensors are present, their configurations may differ. For example, the first image sensor may have a larger field of view (FOV) than the second image sensor, or the first image sensor may have a different sensitivity or a different dynamic range than the second image sensor. In one example, the first image sensor may be a wide-angle image sensor, and the second image sensor may be a long-range image sensor. In another example, the first sensor is configured to acquire an image through a first lens having a first optical axis, and the second sensor is configured to acquire an image through a second lens having a second optical axis different from the first optical axis. Additionally or alternatively, the first lens may have a first magnification, and the second lens may have a second magnification different from the first magnification. Any of these or other configurations may be part of a lens cluster on a mobile device, such as where multiple image sensors and associated lenses are located at offset positions on the front or rear of the mobile device. Additional image sensors with larger, smaller, or the same field of view may be included. The image processing techniques described herein can be applied to image frames captured from any of the image sensors in a multi-sensor device.

[0019] In an additional aspect of this disclosure, an apparatus configured for image processing and / or image capture is disclosed. The apparatus includes components for capturing image frames. The apparatus also includes one or more components for capturing data representing a scene, such as image sensors (including charge-coupled device (CCD), Bayer filter sensors, infrared (IR) detectors, ultraviolet (UV) detectors, complementary metal-oxide-semiconductor (CMOS) sensors) and time-of-flight detectors. The apparatus may further include components for focusing and / or directing light onto one or more image sensors (including simple lenses, compound lenses, spherical lenses, and aspherical lenses). These components can be controlled to capture a first image frame and / or a second image frame input to the image processing techniques described herein.

[0020] Other aspects, features, and specific embodiments will become apparent to those skilled in the art when they review the following description of particular exemplary aspects in conjunction with the accompanying drawings. Although features may be discussed hereinafter with reference to certain aspects and drawings, various aspects may include one or more of the advantageous features discussed herein. In other words, while one or more aspects may be discussed having certain advantageous features, one or more such features may also be used depending on the various aspects. Similarly, although exemplary aspects may be discussed hereinafter as aspects of an apparatus, system, or method, exemplary aspects can be implemented in various apparatuses, systems, and methods.

[0021] This method can be embedded as computer program code in a computer-readable medium, the computer program code including instructions that cause a processor to perform the steps of the method. In some embodiments, the processor may be part of a mobile device including: a first network adapter configured to transmit data, such as recorded images or videos or streaming data, via a first network connection among a plurality of network connections; and a processor coupled to the first network adapter and memory. The processor enables the output image frames described herein to be transmitted via a wireless communication network, such as a 5G NR communication network.

[0022] The features and technical advantages of the examples according to this disclosure have been summarized rather extensively above in order to better understand the detailed description below. Additional features and advantages will be described below. The disclosed concepts and specific examples can be readily used as the basis for modifying or designing other structures for achieving the same purpose of this disclosure. Such equivalent constructions do not depart from the scope of the appended claims. The characteristics of the concepts disclosed herein (both their organization and manner of operation) and their associated advantages will be better understood from the following description when considered in conjunction with the accompanying drawings. Each figure in the drawings is provided for illustrative and descriptive purposes and not as a definition of limitation of the claims.

[0023] While aspects and implementations are described herein by way of example, those skilled in the art will understand that additional implementations and use cases may arise in many different arrangements and scenarios. The innovations described herein can be implemented across many different platform types, devices, systems, shapes, sizes, and package arrangements. For example, aspects and / or implementations may be via integrated chip implementations and other devices based on non-modular components (e.g., end-user devices, vehicles, communication devices, computing devices, industrial equipment, retail / purchasing devices, medical devices, AI-enabled devices, etc.). While some examples may or may not specifically point to use cases or applications, applicability to various types of the described innovations is possible. The scope of implementations can range from chip-level or modular components to non-modular, non-chip-level implementations, and further to aggregated, distributed, or original equipment manufacturer (OEM) devices or systems incorporating one or more aspects of the described innovations. In some practical settings, devices incorporating the described aspects and features may also necessarily include additional components and features for implementing and practicing the claimed and described aspects. For example, the transmission and reception of wireless signals necessarily involve multiple components (e.g., hardware components, including antennas, radio frequency (RF) chains, power amplifiers, modulators, buffers, processors, interleavers, adders / summers, etc.) for analog and digital purposes. The innovations described herein are intended to be implemented in a variety of devices, chip-level components, systems, distributed arrangements, end-user equipment, etc., with different sizes, shapes, and constructions. Attached Figure Description

[0024] A further understanding of the nature and advantages of this disclosure can be achieved by referring to the following figures. In the figures, similar components or features may have the same reference numerals. Furthermore, various components of the same type can be distinguished by adding a dash after the reference numerals and a second reference numeral for differentiation between similar components. If only the first reference numeral is used in the specification, the description applies to any one of the similar components having the same first reference numeral, regardless of the second reference numerals.

[0025] Figure 1 A block diagram of an example device for performing image capture from one or more image sensors is shown.

[0026] Figure 2 This is a block diagram illustrating an example data flow path for image data processing in an image capture device according to one or more embodiments of the present disclosure.

[0027] Figure 3 A flowchart is shown of an example method for processing image data with anchor frame modification prior to multi-frame time processing, according to some embodiments of the present disclosure.

[0028] Figure 4 This is a block diagram illustrating an example method of modifying the anchor frame by changing the subject according to some embodiments of this disclosure.

[0029] Figure 5 A flowchart is shown of an example method for processing image data according to some embodiments of the present disclosure, which involves swapping the subject in an anchor frame prior to multi-frame temporal processing.

[0030] Figure 6 This is a block diagram illustrating some embodiments of image processing according to the present disclosure, which involves swapping the subject in an anchor frame prior to multi-frame time processing.

[0031] Figure 7 This is a block diagram illustrating the determination of the Best Anchor Selector Score (BASS) according to some embodiments of this disclosure.

[0032] The same reference numerals and names in the various figures indicate the same elements. Detailed Implementation

[0033] The detailed description set forth below in conjunction with the accompanying drawings is intended as a description of various configurations and is not intended to limit the scope of this disclosure. Rather, the detailed description includes specific details for providing a thorough understanding of the subject matter of the invention. It will be apparent to those skilled in the art that these specific details are not necessary in every situation, and in some instances, well-known structures and components are shown in block diagram form for clarity.

[0034] This disclosure provides systems, apparatus, methods, and computer-readable media that support image processing, including techniques for multi-frame image processing to provide better noise, tone, and other details of the scene in the photograph. Photographs are improved in various aspects of this disclosure by using statistics determined by an ML model to identify anchor image frames, and / or by modifying anchor image frames to identify replacement pixels in other image frames based on statistics determined by an ML model for the anchor image frames and other image frames.

[0035] Although photographs are described in some embodiments, aspects of this disclosure can be implemented in recording image frames as photographs, or in determining output image frames as part of a camera preview operation. When executed as part of a preview operation, statistics determined by the ML model can be computed during the preview time and stored in a metadata queue along with the image frame buffer. In snapshots, the statistics can be used for zero shutter lag (ZSL) scenes. In other embodiments, the statistics can be determined on demand.

[0036] Specific embodiments of the subject matter described in this disclosure may be implemented to achieve one or more of the following potential advantages or benefits. In some aspects, this disclosure provides techniques for improving the quality of photographs and camera previews involving multi-frame image processing, particularly in scenes with group portraits and / or action portraits. For example, the improved quality may result in the output image frame having better facial expressions for the subjects in the scene by replacing pixels in an anchor image frame with pixels from other image frames, particularly by replacing the face of a subject in an anchor image frame with the face of a subject from another image frame.

[0037] In the description of the embodiments herein, numerous specific details (such as examples of specific components, circuits, and processes) are set forth to provide a thorough understanding of this disclosure. As used herein, the term "coupled" means a direct connection or a connection via one or more intermediate components or circuits. Furthermore, specific terminology is set forth in the following description and for purposes of explanation in order to provide a thorough understanding of this disclosure. However, it will be apparent to those skilled in the art that practicing the teachings disclosed herein may not require these specific details. In other instances, known circuits and devices are illustrated in block diagram form to avoid obscuring the teachings of this disclosure.

[0038] Certain portions of the following detailed description are presented using other symbolic representations of procedures, logic blocks, processes, and data bit operations within computer memory. In this disclosure, procedures, logic blocks, processes, etc., are conceived as a self-consistent sequence of steps or instructions that produce a desired result. These steps are those that require physical operations on physical quantities. Although not strictly necessary, these physical quantities typically take the form of electrical or magnetic signals that can be stored, transferred, combined, compared, and otherwise manipulated within a computer system.

[0039] Example devices (such as smartphones) for capturing image frames using one or more image sensors may include a configuration of one, two, three, four, or more camera modules on the rear side (e.g., the side opposite the main user display) and / or the front side (e.g., the same side as the main user display). These devices may include one or more image signal processors (ISPs), computer vision processors (CVPs) (e.g., AI engines), or other suitable circuitry for processing the images captured by the image sensors. The one or more image signal processors (ISPs) may store the output image frames in memory (e.g., via a bus) and / or provide the output image frames to processing circuitry (e.g., an application processor). The processing circuitry may perform further processing, such as encoding, storing, transmitting, or other manipulation of the output image frames.

[0040] As used herein, a camera module may include an image sensor and certain other components coupled to the image sensor for acquiring a representation of a scene in image data comprising image frames. For example, a camera module may include other components of the camera, including a shutter, buffer, or additional readout circuitry for accessing individual pixels of the image sensor. In some embodiments, a camera module may include one or more components including an image sensor housed in a single package having an interface configured to couple the camera module to an image signal processor or other processor via a bus.

[0041] Figure 1 A block diagram of a device 100 for performing image capture from one or more image sensors is shown. Device 100 may include or be otherwise coupled to an image signal processor (e.g., ISP 112) for processing image frames from one or more image sensors, such as a first image sensor 101, a second image sensor 102, and a depth sensor 140. In some specific embodiments, device 100 may also include or be coupled to a processor 104 and a memory 106 storing instructions 108 (e.g., memory storing processor-readable code or a non-transitory computer-readable medium storing instructions). Device 100 may also include or be coupled to a display 114 and component 116. Component 116 may be used for user interaction, such as a touchscreen interface and / or physical buttons.

[0042] Component 116 may also include network interfaces for communicating with other devices, including a wide area network (WAN) adapter (e.g., WAN adapter 152), a local area network (LAN) adapter (e.g., LAN adapter 153), and / or a personal area network (PAN) adapter (e.g., PAN adapter 154). WAN adapter 152 may be a 4G LTE or 5G NR wireless network adapter. LAN adapter 153 may be an IEEE 802.11 WiFi wireless network adapter. PAN adapter 154 may be a Bluetooth wireless network adapter. Each of WAN adapter 152, LAN adapter 153, and / or PAN adapter 154 may be coupled to an antenna comprising multiple antennas configured for main and diversity reception and / or configured to receive a specific frequency band. In some embodiments, the antennas may be shared by WAN adapter 152, LAN adapter 153, and / or PAN adapter 154 for communication on different networks. In some implementations, WAN adapter 152, LAN adapter 153 and / or PAN adapter 154 may share circuitry and / or be packaged together, such as when LAN adapter 153 and PAN adapter 154 are packaged as a single integrated circuit (IC).

[0043] Device 100 may also include or be coupled to a power source 118 for use with device 100, such as a battery or an adapter for coupling device 100 to an energy source. Device 100 may also include or be coupled to... Figure 1 Additional features or components not shown. In one example, a wireless interface that may include multiple transceivers and a baseband processor in a radio frequency front-end (RFFE) may be coupled to or included in the WAN adapter 152 for use in a wireless communication device. In another example, an analog front-end (AFE) for converting analog image data to digital image data may be coupled between the first image sensor 101 or the second image sensor 102 and the processing circuitry in the device 100. In some embodiments, the AFE may be embedded in the ISP 112.

[0044] The device may include or be coupled to a sensor hub 150, which interfaces with sensors to receive data about the movement of device 100, data about the environment surrounding device 100, and / or other non-camera sensor data. One example non-camera sensor is a gyroscope, a device configured to measure rotation, orientation, and / or angular velocity to generate motion data. Another example non-camera sensor is an accelerometer, a device configured to measure acceleration, which can also be used to determine the speed and distance of travel by appropriately integrating the measured acceleration. In some aspects, a gyroscope in an electronic image stabilization system (EIS) may be coupled to the sensor hub. In another example, the non-camera sensor may be a Global Positioning System (GPS) receiver, a device used to process satellite signals, such as through triangulation and other techniques, to determine the position of device 100. Position can be tracked over time to determine additional motion information, such as velocity and acceleration. Data from one or more sensors may be accumulated by the sensor hub 150 into motion data. One or more of acceleration, velocity, and / or distance may be included in the motion data provided by sensor hub 150 to other components of device 100, including ISP 112 and / or processor 104.

[0045] The ISP 112 can receive captured image data. In one embodiment, a local bus connection couples the ISP 112 to the first image sensor 101 and the second image sensor 102 of the first camera 103 and the second camera 105, respectively. In another embodiment, a wired interface couples the ISP 112 to an external image sensor. In yet another embodiment, a wireless interface couples the ISP 112 to either the first image sensor 101 or the second image sensor 102.

[0046] First image sensor 101 and second image sensor 102 are configured to capture image data representing scenes within the fields of view of first camera 103 and second camera 105, respectively. In some embodiments, first camera 103 and / or second camera 105 output analog data converted by an analog front-end (AFE) and / or analog-to-digital converter (ADC) in device 100 or embedded in ISP 112. In some embodiments, first camera 103 and / or second camera 105 output digital data. The digital image data may be formatted into one or more image frames, whether received from first camera 103 and / or second camera 105 or converted from analog data received from first camera 103 and / or second camera 105.

[0047] The first camera 103 may include a first image sensor 101 and a first lens 131. The second camera may include a second image sensor 102 and a second lens 132. Each of the first lens 131 and the second lens 132 may be controlled by an associated autofocus (AF) algorithm (e.g., AF 133) executed in the ISP 112, which adjusts the first lens 131 and the second lens 132 to focus on a specific focal plane located at a specific scene depth. AF 133 may be assisted by depth data received from the depth sensor 140. The first lens 131 and the second lens 132 focus light onto the first image sensor 101 and the second image sensor 102 respectively through one or more apertures for receiving light, one or more shutters for blocking light when outside the exposure window, and / or one or more color filter arrays (CFAs) for filtering light outside a specific frequency range. The first lens 131 and the second lens 132 may have different fields of view to capture different representations of the scene. For example, the first lens 131 may be an ultra-wide (UW) lens, and the second lens 132 may be a wide (W) lens. Multiple image sensors may include a combination of ultra-wide (high field of view (FOV)) sensors, wide sensors, long-range sensors, and ultra-long-range (low FOV) sensors.

[0048] Each of the first camera 103 and the second camera 105 can be configured through hardware configuration and / or software settings to obtain different but overlapping fields of view. In some configurations, the cameras are configured with different lenses with different magnifications, resulting in different fields of view for capturing different representations of the scene. The cameras can be configured such that the UW camera has a larger FOV than the W camera, the W camera has a larger FOV than the T camera, and the T camera has a larger FOV than the UT camera. For example, a camera configured for a wide FOV can capture a field of view in the range of 64 to 84 degrees, a camera configured for an ultra-side FOV can capture a field of view in the range of 100 to 140 degrees, a camera configured for a long-range FOV can capture a field of view in the range of 10 to 30 degrees, and a camera configured for an ultra-long-range FOV can capture a field of view in the range of 1 to 8 degrees.

[0049] In some implementations, one or more of the first camera 103 and / or the second camera 105 may be variable aperture (VA) cameras, wherein the aperture can be adjusted to set a specific aperture size. Example aperture sizes include f / 2.0, f / 2.8, f / 3.2, f / 8.0, etc. Larger aperture values ​​correspond to smaller aperture sizes, and smaller aperture values ​​correspond to larger aperture sizes. Variable aperture (VA) cameras may have different characteristics that produce different representations of the scene based on the current aperture size. For example, a VA camera may capture image data with a depth of focus (DOF) corresponding to the current aperture size set for the VA camera.

[0050] The ISP 112 processes image frames captured by the first camera 103 and the second camera 105. Although Figure 1 Device 100 is illustrated as including a first camera 103 and a second camera 105, but any number of cameras (e.g., one, two, three, four, five, six, etc.) may be coupled to ISP 112. In some aspects, depth sensors (such as depth sensor 140) may be coupled to ISP 112. The output from depth sensor 140 may be processed in a manner similar to that of the first camera 103 and the second camera 105. Examples of depth sensors 140 include active sensors, including one or more of indirect time-of-flight (iToF), direct time-of-flight (dToF), light detection and ranging (LiDAR), mmWave, radio detection and ranging (radar), and / or hybrid depth sensors (such as structured light sensors). In embodiments without depth sensor 140, similar information about the depth or depth map of an object may be determined based on the parallax between the first camera 103 and the second camera 105, such as by using parallax depth measurement algorithms, stereo depth measurement algorithms, phase detection autofocus (PDAF) sensors, etc. In addition, any number of additional image sensors or image signal processors may be present in device 100.

[0051] In some embodiments, ISP 112 may execute instructions from memory, such as instructions 108 from memory 106, instructions stored in a separate memory coupled to or included in ISP 112, or instructions provided by processor 104. Additionally or alternatively, ISP 112 may include specific hardware (such as one or more integrated circuits (ICs)) configured to perform one or more operations described in this disclosure. For example, ISP 112 may include an image front-end (e.g., IFE 135), an image post-processing engine (e.g., IPE 136), an automatic exposure compensation (AEC) engine (e.g., AEC 134), and / or one or more engines for video analysis (e.g., EVA 137). The image pipeline may be formed by a sequence of one or more of IFE 135, IPE 136, and / or EVA 137. In some embodiments, the image pipeline in ISP 112 may be reconfigured by changing the connections between IFE 135, IPE 136, and / or EVA 137. AF 133, AEC 134, IFE 135, IPE 136 and EVA137 may each include dedicated circuitry and may be embodied as software or firmware executed by ISP 112 and / or a combination of hardware and software or firmware executed on ISP 112.

[0052] Memory 106 may include a non-transient or non-transitory computer-readable medium storing computer-executable instructions (as instructions 108) for performing all or part of one or more of the operations described in this disclosure. Instructions 108 may include a camera application (or other suitable application, such as a messaging application) to be executed by device 100 for taking pictures or videos. Instructions 108 may also include other applications or programs executed by device 100, such as an operating system and applications other than those for image or video generation. Executing a camera application, such as by processor 104, may enable device 100 to record images using the first camera 103 and / or the second camera 105 and the ISP 112.

[0053] In addition to instruction 108, memory 106 may also store image frames. The image frames may be output image frames stored by ISP 112. The output image frames may be accessed by processor 104 for further operation. In some embodiments, device 100 does not include memory 106. For example, device 100 may be circuitry including ISP 112, and the memory may be external to device 100. Device 100 may be coupled to external memory and configured to access that memory to write output image frames for display or long-term storage. In some embodiments, device 100 is a system-on-a-chip (SoC) that integrates ISP 112, processor 104, sensor hub 150, memory 106, and / or component 116 into a single package.

[0054] In some embodiments, at least one of the ISP 112 or processor 104 executes instructions to perform various operations described herein, including replacing portions of anchored image frames. For example, the execution of instructions may instruct the ISP 112 to begin or end capturing image frames or sequences of image frames, wherein the capture includes corrections as described in the embodiments herein. In some embodiments, processor 104 may include one or more general-purpose processor cores 104A to 104N capable of executing instructions to control the operation of the ISP 112. For example, cores 104A to 104N may execute a camera application (or other suitable application for generating images or videos) stored in memory 106 that activates or deactivates the ISP 112 to capture image frames and / or controls the ISP 112 in a portrait swapping application that replaces the face of a subject in an anchored image frame with the face of the same subject from another image frame. The operation of cores 104A to 104N and ISP 112 may be based on user input. For example, a camera application executing on processor 104 may receive a user command to start a video preview display. Upon receiving the user command, video, including a sequence of image frames, is captured and processed via ISP 112 from first camera 103 and / or second camera 105 for display and / or storage. Image processing, such as that described herein, for determining “output” or “corrected” image frames, may be applied to one or more image frames in the sequence.

[0055] In some implementations, processor 104 may include an IC or other hardware (e.g., an artificial intelligence (AI) engine, such as AI engine 124, or other coprocessor) to offload certain tasks from cores 104A through 104N. AI engine 124 may be used to offload tasks related to face detection and / or object recognition, performed, for example, using machine learning (ML) or artificial intelligence (AI). AI engine 124 may be referred to as an artificial intelligence processing unit (AI PU). AI engine 124 may include hardware configured to perform and accelerate convolutional operations involved in performing machine learning algorithms, such as by executing predictive models such as artificial neural networks (ANNs), including multilayer feedforward neural networks (MLFFNNs), recurrent neural networks (RNNs), and / or radial basis functions (RBFs). The ANN executed by AI engine 124 has access to predefined training weights for performing operations on user data. The ANN may optionally be trained during operation of image capture device 100, such as through reinforcement training, supervised training, and / or unsupervised training. In some other embodiments, device 100 does not include processor 104, such as when all the described functionality is configured in ISP 112.

[0056] In some embodiments, display 114 may include one or more suitable displays or screens that allow the user to interact and / or present a preview of an item (such as the output of the first camera 103 and / or the second camera 105) to the user. In some embodiments, display 114 is a touch-sensitive display. Input / output (I / O) components (such as component 116) may be or include any suitable mechanism, interface, or device to receive input (such as commands) from the user and provide output to the user via display 114. For example, component 116 may include (but is not limited to) a graphical user interface (GUI), keyboard, mouse, microphone, speaker, squeezable bezel, one or more buttons (such as a power button), slider, toggle key, switch, etc.

[0057] Although shown as coupled to each other via processor 104, components (such as processor 104, memory 106, ISP 112, display 114, and component 116) may be coupled to each other in various other arrangements, such as via one or more local buses, which are not shown for simplicity. An example of a bus used to interconnect components is a Peripheral Component Interface (PCI) Fast (PCIe) bus.

[0058] Although ISP 112 is illustrated as separate from processor 104, ISP 112 may be the core of processor 104, which is an application processor unit (APU) included in a system-on-a-chip (SoC), or otherwise included in processor 104. While device 100 is referenced in the examples herein to perform aspects of this disclosure, some device components may not be included. Figure 1 The details are shown to prevent obscuring aspects of this disclosure. Additionally, other components, the number of components, or combinations of components may be included in suitable equipment for performing aspects of this disclosure. Therefore, this disclosure is not limited to the configuration of a particular device or component, including device 100.

[0059] Figure 1 An exemplary image capture device can be operated to obtain an improved image by correcting the anchored image frame, including by replacing the low-quality portion of the anchored image frame as part of a preview operation. Figure 2 An example method of operating one or more cameras (such as a first camera 103 and / or a second camera 105) is shown and described below.

[0060] Figure 2 This is a block diagram illustrating an example data flow path for image data processing in an image capture device according to one or more embodiments of the present disclosure. The processor 104 of system 200 communicates with ISP 112 via a bidirectional bus and / or separate control and data lines. The processor 104 can control a first camera 103 via camera control 210. Camera control 210 may be a camera driver executed by processor 104 for configuring the first camera 103 (such as activating or deactivating image capture, configuring exposure settings, and / or configuring aperture size). Camera control 210 may be managed by a camera application 204 executing on processor 104. Camera application 204 provides user-accessible settings, allowing a user to specify individual camera settings or select a profile with corresponding camera settings. Camera control 210 communicates with the first camera 103 to configure the first camera 103 according to commands received from camera application 204. Camera application 204 may be, for example, a photography application, a document scanning application, a messaging application, or other application that processes image data acquired from the first camera 103.

[0061] Camera configuration may include specifying parameters such as frame rate, image resolution, readout duration, exposure level, aspect ratio, aperture size, etc. The first camera 103 may apply the camera configuration and use it to acquire image data representing the scene. In some embodiments, the camera configuration may be adjusted to obtain different representations of the scene. For example, the processor 104 may execute camera application 204 to instruct the first camera 103 to set a first camera configuration via camera control 210, acquire first image data from the first camera 103 operating with the first camera configuration, instruct the first camera 103 to set a second camera configuration, and acquire second image data from the first camera 103 operating with the second camera configuration.

[0062] In some embodiments where the first camera 103 is a variable aperture (VA) camera system, the processor 104 can execute camera application 204 to instruct the first camera 103 to be configured to a first aperture size, acquire first image data from the first camera 103, instruct the first camera 103 to be configured to a second aperture size, and acquire second image data from the first camera 103. The aperture reconfiguration and the acquisition of the first and second image data can occur with little or no change in the scene captured at the first aperture size and the second aperture size. Example aperture sizes are f / 2.0, f / 2.8, f / 3.2, f / 8.0, etc. Larger aperture values ​​correspond to smaller aperture sizes, and smaller aperture values ​​correspond to larger aperture sizes. That is, f / 2.0 corresponds to an aperture size larger than f / 8.0.

[0063] Image data received from the first camera 103 can be processed in one or more blocks of the ISP 112 to determine an output image frame 230 that can be stored in memory 106 and / or otherwise provided to the processor 104. The processor 104 can further process the image data to apply effects to the output image frame 230. Effects may include background blur, lighting, color cast, and / or high dynamic range (HDR) blending. In some embodiments, effects can be applied in the ISP 112.

[0064] The output image frame 230 of ISP 112 may include a representation of the scene improved by various aspects of this disclosure, such that the anchor image frame is improved before multi-frame temporal processing. Processor 104 may display these output image frames 230 to a user, and the improvements provided by the described processing implemented in ISP 112 and / or processor 104 improve image quality and user experience by reducing the appearance of bright and dark areas in the photograph. For example, when determining the output image frame 230, portrait switching 212 in ISP 112 may correct image data received from the first camera 103.

[0065] Figure 2System 200 can be configured to execute the reference. Figure 3 The described operation determines the output image frame 230. Figure 3 A flowchart is shown of an example method for processing image data with anchor frame modification prior to multi-frame time processing, according to some embodiments of the present disclosure. Figure 3 Capturing images in this way allows for an improved digital representation of the scene, resulting in photos or videos with higher image quality (IQ). (Reference) Figure 3 Each operation described may be performed by one or a combination of processor 104 (including cores 104A to 104N or AI engine 124) and / or ISP 112.

[0066] At block 302, first image data is received from the image sensor, such as when the image sensor is configured with a camera configuration. The first image data may be received, for example, from a bus coupled to the first camera 103 or from an analog front-end (AFE) coupled to the first camera 103. Alternatively, the first image data may be received from a wireless camera, wherein the image data is received via one or more of WAN adapter 152, LAN adapter 153, and / or PAN adapter 154. Alternatively, the image data may be received from memory location or network storage location, such as when image data was previously captured and is now retrieved from memory 106 and / or remote location via one or more of the adapters WAN adapter 152, LAN adapter 153, and / or PAN adapter 154. In some embodiments, the capture of image data may be initiated by a camera application executing on processor 104, which causes camera control 210 to activate the first camera 103 to capture image data. The image data retrieved at box 302 can then be processed by the ISP 112 and / or processor 104, or other components for processing the image data according to the operations described in one or more of the following boxes. Multiple image frames received at box 302 may be captured at different times, allowing the subject to have multiple poses in different image frames. Some of these poses may be undesirable, causing subsequent processing in boxes 304, 306, and 308 to replace the subject in the anchored image frame before performing multi-frame temporal processing that combines the anchored image frame with other image frames to improve image quality (e.g., by reducing noise in the MFNR).

[0067] At box 304, an anchor image frame is determined from the image frames received from box 302. The anchor image frame can be selected based on criteria involving statistics calculated by an ML model for the image frames received at box 302.

[0068] At box 306, the first pixel of the anchor frame is modified (e.g., replaced) with the second pixel of the first replacement image frame. The first replacement image frame may be one of a plurality of image frames received at box 302. The first replacement image frame can be identified by analyzing the plurality of image frames. For example, a machine learning (ML) model may use artificial intelligence to identify the first replacement image frame as an image frame having a better pose for at least one subject of the anchor image frame determined at box 304. Example ML models may define computational power for making determinations based on input data, where the determinations are based on patterns identified in the input data. Computational power may be defined by weights and biases. Weights may indicate a relationship between certain input data and certain determinations. Bias may indicate a starting point for the determination. Example ML models that operate on input data may start with a determination defined by biases and subsequently change their determinations based on a combination of input data and weights. Some example ML models include artificial neural networks (ANNs), regression analysis (such as statistical models), decision tree learning (such as predictive models), support vector networks (SVMs), large language models (LLMs), generative models, and probabilistic graphical models (such as Bayesian networks), etc.

[0069] At box 308, the output image frame is determined by merging (e.g., temporal blending) the anchor image frame (as modified at box 306) with one or more additional image frames. Merging may include merging each pixel of the anchor image frame with the corresponding pixel of each of the one or more additional image frames. Alternatively, merging may operate on a subset of pixels of the image frame, such that merging includes combining a first plurality of pixels of the anchor image frame (which may be a subset of all pixels) with a second plurality of pixels of at least one additional image frame. During temporal merging of frames, the contribution weight of each frame may vary between 0% and 100% for each pixel value in the output image frame. Multi-frame merging at box 308 may result in, for example, lower noise in the darker parts of the scene captured by the multiple image frames. Additional image frames may be one or more image frames among the plurality of image frames received at box 302 and / or other image frames not considered part of the image processing in boxes 304 and 306.

[0070] Figure 4 The portrait swapping illustrates an example of pixel replacement in an anchored image frame. Figure 4This is a block diagram illustrating an example method of modifying an anchor frame by swapping subjects according to some embodiments of the present disclosure. Anchor image frame 402 may include multiple subjects. Artificial intelligence (AI) processing of the anchor image frame may determine that one or more of these subjects have an undesirable pose. Replacement image frames 404 and 406 may be identified as having better poses for subjects 414 and 416, respectively. Anchor image frame 402 may be modified by replacing pixels corresponding to subject 414 with a representation of subject 414 from the first replacement image frame 404. Anchor image frame 402 may be further modified by replacing pixels corresponding to subject 416 with a representation of subject 416 from the second replacement image frame 406. Anchor image frame 402 may be processed with additional image frames in multi-frame temporal processing 420. Additional image frames 422 may include replacement image frames 404 and 406. Although subjects and poses are described in some embodiments, portrait swapping features may be applied to swap the representation of an object based on a better representation of the object in the identified image frame. Such objects may include foreground (FG) patches, FG objects at the center of an image frame, and / or semantically segmented regions. For example, non-human subjects may include pets, animals, and / or vehicles.

[0071] The processing of anchored image frames may include user input when identifying the content to be replaced in the anchored image frame, such as... Figure 5 The processing flow is described in the documentation. Figure 5 A flowchart of an example method for processing image data, according to some embodiments of the present disclosure, is shown, which involves swapping the subject in an anchor frame prior to multi-frame temporal processing. Method 500 begins at block 502: a user activates a camera application that begins capturing image frames from an image sensor. At block 504, anchor image frame selection is performed. This selection may be based on statistics about the image frames, such as AI Region Statistics (ARS), determined by an ML model. For example, criteria for selecting anchor image frames may include identifying frames with the highest Best Anchor Selector Score (BASS).

[0072] BASS scores can be like Figure 7As shown, the calculation is based on features present in the image frame. At box 702, the device determines whether the ROI is based on user selection (such as by touching an area on the display). If so, the image frame is analyzed to determine the presence of a person or face. If the ROI is determined to include a person at box 704, the BASS score of the image frame is determined based on average facial brightness, facial sharpness, average skin brightness, skin sharpness, probability of open eyes, and / or probability of maximum or happy and non-neutral expressions. If the ROI is determined to include a detected pet at box 704, the BASS score of the image frame is determined based on average pet brightness and / or average pet sharpness. Otherwise, the BASS score is determined at box 704 based on significantly weighted brightness and / or significantly weighted sharpness. If the ROI is not based on user selection, the device determines at box 706 whether the segmentation includes valid people, pets, and / or vehicles. If yes, then at box 708, the BASS score is determined based on the average facial luminance, facial sharpness, average skin luminance, skin sharpness, average pet luminance, average pet sharpness, probability of open eyes, maximum probability of happy and non-neutral expressions, average vehicle luminance, and / or average vehicle sharpness. If determined as negative at box 706, the ROI region identifier is analyzed to determine whether saliency is a selector for the ROI. If yes, then at box 712, the BASS score is determined based on saliency-weighted luminance and / or saliency-weighted sharpness. Otherwise, at box 714, the BASS score is determined based on average global luminance and / or global sharpness.

[0073] At box 506, subjects in the anchored image frame are identified, for example, by using a face detection algorithm. One or more features can be evaluated and examined against criteria to identify subjects for potential substitutions. For example, image features such as the probability of open eyes, happy expression, and / or non-neutral expression determined by face detection or facial keypoint processing can be evaluated, and / or facial brightness, facial sharpness, skin brightness, skin sharpness, pet brightness, pet sharpness, global brightness, and / or global sharpness evaluated by a semantic segmentation network can be compared with criteria defining a satisfactory pose. These criteria can be configured by the user via a camera application and / or automatically adjusted based on conditions of the image frame (e.g., location, indoor / outdoor, time of day, etc.). If some subjects do not meet the criteria, method 500 can continue further processing at box 508 to modify the anchored image frame. If all subjects meet the criteria or no subjects are identified, method 500 can proceed to box 516 to determine the output image frame using the anchored image frame.

[0074] Additional processing may involve identifying the corresponding region of the modified anchor image frame and the region to be swapped into the replacement image frame in the anchor image frame to improve the appearance of the anchor image frame. At box 508, a subject that does not meet the criteria at box 506 is identified in another image frame, where the subject has better characteristics (such as those indicated by evaluating one or more of the same criteria examined at box 506). When a better representation of the subject is identified, the user may be prompted with the option to swap the subject in the anchor image frame with the subject in the replacement image frame at box 510. At box 512, if the user does not accept the swap, method 500 proceeds to box 516 to determine the output image frame using the unmodified anchor image frame. If the user accepts the swap, method 500 proceeds to box 514 to replace the pixels in the anchor image frame with pixels from the replacement image frame, and then uses the anchor image frame to determine the output image frame.

[0075] Figure 6 As shown Figure 5 The example system for image processing is described in the example. Figure 6 This is a block diagram illustrating image processing according to some embodiments of the present disclosure, which involves subject swapping in an anchor frame prior to multi-frame temporal processing. Input image frame 602 may be provided to an autofocus statistics block 612, a face detection block 614, a facial keypoint block 616, and a semantic segmentation block 618. The autofocus statistics block 612 may determine a region of interest (ROI) provided to the anchor frame selection block 622. The ROI may be used to identify anchor frames, such as anchor frames in focus. The face detection block 614 may determine the presence of a face and / or metadata about the face for the image frame. The face detection block 614 may identify the image frame for further evaluation by the anchor frame selection block 622 and adaptive subject swapping 624. The facial keypoint block 616 may determine statistics determined by an ML model, including the probability of open eyes, the probability of a happy expression, and the probability of a non-neutral expression. In some embodiments, the face detection block 614 may identify faces for processing by the facial keypoint block 616 and the semantic segmentation block 618 to reduce the processing load performed by these blocks. Semantic segmentation block 618 can determine one or more of facial brightness, facial sharpness, skin brightness, skin sharpness, pet brightness, pet sharpness, global brightness, and / or global sharpness. AI metadata from facial keypoint block 616 and semantic segmentation block 618 can be provided to adaptive subject swapping block 624, which determines image frames with better subject representation (e.g., better pose) compared to anchored image frames.

[0076] The subject swapping block 624 can be selectively activated for anchored image frames with detected faces, and this block determines modified anchored image frames to provide to the anchored frame selection block 622. Although some embodiments describe subject swapping based on statistics about faces, other embodiments may operate on image frames involving animals or objects and corresponding characteristics of the animals or objects to modify the anchored image frames based on preferred representations identified by those animals or objects.

[0077] Anchor frame selection block 622 provides the modified anchor image frame to temporal image processing block 626. Block 626 combines multiple image frames captured at different times (but possibly overlapping) to determine an output image frame with improved image quality. For example, multi-frame processing can be used to reduce noise from the image sensor. In some embodiments, the output image frame from block 626 can be further processed in spatial image processing block 628, such as spatial noise suppression, tone mapping, color conversion and correction, sharpening, bokeh, and / or other special effects.

[0078] In one or more aspects, the techniques for supporting image processing may include additional aspects, such as any single aspect or any combination of aspects described below or in conjunction with one or more other processes or devices described elsewhere herein. In a first aspect, supporting image processing may include an apparatus configured to improve an anchor image frame (e.g., by swapping portions of pixels of the anchor image frame with other image frames before temporal merging of the anchor frame with other additional image frames). The apparatus is also configured to perform operations including: receiving, by at least one processor, a plurality of image frames captured at different times; determining, by the at least one processor, an anchor image frame from the plurality of image frames; replacing, by the at least one processor, a first pixel of a region of the anchor image frame with a second pixel of a first replacement image frame from the plurality of image frames; and determining, by the at least one processor, an output image frame based on the anchor image frame and the at least one additional image frame by merging the first plurality of pixels of the anchor image frame with a second plurality of pixels of at least one additional image frame.

[0079] Additionally, the apparatus may perform or operate according to one or more aspects described below. In some embodiments, the apparatus includes a wireless device, such as a UE. In some embodiments, the apparatus includes a remote server (such as a cloud-based computing solution) that receives image data for processing to determine output image frames. In some embodiments, the apparatus may include at least one processor and memory coupled to the processor. The processor may be configured to perform the operations described herein with reference to the apparatus. In some other embodiments, the apparatus may include a non-transitory computer-readable medium having program code recorded thereon, and the program code may be executable by a computer to cause the computer to perform the operations described herein with reference to the apparatus. In some embodiments, the apparatus may include one or more components configured to perform the operations described herein. In some embodiments, a method of wireless communication may include one or more operations described herein with reference to the apparatus.

[0080] In a second aspect, in conjunction with the first aspect, the apparatus is further configured to perform operations including: determining by the at least one processor a first plurality of values, each of the first plurality of values ​​corresponding to the same characteristic of the plurality of image frames; and determining by the at least one processor, based on criteria relating to the first plurality of values, to replace the first pixel of the anchored image frame with the second pixel of the first replacement image frame, wherein the replacement of the first pixel is performed based on: determining to replace the first pixel based on the criteria.

[0081] In a third aspect, in combination with one or more of the first or second aspects, determining the first plurality of values ​​includes determining the first plurality of statistical data based on a machine learning (ML) model.

[0082] In the fourth aspect, in combination with one or more of the first to third aspects, determining the first plurality of statistical data includes determining statistical data about at least one feature of a subject in the plurality of image frames determined using the ML model.

[0083] In the fifth aspect, in combination with one or more of the first to fourth aspects, determining the region to be replaced by the anchored image frame includes determining the pose of the subject in the first replacement image frame and the pose of the subject in the anchored image frame based on the first plurality of values.

[0084] In a sixth aspect, in combination with one or more of the first to fifth aspects, each of the plurality of image frames includes a first subject and a second subject, and / or the replacement of the first pixel of the anchored image frame with the second pixel of the first replacement image frame by the at least one processor includes replacing the first region in the anchored image frame that also corresponds to the first subject with the second region in the first replacement image frame that corresponds to the first subject.

[0085] In a seventh aspect, in combination with one or more of the first to sixth aspects, the device is also configured to perform an operation including displaying the output image frame by the at least one processor as part of a camera preview operation.

[0086] In the eighth aspect, in combination with one or more of the first to seventh aspects, the device is further configured to perform an operation including receiving user input confirming that the first pixel of the anchored image frame is to be replaced with the second pixel of the first replacement image frame, wherein the replacement is performed after receiving the user input confirming that the area is to be replaced.

[0087] In the ninth aspect, in combination with one or more of the first to eighth aspects, the at least one additional image frame is derived from the plurality of image frames.

[0088] In the tenth aspect, in combination with one or more of the first to ninth aspects, receiving the plurality of image frames includes receiving image data from at least one image sensor.

[0089] In the eleventh aspect, in combination with one or more of the first to tenth aspects, the device may include an image sensor coupled to the at least one processor.

[0090] In the accompanying drawings, a single block can be described as performing one or more functions. The one or more functions performed by this block can be performed in a single component or across multiple components, and / or can be performed using hardware, software, or a combination of hardware and software. To clearly illustrate this interchangeability between hardware and software, various exemplary components, blocks, modules, circuits, and steps are described below in general terms of their functionality. Whether such functionality is implemented as hardware or software depends on the specific application and the design constraints imposed on the system as a whole. Those skilled in the art may implement the described functionality in different ways for each specific application, but such implementation decisions should not be construed as departing from the scope of this disclosure. Additionally, the example device may include components other than those shown, including well-known components such as processors, memory, etc.

[0091] The aspects of this disclosure are applicable to any electronic device that includes, is coupled to, or otherwise processes data from one, two, or more image sensors capable of capturing image frames (or “frames”). The terms “output image frame,” “modified image frame,” and “corrected image frame” can refer to an image frame that has been processed by any of the techniques disclosed to adjust the raw image data received from the image sensor. Additionally, aspects of the disclosed techniques can be implemented for processing image data received from image sensors having the same or different capabilities and characteristics, such as resolution, shutter speed, or sensor type. Furthermore, aspects of the disclosed techniques can be implemented in devices for processing image data, whether or not the device includes or is coupled to an image sensor. For example, the disclosed techniques may include operations performed by a processing device in a cloud computing system that retrieves image data previously recorded by a separate device having an image sensor for processing.

[0092] Unless explicitly stated otherwise in the following discussion, it should be understood that throughout this application, the use of terms such as “access,” “receive,” “transmit,” “use,” “select,” “determine,” “normalize,” “multiply,” “average,” “monitor,” “compare,” “apply,” “update,” “measure,” “derive,” “set,” “generate,” etc., refers to the actions and processes of a computer system or similar electronic computing device that manipulate data represented as physical (electronic) quantities in the registers and memories of the computer system and transform them into other data similarly represented as physical quantities in the registers, memories, or other such information storage, transmission, or display devices of the computer system. The use of different terms to refer to actions or processes of a computer system does not necessarily indicate different operations. For example, “determining” data can refer to “generating” data. Similarly, “determining” data can refer to “retrieving” data.

[0093] The terms "device" and "apparatus" are not limited to one or a specific number of physical objects (such as a smartphone, a camera controller, a processing system, etc.). As used herein, a device can be any electronic device having one or more components that can implement at least some parts of this disclosure. Although the description and examples herein use the term "device" to describe various aspects of this disclosure, the term "device" is not limited to a particular configuration, type, or number of objects. As used herein, an apparatus can include a device or part of a device for performing the described operations.

[0094] Certain components in a device or apparatus described as “parts for access,” “parts for receiving,” “parts for transmitting,” “parts for using,” “parts for selecting,” “parts for determining,” “parts for normalizing,” “parts for multiplying,” or other similarly named terms referring to one or more operations on data (such as image data) may refer to processing circuitry (e.g., application-specific integrated circuit (ASIC), digital signal processor (DSP), graphics processing unit (GPU), central processing unit (CPU), computer vision processor (CVP), or neural signal processor (NSP)) configured to perform the described functions by means of hardware, software, or a combination of hardware configured by software.

[0095] Those skilled in the art will understand that information and signals can be represented using any of a variety of different techniques and skills. For example, data, instructions, commands, information, signals, bits, symbols, and chips that may be mentioned throughout the above description can be represented by voltage, current, electromagnetic waves, magnetic fields or magnetic particles, light fields or optical particles, or any combination thereof.

[0096] The components, functional blocks, and modules described herein with respect to the accompanying figures cited above include processors, electronic devices, hardware devices, electronic components, logic circuits, memory, software code, firmware code, and so on, or any combination thereof. Software should be interpreted broadly as instructions, instruction sets, code, code segments, program code, programs, subroutines, software modules, applications, software applications, software packages, routines, subroutines, objects, executable files, threads of execution, procedures, and / or functions, whether referred to as software, firmware, middleware, microcode, hardware description languages, or other terms. Furthermore, the features discussed herein may be implemented via dedicated processor circuitry, via executable instructions, or a combination thereof.

[0097] Those skilled in the art should understand that, with reference to Figures 3 to 6 One or more boxes (or operations) described may be combined with one or more boxes (or operations) described in another figure referring to the figures. For example, Figure 3 One or more boxes (or operations) can be connected with Figures 1 to 2 A combination of one or more boxes (or operations). For example, with... Figure 4 One or more associated boxes can be combined with Figures 1 to 2 A combination of one or more associated boxes (or operations).

[0098] Those skilled in the art will also recognize that the various exemplary logic blocks, modules, circuits, and algorithm steps described in connection with the disclosure herein can all be implemented as electronic hardware, computer software, or a combination of both. To clearly illustrate this interchangeability between hardware and software, various exemplary components, blocks, modules, circuits, and steps have been described above in general terms of their functionality. Whether such functionality is implemented as hardware or software depends on the specific application and the design constraints imposed on the overall system. Those skilled in the art may implement the described functionality in different ways for each specific application, but such implementation decisions should not be construed as departing from the scope of this disclosure. Those skilled in the art will also readily recognize that the order or combination of components, methods, or interactions described herein is merely illustrative, and that components, methods, or interactions of various aspects of this disclosure can be combined or performed in ways other than those illustrated and described herein.

[0099] The various exemplary logics, logic blocks, modules, circuits, and algorithmic processes described in conjunction with the specific implementations disclosed herein can be implemented as electronic hardware, computer software, or a combination of both. The interchangeability of hardware and software has been generally described in terms of functionality and illustrated in the various exemplary components, blocks, modules, circuits, and processes described above. Whether such functionality is implemented in hardware or software depends on the specific application and the design constraints imposed on the overall system.

[0100] Hardware and data processing means for implementing the various exemplary logic units, logic blocks, modules, and circuits described herein can be implemented or executed using general-purpose single-chip or multi-chip processors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs) or other programmable logic devices, discrete gate or transistor logic units, discrete hardware components, or any combination thereof designed to perform the functions described herein. A general-purpose processor can be a microprocessor, or any conventional processor, controller, microcontroller, or state machine. In some embodiments, the processor may be implemented as a combination of computing devices, such as a combination of a DSP and a microprocessor, multiple microprocessors, one or more microprocessors combined with a DSP core, or any other such configuration. In some embodiments, specific processes and methods may be performed by circuitry specific to a given function.

[0101] In one or more aspects, the described functionality may be implemented in hardware, digital electronic circuits, computer software, firmware, including the structures disclosed in this specification and their structural equivalents or any combination thereof. Specific implementations of the subject matter described in this specification may also be implemented as one or more computer programs, i.e., one or more modules of computer program instructions encoded on a computer storage medium for execution by a data processing apparatus or for controlling the operation of a data processing apparatus.

[0102] If implemented in software, the functionality can be stored as one or more instructions or code on or transmitted through a computer-readable medium. The processes of the methods or algorithms disclosed herein can be implemented in a processor-executable software module that can reside on a computer-readable medium. Computer-readable media include both computer storage media and communication media, including any medium that can be implemented to transfer a computer program from one location to another. Storage media can be any available medium accessible to a computer. By way of example and not limitation, such computer-readable media may include random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), CD-ROM or other optical disc storage devices, magnetic disk storage devices or other magnetic storage devices, or any other medium capable of storing desired program code in the form of instructions or data structures and accessible to a computer. Additionally, any connection may be appropriately referred to as a computer-readable medium. As used herein, disks and optical discs include compact optical discs (CDs), laser discs, optical discs, digital versatile optical discs (DVDs), floppy disks, and Blu-ray discs, wherein disks typically magnetically reproduce data, while optical discs optically reproduce data using lasers. Combinations of the above should also be included within the scope of computer-readable media. Additionally, the operation of a method or algorithm may reside as a set of code and instructions or any combination of code and instructions on a machine-readable medium and a computer-readable medium that may be incorporated into a computer program product.

[0103] Various modifications to the specific embodiments described in this disclosure will be apparent to those skilled in the art, and the general principles defined herein may be applied to other specific embodiments without departing from the spirit or scope of this disclosure. Therefore, the claims are not intended to be limited to the specific embodiments shown herein, but are to be granted the broadest scope consistent with this disclosure, the principles disclosed herein, and the novel features.

[0104] Additionally, those skilled in the art will readily recognize that, for the convenience of describing the accompanying drawings, contrasting terms such as “upper” and “lower” or “front” and “back” or “top” and “bottom” or “forward” and “backward” are sometimes used, indicating relative positions on a correctly oriented page corresponding to the orientation of the drawings, and may not reflect the correct orientation of any device as implemented.

[0105] Certain features described in this specification in the context of a single embodiment may also be implemented in combination in a single embodiment. Conversely, various features described in the context of a single embodiment may also be implemented individually or in any suitable sub-combination in multiple embodiments. Furthermore, although features may be described above as operating in certain combinations and even originally claimed in this way, one or more features from the claimed combination may be removed from that combination in some cases, and the claimed combination may be for sub-combinations or variations thereof.

[0106] Similarly, although operations are depicted in a specific order in the figures, this should not be construed as requiring such operations to be performed in the indicated specific order or sequential order, or to perform all illustrated operations to achieve the desired result. Furthermore, the figures may schematically depict one or more example processes in the form of flowcharts. However, other operations not depicted may be combined with the schematically illustrated example processes. For example, one or more additional operations may be performed before, after, simultaneously with, or between any illustrated operation. In some contexts, multitasking and parallel processing are advantageous. Moreover, the separation of the various system components in the embodiments described above should not be construed as requiring such separation in all embodiments, and it should be understood that the described program components and systems can generally be integrated together in a single software product or packaged into multiple software products. Additionally, some other embodiments also fall within the scope of the appended claims. In some cases, the actions recited in the claims can be performed in a different order and still achieve the desired result.

[0107] As used herein (including the claims), the term "or" in a list of two or more items means that any one of the listed items may be used alone, or any combination of two or more listed items may be used. For example, if a composition is described as containing component A, B, or C, the composition may contain A alone; B alone; C alone; a combination of A and B; a combination of A and C; a combination of B and C; or a combination of A, B, and C. Additionally, as used herein (including the claims), "or" in a list of items beginning with "at least one of" indicates a separate list, such that a list such as "at least one of A, B, or C" refers to A or B or C or AB or AC or BC or ABC (i.e., A and B and C) or any combination of any of these items.

[0108] The term “substantially” is defined as being largely but not necessarily entirely what is specified (and includes what is specified; for example, substantially 90 degrees includes 90 degrees and substantially parallel includes parallel), as understood by one of ordinary skill in the art. In any specific implementation of the disclosure, the term “substantially” may be used in place of the “[percentage]” of the specified content, where the percentage includes 0.1%, 1%, 5%, or 10%.

[0109] The prior description of this disclosure is provided to enable any person skilled in the art to make or use this disclosure. Various modifications to this disclosure will be apparent to those skilled in the art, and the general principles defined herein may be applied to other variations without departing from the spirit or scope of this disclosure. Therefore, this disclosure is not intended to be limited to the examples and designs described herein, but is to be accorded the widest scope consistent with the principles and novel features disclosed herein.

Claims

1. A method, the method comprising: Multiple image frames captured at different times are received by at least one processor; The anchor image frame is determined from the plurality of image frames by the at least one processor; The at least one processor replaces the first pixel of the region of the anchored image frame with the second pixel of the first replacement image frame in the plurality of image frames; as well as The at least one processor determines an output image frame based on the anchor image frame and the at least one additional image frame by merging a first plurality of pixels of the anchor image frame with a second plurality of pixels of at least one additional image frame.

2. The method according to claim 1, further comprising: The at least one processor determines a first plurality of values, each of the first plurality of values ​​corresponding to the same characteristic of the plurality of image frames; as well as The at least one processor determines, based on criteria involving the first plurality of values, which pixel of the anchored image frame should be replaced with the second pixel of the first replacement image frame. The replacement of the first pixel is performed based on the following criteria: determining which first pixel to replace is based on the criteria.

3. The method of claim 2, wherein determining the first plurality of values ​​comprises determining the first plurality of statistical data based on a machine learning (ML) model.

4. The method of claim 3, wherein determining the first plurality of statistical data comprises determining statistical data about at least one feature of a subject in the plurality of image frames, determined using the ML model.

5. The method of claim 4, wherein determining the region to be replaced in the anchored image frame comprises determining the pose of the subject in the first replacement image frame and the pose of the subject in the anchored image frame based on the first plurality of values.

6. The method according to claim 1, wherein: Each of the plurality of image frames includes a first subject and a second subject; and Replacing the first pixel of the anchored image frame with the second pixel of the first replacement image frame by the at least one processor includes replacing the first region in the anchored image frame that also corresponds to the first subject with the second region in the first replacement image frame that corresponds to the first subject.

7. The method of claim 1, further comprising having the at least one processor display the output image frame as part of a camera preview operation.

8. The method of claim 1, further comprising receiving user input confirming that the first pixel of the anchored image frame is to be replaced with the second pixel of the first replaced image frame, wherein the replacement is performed after receiving the user input confirming that the region is to be replaced.

9. The method of claim 1, wherein the at least one additional image frame is derived from the plurality of image frames.

10. The method of claim 1, wherein receiving the plurality of image frames comprises receiving image data from at least one image sensor.

11. An apparatus comprising: Memory, the memory storing processor-readable code; and At least one processor, coupled to the memory, is configured to execute processor-readable code to cause the at least one processor to perform operations, the operations including: Multiple image frames captured at different times are received by at least one processor; The anchor image frame is determined from the plurality of image frames by the at least one processor; The at least one processor replaces the first pixel of the region of the anchored image frame with the second pixel of the first replacement image frame from the plurality of image frames; and The at least one processor determines an output image frame based on the anchor image frame and the at least one additional image frame by merging a first plurality of pixels of the anchor image frame with a second plurality of pixels of at least one additional image frame.

12. The apparatus of claim 11, wherein the operation further comprises: The at least one processor determines a first plurality of values, each of the first plurality of values ​​corresponding to the same characteristic of the plurality of image frames; The at least one processor determines, based on criteria involving the first plurality of values, which pixel of the anchored image frame should be replaced with the second pixel of the first replacement image frame. The replacement of the first pixel is performed based on the following criteria: determining which first pixel to replace is based on the criteria.

13. The apparatus of claim 12, wherein determining the first plurality of values ​​comprises determining the first plurality of statistical data based on a machine learning (ML) model.

14. The apparatus of claim 13, wherein determining the first plurality of statistical data includes determining statistical data about at least one feature of a subject in the plurality of image frames, determined using the ML model.

15. The apparatus of claim 14, wherein determining the region to be replaced in the anchored image frame comprises determining the pose of the subject in the first replacement image frame and the pose of the subject in the anchored image frame based on the first plurality of values.

16. The apparatus according to claim 11, wherein: Each of the plurality of image frames includes a first subject and a second subject; and Replacing the first pixel of the anchored image frame with the second pixel of the first replacement image frame by the at least one processor includes replacing the first region in the anchored image frame that also corresponds to the first subject with the second region in the first replacement image frame that corresponds to the first subject.

17. The apparatus of claim 11, wherein the operation further comprises the at least one processor displaying the output image frame as part of a camera preview operation.

18. The apparatus of claim 11, wherein the operation further comprises receiving user input confirming that the first pixel of the anchored image frame is to be replaced with the second pixel of the first replaced image frame, wherein the replacement is performed after receiving the user input confirming that the region is to be replaced.

19. The apparatus of claim 11, wherein the at least one additional image frame is derived from the plurality of image frames.

20. The apparatus of claim 11, wherein receiving the plurality of image frames includes receiving image data from at least one image sensor.

21. A non-transitory computer-readable medium storing instructions that, when executed by a processor, cause the processor to perform operations, the operations including: Multiple image frames captured at different times are received by at least one processor; The anchor image frame is determined from the plurality of image frames by the at least one processor; The at least one processor replaces the first pixel of the region of the anchored image frame with the second pixel of the first replacement image frame in the plurality of image frames; as well as The at least one processor determines an output image frame based on the anchor image frame and the at least one additional image frame by merging a first plurality of pixels of the anchor image frame with a second plurality of pixels of at least one additional image frame.

22. The non-transitory computer-readable medium of claim 21, wherein the operation further comprises: The at least one processor determines a first plurality of values, each of the first plurality of values ​​corresponding to the same characteristic of the plurality of image frames; The at least one processor determines, based on criteria involving the first plurality of values, which pixel of the anchored image frame should be replaced with the second pixel of the first replacement image frame. The replacement of the first pixel is performed based on the following criteria: determining which first pixel to replace is based on the criteria.

23. The non-transitory computer-readable medium of claim 22, wherein determining the first plurality of values ​​comprises determining the first plurality of statistical data based on a machine learning (ML) model.

24. The non-transitory computer-readable medium of claim 23, wherein determining the first plurality of statistical data includes determining statistical data about at least one feature of a subject in the plurality of image frames, determined using the ML model.

25. The non-transitory computer-readable medium of claim 24, wherein determining the region to be replaced in the anchored image frame comprises determining the pose of the subject in the first replacement image frame and the pose of the subject in the anchored image frame based on the first plurality of values.

26. An image capturing device, the image capturing device comprising: Image sensor; Memory, the memory storing processor-readable code; and At least one processor, coupled to the memory and the image sensor, is configured to execute processor-readable code to cause the at least one processor to: Multiple image frames captured at different times are received by at least one processor; The anchor image frame is determined from the plurality of image frames by the at least one processor; The at least one processor replaces the first pixel of the region of the anchored image frame with the second pixel of the first replacement image frame in the plurality of image frames; as well as The at least one processor determines an output image frame based on the anchor image frame and the at least one additional image frame by merging a first plurality of pixels of the anchor image frame with a second plurality of pixels of at least one additional image frame.

27. The image capture device of claim 26, wherein the operation further comprises: The at least one processor determines a first plurality of values, each of the first plurality of values ​​corresponding to the same characteristic of the plurality of image frames; The at least one processor determines, based on criteria involving the first plurality of values, which pixel of the anchored image frame should be replaced with the second pixel of the first replacement image frame. The replacement of the first pixel is performed based on the following criteria: determining which first pixel to replace is based on the criteria.

28. The image capture device of claim 27, wherein determining the first plurality of values ​​comprises determining the first plurality of statistical data based on a machine learning (ML) model.

29. The image capture device of claim 28, wherein determining the first plurality of statistical data includes determining statistical data about at least one feature of a subject in the plurality of image frames, determined using the ML model.

30. The image capture device of claim 29, wherein determining the region to be replaced in the anchored image frame comprises determining the pose of the subject in the first replacement image frame and the pose of the subject in the anchored image frame based on the first plurality of values.