Selective motion distortion correction within image frame

By detecting moving hotspots within the image frame and applying only McTF corrections to them, the problem of image capture devices is solved, extending battery life and improving processing efficiency.

CN120303687APending Publication Date: 2025-07-11QUALCOMM INC
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202380083695.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2022-12-12
Filing Date
2023-11-21
Publication Date
2025-07-11

AI Technical Summary

Technical Problem

Existing image capture devices are intensive in computing resources when processing high-resolution image frames, resulting in high power consumption, short battery life, and limited video processing pipelines.

Method used

By detecting motion hotspots in the image frame, only motion compensation time filtering (McTF) is applied to the areas that meet the motion threshold for correction, reducing the use of computing resources.

Benefits of technology

Effectively reduces the computing resources required for correction of motion distortion in the image frame, extends the battery life of the device and improves the efficiency of the image processing pipeline.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120303687A_ABST
    Figure CN120303687A_ABST
Patent Text Reader

Abstract

This disclosure provides systems, methods, and devices for image signal processing that support improved correction of motion artifacts within an image frame. In a first aspect, an image processing method includes receiving a first image frame and a second image frame, determining a motion map indicative of motion of an object within the first image frame and the second image frame. Additionally, motion hotspots may be identified within the second image frame based on the motion map. A temporal filtering process may be applied to a portion of the second image frame located within the motion hotspot to generate a corrected image frame. Other aspects and features are also claimed and described.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] Cross - Reference to Related Applications

[0002] This application claims the benefit of U.S. Patent Application No. 18 / 064,542, filed on December 12, 2022, entitled "SELECTIVE MOTION DISTORTION CORRECTION WITHIN IMAGE FRAMES", which is hereby incorporated by reference in its entirety. Technical Field

[0003] Aspects of the present disclosure generally relate to image processing, and more particularly to image processing performed on image frames to correct movement within the image frames. Some features enable and provide improved image processing, including using motion - compensated temporal filtering (McTF) in an improved manner to reduce the overall computational resources required to correct image frames using McTF techniques, especially for image frames of captured video. Background Art

[0004] An image - capture device is a device that can capture one or more digital images (either still images for photos or sequences of images for video). The capture device can be incorporated into various devices. By way of example, the image - capture device can include a standalone digital camera or digital video camera, a wireless communication device phone equipped with a camera (such as a mobile phone, cellular or satellite radiotelephone), a personal digital assistant (PDA), a panel or tablet device, a gaming device, a computing device (such as a webcam, video surveillance camera), or other devices having digital imaging or video capabilities.

[0005] The amount of image data captured by an image sensor has increased through successive generations of image - capture devices. The amount of information captured by an image sensor is related to the number of pixels in the image sensor of the image - capture device, and the number of pixels can be measured as the number of megapixels indicating the number of millions of sensors in the image sensor. For example, a 12 - megapixel image sensor has 12 million pixels. Higher megapixel values generally represent higher - resolution images that are more suitable for user viewing.

[0006] The increased amount of image data captured by an image capture device can have some negative impacts as the resolution obtained from the additional image data increases. The additional image data increases the amount of processing performed by the image capture device when determining image frames and videos from the image data and performing other operations related to the image data. For example, before displaying the image data to a user on a display or sending the image data to a recipient in a message, the image data can be processed through several processing blocks to enhance the image. Each of the processing blocks consumes additional power proportional to the amount of image data or the number of megapixels in the image capture. The additional power consumption can shorten the operating time of an image capture device that uses battery power, such as a mobile phone. SUMMARY OF THE INVENTION

[0007] Some aspects of the present disclosure are summarized below to provide a basic understanding of the technology being discussed. This summary is not an exhaustive overview of all the expected features of the present disclosure, and is neither intended to identify the key or important elements of all aspects of the present disclosure, nor to depict the scope of any or all aspects of the present disclosure. The sole purpose of this summary of the invention is to present some concepts of one or more aspects of the present disclosure in a general form as a prelude to the more detailed embodiments that are presented later.

[0008] In some aspects of the present disclosure, temporal filtering can be applied to or not applied to portions of an image frame based on motion detected within the image frame. The selective operation of temporal filtering on portions of the image data reduces the amount of processing performed by the processor applying the temporal filtering, which can reduce power consumption and increase operating speed. Temporal filtering can be applied to remove distortions within an image frame caused by the movement of an object when capturing image frames for synthesis into a single still image or as part of a video sequence. For example, temporal filtering selectively applied to a region of an image frame can reduce distortions caused by blurring within the image frame and / or the movement of one or more objects blended together.

[0009] The motion map can indicate the local movement of an object depicted within a second image frame in a first image frame. Additionally or alternatively, the motion map can reflect the global movement of a device used to capture the image frames. Motion hotspots can also be identified based on the motion map. The motion hotspots can identify corresponding portions of the second image frame that contain or may contain motion distortion and / or motion artifacts. Specifically, motion hotspots can be determined to identify portions of the second image frame that have motion greater than or equal to a predetermined threshold. A temporal filtering process can then be applied to the portions of the second image frame that are within the motion hotspots to generate a corrected image frame. Applying the temporal filtering process can include determining an image transformation matrix, such as a motion compensation transformation matrix. The image transformation matrix can be determined based on the motion map and / or the first and second image frames. The image transformation matrix can be determined such that when applied to pixels or other portions (e.g., blocks, superblocks, macroblocks, regions) of the second image frame, the motion distortion and / or motion artifacts are corrected (e.g., removed, reduced). The corrected image frame can then be added to an output file. This process can be repeated to process multiple image frames. For example, this process can be repeated sequentially across all received image frames (e.g., in the order in which the image frames are captured and / or received).

[0010] In some aspects, the temporal filtering process applied to the corrected image frame is a motion compensated temporal filtering (MCTF) process. MCTF reduces noise and / or motion artifacts in video by filtering motion regions based on the global and / or local movement of the current frame relative to the previous frame. When processing a scene with limited motion within the field of view (FOV), selectively performing MCTF filtering only on identified regions that meet certain criteria can reduce core power consumption.

[0011] In some aspects, the techniques described herein relate to a method that includes: receiving a first image frame and a second image frame; determining a motion map based on the second image frame and the first image frame; determining a region of the second image frame that has motion meeting at least one criterion based on the motion map; and determining a corrected image frame by applying a temporal filtering process to the region.

[0012] In some aspects, the techniques described herein relate to an apparatus that includes: a memory that stores processor-readable code; and at least one processor coupled to the memory, the at least one processor being configured to execute the processor-readable code to cause the at least one processor to perform operations including: receiving a first image frame and a second image frame; determining a motion map based on the second image frame and the first image frame; determining a region of the second image frame that has motion meeting at least one criterion based on the motion map; and determining a corrected image frame by applying a temporal filtering process to the region.

[0013] In some aspects, the techniques described herein relate to a non-transitory computer-readable medium storing instructions that, when executed by a processor, cause the processor to perform operations including: receiving a first image frame and a second image frame; determining a motion map based on the second image frame and the first image frame; determining, based on the motion map, a region of the second image frame having motion that satisfies at least one criterion; and determining a corrected image frame by applying a temporal filtering process to the region.

[0014] In some aspects, the techniques described herein relate to an image capture device including: an image sensor; a memory storing processor-readable code; and at least one processor coupled to the memory and the image sensor, the at least one processor configured to execute the processor-readable code to cause the at least one processor to: receive a first image frame and a second image frame in image data received from the image sensor; determine a motion map based on the second image frame and the first image frame; determine, based on the motion map, a region of the second image frame having motion that satisfies at least one criterion; and determine a corrected image frame by applying a temporal filtering process to the region.

[0015] The image processing methods described herein may be performed by an image capture device and / or on image data captured by one or more image capture devices. An image capture device (a device that can capture one or more digital images, whether a still image photograph or a sequence of images of a video) may be incorporated into a variety of devices. By way of example, an image capture device may include a standalone digital camera or digital video camera, a wireless communication device phone equipped with a camera (such as a mobile phone, a cellular or satellite radiotelephone), a personal digital assistant (PDA), a panel or tablet device, a gaming device, a computing device (such as a webcam, a video surveillance camera), or other devices having digital imaging or video capabilities.

[0016] The image processing techniques described herein may relate to a digital camera having an image sensor and processing circuitry (e.g., an application specific integrated circuit (ASIC), a digital signal processor (DSP), a graphics processing unit (GPU), or a central processing unit (CPU)). An image signal processor (ISP) may include one or more of these processing circuits and be configured to perform operations to obtain image data for processing according to the image processing techniques described herein and / or the image processing techniques involved in the techniques described herein. The ISP may be configured to control the capture of image frames from one or more image sensors and determine one or more image frames from the one or more image sensors to generate a view of a scene in an output image frame. The output image frame may be part of a sequence of image frames forming a video sequence. The video sequence may include other image frames received from the image sensor or other image sensors.

[0017] In an example application, an Image Signal Processor (ISP) may receive instructions for capturing a sequence of image frames in response to the loading of software, such as a camera application, to generate a preview display from an image capture device. The image signal processor may be configured to generate a single output image frame stream based on the image frames received from one or more image sensors. The single output image frame stream may include raw image data from the image sensors, merged image data from the image sensors, or corrected image data processed by one or more algorithms within the image signal processor. For example, the image frames may be processed by an Image Post-Processing Engine (IPE) and / or other image processing circuitry to process the image frames obtained from the image sensors in the image signal processor (the image frames may have had some processing of the data performed on them before being output to the image signal processor), thereby performing one or more of tone mapping, portrait illumination, contrast enhancement, gamma correction, etc. The output image frames from the ISP may be stored in memory and retrieved by an application processor executing the camera application, which may perform further processing on the output image frames to adjust the appearance of the output image frames and reproduce the output image frames on a display for user viewing.

[0018] After an output frame representing a scene is determined by the image signal processor and / or the application processor (such as through the image processing techniques described in various embodiments herein), the output image frames may be displayed on a device display as a single still image and / or as part of a video sequence, saved to a storage device as a picture or video sequence, sent over a network, and / or printed to an output medium. For example, an Image Signal Processor (ISP) may be configured to obtain an input frame of image data (e.g., pixel values) from one or more image sensors and, in turn, generate a corresponding output image frame (e.g., a preview display frame, a still image capture, a frame for video, a frame for object tracking, etc.). In other examples, the image signal processor may output the image frames to various output devices and / or camera modules for further processing, such as for 3A parameter synchronization (e.g., Auto Focus (AF), Auto White Balance (AWB), and Auto Exposure Control (AEC)), generating a video file via the output frames, configuring the frames for display, configuring the frames for storage, sending the frames over a network connection, etc. Generally, an Image Signal Processor (ISP) may obtain incoming frames from one or more image sensors, generate an output frame stream, and output the output frame stream to various output destinations.

[0019] In some aspects, an output image frame can be generated by combining aspects of the image correction of the present disclosure with other computational photography techniques such as high dynamic range (HDR) photography or multi-frame noise reduction (MFNR). In the case of HDR photography, a first image frame and a second image frame are captured using different exposure times, different apertures, different lenses, and / or other characteristics that can improve the dynamic range of the fused image when combining the two image frames. In some aspects, the method can be performed for MFNR photography, where the first image frame and the second image frame are captured using the same or different exposure times, and the first image frame and the second image frame are fused to generate a corrected first image frame that has reduced noise compared to the captured first image frame.

[0020] In some aspects, the device can include an image signal processor or a processor (e.g., an application processor) that includes specific functionality for camera control and / or processing, such as enabling or disabling a merging module or otherwise controlling aspects of the image correction. The methods and techniques described herein can be performed entirely by the image signal processor or the processor, or the various operations can be split between the image signal processor and the processor, and in some aspects across additional processors.

[0021] The device can include one, two, or more image sensors, such as a first image sensor. When there are multiple image sensors, the configurations of these image sensors can be different. For example, the first image sensor can have a larger field of view (FOV) than the second image sensor, or the first image sensor can have a different sensitivity or a different dynamic range than the second image sensor. In one example, the first image sensor can be a wide-angle image sensor, and the second image sensor can be a telephoto image sensor. In another example, the first sensor is configured to obtain an image through a first lens having a first optical axis, and the second sensor is configured to obtain an image through a second lens having a second optical axis different from the first optical axis. Additionally or alternatively, the first lens can have a first magnification, and the second lens can have a second magnification different from the first magnification. Any of these or other configurations can be part of a lens cluster on a mobile device, such as where multiple image sensors and associated lenses are located at offset positions on the front or back side of the mobile device. Additional image sensors with larger, smaller, or the same field of view can be included. The image processing techniques described herein can be applied to image frames captured from any of the image sensors in a multi-sensor device.

[0022] In additional aspects of the present disclosure, a device configured for image processing and / or image capture is disclosed. The device includes components for capturing image frames. The device also includes one or more components for capturing data representative of a scene, such as image sensors (including charge-coupled devices (CCDs), Bayer filter sensors, infrared (IR) detectors, ultraviolet (UV) detectors, complementary metal-oxide semiconductor (CMOS) sensors) and time-of-flight detectors. The device may also include one or more components for focusing and / or concentrating light onto one or more of the image sensors (including simple lenses, compound lenses, spherical lenses, and aspherical lenses). These components can be controlled to capture a first image frame and / or a second image frame input to the image processing techniques described herein.

[0023] For those of ordinary skill in the art, other aspects, features, and specific implementations will become apparent upon reviewing the following description of specific exemplary aspects in conjunction with the accompanying drawings. Although the features may be discussed below with respect to certain aspects and drawings, various aspects may include one or more of the advantageous features discussed herein. In other words, although one or more aspects may be discussed as having certain advantageous features, one or more of such features may also be used according to various aspects. In a similar manner, although the exemplary aspects may be discussed below as device, system, or method aspects, the exemplary aspects may be implemented in various devices, systems, and methods.

[0024] The method may be embedded in a computer-readable medium as computer program code, the computer program code including instructions that cause a processor to perform the steps of the method. In some embodiments, the processor may be part of a mobile device that includes: a first network adapter configured to send data, such as an image or video as recorded data or as streaming data, over a first network connection of a plurality of network connections; and a processor coupled to the first network adapter and a memory. The processor may cause the output image frames described herein to be sent over a wireless communication network, such as a 5G NR communication network.

[0025] The features and technical advantages of examples in accordance with the present disclosure have been outlined above rather broadly so that the detailed description below may be better understood. Additional features and advantages will be described below. The disclosed concepts and specific examples may be readily utilized as a basis for modifying or designing other structures for achieving the same purposes of the present disclosure. Such equivalent constructions do not depart from the scope of the appended claims. When considered in conjunction with the accompanying drawings, the characteristics (both the organization and method of operation) of the concepts disclosed herein, as well as the associated advantages, will be better understood. Each of the drawings provided is for purposes of illustration and description and not as a definition of the limits of the claims.

[0026] While aspects and specific implementations are described herein by way of some examples, those skilled in the art will understand that additional specific implementations and use cases may arise in many different arrangements and scenarios. The innovations described herein can be implemented across many different platform types, devices, systems, shapes, sizes, and packaging arrangements. For example, aspects and / or uses can be implemented via integrated chips and other non-module component-based devices (e.g., end-user devices, vehicles, communication devices, computing devices, industrial equipment, retail / purchase devices, medical devices, artificial intelligence (AI)-enabled devices, etc.). Although some examples may or may not specifically point to use cases or applications, applicability of various types of the described innovations may occur. The scope of specific implementations can range from chip-level or module components to non-module, non-chip-level implementations and further to aggregated, distributed, or original equipment manufacturer (OEM) devices or systems incorporating one or more aspects of the described innovations. In some practical environments, devices incorporating the described aspects and features may also necessarily include additional components and features for implementing and practicing the claimed and described aspects. For example, the transmission and reception of wireless signals necessarily includes multiple components for analog and digital purposes (e.g., hardware components including antennas, radio frequency (RF) chains, power amplifiers, modulators, buffers, processors, interleavers, adders / summers, etc.). The innovations described herein are intended to be practiced in a variety of devices, chip-level components, systems, distributed arrangements, end-user devices, etc., having different sizes, shapes, and configurations. BRIEF DESCRIPTION OF THE DRAWINGS

[0027] A further understanding of the nature and advantages of the present disclosure can be realized by referring to the following drawings. In the drawings, like components or features may have the same reference numeral. Additionally, various components of the same type can be distinguished by adding a dash and a second numeral used to differentiate between like components after the reference numeral. If only the first reference numeral is used in the specification, the description applies to any one of the like components having the same first reference numeral, regardless of the second reference numeral.

[0028] Figure 1 A block diagram illustrating an example device for performing image capture from one or more image sensors.

[0029] Figure 2 A block diagram illustrating an example data flow path for image data processing in an image capture device in accordance with one or more embodiments of the present disclosure.

[0030] Figure 3 A block diagram of an example implementation of an engine for video analysis and an image processing engine in accordance with an exemplary embodiment of the present disclosure.

[0031] Figure 4 A flowchart illustrating an example method for processing an image frame to correct for motion in accordance with some embodiments of the present disclosure.

[0032] Figure 5 A block diagram illustrating an example processor configuration for image data processing in an image capture device in accordance with one or more embodiments of the present disclosure.

[0033] Like reference numerals and names in the various figures indicate like elements. Detailed Description

[0034] The detailed description set forth below in connection with the appended drawings is intended as a description of various configurations and is not intended to limit the scope of the present disclosure. On the contrary, the detailed description includes specific details for providing a thorough understanding of the subject matter of the present invention. It will be apparent to those skilled in the art that these specific details are not required in every instance and that in some instances, well-known structures and components are shown in block diagram form for clarity of presentation.

[0035] There are various techniques for correcting motion distortion within a captured image frame. For example, motion compensated temporal filtering (McTF) processing allows a video system to reduce noise and motion artifacts within a captured image. However, these techniques are typically computationally resource intensive. Additionally, a video processing pipeline typically requires real-time or near real-time processing of captured image frames to ensure that the video pipeline can keep up with incoming image frames to continue capturing video data. Thus, the intensive level of computational resources required by McTF and other techniques can reduce the device battery life when capturing video. Additionally, the processing constraints of a mobile computing device can limit the total amount of computational resources that can be provided to the video processing pipeline. Thus, the limited resources can also limit the amount of video that can be captured (e.g., can limit video capture resolution, video capture frame rate, bit rate, etc.).

[0036] The disadvantages mentioned here are merely representative and are included to emphasize the problems that the inventors have identified and sought to improve with respect to existing devices. Aspects of the devices described below may address some or all of these disadvantages as well as other disadvantages known in the art. Aspects of the improved devices described herein may present other benefits different from those described above and may be used in other applications different from those described above.

[0037] The present disclosure provides systems, apparatuses, methods, and computer-readable media for supporting image processing, including techniques for improved processing and correction of motion distortion in a captured image frame caused by movement of an object within the image frame. Specifically, an image transformation matrix may be generated to correct motion distortion within the received image frame. Additionally, a motion map that is calculated to indicate movement of an object within the image frame may be analyzed to identify one or more motion hotspots within the image frame. The motion hotspots may then be used to constrain the application of a temporal filtering process to the image frame. For example, the temporal filtering process may be applied only to portions of the image frame where the motion hotspots indicate movement greater than a predetermined threshold.

[0038] Certain specific implementations of the subject matter described in this disclosure can be realized to achieve one or more of the following potential advantages or benefits. In some aspects, the present disclosure provides techniques for reducing the total amount of computing resources required to detect and correct motion distortion in a captured image frame. Specifically, by restricting the application of one or more motion distortion correction techniques to only include regions where sufficient motion has occurred, the techniques can reduce the total proportion of the image that needs to be modified or transformed to correct motion distortion. This can result in improved device battery life for image capture devices implementing these techniques. Similarly, these techniques can reduce overall device heat and associated cooling requirements. Relatedly, by reducing the amount of image processing that needs to be performed, the techniques can improve the capabilities of an image capture device to capture processed video data, for example, at a higher resolution, higher frame rate, higher bit rate, etc.

[0039] Example devices for capturing image frames using one or more image sensors, such as smart phones, may include a configuration of one, two, three, four, or more cameras on the back side (e.g., the side opposite the main user display) and / or the front side (e.g., the side same as the main user display) of the device. These devices may include one or more image signal processors (ISPs), computer vision processors (CVPs) (e.g., AI engines), or other suitable circuitry for processing images captured by the image sensors. The one or more image signal processors (ISPs) may store the output image frames in a memory and / or otherwise provide the output image frames to the processing circuitry (such as via a bus). The processing circuitry may perform further processing, such as encoding, storing, transmitting, or other manipulation of the output image frames.

[0040] As used herein, an image sensor may refer to the image sensor itself and any particular other components coupled to the image sensor for generating image frames for processing by an image signal processor or other logic circuitry or for storage in a memory, whether a short-term buffer or a long-term non-volatile memory. For example, an image sensor may include other components of a camera, including a shutter, a buffer, or other readout circuitry for accessing the individual pixels of the image sensor. An image sensor may also refer to an analog front end or other circuitry for converting an analog signal to a digital representation of an image frame, which digital representation is provided to digital circuitry coupled to the image sensor.

[0041] In the description of the embodiments herein, numerous specific details (such as examples of specific components, circuits, and processes) are set forth to provide a thorough understanding of the present disclosure. As used herein, the term "coupled" means directly connected or connected through one or more intermediate components or circuits. Additionally, in the following description and for purposes of explanation, specific terms are set forth to provide a thorough understanding of the present disclosure. However, those skilled in the art will appreciate that implementing the teachings disclosed herein may not require these specific details. In other instances, well-known circuits and devices are shown in block diagram form to avoid obscuring the teachings of the present disclosure.

[0042] Certain portions of the detailed description that follows are presented in terms of procedures, logic blocks, processing, and other symbolic representations of operations on data bits within a computer memory. In the present disclosure, procedures, logic blocks, processes, etc. are conceived of as a self-consistent sequence of steps or instructions leading to a desired result. These steps are those requiring physical manipulation of physical quantities. Although not necessarily, typically, these physical quantities take the form of electrical or magnetic signals capable of being stored, transferred, combined, compared, and otherwise manipulated in a computer system.

[0043] In the figures, a single box may be described as performing one or more functions. The one or more functions performed by the box may be performed in a single component or across multiple components and / or may be performed using hardware, software, or a combination of hardware and software. To clearly illustrate this interchangeability of hardware and software, various illustrative components, blocks, modules, circuits, and steps are described hereinafter in terms of their functionality. Whether such functionality is implemented as hardware or software depends on the particular application and the design constraints imposed on the overall system. Those skilled in the art may implement the described functionality in different ways for each particular application, but such specific implementation decisions should not be construed as causing a departure from the scope of the present disclosure. Moreover, an example device may include components other than those shown, including well-known components such as processors, memories, and the like.

[0044] Aspects of the present disclosure are applicable to any electronic device that includes, is coupled to, or otherwise processes data from one, two, or more image sensors capable of capturing image frames (or "frames"). The terms "output image frame" and "corrected image frame" may refer to an image frame that has been processed by any of the techniques discussed herein. Additionally, aspects of the present disclosure may be implemented in image sensors or devices coupled to the image sensors having the same or different capabilities and characteristics, such as resolution, shutter speed, sensor type, etc. Further, aspects of the present disclosure may be implemented in a device for processing image frames, whether or not the device includes or is coupled to an image sensor, such as a processing device that can retrieve stored images for processing, including processing devices present in a cloud computing system.

[0045] Unless otherwise specifically stated, it should be understood from the following discussion that, throughout this application, discussions using terms such as "access", "receive", "transmit", "use", "select", "determine", "normalize", "multiply", "average", "monitor", "compare", "apply", "update", "measure", "derive", "set", "generate", etc. refer to actions and processes of a computer system or similar electronic computing device that manipulate and transform data represented as physical (electronic) quantities within the registers and memories of the computer system into other data similarly represented as physical quantities within the registers, memories, or other such information storage, transmission, or display devices of the computer system.

[0046] The terms "device" and "apparatus" are not limited to one or a specific number of physical objects (e.g., a smart phone, a camera controller, a processing system, etc.). As used herein, a device can be any electronic device having one or more parts capable of implementing at least some portions of the present disclosure. Although the description and examples herein use the term "device" to describe various aspects of the present disclosure, the term "device" is not limited to a particular configuration, type, or number of objects. As used herein, an apparatus can include a device or a part of a device for performing the described operations.

[0047] Certain components in a device or apparatus described as "components for accessing", "components for receiving", "components for transmitting", "components for using", "components for selecting", "components for determining", "components for normalizing", "components for multiplying", or other similarly named terms referring to one or more operations on data (such as image data) may refer to processing circuitry (e.g., an application specific integrated circuit (ASIC), a digital signal processor (DSP), a graphics processing unit (GPU), a central processing unit (CPU)) configured to perform the recited functions by a combination of hardware, software, or hardware configured by software.

[0048] Figure 1 FIG. 2 shows a block diagram of an example device 100 for performing image capture from one or more image sensors. Device 100 may include or otherwise be coupled to an image signal processor 112 for processing image frames from one or more image sensors such as a first image sensor 101, a second image sensor 102, and a depth sensor 140. In some embodiments, device 100 also includes or is coupled to a processor 104 and a memory 106 storing instructions 108. Device 100 may also include or be coupled to a display 114 and input / output (I / O) components 116. The I / O components 116 may be used to interact with a user, such as a touchscreen interface and / or physical buttons.

[0049] The I / O components 116 may also include a network interface for communicating with other devices including a wide area network (WAN) adapter 152, a local area network (LAN) adapter 153, and / or a personal area network (PAN) adapter 154. An example WAN adapter is a 4G LTE or 5G NR wireless network adapter. An example LAN adapter 153 is an IEEE 802.11 WiFi wireless network adapter. An example PAN adapter 154 is a Bluetooth wireless network adapter. Each of the adapters 152, 153, and / or 154 may be coupled to an antenna that includes multiple antennas configured for main set reception and diversity reception and / or configured to receive a specific frequency band.

[0050] Device 100 may also include or be coupled to a power source 118 for device 100, such as a battery or a component that couples device 100 to an energy source. Device 100 may also include or be coupled to Figure 1 additional features or components not shown. In one example, a wireless interface that may include multiple transceivers and a baseband processor may be coupled to or included in the WAN adapter 152 for a wireless communication device. In another example, an analog front end (AFE) for converting analog image frame data to digital image frame data may be coupled between the image sensors 101 and 102 and the image signal processor 112.

[0051] The device may include or be coupled to a sensor hub 150 that interfaces with sensors to receive data about the movement of the device 100, data about the environment surrounding the device 100, and / or other non-camera sensor data. An example non-camera sensor is a gyroscope, i.e., a device configured to measure rotation, orientation, and / or angular velocity to generate motion data. Another example non-camera sensor is an accelerometer, i.e., a device configured to measure acceleration, which can also be used to determine velocity and distance traveled by appropriately integrating the measured acceleration, and one or more of acceleration, velocity, and / or distance may be included in the generated motion data. In some aspects, the gyroscope in an electronic image stabilization system (EIS) may be coupled to the sensor hub or directly to the image signal processor 112. In another example, the non-camera sensor may be a global positioning system (GPS) receiver.

[0052] The image signal processor 112 may receive image data such as for forming image frames. In one implementation, a local bus connection couples the image signal processor 112 to the image sensors 101 and 102 of the first camera 103 and the second camera 105, respectively. In another implementation, a wired interface couples the image signal processor 112 to an external image sensor. In yet another implementation, a wireless interface couples the image signal processor 112 to the image sensors 101, 102.

[0053] The first camera 103 may include a first image sensor 101 and a corresponding first lens 131. The second camera may include a second image sensor 102 and a corresponding second lens 132. Each of the lenses 131 and 132 may be controlled by an associated autofocus (AF) algorithm 133 executed in the ISP 112, which adjusts the lenses 131 and 132 to focus on a specific focal plane at a certain scene depth from the image sensors 101 and 102. The AF algorithm 133 may be assisted by a depth sensor 140.

[0054] The first image sensor 101 and the second image sensor 102 are configured to capture one or more image frames. Lenses 131 and 132 focus light at image sensors 101 and 102 respectively through one or more apertures for receiving light, one or more shutters for blocking light when outside the exposure window, one or more color filter arrays (CFAs) for filtering light outside a specific frequency range, one or more analog front ends for converting analog measurements to digital information, and / or other suitable components for imaging. The first lens 131 and the second lens 132 may have different fields of view to capture different representations of a scene. For example, the first lens 131 may be an ultra-wide (UW) lens and the second lens 132 may be a wide (W) lens. The plurality of image sensors may include a combination of ultra-wide (high field of view (FOV)) sensors, wide sensors, tele sensors, and ultra-tele (low FOV) sensors.

[0055] That is, each image sensor can be configured through hardware configuration and / or software settings to obtain different but overlapping fields of view. In one configuration, the image sensor is configured with different lenses having different magnifications, which results in different fields of view. The sensors can be configured such that the UW sensor has a larger FOV than the W sensor, the W sensor has a larger FOV than the T sensor, and the T sensor has a larger FOV than the UT sensor. For example, a sensor configured for a wide FOV may capture a field of view in the range of 64 degrees to 84 degrees, a sensor configured for an ultra-side FOV may capture a field of view in the range of 100 degrees to 140 degrees, a sensor configured for a tele FOV may capture a field of view in the range of 10 degrees to 30 degrees, and a sensor configured for an ultra-tele FOV may capture a field of view in the range of 1 degree to 8 degrees.

[0056] The camera 103 can be a variable aperture (VA) camera, where the aperture can be controlled to a specific size. Example aperture sizes are f / 2.0, f / 2.8, f / 3.2, f / 8.0, etc. Larger aperture values correspond to smaller aperture sizes, and smaller aperture values correspond to larger aperture sizes. The camera 103 may have different characteristics based on the current aperture size, such as different depths of field (DOF) at different aperture sizes.

[0057] The image signal processor 112 processes the image frames captured by the image sensors 101 and 102. Although Figure 1Exemplarily, device 100 includes two image sensors 101 and 102 coupled to image signal processor 112, but any number (e.g., one, two, three, four, five, six, etc.) of image sensors may be coupled to image signal processor 112. In some aspects, a depth sensor such as depth sensor 140 may be coupled to image signal processor 112 and process the output from the depth sensor in a manner similar to that of image sensors 101 and 102. Example depth sensors include active sensors, including one or more of indirect time-of-flight (iToF), direct time-of-flight (dToF), light detection and ranging (LiDAR), mmWave, radio detection and ranging (radar), and / or hybrid depth sensors (such as structured light). In embodiments without depth sensor 140, similar information regarding the depth or depth map of an object can be generated passively from the disparity between two image sensors (e.g., using disparity depth measurement or stereo depth measurement), phase detection autofocus (PDAF) sensors, etc. Additionally, there may be any number of additional image sensors or image signal processors for device 100.

[0058] In some embodiments, image signal processor 112 may execute instructions from a memory, such as instructions 108 from memory 106, instructions stored in a separate memory coupled to or included in image signal processor 112, or instructions provided by processor 104. Additionally or alternatively, image signal processor 112 may include specific hardware (such as one or more integrated circuits (ICs)) configured to perform one or more operations described in this disclosure. For example, image signal processor 112 may include one or more image front ends (IFE) 135, one or more image post-processing engines 136 (IPE), one or more automatic exposure compensation (AEC) 134 engines, and / or one or more engines for video analysis (EVA). AF 133, AEC 134, IFE 135, IPE 136, and EVA 137 may each include dedicated circuitry, may be embodied as software code executed by ISP 112, and / or a combination of hardware and software code executed on ISP 112.

[0059] In some embodiments, the memory 106 may include a non-transitory or non-volatile computer-readable medium storing computer-executable instructions 108 to perform all or a portion of one or more operations described in the present disclosure. In some embodiments, the instructions 108 include a camera application (or other suitable application) for generating an image or video to be executed by the device 100. The instructions 108 may also include other applications or programs to be executed by the device 100, such as an operating system and specific applications other than those for image or video generation. A camera application, such as executed by the processor 104, may cause the device 100 to generate an image using the image sensors 101 and 102 and the image signal processor 112. The memory 106 may also be accessed by the image signal processor 112 to store processed frames or may be accessed by the processor 104 to obtain the processed frames. In some embodiments, the device 100 does not include the memory 106. For example, the device 100 may be a circuit including the image signal processor 112, and the memory may be external to the device 100. The device 100 may be coupled to an external memory and configured to access the memory to write output frames for display or long-term storage. In some embodiments, the device 100 is a system-on-chip (SoC) that combines the image signal processor 112, the processor 104, the sensor hub 150, the memory 106, and the input / output component 116 into a single package.

[0060] In some embodiments, at least one of the image signal processor 112 or the processor 104 executes instructions to perform the various operations described herein, including motion distortion detection and correction operations. For example, the execution of the instructions may instruct the image signal processor 112 to start or end capturing an image frame or a sequence of image frames, where the capture includes motion distortion or other motion errors as described in the embodiments herein. In some embodiments, the processor 104 may include one or more general-purpose processor cores 104A capable of executing a script or instructions of one or more software programs (such as the instructions 108 stored in the memory 106). For example, the processor 104 may include one or more application processors configured to execute a camera application (or other suitable application for generating an image or video) stored in the memory 106.

[0061] When executing a camera application, the processor 104 may be configured to instruct the image signal processor 112 to perform one or more operations with reference to the image sensor 101 or 102. For example, a camera application executed on the processor 104 may receive a user command to start a video preview display. Upon receiving the user command, the image signal processor 112 is used to capture and process a video including a sequence of image frames from one or more of the image sensors 101 or 102. Image processing such as generating an "output" or "corrected" image frame according to the techniques described herein may be applied to one or more of the image frames in the sequence. Executing instructions 108 outside the camera application by the processor 104 may also cause the device 100 to perform any number of functions or operations. In some embodiments, the processor 104 may include an IC or other hardware (e.g., an artificial intelligence (AI) engine 124 or other coprocessor) to offload certain tasks from the core 104A. The AI engine 124 may be used to offload tasks related to, for example, face detection and / or object recognition. In some other embodiments, the device 100 does not include the processor 104, such as when all the described functionality is configured in the image signal processor 112.

[0062] In some embodiments, the display 114 may include one or more suitable displays or screens that allow user interaction and / or present items to the user, such as a preview of the image frames captured by the image sensors 101 and 102. In some embodiments, the display 114 is a touch-sensitive display. The I / O component 116 may be or include any suitable mechanism, interface, or device to receive input (such as commands) from the user and provide output to the user via the display 114. For example, the I / O component 116 may include (but is not limited to) a graphical user interface (GUI), a keyboard, a mouse, a microphone, a speaker, a squeezable bezel, one or more buttons (such as a power button), a slider, a switch, etc.

[0063] Although shown as being coupled to each other via the processor 104, components such as the processor 104, the memory 106, the image signal processor 112, the display 114, and the I / O component 116 may be coupled to each other in various other arrangements, such as being coupled to each other via one or more local buses, which are not shown for simplicity. Although the image signal processor 112 is illustrated as being separate from the processor 104, the image signal processor 112 may be a core of the processor 104, which is an application processor unit (APU), included in a system-on-chip (SoC), or otherwise included with the processor 104. Although reference is made to the device 100 in the examples herein to perform aspects of the present disclosure, some device components may not be Figure 1is shown to prevent obscuring aspects of the present disclosure. Additionally, other components, the number of components, or combinations of components may be included in a suitable device for performing aspects of the present disclosure. Accordingly, the present disclosure is not limited to a particular device or component configuration, including device 100.

[0064] Figure 1 An exemplary image capture device may be operated to obtain improved images and / or video by adjusting the application of one or more motion distortion correction techniques (e.g., McTF) based on motion hotspots detected within a received image frame. A motion hotspot refers to a portion of an image frame that meets at least one criterion regarding motion, such as a determined amount of motion above a threshold, which may be determined based on one or more motion vectors corresponding to the portion of the image frame. Figure 2 An example method of operating one or more cameras (such as camera 103) is shown and described below.

[0065] Figure 2 is a block diagram illustrating an example data flow path for image data processing in an image capture device according to one or more embodiments of the present disclosure. The processor 104 of system 200 may communicate with an image signal processor (ISP) 112 via a bidirectional bus and / or separate control and data lines. The processor 104 may control the camera 103 via the camera control 210, such as for configuring the camera 103 via a driver executed on the processor 104. The camera control 210 may be managed by a camera application 204 executed on the processor 104, which provides user-accessible settings such that a user may specify individual camera settings or select a profile with corresponding camera settings. The camera control 210 communicates with the camera 103 to configure the camera 103 according to commands received from the camera application 204. The camera application 204 may be, for example, a photography application, a document scanning application, a messaging application, or other application that processes image data obtained from the camera 103.

[0066] Camera configuration may be parameters that specify, for example, frame rate, image resolution, readout duration, exposure level, aspect ratio, aperture size, etc. The camera 103 may obtain image data based on the camera configuration. For example, the processor 104 may execute the camera application 204 to instruct the camera 103 via the camera control 210 to set a first camera configuration for the camera 103, obtain first image data from the camera 103 operating in the first camera configuration, instruct the camera 103 to set a second camera configuration for the camera 103, and obtain second image data from the camera 103 operating in the second camera configuration.

[0067] In some embodiments in which camera 103 is a variable aperture (VA) camera system, processor 104 may execute camera application 204 to instruct camera 103 to be configured to a first aperture size, obtain first image data from camera 103, instruct camera 103 to be configured to a second aperture size, and obtain second image data from camera 103. Reconfiguration of the aperture and the obtaining of the first and second image data may occur with little or no change in the scene that can be captured at the first aperture size and at the second aperture size. Example aperture sizes are f / 2.0, f / 2.8, f / 3.2, f / 8.0, etc. Larger aperture values correspond to smaller aperture sizes, and smaller aperture values correspond to larger aperture sizes. That is, f / 2.0 is a larger aperture size than f / 8.0.

[0068] The image data received from camera 103 may be processed in one or more blocks of ISP 112 to form image frame 230 stored in memory 106 and / or provided to processor 104. Image frame 230 output by ISP 112 may have motion compensated temporal filtering (McTF) selectively applied to the motion hotspots and / or other regions of the image frame. The McTF applied by ISP 112 may be applied to determine a corrected image frame based on a first image frame and a second image frame received in the image data from first camera 103. Processor 104 may further process the image data to apply effects to image frame 230. The effects may include Bokeh, lighting, color cast, and / or high dynamic range (HDR) merging. In some embodiments, the functionality may be embedded in different components, such as ISP 112, DSP, ASIC, or other custom logic circuitry for performing additional image processing.

[0069] Figure 3FIG. 300 is a block diagram of an example implementation of a system 300 that includes an Engine for Video Analysis (EVA) 137 and an Image Processing Engine (IPE) 136 according to an exemplary implementation of the present disclosure. In system 300, EVA 137 receives a first image frame 302, a second image frame 304, and sensor data 306. The first image frame 302 and the second image frame 304 may be received as image data from an image sensor (such as image sensors 101, 102). In some instances, EVA 137 may be configured to receive and sequentially process the image frames (e.g., as part of an image processing pipeline and / or a video processing pipeline). For example, EVA 137 may initially receive the first image frame 302 (e.g., after being captured by image sensors 101, 102), and it may subsequently receive the image frame 304 (e.g., after being subsequently captured by image sensors 101, 102). In some instances, the second image frame 304 may be considered the current image frame, or the target image frame of the image processing or video processing pipeline, and the first image frame 302 may be considered the reference frame of the pipeline. In additional or alternative implementations, EVA 137 may receive a plurality of image frames including image frames 302, 304. For example, the plurality of image frames may have been previously stored and may be retrieved by EVA 137 for further processing. The sensor data 306 may be received from a sensor hub (such as sensor hub 150). The sensor data 306 may include data from one or more motion sensors (such as an accelerometer and a gyroscope). The sensor data 306 may reflect the motion of the device (such as device 100) that captured the image frames 302, 304.

[0070] EVA 137 includes a motion map engine 308 that can be configured to generate a motion map 316. The motion map 316 can be generated to indicate motion within a second image frame 304 relative to a first image frame 302. In some embodiments, the motion map 316 can be calculated based on the difference between the second image frame 304 and the first image frame 302. In additional or alternative embodiments, the motion map 316 can be calculated based on sensor data 306 (e.g., motion sensor data) from a device 100 (e.g., an image capture device) that captured the first image frame 302 and the second image frame 304. For example, the motion map engine 308 can be configured to calculate one or more global motion estimates that reflect the movement of the image capture device and local motion estimates that reflect the movement of one or more objects depicted within the image frames 302, 304. In some instances, the global motion estimate can be calculated based on sensor data 306 such as gyroscope or accelerometer data that indicates the movement of the image capture device. The global motion estimate can include one or both of the magnitude and direction of the movement of the image capture device. The local motion estimate can be captured by comparing the first image frame 302 and the second image frame 304. For example, the local motion estimate can be determined by comparing the positions of one or more objects within each of the first image frame 302 and the second image frame 304. The local motion estimate can be calculated as the difference between the image frames 302, 304. As another example, the local motion estimate can be calculated based on texture processing using Harris corner detection and related techniques. The local motion estimate and / or the global motion estimate can then be combined to generate the motion map 316.

[0071] In some embodiments, the motion map 316 can be implemented separately from the second image frame 304 (e.g., implemented as a separate data structure corresponding to the second image frame 304). In additional or alternative embodiments, the motion map 316 can be implemented as part of the second image frame 304 (e.g., implemented as a data layer or metadata layer of the second image frame 304). Additionally, the content of the motion map 316 can correspond to a specific portion of the second image frame 304. For example, each pixel of the second image frame 304 can have a corresponding entry in the motion map 316. As another example, each entry in the motion map 316 can correspond to multiple pixels (e.g., 4 pixels, 9 pixels, 16 pixels, or more). The content of the motion map 316 can indicate motion within the corresponding portion of the second image frame 304. For example, an entry in the motion map 316 can indicate the magnitude of the movement of an object or feature depicted within the corresponding portion of the second image frame 304 (e.g., relative to the first image frame 302). In various embodiments, the magnitude of the movement can be indicated as one or more of the number of pixels moved, the distance moved, the speed of movement, etc. In some embodiments, the entry can also indicate the direction of the movement.

[0072] In a particular implementation, the motion map 316 can be calculated by the motion map engine 308 by comparing the second image frame 304 with the first image frame 302 to generate an estimate of the motion vectors between the image frames 302, 304. The motion vectors and the sensor data 306 can then be analyzed together to determine the alignment of the image capture device (e.g., to separate the global motion of the image capture device from the local motion of the objects depicted within the image frames 302, 304). The motion map engine 308 can then perform a matching process based on the alignment and the image frames 302, 304 to generate the motion map 316. In some implementations, the matching process can be performed as a semi-global matching (SGM) process.

[0073] The motion map 316 can then be used to generate an image transformation matrix 318. For example, the image transformation matrix 318 can be generated to indicate the image transformation that should be applied to the second image frame 304 to correct for motion distortion within the second image frame 304. The image transformation matrix 318 can include an image transformation based on individual pixels within the second image frame 304 and / or one or more adjacent pixels within the image frame 304. Additionally or alternatively, the image transformation matrix 318 can include a transformation based on other image frames (e.g., the image frame 302 captured before the image frame 304 and / or an image frame captured after the image frame 304). In various implementations, the image transformation matrix 318 can include the same or similar transformation for each pixel and / or portion of the second image frame 304. In additional or alternative implementations, the image transformation matrix 318 can indicate different transformations for different pixels and / or different portions of the second image frame 304. In some implementations, the motion map 316 can be generated according to one or more motion distortion correction techniques such as McTF. In some implementations, the image transformation matrix 318 can be generated to reverse the movement of a pixel block (such as an 8-pixel by 8-pixel block) or other image features from the second image frame relative to the first image frame 302.

[0074] The IPE 136 can receive both the motion map 316 and the image transformation matrix 318. Specifically, the IPE 136 includes a motion hot spot engine 312 that can receive the motion map 316 and an McTF engine 314 that can receive the image transformation matrix 318. The motion hot spot engine 312 can be configured to identify motion hot spots 320 within the second image frame 304. The motion hot spots 320 can represent regions within the second image frame 304 that have significant motion and / or motion distortion. Additionally or alternatively, the motion hot spots 320 can represent regions within the second image frame 304 where motion distortion may exist. In some embodiments, the motion hot spots 320 can include or otherwise identify the locations within the second image frame 304 that have motion greater than or equal to a predetermined threshold. Specifically, in some embodiments, the motion map 316 can indicate the motion within the second image frame 304 as the number of pixels of movement of the corresponding portion of the second image frame 304 between the first image frame 302 and the second image frame 304. In such embodiments, the motion hot spots 320 can be implemented as regions where the movement indicated by the motion map 316 is greater than a predetermined number of pixels (e.g., 1 pixel, 2 pixels, 4 pixels, 10 pixels, etc.).

[0075] In various embodiments, the motion hot spots 320 can be stored in a format similar to the format of the motion map 316. In some embodiments, the motion hot spots 320 can be implemented separately from the second image frame 304 (e.g., implemented as a separate data structure corresponding to the second image frame 304). In additional or alternative embodiments, the motion hot spots 320 can be implemented as part of the second image frame 304 (e.g., implemented as a data layer or metadata layer of the second image frame 304). Additionally, the content of the motion hot spots 320 can correspond to a specific portion of the second image frame 304. For example, each pixel of the second image frame 304 can have a corresponding entry in the motion characteristics 320. As another example, each entry in the motion hot spots 320 can correspond to multiple pixels (e.g., 4 pixels, 9 pixels, 16 pixels, or more). In additional or alternative embodiments, the motion hot spots 320 can be generated only to include an indication of the locations of the motion hot spots 320 within the second image frame 304. For example, the motion hot spots 320 can include the coordinates or other identifiers of a polygon or other bounded region that includes the motion hot spots within the second image frame 304.

[0076] The McTF engine 314 can be configured to receive a motion hotspot 320 and an image transformation matrix 318, and generate a corrected image frame 324. Specifically, the McTF 314 can be configured to apply a temporal filtering process (such as the McTF process) to the second image frame 304 to generate the corrected image frame 324. In some embodiments, applying the temporal filtering process may include applying the image transformation matrix 318 to all or part of the second image frame 304. As explained above, applying the temporal filtering process to the entire second image frame 304 may utilize excessive computing resources, thereby reducing the device battery life and the image frame processing ability. Instead, the McTF engine 314 can be configured to apply the temporal filtering process only to the pixels of the second image frame 304 that are located within the motion hotspot 320. For example, the McTF engine 314 can analyze each pixel within the second image frame 304 and can determine whether the pixel is included within the motion hotspot based on the corresponding portion of the motion hotspot 320. If the pixel is included within the motion hotspot, the McTF engine 314 can apply the temporal filtering process to the pixel. The McTF engine 314 can traverse all the pixels within the second image frame 304 to generate the corrected image frame 324. As a specific example, Figure 3 depicts an exemplary image frame with the identified exemplary motion hotspot 332. The McTF engine 314 can apply the temporal filtering process to these regions within the second image frame to generate the corrected image frame 324.

[0077] In additional or alternative embodiments, the McTF engine 314 can alternatively analyze the motion hotspot 320 and can determine the corresponding locations (e.g., corresponding pixels) of the second image frame 304 that are included within the motion hotspot 320. The McTF engine 314 can then apply the image transformation matrix 318 to the corresponding locations to generate the corrected image frame 324. Once generated, the IPE 136 can add the corrected image frame 324 to the output image frame 330 (e.g., the output image frame of a video and / or a composite image). In some embodiments, the corrected image frame 324 can also be used as a reference image frame for correcting future image frames. For example, the first image frame 302 can be an image frame that was previously corrected by the EVA 137 and the IPE 136 using the techniques discussed above.

[0078] Figure 2 of the system 200 and / or Figure 3 of the system 300 can be configured to perform the operations described in reference Figure 4 to determine the output image frames 230, 330 (e.g., for a video sequence). Figure 4 shows a flowchart of an example method 400 for processing image frames to correct motion according to some embodiments of the present disclosure. Figure 4The capture in [the device] can obtain an improved digital representation of the scene, thereby producing a photo or video with higher image quality (IQ).

[0079] At block 402, a first image frame and a second image frame are received. The first image frame 302 and the second image frame 304 can be received from image sensors 101, 102, such as when the image sensors are configured with a camera configuration. The first image frame and the second image frame can be received at the ISP 112, processed by the image front end (IFE) and / or the image post - processing engine (IPE) of the ISP 112, and stored in the memory. In some embodiments, the capture of the image data can be initiated by a camera application executing on the processor 104, which causes the camera control 210 to activate the camera 103 to capture the image data and causes the image data to be provided to a processor, such as the processor 104 or the ISP 112. In certain instances, the first image frame 302 and the second image frame 304 can be received in sequence. For example, the first image frame 302 can be received before the second image frame 304.

[0080] At block 404, the second image frame is compared with the first image frame to calculate a motion map. For example, the second image frame 304 can be compared with the first image frame 302 to calculate the motion map 316. As further explained above, the motion map 316 can be calculated based on one or both of local motion estimation and global motion estimation. For example, the local motion estimation can be calculated based on the change in the position of one or more objects and / or features within the second image frame 304 and the first image frame 302. Additionally or alternatively, the global motion estimation can be calculated based on the sensor data 306 reflecting the movement of the device 100.

[0081] At block 406, an image transformation matrix for the second image frame is determined. For example, the image transformation matrix 318 for the second image frame 304 can be determined by the EVA 137 and / or the transformation matrix engine 310. Specifically, the image transformation matrix 318 can be determined to correct one or more motion - artifact falsehoods or other errors within the second image frame 304. In various embodiments, the image transformation matrix 318 can be calculated according to the McTF technique.

[0082] At block 408, motion hotspots are identified within the second image frame. For example, the motion hotspot engine 312 and / or the IPE 136 can identify motion hotspots 320 within the second image frame 304. The motion hotspots 320 can be identified as the positions within the second image frame 304 that contain significant motion and / or motion - artifact falsehoods. As further explained above, in various embodiments, the motion hotspots 320 can be identified as the positions within the second image frame 304 corresponding to the portions of the motion map 316 that indicate motion exceeding a predetermined threshold.

[0083] At block 410, a temporal filtering process is applied to pixels located within a motion hotspot. For example, the McTF engine 314 and / or the IPE 136 may apply a temporal filtering process, such as the McTF process, to pixels of the second image frame 304 that are located within the motion hotspot 320 to generate a corrected image frame 324. In some embodiments, applying the temporal filtering process may include applying an image transformation matrix 318 to the second image frame 302 (such as pixels of the second image frame 304 that are located within the motion hotspot 320).

[0084] At block 412, the corrected image frame is added to an output file. For example, the IPE 136 and / or the device 100 may add the corrected image frame 324 to an output file that includes one or more output image frames 330. The output image frames 230, 330 may be determined by the processor 104 or the ISP 112 and stored in the memory 106. The stored image frames may be read by the processor 104 and used to form a preview display on a display of the device 100 and / or processed to form a photograph for storage in the memory 106 and / or sent to another device. In some embodiments, the output image frame 330 may correspond to a video file. In such instances, the corrected image frame 324 may be appended as a frame of the video file (e.g., as the next sequential image frame of the video file).

[0085] Accordingly, method 400 implements improved detection and correction of motion distortion errors and artifacts within image frames captured by a computing device. Specifically, method 400 reduces the overall computational resources required to correct motion distortion artifacts and other errors. These techniques may correspondingly improve device battery life and increase the capacity of the image processing pipeline within the computing device, thereby increasing the resolution, size, and / or frame rate of image frames and / or video frames that can be processed and corrected.

[0086] Figure 5 is a block diagram illustrating an example processor configuration for image data processing in an image capture device according to one or more embodiments of the present disclosure. The processor 104 or other processing circuitry may be configured to operate on image data to perform Figure 4 one or more operations of the method. The image data may be processed to determine one or more output image frames 510. In Figure 5 the processor 104 implements a motion mapper 502, an image transformation generator 504, a motion hotspot identifier 506, and an image transformer 508.

[0087] The processor 104 is configured to receive first image data and second image data. The first image data may represent a first image frame, and the second image data may represent a second image frame. The image data may be captured by an image sensor. The motion mapper 502 may be configured to compare the second image frame with the first image frame to calculate a motion map of the second image frame. Specifically, the motion map may be calculated to indicate the movement of one or more objects within the second image frame relative to the first image frame. Additionally or alternatively, the motion map may be calculated to indicate the movement of the device 100 that captured the first image frame and the second image frame. The image transformation generator 504 may be configured to determine an image transformation matrix based on the motion map to correct the second image frame. Specifically, the image transformation matrix may be determined to correct motion distortion and / or motion artifacts within the second image frame. In certain embodiments, the image transformation matrix 318 may be calculated according to the McTF technique. The motion hot spot identifier 506 may be configured to identify motion hot spots within the second image frame based on the motion map. The image transformer 508 may be configured to generate a corrected image frame based on the second image frame, the image transformation matrix, and the motion hot spots. Specifically, the image transformer 508 may be configured to apply a temporal filtering process to a portion of the second image frame that is within the motion hot spots to generate the corrected image frame. The corrected image frame may then be added to the output image frame 510.

[0088] In one or more aspects, techniques for supporting image processing may include additional aspects, such as any single aspect or any combination of aspects described below or in connection with one or more other processes or devices described elsewhere herein. In a first aspect, the techniques described herein relate to a method that includes: receiving a first image frame and a second image frame; determining a motion map based on the second image frame and the first image frame; determining a region of the second image frame having motion that meets at least one criterion; and determining a corrected image frame by applying a temporal filtering process to the region.

[0089] In a second aspect according to the first aspect, the at least one criterion includes movement exceeding a predetermined threshold.

[0090] In a third aspect according to at least one of the first aspect to the second aspect, the method is performed by an image signal processor that includes an image processing engine and an engine for video analysis.

[0091] In a fourth aspect according to the third aspect, applying the temporal filtering process to the region includes determining an image transformation matrix, and the image transformation matrix is determined by the image processing engine, and the region is determined by the engine for video analysis.

[0092] In a fifth aspect according to at least one of the first to fourth aspects, the method is performed as part of an image processing pipeline, and the first image frame is a reference image frame of the image processing pipeline, and the second image frame is a current image frame of the image processing pipeline.

[0093] In a sixth aspect according to at least one of the first to fifth aspects, the motion map reflects movement within the second image frame relative to the first image frame.

[0094] In a seventh aspect according to the sixth aspect, determining the motion map includes determining the difference between the second image frame and the first image frame.

[0095] In an eighth aspect according to at least one of the sixth to seventh aspects, determining the motion map is based on motion sensor data from an image capture device that captured the first image frame and the second image frame.

[0096] In a ninth aspect according to at least one of the first to eighth aspects, the first image frame was previously transformed to correct for motion errors before being compared with the second image frame to calculate the motion map of the second image frame.

[0097] In a tenth aspect according to at least one of the first to ninth aspects, the image transformation matrix is generated according to a motion compensation temporal filtering process.

[0098] In an eleventh aspect, there is provided an apparatus including: a memory that stores processor-readable code; and at least one processor coupled to the memory, the at least one processor being configured to execute the processor-readable code to cause the at least one processor to perform operations including: receiving a first image frame and a second image frame; determining a motion map based on the second image frame and the first image frame; determining a region of the second image frame having motion that satisfies at least one criterion based on the motion map; and determining a corrected image frame by applying a temporal filtering process to the region.

[0099] Additionally, the apparatus may perform or operate according to one or more aspects as described hereinbelow. In some specific implementations, the apparatus includes a wireless device, such as a UE. In some specific implementations, the apparatus includes a remote server (such as a cloud-based computing solution) that receives image data for processing to determine an output image frame. In some specific implementations, the apparatus may include at least one processor and a memory coupled to the processor. The processor may be configured to perform the operations described herein for the apparatus. In some other specific implementations, the apparatus may include a non-transitory computer-readable medium having program code recorded thereon, and the program code may be executable by a computer to cause the computer to perform the operations described herein with reference to the apparatus. In some specific implementations, the apparatus may include one or more components configured to perform the operations described herein. In some specific implementations, a method of wireless communication may include one or more operations described herein with reference to the apparatus.

[0100] In a twelfth aspect according to the eleventh aspect, at least one criterion includes movement exceeding a predetermined threshold.

[0101] In a thirteenth aspect according to at least one of the eleventh aspect to the twelfth aspect, at least one processor includes an image signal processor, and the image signal processor includes an image processing engine and an engine for video analysis.

[0102] In a fourteenth aspect according to the thirteenth aspect, applying a temporal filtering process to a region includes determining an image transformation matrix, and the image transformation matrix is determined by the image processing engine, and the region is determined by the engine for video analysis.

[0103] In a fifteenth aspect according to at least one of the eleventh aspect to the fourteenth aspect, the apparatus is executed as part of an image processing pipeline, and the first image frame is a reference image frame of the image processing pipeline, and the second image frame is a current image frame of the image processing pipeline.

[0104] In a sixteenth aspect according to at least one of the eleventh aspect to the fifteenth aspect, the motion map reflects movement within the second image frame relative to the first image frame.

[0105] In a seventeenth aspect according to the sixteenth aspect, determining the motion map includes determining the difference between the second image frame and the first image frame.

[0106] In an eighteenth aspect according to at least one of the sixteenth aspect to the seventeenth aspect, determining the motion map is based on motion sensor data from an image capture device that captures the first image frame and the second image frame.

[0107] In a nineteenth aspect according to at least one of the eleventh to eighteenth aspects, the first image frame was previously transformed to correct for motion errors before being compared with a second image frame to calculate a motion map of the second image frame.

[0108] In a twentieth aspect according to at least one of the eleventh to nineteenth aspects, the image transformation matrix is generated according to a motion compensation temporal filtering process.

[0109] In a twenty - first aspect, the techniques described herein relate to a non - transitory computer - readable medium storing instructions that, when executed by a processor, cause the processor to perform operations including: receiving a first image frame and a second image frame; determining a motion map based on the second image frame and the first image frame; determining a region of the second image frame having motion that satisfies at least one criterion; and determining a corrected image frame by applying a temporal filtering process to the region.

[0110] In a twenty - second aspect according to the twenty - first aspect, at least one criterion includes movement exceeding a predetermined threshold.

[0111] In a twenty - third aspect according to at least one of the twenty - first to twenty - second aspects, the instructions further cause the processor to implement an image signal processor that includes an image processing engine and an engine for video analysis.

[0112] In a twenty - fourth aspect according to the twenty - third aspect, applying the temporal filtering process to the region includes determining an image transformation matrix, and the image transformation matrix is determined by the image processing engine, and the region is determined by the engine for video analysis.

[0113] In a twenty - fifth aspect according to at least one of the twenty - first to twenty - fourth aspects, the motion map reflects movement within the second image frame relative to the first image frame.

[0114] In a twenty - sixth aspect, the techniques described herein relate to an image capture device that includes: an image sensor; a memory that stores processor - readable code; and at least one processor coupled to the memory and the image sensor, the at least one processor being configured to execute the processor - readable code to cause the at least one processor to: receive a first image frame and a second image frame in image data received from the image sensor; determine a motion map based on the second image frame and the first image frame; determine a region of the second image frame having motion that satisfies at least one criterion; and determine a corrected image frame by applying a temporal filtering process to the region.

[0115] In a twenty - seventh aspect according to the twenty - sixth aspect, at least one criterion includes movement exceeding a predetermined threshold.

[0116] In a twenty-eighth aspect according to at least one of the twenty-sixth to twenty-seventh aspects, the processor-readable code further causes the processor to implement an image signal processor, the image signal processor including an image processing engine and an engine for video analysis.

[0117] In a twenty-ninth aspect according to the twenty-ninth aspect, applying a temporal filtering process to a region includes determining an image transformation matrix, and the image transformation matrix is determined by the image processing engine, and the region is determined by the engine for video analysis.

[0118] In a thirtieth aspect according to at least one of the twenty-sixth to twenty-ninth aspects, the motion map reflects movement within a second image frame relative to a first image frame.

[0119] Those skilled in the art should understand that any of a variety of different techniques and arts can be used to represent information and signals. For example, data, instructions, commands, information, signals, bits, symbols, and chips that may have been mentioned throughout the above description can be represented by voltage, current, electromagnetic waves, magnetic fields or magnetic particles, optical fields or optical particles, or any combination thereof.

[0120] As used herein Figures 1 to 5 The described components, functional blocks, and modules include processors, electronic devices, hardware devices, electronic components, logic circuits, memories, software code, firmware code, etc., or any combination thereof. Software should be construed broadly to mean instructions, instruction sets, code, code segments, program code, programs, subroutines, software modules, applications, software applications, software packages, routines, subroutines, objects, executables, execution threads, procedures, and / or functions, etc., regardless of whether it is referred to as software, firmware, middleware, microcode, hardware description language, or other terms. Additionally, the features discussed herein can be implemented via dedicated processor circuitry, via executable instructions, or a combination thereof.

[0121] Those skilled in the art should understand that: reference Figure 4 and Figure 5 One or more of the boxes (or operations) described with reference to Figure 4 can be combined with one or more of the boxes (or operations) described with reference to another figure in the reference drawings. For example, Figures 1 to 3 One or more of the boxes (or operations) of Figure 5 can be combined with one or more of the boxes (or operations) of Figures 1 to 3 . As another example, one or more of the boxes associated with Figure 5 can be combined with one or more of the boxes (or operations) associated with Figures 1 to 3 .

[0122] Those of ordinary skill in the art should also recognize that: all of the various illustrative logical blocks, modules, circuits, and algorithmic steps described in connection with the disclosure herein can be implemented as electronic hardware, computer software, or a combination of the two. To clearly illustrate this interchangeability of hardware and software, the various illustrative components, blocks, modules, circuits, and steps have been described above generally in terms of their functionality. Whether such functionality is implemented as hardware or software depends upon the particular application and design constraints imposed on the overall system. Those skilled in the art can implement the described functionality in different ways for each particular application, but such specific implementation decisions should not be construed as causing a departure from the scope of the present disclosure. Those skilled in the art will also readily recognize that the order or combination of the components, methods, or interactions described herein are merely examples, and the components, methods, or interactions of the various aspects of the present disclosure can be combined or performed in ways other than those illustrated and described herein.

[0123] The various illustrative logical components, logical blocks, modules, circuits, and algorithmic processes described in connection with the specific implementations disclosed herein can be implemented as electronic hardware, computer software, or a combination of the two. The interchangeability of hardware and software has been generally described in terms of functionality and illustrated in the various illustrative components, blocks, modules, circuits, and processes described above. Whether such functionality is implemented as hardware or software depends upon the particular application and design constraints imposed on the overall system.

[0124] The hardware and data processing means for implementing or performing the various illustrative logics, logic blocks, modules, and circuits described in connection with the aspects disclosed herein can be realized using a general single-chip or multi-chip processor, a digital signal processor (DSP), an application specific integrated circuit (ASIC), a field programmable gate array (FPGA), or other programmable logic device, discrete gate or transistor logic components, discrete hardware components, or any combination thereof, which are designed to perform the functions described herein. A general purpose processor may be a microprocessor, or any conventional processor, controller, microcontroller, or state machine. In some specific implementations, the processor may be implemented as a combination of computing devices, such as a combination of a DSP and a microprocessor, multiple microprocessors, one or more microprocessors in combination with a DSP core, or any other such configuration. In some specific implementations, specific processes and methods may be performed by circuitry specific to a given function.

[0125] In one or more aspects, the described functionality can be implemented in hardware, digital electronic circuits, computer software, firmware, including the structures disclosed in this specification and structural equivalents thereof, or any combination thereof. The specific implementations of the subject matter described in this specification can also be implemented as one or more computer programs, i.e., one or more modules of computer program instructions encoded on a computer storage medium for execution by, or to control the operation of, a data processing apparatus.

[0126] If implemented in software, the functions may be stored on or transmitted through a computer-readable medium as one or more instructions or code. The processes of the methods or algorithms disclosed herein may be implemented in a processor-executable software module that may reside on a computer-readable medium. Computer-readable media includes both computer storage media and communication media, which communication media includes any medium that can be implemented to transfer a computer program from one place to another. The storage media can be any available medium accessible by a computer. By way of example and not limitation, such computer-readable media may include random access memory (RAM), read only memory (ROM), electrically erasable programmable read only memory (EEPROM), CD-ROM or other optical disk storage, magnetic disk storage or other magnetic storage devices, or any other medium that can be used to store the desired program code in the form of instructions or data structures and that can be accessed by a computer. Additionally, any connection may be properly termed a computer-readable medium. As used herein, disk and disc include compact disc (CD), laser disc, optical disc, digital versatile disc (DVD), floppy disk and Blu-ray disc, where disks usually magnetically reproduce data, while discs optically reproduce data with lasers. Combinations of the above should also be included within the scope of computer-readable media. Additionally, the operations of a method or algorithm may reside as one set of code and instructions or any combination of sets of code and instructions on a machine-readable medium and a computer-readable medium, which may be incorporated into a computer program product.

[0127] Various modifications to the specific implementations described in this disclosure will be apparent to those of ordinary skill in the art, and the general principles defined herein may be applied to some other specific implementations without departing from the spirit or scope of this disclosure. Thus, the claims are not intended to be limited to the specific implementations shown herein, but rather to conform to the broadest scope consistent with this disclosure, the principles disclosed herein, and novel features.

[0128] Additionally, those of ordinary skill in the art will readily recognize that, for ease of description of the drawings, terms such as "above" and "below" or "front" and "back" or "top" and "bottom" or "forward" and "backward" may sometimes be used, and indicate relative positions corresponding to the orientation of the drawing on the correctly oriented page, and may not reflect the correct orientation of any device as implemented.

[0129] Certain features that are described in the context of separate embodiments in this specification can also be implemented in a single embodiment in combination. Conversely, the various features that are described in the context of a single embodiment can also be implemented separately or in any suitable sub-combination in multiple embodiments. Additionally, although features may have been described above as acting in certain combinations and even initially claimed as such, one or more features from the claimed combination may in some cases be excluded from the combination, and the claimed combination may be directed to a sub-combination or variation of a sub-combination.

[0130] Similarly, although operations are depicted in the figures in a particular order, this should not be construed as requiring that the operations be performed in the particular order shown or in a sequential order, or that all of the illustrated operations be performed to achieve the desired result. Additionally, the figures may schematically depict one or more example processes in the form of a flowchart. However, other operations not depicted may be incorporated into the example processes that are schematically illustrated. For example, one or more additional operations may be performed before, after, simultaneously with, or between any of the illustrated operations. In certain environments, multitasking and parallel processing are advantageous. Moreover, the separation of the various system components in the embodiments described above should not be construed as requiring such separation in all embodiments, and it should be understood that the described program components and systems can generally be integrated together in a single software product or packaged into multiple software products. Additionally, some other embodiments also fall within the scope of the appended claims. In some cases, the acts recited in the claims can be performed in a different order and still achieve the desired result.

[0131] As used herein (including in the claims), the term "or" as used in a list of two or more items means that any one of the listed items can be employed alone, or any combination of two or more of the listed items can be employed. For example, if a composition is described as containing components A, B, or C, then the composition can contain A alone; B alone; C alone; a combination of A and B; a combination of A and C; a combination of B and C; or a combination of A, B, and C. Additionally, as used herein (including in the claims), "or" as used in a list of items beginning with "at least one of" indicates a disjunctive list, such that a list of, for example, "at least one of A, B, or C" means A or B or C or AB or AC or BC or ABC (i.e., A and B and C) or any combination of any of these items.

[0132] The term "substantially" is defined as being largely but not necessarily wholly that which is specified (and includes that which is specified; for example, substantially 90 degrees includes 90 degrees, and substantially parallel includes parallel), as understood by one of ordinary skill in the art. In any of the disclosed specific embodiments, the term "substantially" may be replaced by "[percentage] within" that which is specified, where the percentage includes 0.1%, 1%, 5%, or 10%.

[0133] The foregoing description of the disclosure has been provided to enable any person skilled in the art to make or use the disclosure. Various modifications to the disclosure will be readily apparent to those skilled in the art, and the general principles defined herein may be applied to other variations without departing from the spirit or scope of the disclosure. Thus, the disclosure is not intended to be limited to the examples and designs described herein but is to be accorded the widest scope consistent with the principles and novel features disclosed herein.

Claims

1. A method, the method comprising: Receiving a first image frame and a second image frame; Determining a motion map based on the second image frame and the first image frame; Determining, based on the motion map, a region of the second image frame having motion that satisfies at least one criterion; And Determining a corrected image frame by applying a temporal filtering process to the region.

2. The method according to claim 1, wherein the at least one criterion includes movement exceeding a predetermined threshold.

3. The method according to claim 1, wherein the method is performed by an image signal processor that includes an image processing engine and an engine for video analysis.

4. The method according to claim 3, wherein applying the temporal filtering process to the region includes determining an image transformation matrix, and wherein the image transformation matrix is determined by the image processing engine and the region is determined by the engine for video analysis.

5. The method according to claim 1, wherein the method is performed as part of an image processing pipeline, and wherein the first image frame is a reference image frame of the image processing pipeline and the second image frame is a current image frame of the image processing pipeline.

6. The method according to claim 1, wherein the motion map reflects movement within the second image frame relative to the first image frame.

7. The method according to claim 6, wherein determining the motion map includes determining a difference between the second image frame and the first image frame.

8. The method according to claim 6, wherein determining the motion map is based on motion sensor data from an image capture device that captured the first image frame and the second image frame.

9. The method according to claim 1, wherein the first image frame was previously transformed to correct for motion errors prior to being compared with the second image frame to calculate the motion map of the second image frame.

10. The method according to claim 1, wherein the image transformation matrix is generated according to a motion compensation temporal filtering process.

11. An apparatus, the apparatus comprising: A memory that stores processor-readable code; And At least one processor coupled to the memory, the at least one processor configured to execute the processor-readable code to cause the at least one processor to perform operations including the following: Receiving a first image frame and a second image frame; Determining a motion map based on the second image frame and the first image frame; Determining, based on the motion map, a region of the second image frame having motion that satisfies at least one criterion; And Determining a corrected image frame by applying a temporal filtering process to the region.

12. The apparatus according to claim 11, wherein the at least one criterion includes movement exceeding a predetermined threshold.

13. The apparatus according to claim 11, wherein the at least one processor includes an image signal processor that includes an image processing engine and an engine for video analysis.

14. The apparatus according to claim 13, wherein applying the temporal filtering process to the region includes determining an image transformation matrix, and wherein the image transformation matrix is determined by the image processing engine, and the region is determined by the engine for video analysis.

15. The apparatus according to claim 11, wherein the apparatus is executed as part of an image processing pipeline, and wherein the first image frame is a reference image frame of the image processing pipeline, and the second image frame is a current image frame of the image processing pipeline.

16. The apparatus according to claim 11, wherein the motion map reflects movement within the second image frame relative to the first image frame.

17. The apparatus according to claim 16, wherein determining the motion map includes determining the difference between the second image frame and the first image frame.

18. The apparatus according to claim 16, wherein determining the motion map is based on motion sensor data from an image capture device that captures the first image frame and the second image frame.

19. The apparatus according to claim 11, wherein the first image frame was previously transformed to correct for motion errors before being compared with the second image frame to calculate the motion map of the second image frame.

20. The apparatus according to claim 11, wherein the image transformation matrix is generated according to a motion compensated temporal filtering process.

21. A non-transitory computer-readable medium storing instructions that, when executed by a processor, cause the processor to perform operations including the following: Receive a first image frame and a second image frame; Determine a motion map based on the second image frame and the first image frame; Determine a region of the second image frame having motion that meets at least one criterion based on the motion map; And Determine a corrected image frame by applying a temporal filtering process to the region.

22. The non-transitory computer-readable medium according to claim 21, wherein the at least one criterion includes movement exceeding a predetermined threshold.

23. The non-transitory computer-readable medium according to claim 21, wherein the instructions further cause the processor to implement an image signal processor that includes an image processing engine and an engine for video analysis.

24. The non-transitory computer-readable medium according to claim 23, wherein applying the temporal filtering process to the region includes determining an image transformation matrix, and wherein the image transformation matrix is determined by the image processing engine, and the region is determined by the engine for video analysis.

25. The non-transitory computer-readable medium according to claim 21, wherein the motion map reflects movement within the second image frame relative to the first image frame.

26. An image capture device, the image capture device comprising: An image sensor; A memory that stores processor-readable code; And at least one processor coupled to the memory and the image sensor, the at least one processor configured to execute the processor-readable code to cause the at least one processor to: receive a first image frame and a second image frame in image data received from the image sensor; determine a motion map based on the second image frame and the first image frame; determine a region of the second image frame having motion that satisfies at least one criterion based on the motion map; and determine a corrected image frame by applying a temporal filtering process to the region.

27. The image capture device according to claim 26, wherein the at least one criterion includes movement exceeding a predetermined threshold.

28. The image capture device according to claim 26, wherein the processor-readable code further causes the processor to implement an image signal processor, the image signal processor including an image processing engine and an engine for video analysis.

29. The image capture device according to claim 28, wherein applying the temporal filtering process to the region includes determining an image transformation matrix, and wherein the image transformation matrix is determined by the image processing engine and the region is determined by the engine for video analysis.

30. The image capture device according to claim 26, wherein the motion map reflects movement within the second image frame relative to the first image frame.