Image enhancement for region of interest of image

By identifying the region of interest in the image and deciding whether to perform upsampling based on its characteristics, and using machine learning models for super-resolution processing, the problem of image quality improvement in the prior art is solved, and efficient image quality improvement for the region of interest is achieved.

CN120112935APending Publication Date: 2025-06-06QUALCOMM INC
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202380074363.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2022-10-26
Filing Date
2023-10-04
Publication Date
2025-06-06

AI Technical Summary

Technical Problem

The prior art is difficult to effectively improve the image quality of the region of interest in an image, especially when the object is small or out of focus.

Method used

By determining the region of interest (ROI) in the image and based on the characteristics of the ROI, such as size and sharpness, it is decided whether to upsample the image data in the ROI. The system includes identifying ROIs, determining its characteristics, and super-resolution processing using machine learning models such as GANs.

Benefits of technology

Efficient image quality improvement for areas of interest in the image is achieved, especially when the object is small or out of focus, improving the image clarity and visual fidelity.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120112935A_ABST
    Figure CN120112935A_ABST
Patent Text Reader

Abstract

Systems, apparatuses, processes, and computer-readable media for capturing images are disclosed. A method of processing image data includes determining a first region of interest (ROI) in an image. The first ROI is associated with a first object. The method may include determining one or more image characteristics of the first ROI. The method may further include determining whether to perform an up-sampling process on the image data in the first ROI based on the one or more image characteristics of the first ROI.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application generally relates to processing image data. For example, aspects of the present disclosure relate to systems and techniques for selectively enhancing regions of an image based on characteristics of the region. Background Art

[0002] A camera is a device that uses an image sensor to receive light and capture image frames (such as still images or video frames). A camera may include a processor that can receive one or more image frames and process one or more image frames, such as an image signal processor (ISP). For example, a raw image frame captured by a camera sensor may be processed by an ISP to generate a final image. A camera may be configured with various image capture and image processing settings to change the appearance of an image. The application of different settings may result in frames or images with different appearances. Summary of the invention

[0003] In some examples, systems and techniques are described for selectively enhancing region(s) of interest (e.g., corresponding to one or more objects, such as people, faces, vehicles, etc.) in an image based on characteristics of image data in the region of interest. The systems and techniques can improve image quality of objects of interest in captured images.

[0004] According to at least one example, a method for processing one or more images is provided. The method includes: determining a first region of interest (ROI) in the image, wherein the first ROI is associated with a first object; determining one or more image characteristics of the first ROI; and determining whether to perform an upsampling process on image data in the first ROI based on the one or more image characteristics of the first ROI.

[0005] In another example, an apparatus for processing one or more images is provided, comprising at least one memory and at least one memory coupled to at least one processor. The at least one processor is configured to: determine a first ROI in the image, wherein the first ROI is associated with a first object; determine one or more image characteristics of the first ROI; and determine whether to perform an upsampling process on image data in the first ROI based on the one or more image characteristics of the first ROI.

[0006] In another example, a non-transitory computer-readable medium having instructions stored thereon is provided, which, when executed by one or more processors, causes the one or more processors to: determine a first ROI in an image, wherein the first ROI is associated with a first object; determine one or more image characteristics of the first ROI; and determine whether to perform an upsampling process on image data in the first ROI based on the one or more image characteristics of the first ROI.

[0007] In another example, an apparatus for processing one or more images is provided. The apparatus includes: a component for determining a first ROI in the image, wherein the first ROI is associated with a first object; a component for determining one or more image characteristics of the first ROI; and a component for determining whether to perform an upsampling process on image data in the first ROI based on the one or more image characteristics of the first ROI.

[0008] In some aspects, the apparatus is a mobile device (e.g., a mobile phone and / or mobile handset and / or so-called "smartphone" or other mobile device), an extended reality (XR) device (e.g., a virtual reality (VR) device, an augmented reality (AR) device, or a mixed reality (MR) device), a head mounted device (HMD) device, a vehicle or a computing device or component of a vehicle, a wearable device (e.g., a network-connected watch or other wearable device), a wireless communication device, a camera, a personal computer, a laptop computer, a server computer, another device, or a combination thereof, and is part of and / or includes a part thereof. In some aspects, the apparatus includes a camera or multiple cameras for capturing one or more images. In some aspects, the apparatus also includes a display for displaying one or more images, notifications, and / or other displayable data. In some aspects, the apparatus may include one or more sensors (e.g., one or more inertial measurement units (IMUs), such as one or more gyroscopes, one or more gyroscopes, one or more accelerometers, any combination thereof, and / or other sensors).

[0009] This summary is not intended to identify key or essential features of the claimed subject matter, nor is it intended to be used alone to determine the scope of the claimed subject matter. The subject matter should be understood by reference to appropriate portions of the entire specification of this patent, any or all drawings, and each claim.

[0010] The foregoing and other features and aspects will become more fully apparent by reference to the following description, claims and accompanying drawings. BRIEF DESCRIPTION OF THE DRAWINGS

[0011] Illustrative aspects of the present application are described in detail below with reference to the following drawings:

[0012] Figure 1A , Figure 1B and Figure 1C is a diagram illustrating an example configuration of an image sensor of an image capturing device according to aspects of the present disclosure.

[0013] Figure 2 is a block diagram illustrating the architecture of an image capture and processing device according to aspects of the present disclosure.

[0014] Figure 3 is a block diagram illustrating an example of an image capture system according to aspects of the present disclosure.

[0015] Figure 4 A block diagram of an image processing apparatus 400 for enhancing at least a portion of an image using super-resolution techniques according to aspects of the present disclosure is shown.

[0016] Figure 5A An example of an image obtained by an image processing device according to some aspects of the present disclosure is shown.

[0017] Figure 5B Shows Figure 5A An upsampled version of the region of interest 502 is shown in FIG.

[0018] Figure 5C The modified image based on super-resolution technology according to some aspects of the present disclosure is shown. Figure 5A An upsampled version of the region of interest 502 is shown in FIG.

[0019] Fig. 6A An example of a region of interest 602 detected by an image processing device is shown.

[0020] Figure 6B Example keypoints are shown that may be detected by a ROI pre-processor to determine a transformation to apply to the ROI in accordance with some aspects of the present disclosure.

[0021] Figure 6C An example transformation applied to a region of interest 602 is shown.

[0022] Fig. 7A The results of image synthesis after post-processing by a ROI post-processor according to some aspects of the present disclosure are shown.

[0023] Figure 7B The results of image synthesis after post-processing by a ROI post-processor after adjusting the ROI according to some aspects of the present disclosure are shown.

[0024] Figure 8 An example of an image showing a ROI analyzer that can identify at least one ROI that can cause the ROI analyzer to control an image sensor in accordance with some aspects of the present disclosure.

[0025] Fig. 9 An example of a trigger module 900 configured to enable or disable SR functionality according to some aspects of the present disclosure is shown.

[0026] Fig.10 is a flow chart illustrating another example of a method for performing a super-resolution function according to aspects of the present disclosure.

[0027] Fig.11 An example diagram illustrating an implementation of a generative adversarial network (GAN) to perform super-resolution (SR) functionality in accordance with certain aspects of the present disclosure.

[0028] Fig.12 An example generator portion of a GAN according to certain aspects of the present disclosure is shown.

[0029] Fig.13 An example discriminator portion of a GAN in accordance with certain aspects of the present disclosure is shown.

[0030] Fig.14 is a diagram illustrating an example of a system for implementing certain aspects described herein. DETAILED DESCRIPTION

[0031] Certain aspects of the present disclosure are provided below. Some of these aspects can be applied independently, and some of them can be applied in combination, which is clear to those skilled in the art. In the following description, for the purpose of explanation, specific details are set forth in order to provide a thorough understanding of various aspects of the application. However, it is clear that various aspects can be practiced without these specific details. The drawings and description are not intended to be limiting.

[0032] The following description provides only example aspects and is not intended to limit the scope, applicability or configuration of the present disclosure. On the contrary, the following description of the example aspects will provide an enabling description for implementing the example aspects for those skilled in the art. It should be understood that various changes may be made to the function and arrangement of the elements without departing from the spirit and scope of the present application as set forth in the appended claims.

[0033] The following description provides only exemplary aspects and is not intended to limit the scope, applicability or configuration of the present disclosure. On the contrary, the following description of exemplary aspects will provide an enabling description for implementing aspects of the present disclosure to those skilled in the art. It should be understood that various changes may be made to the function and arrangement of elements without departing from the spirit and scope of the present application as set forth in the appended claims.

[0034] The terms "exemplary" and / or "example" are used herein to mean "serving as an example, instance, or illustration." Any aspect described herein as "exemplary" and / or "example" is not necessarily to be construed as preferred or advantageous over other aspects. Likewise, the term "aspects of the disclosure" does not require that all aspects of the disclosure include the discussed feature, advantage, or mode of operation.

[0035] A camera is a device that uses an image sensor to receive light and capture an image frame (such as a still image or a video frame). The terms "image", "image frame" and "frame" are used interchangeably herein. A camera can be configured with various image capture and image processing settings. Different settings result in images with different presentations. Some camera settings, such as ISO, exposure time, aperture size, aperture value (f / stop), shutter speed, focus and gain, are determined and applied before or during the capture of one or more image frames. For example, settings or parameters can be applied to an image sensor for capturing one or more image frames. Other camera settings can configure post-processing of one or more image frames, such as changes to contrast, brightness, saturation, sharpness, level, curve or color. For example, settings or parameters can be applied to a processor (e.g., an image signal processor (ISP)) for processing one or more image frames captured by an image sensor.

[0036] Digital image upscaling operations may sometimes be referred to as super-resolution (SR) operations. For example, using a super-resolution operation, a 4k (3840×2160 pixel) resolution display output may be scaled up from a 1080p (1920×1080 pixel) resolution input. Such super-resolution operations may be applied to different resolutions, depending on the specific application (e.g., from 4k to 7680×4320 pixels). There are different algorithms that may be performed to upscale an image to a higher resolution, each achieving a different level of restoration (e.g., if previously compressed or scaled down), realism (e.g., if artificially generated), and / or accuracy (e.g., if a real-life scene is available for comparison).

[0037] Machine learning (such as employing a deep learning neural network) can often achieve better results in terms of restoration, realism, or accuracy than predefined algorithms (e.g., bicubic). In some cases, part of the deep learning network generates one or more scaled-up versions of the input, and another part compares and rates the scaled-up versions with a reference image set to identify a processing regime (e.g., learn one or more parameters) that produces a desired output. Thus, the processing regime is "learned" by the super-resolution network rather than being pre-programmed.

[0038] In some cases, the captured image may be limited based on the resolution of the image and based on whether different portions of the image are in focus or out of focus. As an example, a user of the device may want to enhance an image that includes multiple regions of interest (ROIs) (e.g., corresponding to a group of faces) so that the resulting image more clearly identifies at least one face of a person from the group of faces. This scenario may include many issues, such as the ROI associated with the person's face having a small size relative to the entire image. In other cases, the person's face may be out of focus based on the depth of field (DOF) associated with the camera settings. In these cases, the SR function cannot be applied to the image.

[0039] In some aspects, systems, apparatus, processes (also referred to as methods), and computer-readable media (collectively referred to herein as "systems and techniques") for identifying and improving an image that includes at least one ROI are described. For example, the systems and techniques may determine or identify an ROI and determine whether to perform an upsampling process (e.g., SR techniques) if a person is located within the ROI. For example, based on determining the ROI, the systems and techniques may determine one or more image characteristics of the ROI. An illustrative example of an image characteristic is the size of the ROI (or the size of the ROI relative to the entire image). For example, the ROI may be less than a size threshold suitable for an SR process. In some cases, the size threshold may be any suitable size, such as 100 pixels (px)×100px, 50px×50px, or other suitable sizes. Another illustrative example of an image characteristic is the sharpness of the image, which may be measured based on a signal-to-noise ratio (SNR). Another example of an image characteristic is the distance of the first ROI from the focal point relative to a threshold distance.

[0040] After detecting one or more image characteristics of the ROI, the systems and techniques can determine whether to perform an upsampling process (e.g., an SR technique) on the image data in the ROI based on the one or more image characteristics of the ROI. For example, if the size of the ROI is less than a size threshold, if the sharpness of the image data in the ROI is less than a sharpness threshold, if the distance of the first ROI from the focus is greater than a threshold distance, and / or based on other characteristics of the ROI, the systems and techniques can perform an upsampling process.

[0041] In some aspects, images with limited quality and within range can be pre-processed prior to the upsampling process. For example, in the case where an object identified in an image is rotated based on the position of a person (e.g., tilt) or the position of a camera (e.g., rotation such that the object is slightly rotated), the systems and techniques can pre-process the ROI to align the object to improve image enhancement operations. Other aspects of the systems and techniques include implementing the upsampling process based on gaze detection, device orientation, facial recognition, and other factors.

[0042] Additional details and aspects of the disclosure are described in more detail below with reference to the accompanying drawings.

[0043] Image sensors include one or more arrays of photodiodes or other light-sensitive elements. Each photodiode measures the amount of light that ultimately corresponds to a specific pixel in the image produced by the image sensor. In some cases, different photodiodes can be covered by different color filters of a color filter array and can therefore measure light that matches the color of the color filter covering the photodiode.

[0044] Various color filter arrays may be used, including a Bayer color filter array, a quad color filter array (also referred to as a quad Bayer color filter or QCFA), and / or other color filter arrays. Figure 1A An example of a Bayer color filter array 100 is shown in FIG. As shown, the Bayer color filter array 100 includes a repeating pattern of red, blue, and green filters. Figure 1B As shown, QCFA 110 includes a 2×2 (or "quad") pattern of color filters, including a 2×2 pattern of red (R) color filters, a pair of 2×2 patterns of green (G) color filters, and a 2×2 pattern of blue (B) color filters. Figure 1B 1. Using the QCFA 110 or Bayer color filter array 100, each pixel of the image is generated based on red light data from at least one photodiode covered in the red filter of the color filter array, blue light data from at least one photodiode covered in the blue filter of the color filter array, and green light data from at least one photodiode covered in the green filter of the color filter array. Other types of color filter arrays may use yellow, magenta, and / or cyan (also known as "emerald") filters instead of or in addition to the red, blue, and / or green filters. Different photodiodes throughout the pixel array may have different spectral sensitivity curves and therefore respond to different wavelengths of light. Monochrome image sensors may also lack color filters and therefore lack color depth.

[0045] In some cases, a subset of adjacent photodiodes (e.g., when using Figure 1B , a 2×2 patch of photodiodes can measure light of the same color for substantially the same area of ​​a scene. For example, when the photodiodes included in each of the subgroups of photodiodes are in close physical proximity, the light incident on each photodiode of the subgroup can originate from substantially the same location in the scene (e.g., a portion of leaves on a tree, a small portion of the sky, etc.).

[0046] In some examples, the brightness range of the light from the scene may significantly exceed the brightness levels that the image sensor can capture. For example, a digital single-lens reflex (DSLR) camera may be able to capture a 1:30,000 contrast ratio of the light from the scene, while the brightness levels of an HDR scene may exceed a 1:1,000,000 contrast ratio.

[0047] In some cases, an HDR sensor can be used to enhance the contrast of an image captured by an image capture device. In some examples, an HDR sensor can be used to obtain multiple exposures within an image or frame, where such multiple exposures can include shorter (e.g., 5 ms) and longer (e.g., 15 ms or more) exposure times. As used herein, a longer exposure time generally refers to any exposure time that is longer than a shorter exposure time.

[0048] In some embodiments, an HDR sensor may be capable of isolating individual photodiodes within a photodiode subgroup (e.g., from Figure 1B The four individual R photodiodes, four individual B photodiodes, and four individual G photodiodes in each of the two 2×2G patches in the QCFA 110 shown in FIG. 1 are configured with different exposure settings. A collection of photodiodes with matching exposure settings is also referred to herein as a photodiode exposure group. Figure 1C FIG. 4 shows a portion of an image sensor array with a QCFA filter configured with four different photodiode exposure groups 1 to 4. Figure 1C As shown in the example photodiode exposure group array 120 in FIG. 1 , each 2×2 tile may include photodiodes from each different photodiode exposure group for a particular image sensor. Figure 1C Four groups are shown in the specific grouping in FIG, but one of ordinary skill will recognize that different numbers of photodiode exposure groups, different arrangements of photodiode exposure groups within a subgroup, and any combination thereof may be used without departing from the scope of the present disclosure.

[0049] As about Figure 1CAs described, in some HDR image sensor embodiments, the exposure settings corresponding to different photodiode exposure groups may include different exposure times (also referred to as exposure lengths), such as short exposure, medium exposure, and long exposure. In some cases, different images of scenes associated with different exposure settings can be formed from light captured by the photodiodes of each photodiode exposure group. For example, a first image may be formed by light captured by the photodiodes of photodiode exposure group 1, a second image may be formed by light captured by the photodiodes of photodiode exposure group 2, a third image may be formed by light captured by the photodiodes of photodiode exposure group 3, and a fourth image may be formed by light captured by the photodiodes of photodiode exposure group 4. Based on the difference in exposure settings corresponding to each group, the brightness of objects in the scene captured by the image sensor may be different in each image. For example, a well-lit object captured by a photodiode with a long exposure setting may appear saturated (e.g., completely white). In some cases, the image processor may select between pixels of images corresponding to different exposure settings to form a combined image.

[0050] In one illustrative example, the first image corresponds to a short exposure time (also referred to as a short exposure image), the second image corresponds to a medium exposure time (also referred to as a medium exposure image), and the third image and the fourth image correspond to long exposure times (also referred to as long exposure images). In such an example, pixels of the combined image corresponding to a portion of the scene with a lower illumination (e.g., a portion of the scene in shadow) may be selected from the long exposure image (e.g., the third image or the fourth image). Similarly, pixels of the combined image corresponding to a portion of the scene with a higher illumination (e.g., a portion of the scene in direct sunlight) may be selected from the short exposure image (e.g., the first image).

[0051] In some cases, the image sensor can also utilize a photodiode exposure group to capture objects in motion without blurring. The length of the exposure time of the photodiode group can correspond to the distance that the object in the scene moves during the exposure time. If the light from the object in motion is captured by a photodiode corresponding to multiple image pixels during the exposure time, the object in motion may appear blurred across multiple image pixels (also known as motion blur). In some embodiments, motion blur can be reduced by configuring one or more photodiode groups with a short exposure time. In some embodiments, an image capture device (e.g., a camera) can determine the amount of local motion (e.g., a motion gradient) within a scene by comparing the position of an object between two consecutive captured images. For example, motion can be detected in a preview image captured by an image capture device to provide a preview function to a user on a display. In some cases, a machine learning model can be trained to detect local motion between consecutive images.

[0052] Various aspects of the technology described herein will be discussed below with reference to the accompanying drawings. Figure 2 2 is a block diagram illustrating the architecture of the image capture and processing system 200. The image capture and processing system 200 includes various components for capturing and processing an image of a scene (e.g., an image of the scene 210). The image capture and processing system 200 may capture an independent image (or photograph) and / or may capture a video including a plurality of images (or video frames) of a particular sequence. In some cases, the lens 215 and the image sensor 230 may be associated with an optical axis. In an illustrative example, the photosensitive area (e.g., a photodiode) of the image sensor 230 and the lens 215 may both be centered on the optical axis. The lens 215 of the image capture and processing system 200 faces the scene 210 and receives light from the scene 210. The lens 215 bends the incident light from the scene toward the image sensor 230. The light received by the lens 215 passes through an aperture. In some cases, the aperture (e.g., the aperture size) is controlled by one or more control mechanisms 220 and received by the image sensor 230. In some cases, the aperture may have a fixed size.

[0053] The one or more control mechanisms 220 may control exposure, focus, and / or zoom based on information from the image sensor 230 and / or based on information from the image processor 250. The one or more control mechanisms 220 may include multiple mechanisms and components; for example, the control mechanisms 220 may include one or more exposure control mechanisms 225A, one or more focus control mechanisms 225B, and / or one or more zoom control mechanisms 225C. The one or more control mechanisms 220 may also include additional control mechanisms other than those shown, such as control mechanisms to control analog gain, flash, HDR, depth of field, and / or other image capture attributes.

[0054] The focus control mechanism 225B of the control mechanism 220 can obtain the focus setting. In some examples, the focus control mechanism 225B stores the focus setting in a memory register. Based on the focus setting, the focus control mechanism 225B can adjust the position of the lens 215 relative to the position of the image sensor 230. For example, based on the focus setting, the focus control mechanism 225B can move the lens 215 closer to the image sensor 230 or farther away from the image sensor 230 by actuating a motor or servo mechanism (or other lens mechanism), thereby adjusting the focus. In some cases, additional lenses may be included in the image capture and processing system 200, such as one or more microlenses above each photodiode of the image sensor 230, each microlens bending the light received from the lens 215 toward the corresponding photodiode before the light reaches the photodiode. The focus setting may be determined via contrast detection autofocus (CDAF), phase detection autofocus (PDAF), hybrid autofocus (HAF), or some combination thereof. The focus setting may be determined using the control mechanism 220, the image sensor 230, and / or the image processor 250. The focus setting may be referred to as an image capture setting and / or an image processing setting. In some cases, lens 215 may be fixed relative to the image sensor and focus control mechanism 225B may be omitted without departing from the scope of the present disclosure.

[0055] The exposure control mechanism 225A of the control mechanism 220 can obtain an exposure setting. In some cases, the exposure control mechanism 225A stores the exposure setting in a memory register. Based on the exposure setting, the exposure control mechanism 225A can control the size of the aperture (e.g., aperture size or aperture value (f / stop)), the duration that the aperture is open (e.g., exposure time or shutter speed), the duration that the sensor collects light (e.g., exposure time or electronic shutter speed), the sensitivity of the image sensor 230 (e.g., ISO speed or film speed), the analog gain applied by the image sensor 230, or any combination thereof. The exposure setting can be referred to as an image capture setting and / or an image processing setting.

[0056] The zoom control mechanism 225C of the control mechanism 220 can obtain the zoom setting. In some examples, the zoom control mechanism 225C stores the zoom setting in a memory register. Based on the zoom setting, the zoom control mechanism 225C can control the focal length of the lens element assembly (lens assembly) including the lens 215 and one or more additional lenses. For example, the zoom control mechanism 225C can control the focal length of the lens assembly by actuating one or more motors or servo mechanisms (or other lens mechanisms) to move one or more of the lenses relative to each other. The zoom setting can be referred to as an image capture setting and / or an image processing setting. In some examples, the lens assembly can include a parfocal zoom lens or a varifocal zoom lens. In some examples, the lens assembly can include a focusing lens (which can be lens 215 in some cases) that first receives light from the scene 210, and then the light passes through an afocal zoom system between the focusing lens (e.g., lens 215) and the image sensor 230 before the light reaches the image sensor 230. In some cases, the afocal zoom system may include two positive (e.g., converging, convex) lenses having equal or similar focal lengths (e.g., within a threshold difference of each other) with a negative (e.g., diverging, concave) lens between them. In some cases, zoom control mechanism 225C moves one or more lenses in the afocal zoom system, such as a negative lens and one or two positive lenses. In some cases, zoom control mechanism 225C may control the zoom by capturing images from an image sensor in a plurality of image sensors (e.g., including image sensor 230) at a zoom corresponding to the zoom setting. For example, image processing system 200 may include a wide-angle image sensor with a relatively low zoom and a telephoto image sensor with a greater zoom. In some cases, based on the selected zoom setting, zoom control mechanism 225C may capture images from the corresponding sensors.

[0057] Image sensor 230 includes one or more arrays of photodiodes or other light sensitive elements. Each photodiode measures the amount of light that ultimately corresponds to a particular pixel in the image produced by image sensor 230. In some cases, different photodiodes may be covered by different filters. In some cases, different photodiodes may be covered in color filters and thus may measure light that matches the color of the color filter covering the photodiode. Various color filter arrays may be used, including Bayer color filter arrays (such as Figure 1A shown), QCFA (see Figure 1B ) and / or any other color filter array.

[0058] Back to Figure 1A and Figure 1B, other types of color filters may use yellow, magenta, and / or cyan (also referred to as "emerald") filters instead of or in addition to red, blue, and / or green filters. In some cases, some photodiodes may be configured to measure infrared (IR) light. In some embodiments, the photodiode measuring IR light may not be covered by any filter, thereby allowing the IR photodiode to measure both visible (e.g., color) light and IR light. In some examples, the IR photodiode may be covered by an IR filter, thereby allowing IR light to pass through and blocking light from other parts of the spectrum (e.g., visible light, color). Some image sensors (e.g., image sensor 230) may lack filters (e.g., color, IR, or any other part of the spectrum) entirely, and may instead use different photodiodes (in some cases vertically stacked) throughout the pixel array. Different photodiodes throughout the pixel array may have different spectral sensitivity curves, and therefore respond to light of different wavelengths. Monochrome image sensors may also lack filters, and therefore lack color depth.

[0059] In some cases, the image sensor 230 may alternatively or additionally include an opaque and / or reflective mask that blocks light from reaching certain photodiodes or portions of certain photodiodes at certain times and / or from certain angles. In some cases, the opaque and / or reflective mask may be used for PDAF. In some cases, the opaque and / or reflective mask may be used to block portions of the electromagnetic spectrum from reaching the photodiodes of the image sensor (e.g., IR cutoff filters, ultraviolet (UV) cutoff filters, bandpass filters, low-pass filters, high-pass filters, etc.). The image sensor 230 may also include an analog gain amplifier for amplifying the analog signal output by the photodiode and / or an analog-to-digital converter (ADC) for converting the analog signal output of the photodiode (and / or amplified by the analog gain amplifier) ​​into a digital signal. In some cases, certain components or functions discussed with respect to one or more of the control mechanisms 220 may be included in the image sensor 230 alternatively or additionally. Image sensor 230 may be a charge coupled device (CCD) sensor, an electron multiplying CCD (EMCCD) sensor, an active pixel sensor (APS), a complementary metal oxide semiconductor (CMOS), an N-type metal oxide semiconductor (NMOS), a hybrid CCD / CMOS sensor (e.g., sCMOS), or some other combination thereof.

[0060] The image processor 250 may include one or more processors, such as one or more ISPs (e.g., ISP 254), one or more host processors (e.g., host processor 252), and / or Fig.14The host processor 252 may be a digital signal processor (DSP) and / or other type of processor. In some embodiments, the image processor 250 is a single integrated circuit or chip (e.g., referred to as a system on a chip or SoC) that includes the host processor 252 and the ISP 254. In some cases, the chip may also include one or more input / output ports (e.g., input / output (I / O) port 256), a central processing unit (CPU), a graphics processing unit (GPU), a broadband modem (e.g., 3G, 4G or LTE, 5G, etc.), memory, connectivity components (e.g., Bluetooth TM , global positioning system (GPS), etc.), any combination thereof, and / or other components. The I / O port 256 may include any suitable input / output port or interface according to one or more protocols or specifications, such as an integrated circuit 2 (I2C) interface, an integrated circuit 3 (I3C) interface, a serial peripheral interface (SPI) interface, a serial general purpose input / output (GPIO) interface, a mobile industry processor interface (MIPI) (such as a MIPI CSI-2 physical (PHY) layer port or interface), an advanced high-performance bus (AHB) bus, any combination thereof, and / or other input / output ports. In an illustrative example, the host processor 252 may communicate with the image sensor 230 using an I2C port, and the ISP 254 may communicate with the image sensor 230 using a MIPI port.

[0061] The image processor 250 may perform a number of tasks, such as demosaicing, color space conversion, image frame downsampling, pixel interpolation, automatic exposure (AE) control, automatic gain control (AGC), CDAF, PDAF, automatic white balance, merging image frames to form an HDR image, image recognition, object recognition, feature recognition, receiving input, managing output, managing memory, or some combination thereof. The image processor 250 may store the image frames and / or processed images in a random access memory (RAM) 240, a read-only memory (ROM) 245, a cache memory, a memory unit, another storage device, or some combination thereof.

[0062] Various input / output (I / O) devices 260 may be connected to the image processor 250. The I / O devices 260 may include a display screen, a keyboard, a keypad, a touch screen, a touch pad, a touch-sensitive surface, a printer, any other output device (e.g., Fig.14 output device 1435), any other input device (e.g. Fig.141445) or some combination thereof. In some cases, subtitles may be entered into the image processing device 205B via a physical keyboard or keypad of the I / O device 260 or via a virtual keyboard or keypad of a touch screen of the I / O device 260. The I / O 260 may include one or more ports, sockets, or other connectors that enable a wired connection between the image capture and processing system 200 and one or more peripheral devices, through which the image capture and processing system 200 may receive data from and / or send data to the one or more peripheral devices. The I / O 260 may include one or more wireless transceivers that enable a wireless connection between the image capture and processing system 200 and one or more peripheral devices, through which the image capture and processing system 200 may receive data from and / or send data to the one or more peripheral devices. The peripheral devices may include any of the previously discussed types of I / O devices 260, and once they are coupled to ports, sockets, wireless transceivers, or other wired and / or wireless connectors, the peripheral devices themselves may be considered to be I / O devices 260.

[0063] In some cases, the image capture and processing system 200 can be a single device. In some cases, the image capture and processing system 200 can be two or more separate devices, including an image capture device 205A (e.g., a camera) and an image processing device 205B (e.g., a computing device coupled to the camera). In some embodiments, the image capture device 205A and the image processing device 205B can be wirelessly coupled together, for example, via one or more wires, cables, or other electrical connectors and / or via one or more wireless transceivers. In some embodiments, the image capture device 205A and the image processing device 205B can be disconnected from each other.

[0064] like Figure 2 As shown, the vertical dashed line will Figure 2 The image capture and processing system 200 is divided into two parts, respectively representing the image capture device 205A and the image processing device 205B. The image capture device 205A includes a lens 215, a control mechanism 220, and an image sensor 230. The image processing device 205B includes an image processor 250 (including an ISP 254 and a host processor 252), a RAM 240, a ROM 245, and an I / O 260. In some cases, some of the components shown in the image capture device 205A (such as the ISP 254 and / or the host processor 252) can be included in the image capture device 205A.

[0065] The image capture and processing system 200 may include an electronic device, such as a mobile or fixed telephone handset (e.g., a smart phone, a cellular phone, etc.), a desktop computer, a laptop or notebook computer, a tablet computer, a set-top box, a television, a camera, a display device, a digital media player, a video game console, a video streaming device, an Internet Protocol (IP) camera, or any other suitable electronic device. In some examples, the image capture and processing system 200 may include one or more wireless transceivers for wireless communications (such as cellular network communications, 802.11 wi-fi communications, wireless local area network (WLAN) communications, or some combination thereof). In some embodiments, the image capture device 205A and the image processing device 205B may be different devices. For example, the image capture device 205A may include a camera device, and the image processing device 205B may include a computing device, such as a mobile phone, a desktop computer, or other computing devices.

[0066] Although the image capture and processing system 200 is shown as including certain components, one of ordinary skill will appreciate that the image capture and processing system 200 may include more than Figure 2 . The components of the image capture and processing system 200 may include software, hardware, or one or more combinations of software and hardware. For example, in some embodiments, the components of the image capture and processing system 200 may include and / or may be implemented using electronic circuits or other electronic hardware, which may include one or more programmable electronic circuits (e.g., microprocessors, GPUs, DSPs, CPUs, and / or other suitable electronic circuits), and / or may include and / or be implemented using computer software, firmware, or any combination thereof to perform the various operations described herein. The software and / or firmware may include one or more instructions stored on a computer-readable storage medium and executable by one or more processors of an electronic device that implements the image capture and processing system 200.

[0067] Figure 3 3 is a block diagram illustrating an example of an image capture system 300. The image capture system 300 includes various components for processing input images or frames to produce output images or frames. As shown, the components of the image capture system 300 include one or more image capture devices 302, an image processing engine 310, and an output device 312. The image processing engine 310 can produce a high dynamic range depiction of a scene, as described in more detail herein.

[0068] The image capture system 300 may include or be part of an electronic device or system. For example, the image capture system 300 may include or be part of an electronic device or system, such as a mobile or fixed telephone handset (e.g., a smart phone, a cellular phone, etc.), an extended reality (XR) device (e.g., a virtual reality (VR) device, an augmented reality (AR) device, or a mixed reality (MR) device), a vehicle or a vehicle computing device / system, a server computer (e.g., communicating with another device or system, such as a mobile device, an XR system / device, a vehicle computing system / device, etc.), a desktop computer, a laptop or notebook computer, a tablet computer, a set-top box, a television, a camera device, a display device, a digital media player, a video streaming device, or any other suitable electronic device. In some examples, the image capture system 300 may include one or more wireless transceivers (or separate wireless receivers and transmitters) for wireless communications, such as cellular network communications, 802.11 Wi-Fi communications, WLAN communications, Bluetooth or other short-range communications, any combination thereof, and / or other communications. In some embodiments, the components of the image capture system 300 may be part of the same computing device. In some implementations, components of image capture system 300 may be part of two or more separate computing devices.

[0069] Although image capture system 300 is shown as including certain components, one of ordinary skill will appreciate that image capture system 300 may include more than Figure 3 In some cases, additional components of image capture system 300 may include software, hardware, or one or more combinations of software and hardware. For example, in some cases, image capture system 300 may include one or more other sensors (e.g., one or more inertial measurement units (IMUs), radars, light detection and ranging (LIDAR) sensors, audio sensors, etc.), one or more display devices, one or more other processing engines, one or more other hardware components, and / or Figure 3One or more other software and / or hardware components not shown in the figure. In some embodiments, the additional components of the image capture system 300 may include and / or may be implemented using electronic circuits or other electronic hardware, which may include one or more programmable electronic circuits (e.g., DSPs, microprocessors, microcontrollers, GPUs, CPUs, any combination thereof, and / or other suitable electronic circuits), and / or may include and / or may be implemented using computer software, firmware, or any combination thereof to perform the various operations described herein. The software and / or firmware may include one or more instructions stored on a computer-readable storage medium and executable by one or more processors of an electronic device implementing the image capture system 300.

[0070] One or more image capture devices 302 may capture image data and generate images (or frames) based on the image data and / or may provide the image data to an image processing engine 310 for further processing. One or more image capture devices 302 may also provide the image data to an output device 312 for output (e.g., on a display). In some cases, the output device 312 may also include a storage device. An image or frame may include an array of pixels representing a scene. For example, the image may be a red-green-blue (RGB) image having red, green, and blue components per pixel; a luminance, chrominance-red, chrominance-blue (YCbCr) image having a luminance component and two chrominance (color) components (chrominance-red and chrominance-blue) per pixel; or any other suitable type of color or monochrome image. In addition to the image data, the image capture device may also generate supplemental information such as the amount of time between consecutively captured images, a timestamp of image capture, and the like.

[0071] Figure 4 A block diagram of an image processing device 400 that uses an upsampling process to enhance at least a portion of an image is shown. An illustrative example of an upsampling process is a super-resolution (SR) process. The image processing device 400 includes an image sensor 402 (e.g., image processor 250) configured to capture an image by exposing light to a sensor array. The image sensor 402 is configured to provide the captured image to a region of interest (ROI) preprocessor 404. The ROI preprocessor 404 is configured to use various techniques to identify various ROIs. An illustrative example of an ROI is an object of interest in an image, such as a face. Other examples of ROIs include a facial region of a person, a stationary object (e.g., a landmark such as a sculpture, a building, etc.), an animated object (e.g., a moving vehicle), vegetation, a natural geographic structure, and one of an animal. In some examples, the ROI preprocessor 404 can be configured to detect multiple ROIs, such as different objects in a scene, including faces, animals, landmarks, etc.

[0072] In some aspects, the ROI preprocessor 404 may include a feature detection module 406 that is configured to identify an ROI that can be enhanced using an upsampling process (e.g., an SR process). Examples of ROIs identified for the upsampling process include areas that are smaller than a size threshold (e.g., a resolution of 100×100 pixels). Another example of an area for the upsampling process includes an area with a sharpness within a specific range or less than a sharpness threshold. For example, an object within a captured image may be slightly defocused based on the distance from the image sensor to the object and the focal length. For example, the focal length includes a depth of field (DOF), and objects at the very edge of the DOF or outside the DOF may appear blurred. The sharpness of the area can be determined to determine whether an upsampling process can be applied to the image. In some aspects, at least two features of the feature detection module 406 can be combined to determine whether an upsampling process can be implemented to improve the quality of the image.

[0073] The ROI preprocessor 404 may include a trigger module 408 that can selectively enable an upsampling process (e.g., an SR process) based on various aspects. Illustrative examples of triggers include gaze detection, voice detection, facial recognition, touch detection, device orientation detection, and brightness / occlusion detection. For example, the ROI preprocessor 404 can be triggered by gaze detection, or when a face in an image is determined to be looking at the image sensor 402. As a result of detecting a face that directs gaze at the image sensor 402, the ROI preprocessor 404 can determine that the face is an ROI and can be enhanced using a super-resolution function. This article refers to Fig. 9 Various examples of triggers are further described.

[0074] Other aspects of the ROI preprocessor 404 include additional processing steps to provide various corrections to improve SR techniques. In one illustrative aspect, the ROI preprocessor 404 can extract a portion of the image corresponding to the ROI and perform various corrections to improve SR systems and techniques that can be applied to the ROI. Figure 6B As further described in , the ROI preprocessor 404 can identify key points of various features within the ROI and transform the ROI (e.g., a portion of the image) to align the portion of the key points. As an example, the key points can correspond to the edges of a person's eyes and can be slightly rotated when the image is obtained. In this case, the ROI preprocessor 404 can rotate the ROI (e.g., a portion of the image) to vertically align the key points associated with the person's eyes. In other aspects, other types of transformations can be applied, such as three-dimensional skew, or corrections caused by distortion caused by the lens at the peripheral edge.

[0075] At least one pre-processed ROI (e.g., based on calculating the rotated portion of the image that is considered blurry) can be provided to a deep learning model 410 configured to improve the image quality of the at least one pre-processed ROI. In some aspects, the deep learning model 410 can be a neural network configured to improve the image quality using various techniques. An illustrative deep learning model 410 can be a convolutional neural network (CNN), which is designed to adaptively learn the spatial hierarchy of features through back propagation using multiple building blocks (such as convolutional layers, pooling layers, and fully connected layers). In some cases, CNN can be used to improve the image. Another illustrative deep learning model 410 is a generative adversarial network (GAN). The GAN can be configured to infer content and then apply the inferred content to the original content to improve the quality of the original content.

[0076] In some aspects, the deep learning model 410 can be configured to perform the SR function and create an upsampled version of at least one pre-processed ROI having a higher resolution (e.g., 400px×400px) than the original ROI (e.g., 100px×100px) and having enhancements applied to fill in additional details. In one illustrative example, the GAN can be configured to create content based on a training set (such as a collection of facial images) and learn how to infer details within the upsampled version of the original image. Reference Figure 5B and Figure 5C An example of SR functionality is shown. The deep learning model 410 may output an enhanced ROI for each of the at least one pre-processed ROI to the ROI analyzer 412 and the post-processor 414 .

[0077] Various aspects of the present disclosure include a ROI analyzer 412 configured to analyze enhanced ROIs to improve image capture of an image processing device 400. In one illustrative aspect, the image processing device 400 may be configured to capture an image or video, and the ROI analyzer 412 may identify at least one ROI from a plurality of ROIs as having a characteristic indicating lower quality. Examples of characteristics indicating lower quality include sharpness factors. In this case, the ROI analyzer 412 may control the image sensor 402 to move the focus of a lens module (not shown) to improve the sharpness of at least one ROI. For example, if there are four ROIs in the image and the background ROI is out of focus based on being outside the DOF, the ROI analyzer 412 may be configured to move the focus between the midpoint between the foreground ROI and the background ROI with higher sharpness to improve the sharpness of the background ROI. In this case, the sharpness of the foreground ROI may be reduced by a negligible amount while increasing the sharpness of the background ROI. The image sensor 402 may thereby capture subsequent images and apply the deep learning model 410 to subsequent images to improve image quality. In particular, the SR function applied to the background ROI will have a larger effect.

[0078] The at least one enhanced ROI is also provided to a post-processor 414 for merging the at least one enhanced ROI into the original image or an upsampled version of the original image. For example, bicubic interpolation may be applied to the original image to produce an upsampled image, and the at least one enhanced ROI may be inserted or blended into the upsampled image. In another aspect of the post-processor 414, the post-processor 414 may be configured to reverse the transformation applied by the ROI pre-processor 404, such as rotating the at least one enhanced ROI. In yet another aspect of the post-processor 414, the post-processor 414 may control a boundary region associated with the ROI pre-processor 404 when inserting or blending the at least one enhanced ROI into the upsampled image. Reference Fig. 7A and Figure 7B An example of a control boundary region is further shown.

[0079] Figure 5A An example of an image obtained by an image processing device according to some aspects of the present disclosure is shown. In one illustrative example, a person located within an image may want to zoom in on the image for various purposes, such as for use as a head shot in various profiles. In this case, the boundary region of the region of interest 502 identifying the face of the person may have a resolution less than a threshold value (e.g., 100×100). Due to the lower resolution of the face of the person, the image may have less utility, and the person may desire to enhance the version of the image by zooming in on the image.

[0080] Figure 5B Shows Figure 5A In particular, Figure 5B Conventional bicubic interpolation is shown for upsampling the region of interest 502 and blurring the subject due to conventional techniques. In bicubic interpolation, the image is enhanced based only on the internal data and the result in the blurred result.

[0081] Figure 5C The use of super-resolution techniques according to some aspects of the present disclosure is shown. Figure 5A 502 is an upsampled version of the ROI 502 shown in . For example, the region of interest may be identified and enhanced using the image processing device 400, which adds content to the ROI 502 to create a clearer image. For example, the ROI 502 may be input into a deep learning model (e.g., deep learning model 410), which infers the content based on learning performed during training, and creates Figure 5C In this case, Figure 5C The images in can have more practicality.

[0082] Fig. 6A An example of a region of interest 602 detected by an image processing device is shown. In this case, objects in the region of interest 602 are misaligned due to various factors, such as the image processing device being slightly rotated. Figure 6B Example key points that may be detected by an ROI preprocessor to determine a transformation to apply to an ROI are shown in accordance with some aspects of the present disclosure. In this case, the preprocessor may include a model for detecting various features, such as key points associated with a person's face. Non-limiting examples of key points detected by the preprocessor include the outer corner 612 of the right eye, the inner corner 614 of the right eye, the outer corner 616 of the left eye, the inner corner 618 of the left eye, and the edge 620 of the nose. In some aspects, the preprocessor may identify alignment issues, such as rotation angles, based on the various key points.

[0083] Figure 6C An example transformation is shown applied to the region of interest 602. In this case, the ROI is rotated to align key points associated with the right eye 612, the inner corner of the right eye 614, the outer corner of the left eye 616, and the inner corner of the left eye 618. As mentioned above, aligning the ROI can increase the effectiveness of the deep learning model.

[0084] Fig. 7AThe result of image synthesis after post-processing by the ROI post-processor according to some aspects of the present disclosure is shown. In this case, the boundary region 702 associated with the ROI is enlarged and enhanced by the deep learning model and blended back to an enlarged version of the original image (e.g., bicubic scaling). The boundary region 702 omits the first portion 704 and the second portion 706 of the person's face, and the blending creates a noise region 708 around the peripheral area of ​​the boundary region 702. As a result of the noise region 708, artifacts are added that reduce the image quality of the enlarged image.

[0085] In some aspects, the ROI post-processor can be configured to perform an iterative check to identify a noise region based on an analysis of a peripheral region associated with a skin region. For example, the ROI post-processor can analyze a region at the peripheral region to determine whether the region corresponds to a skin region based on color (e.g., by looking at a color profile, etc.). If the ROI post-processor identifies a skin region at the peripheral region, the ROI post-processor can then determine whether a region outside the ROI but adjacent to the peripheral region also corresponds to a skin region. In the event that the ROI post-processor identifies a skin region in the ROI and a skin region outside the ROI, the ROI post-processor can then feed the information back to the pre-processor to modify the ROI.

[0086] Figure 7B The result of image synthesis after post-processing by the ROI post-processor after adjusting the ROI according to some aspects of the present disclosure is shown. In this case, the boundary region 710 is modified to include the first portion 704 and the second portion 706, and the blending of the images omits the noise region 708.

[0087] Figure 8 An example of an image in which a ROI analyzer according to some aspects of the present disclosure can identify at least one ROI that can cause the ROI analyzer to control an image sensor is shown. As described above, the ROI analyzer is configured to analyze each ROI and control the image processing system based on the ROI. For example, a first ROI 802 is located near an image sensor that obtains an image with a larger aperture ratio to control the DOF to capture the first ROI 802 with sufficient sharpness. A larger aperture (e.g., a smaller f-stop number such as f / 4.0) produces a very shallow DOF, and a smaller aperture (e.g., a larger f-stop number) produces an image with a larger DOF.

[0088] Due to the large aperture ratio, the DOF is limited, and the second ROI 804 in the image is outside the DOF, which causes the person in the second ROI 804 to be blurry and not sharp enough. In some aspects, the ROI analyzer 412 can control the image sensor 402 to move the focus of the lens module (not shown) to improve the sharpness of at least one ROI. For example, if there are four ROIs in the image and the background ROI is out of focus based on being outside the DOF, the ROI analyzer 412 can be configured to modify the lens focus to move the focus to a midpoint between the background ROI and the four ROIs to ensure that each ROI has a sharpness that can be corrected using iterative techniques, one or more machine learning models, and / or other techniques.

[0089] Fig. 9 An example of a trigger module 900 configured to enable or disable SR functionality in accordance with some aspects of the present disclosure is shown. In some aspects, the trigger module 900 can be used in a variety of different devices. In one aspect, the trigger module can use a gaze detection module 902 to determine whether to activate SR. For example, the image capture system can be configured to enable the SR functionality based on the subject of the captured image. As an example, a viewfinder function can identify a person's gaze during the viewfinder function and can identify that the person is not in focus, as described above. Therefore, the focus of the image sensor can be modified by moving the focus between a foreground target and a background target to minimize focus loss (e.g., minimize sharpness loss).

[0090] In other aspects, the trigger module 900 may include a voice detection module 904 configured to detect the voice of a person within the captured image. In some aspects, the voice detection module may use voice detection detected by an auxiliary device (e.g., a smart speaker) to identify the speaker. For example, a user device including a trigger module may use voice to authenticate that the speaker corresponds to an authenticated user of the user device and then enable SR functionality associated with a person known to the user device. The user device may be notified of known persons to capture images based on frequent use of the user device, or other inferences corresponding to the relationship of the person within the captured image may be made.

[0091] The trigger module 900 may also include a facial recognition module 906 for identification to enable SR functionality. For example, in a situation where a person is browsing their photos with their user device, the user device may identify certain parameters that selectively enable the user device to perform SR functionality in captured images. In the event that a person is identified in at least two previous images, the disclosed systems and techniques may then implement SR for the user without user input to enhance the image quality associated with the ROI in the image. In some aspects, the facial recognition module 906 may be associated and / or combined with biometric functionality associated with the user device. For example, if an authenticated user of a user device is identified in the lock screen, the facial recognition module 906 may use an internal application or an external third-party application to selectively authenticate the user of the user device for image capture functionality.

[0092] Additional aspects include a touch detection module 908 to enable the SR function based on touch input. As an example, a user may purposefully or accidentally select a person in a user device to enable the SR function. The trigger module 900 may also include a brightness / occlusion module 910, which is configured to detect whether the SR function can be performed based on the detection of brightness and / or occlusion caused by changing lighting conditions. For example, some shadow conditions may prevent the user device from correctly applying a deep learning model to improve the quality of the upsampled image. The brightness / occlusion module 910 may also be configured to detect exposed faces or other characteristics indicating facial information, and enable and / or disable the SR function based on the detected information. In additional aspects, the trigger module 900 may also include a device orientation module 912, which is configured to determine whether the orientation (e.g., yaw, pitch, and / or roll) indicates that the SR function is enabled or disabled. For example, if the user device is rotated at an abnormal (e.g., non-representative) angle, the user device may determine that there is no face in the captured image, and determine that the input is received unintentionally and the SR function should be disabled.

[0093] Fig.10 1 is a flowchart illustrating an example of a method 1000 for processing image data according to certain aspects of the present disclosure. The method 1000 may be performed by a computing system or device (or component thereof, such as a chipset) having an image sensor, such as a mobile wireless communication device, a camera, an XR device, a wireless-enabled vehicle, or another computing device. In one illustrative example, the computing system 1400 may be configured to perform all or part of the method 1000. In one illustrative example, an ISP such as the ISP 254 may be configured to perform all or part of the method.

[0094] At block 1002, a computing system (or a component thereof) may be configured to determine a first region of interest (ROI) in an image, wherein the first ROI is associated with a first object. Examples of the first object include a person (e.g., a face of a person), but the first object may be any other object, such as a landmark, an animal, vegetation, etc.

[0095] At block 1004, the computing system (or components thereof) may be configured to determine one or more image characteristics of the first ROI. In some aspects, the one or more image characteristics include at least one of a size of the first ROI or a distance of the first ROI from a focal point associated with the image.

[0096] In one illustrative aspect, the computing system (or a component thereof) may determine that the size of the first ROI is less than a size threshold (e.g., 100 pixels by 100 pixels). The computing system may determine to perform an upsampling process on the image data in the ROI based on the size of the first ROI being less than the size threshold. In some aspects, the ROI may be associated with an ML model, such as an ML model trained to perform a super-resolution function based on a downsampled image having a resolution of 100 pixels by 100 pixels. For example, a GAN may receive an original image (e.g., 400 pixels by 400 pixels) and a downsampled version of the original image (e.g., 100 pixels by 100 pixels), and the GAN learns a technique to upsample the downsampled image to correspond to the original image.

[0097] In another illustrative aspect, the computing system (or a component thereof) may be a component of an imaging system (e.g., a camera) that may use the ROI to improve image quality. According to this aspect, the computing system (or a component thereof) or the imaging system may determine that the distance of the first ROI from the focal point is greater than a threshold distance. The computing system (or a component thereof) or the imaging system may determine to perform an upsampling process on the image data in the first ROI based on the distance of the first ROI from the focal point being greater than the threshold distance.

[0098] Another illustrative aspect is related to an object corresponding to an ROI located outside the DOF, and the ROI may be blurred. In this aspect, the computing system (or its components) or the imaging system can determine a second ROI in the image, and the second ROI is associated with the second object. The computing system (or its components) or the imaging system can determine not to perform an upsampling process on the image data in the second ROI based on one or more image characteristics of the second ROI. For example, the size of the ROI can be greater than a threshold. In another example, the imaging system can determine that the sharpness of the image does not meet the threshold sharpness value. In some cases, the first ROI and the second ROI are associated with a common object type. For example, the computing system (or its components) or the imaging system can determine that the first ROI and the second ROI are associated with a common object type. In one aspect, the common object type can be a facial area of ​​a person. Other example object types include vehicles, people, landmarks, stationary objects (e.g., street signs), bicycles, animated objects, vegetation, geographic formations, animals, and / or other objects.

[0099] At block 1006, the computing system (or a component thereof) may be configured to determine whether to perform an upsampling process on the image data in the first ROI based on one or more image characteristics of the first ROI. In one illustrative aspect, the computing system is configured to perform an upsampling process or a super-resolution process (or function) on an image having a resolution of 100 pixels x 100 pixels as described above using the ML model.

[0100] In some aspects, the computing system may be configured to align the ROI to improve the upsampling process when the computing system determines to perform an upsampling process. In one illustrative aspect, the computing system may detect key points associated with a first ROI and transform the partial image based on aligning the key points associated with the first ROI. For example, the first ROI may be a person's face, and the computing system may detect key points associated with the eyes, nose, and mouth, and use the key points to align the first ROI. Alignment of the key points may improve the upsampling process. Based on the transformed partial image, the computing system may obtain an output image from an ML model that is trained to increase resolution and enhance the face at least in part by inputting the first ROI.

[0101] The computing system (or a component thereof) may be configured to overlay an output image (e.g., from an ML model) on an upsampled version of the image. For example, the computing system may upsample the entire image and may overlay the first ROI that has been upsampled by the ML model onto the upsampled version of the image. In other aspects, the computing system may overlay the upsampled first ROI onto the original image. In this case, when the user zooms in on the first ROI, the image quality of the first ROI increases.

[0102] In some aspects, the computing system may be configured to determine that the ROI creates visual fidelity issues when the upsampled ROI is superimposed into the original (or resized) image. In one illustrative aspect, the computing system may be configured to resize a bounding box associated with a first ROI and determine that the resized bounding box crops the skin information based on a boundary area of ​​the resized bounding box. By cropping the skin information, the superposition of the ROI produces noise and reduces visual fidelity. In some aspects, based on determining that the resized bounding box crops the skin information, the computing system may modify the resized bounding box to include an area corresponding to the skin information outside the resized bounding box.

[0103] In some other aspects, the computing system can be a component of the imaging system, and the computing system can be configured to adjust the image capture settings to improve the visual fidelity of the captured image. In one aspect, the computing system can detect a second ROI in the image and determine the difference in sharpness between the first ROI and the second ROI. For example, the computing system can determine that the second ROI is blurrier than the first ROI, and then adjust the focus of the lens of the image sensor to increase the sharpness of the first ROI and reduce the sharpness of the second ROI for the additional image. The computing system can then obtain the additional image based on adjusting the focus of the lens. In this case, the computing system improves the image quality of the second ROI, and the upsampling process can be applied to the first ROI and the second ROI to enhance the image quality associated with the object.

[0104] As described above, the processes or methods described herein (e.g., method 1000 and / or other processes described herein) can be performed by a computing system (or device or apparatus). In one example, method 1000 can be performed by a computer having Fig.14 The computing architecture of the computing system 1400 shown is a computing device (e.g., Figure 2 The image capture and processing system 200 is executed.

[0105] The computing device may include any suitable device, such as a mobile device (e.g., a mobile phone), a desktop computing device, a tablet computing device, a wearable device (e.g., a VR headset, an AR headset, AR glasses, a network-connected watch or smart watch or other wearable device), a server computer, an autonomous vehicle or a computing device of an autonomous vehicle, a robotic device, a television, and / or any other computing device having resource capabilities to perform the methods described herein (including method 1000). In some cases, the computing device or apparatus may include various components, such as one or more input devices, one or more output devices, one or more processors, one or more microprocessors, one or more microcomputers, one or more cameras, one or more sensors, and / or other components configured to perform the steps of the methods described herein. In some examples, the computing device may include a display, a network interface configured to transmit and / or receive data, any combination thereof, and / or other components. The network interface may be configured to transmit and / or receive IP-based data or other types of data.

[0106] The components of the computing device may be implemented in circuits. For example, the components may include and / or may be implemented using electronic circuits or other electronic hardware, which may include one or more programmable electronic circuits (e.g., microprocessors, GPUs, DSPs, CPUs, and / or other suitable electronic circuits), and / or may include and / or may be implemented using computer software, firmware, or any combination thereof to perform the various operations described herein.

[0107] Method 1000 is illustrated as a logical flow diagram, the operations of which represent a series of operations that can be implemented in hardware, computer instructions, or a combination thereof. In the context of computer instructions, the operations represent computer-executable instructions stored on one or more computer-readable storage media that, when executed by one or more processors, perform the enumerated operations. Typically, computer-executable instructions include routines, programs, objects, components, data structures, etc. that perform specific functions or implement specific data types. The order in which the operations are described is not intended to be construed as a limitation, and any number of the described operations may be combined in any order and / or in parallel to implement the method.

[0108] Method 1000 and / or other methods or processes described herein may be performed under the control of one or more computer systems configured with executable instructions, and may be implemented as code (e.g., executable instructions, one or more computer programs, or one or more applications) that is executed together on one or more processors, by hardware, or a combination thereof. As described above, the code may be stored on a computer-readable or machine-readable storage medium, for example, in the form of a computer program including a plurality of instructions executable by one or more processors. The computer-readable or machine-readable storage medium may be non-transitory.

[0109] Fig.11 An example diagram 1100 is shown that implements super-resolution of an image that can be performed using a GAN in accordance with certain aspects of the present disclosure. As shown, in certain embodiments, a neural network processor can be implemented as a GAN to generate content based on training inferences. The modified GAN employs color features 1102 (such as RGB values ​​of a low-resolution frame) and non-color features. As shown, the non-color features include one or more of depth information, normal information, or texture information. In some aspects, the GAN can implement a discriminator function that employs a relativistic loss function to identify a desired version of a high-resolution output, thereby learning applicable parameters implemented by the neural network processor.

[0110] A quantization process 1122 is performed before outputting one or more high-resolution frames 1130 (also referred to as super-resolution frames or images). The quantization process 1122 evaluates the effect of quantization. For example, the modified GAN uses a regular quantization method available in a standard framework to quantize the network. Such as Tensorflow (or PyTorch) or a vendor-specific quantization method (e.g., quantization provided by Qualcomm's SNPE SDK). In some aspects, quantization using the described method for 8-bit / 16-bit utilizes a high-speed accelerator specifically designed for 8-bit / 16-bit operations to improve performance with a marginal quality tradeoff.

[0111] As shown in the figure, the GAN can accept additional various features, such as one or more of depth information, normal information, texture information, etc. The GAN also includes more learning parameters for several subsequent layers than the known GAN examples, such as 22 subsequent layers instead of 10 layers of some GAN examples. This example is Fig.12 and Fig.13 Further shown in .

[0112] In some ways, the intermediate blocks used in GANs ( Fig.12 and Fig.13The number of (as shown in ) is further tuned to balance between output quality (e.g., the higher the quality, the greater the computational requirements, which may or may not be required) and computational requirements (e.g., gigaFLOPS). In some aspects, this variation improves the performance and power budget of the computation.

[0113] Fig.12 According to certain aspects of the present disclosure Fig.11 An example generator portion 1200 of a modified GAN. As shown, the generator portion 1200 can receive depth features 1202 and color features 1204 at a concatenation layer 1210. The concatenation layer 1210 is followed by a depth and point convolution layer 1212. The convolution layer 1212 is followed by a parametric rectified linear unit (PReLU) layer 1214. The PReLu layer 1214 can learn parameters that control the shape and leakiness of the function.

[0114] A plurality of intermediate blocks (MID blocks) 1220 follow the PReLU layer 1214. As shown, 22 intermediate blocks 1220 are included in this example, but the total number of intermediate blocks may vary depending on the output quality requirements. The feature map (e.g., depth information, texture information, and normal information to be included at the concatenated layer 1210) and the exact number of intermediate blocks 1220 may be adjusted or tuned depending on the desired quality and performance. Each intermediate block 1220 includes a depth-wise and point-wise convolution layer 1221, a batch normalization layer 1222, a PReLU layer 1223, another layer 1224 of depth-wise and point-wise convolution, another layer 1225 of batch normalization, and a layer 1226 for element-wise addition.

[0115] After the series of intermediate blocks 1220, two or more blocks 1230 including depth-wise and point-wise convolutions, pixel shuffling, and PReLU follow. One or more depth-wise and point-wise convolutions may follow before outputting (one or more) high-resolution frames 1130 at the last layer. The (one or more) output high-resolution frames 1130 may include two or more versions of the predicted super-resolution images of the low-resolution input image.

[0116] Fig.13An example discriminator portion 1300 of a GAN according to certain aspects of the present disclosure is shown. When the super-resolution image 1030 generated by the generator portion 1300 is compared to a database of reference images, the discriminator portion 1300 outputs false or true information (to identify whether the material is generated or real). The true output can be (e.g., the most) accurate or true to the final output of the neural network of the processor.

[0117] The discriminator part 1300 can access the reference image and receive the super-resolution image 1030 from the generator part 1300. The modified GAN can compare the predicted image and the reference image using features generated by the loss network (e.g., the VGG-19 network) to learn the corresponding parameters of the desired output. For example, the comparison is on the ReLU activated feature output. ReLU inherently suppresses all negative values ​​and only considers positive values. GAN enhances ReLU by removing the ReLU layer at the end of the known GAN example and by involving full-size features in the comparison.

[0118] The reference image and the super-resolution image 1030 are input to the convolution layer 1304. A leaky ReLU layer 1306 follows the convolution layer 1304. The leaky ReLU layer 1306 can modify the ReLU function to allow small negative values ​​when the input is less than zero. A series of intermediate blocks 1310 follow the leaky ReLU layer 1306. Each of the intermediate blocks 1310 includes a convolution layer, a batch normalization layer, and a leaky ReLU layer. Similar to the intermediate blocks 1120, the number of intermediate blocks 1310 can vary depending on the specific application.

[0119] Fig.14 is a diagram illustrating an example of a system for implementing certain aspects of the present technology. In particular, Fig.14 An example of a computing system 1400 is shown, which may be any computing device, for example, constituting an internal computing system, a remote computing system, a camera, or any component thereof, wherein the components of the system communicate with each other using connection 1405. Connection 1405 may be a physical connection using a bus, or a direct connection into processor 1410, such as in a chipset architecture. Connection 1405 may also be a virtual connection, a networked connection, or a logical connection.

[0120] In some aspects, computing system 1400 is a distributed system, wherein the functionality described in the present disclosure may be distributed within a data center, multiple data centers, a peer-to-peer network, etc. In some aspects, one or more of the described system components represent a number of such components, each component performing some or all of the functionality for which the component is described. In some aspects, a component may be a physical or virtual device.

[0121] The example computing system 1400 includes at least one processing unit (CPU or processor) 1410 and connections 1405 coupling various system components including system memory 1415, such as ROM 1420 and RAM 1425, to the processor 1410. The computing system 1400 may include a cache 1412 of high-speed memory directly connected to, immediately adjacent to, or integrated as part of the processor 1410.

[0122] Processor 1410 may include any general purpose processor and hardware services or software services, such as services 1432, 1434, and 1436 stored in storage device 1430, configured to control processor 1410 as well as special purpose processors where the software instructions are incorporated into the actual processor design. Processor 1410 may be essentially a completely independent computing system containing multiple cores or processors, buses, memory controllers, caches, etc. Multi-core processors may be symmetric or asymmetric.

[0123] To enable user interaction, the computing system 1400 includes an input device 1445, which can represent any number of input mechanisms, such as a microphone for voice, a touch-sensitive screen for gesture or graphical input, a keyboard, a mouse, motion input, voice, and the like. The computing system 1400 can also include an output device 1435, which can be one or more of a plurality of output mechanisms. In some cases, a multimodal system can enable a user to provide multiple types of input / output to communicate with the computing system 1400. The computing system 1400 can include a communication interface 1440, which can generally govern and manage user input and system output. The communication interface can use a wired and / or wireless transceiver to perform or facilitate receiving and / or sending wired or wireless communications, including utilizing an audio jack / plug, a microphone jack / plug, a universal serial bus (USB) port / plug, Ports / plugs, Ethernet ports / plugs, Fiber optic ports / plugs, Proprietary wired ports / plugs, Wireless signal transmission, BLE wireless signal transmission, Wireless signal transmission, RFID wireless signal transmission, near field communication (NFC) wireless signal transmission, dedicated short range communication (DSRC) wireless signal transmission, 802.11 WiFi wireless signal transmission, WLAN signal transmission, visible light communication (VLC), World Interoperability for Microwave Access (WiMAX), IR communication wireless signal transmission, public switched telephone network (PSTN) signal transmission, integrated services digital network (ISDN) signal transmission, 3G / 4G / 5G / LTE cellular data network wireless signal transmission, ad hoc network signal transmission, radio wave signal transmission, microwave signal transmission, infrared signal transmission, visible light signal transmission, ultraviolet light signal transmission, wireless signal transmission along the electromagnetic spectrum, or some combination thereof. The communication interface 1440 may also include one or more global navigation satellite system (GNSS) receivers or transceivers for determining the location of the computing system 1400 based on receiving one or more signals from one or more satellites associated with one or more GNSS systems. GNSS systems include, but are not limited to, GPS based in the United States, Global Navigation Satellite System (GLONASS) based in Russia, BeiDou Navigation Satellite System (BDS) based in China, and Galileo GNSS based in Europe. There is no restriction to operating on any particular hardware arrangement, and thus the basic features herein may be readily substituted for improved hardware or firmware arrangements as such are developed.

[0124] The storage device 1430 may be a non-volatile and / or non-transitory and / or computer-readable memory device and may be a hard disk or other type of computer-readable medium that can store data accessible by a computer, such as a magnetic tape cartridge, a flash memory card, a solid-state memory device, a digital versatile disk, a magnetic cassette, a floppy disk, a flexible disk, a hard disk, a magnetic tape, a magnetic stripe / strip, any other magnetic storage medium, a flash memory, a memristor memory, any other solid-state memory, a compact disk read-only memory (CD-ROM) disc, a rewritable compact disk (CD) disc, a digital video disk (DVD) disc, a Blu-ray disc (BDD) disc, a holographic disc, another optical medium, a secure digital (SD) card, a micro secure digital (microSD) card, a Memory Card, or a memory card. card, a smart card chip, an EMV chip, a subscriber identity module (SIM) card, a mini / micro / nano / pico SIM card, another integrated circuit (IC) chip / card, a RAM, a static RAM (SRAM), a dynamic RAM (DRAM), a ROM, a programmable read-only memory (PROM), an erasable programmable read-only memory (EPROM), an electrically erasable programmable read-only memory (EEPROM), a flash EPROM (FLASH EPROM), a cache memory (L1 / L2 / L3 / L4 / L5 / L#), a resistive random access memory (RRAM / ReRAM), a phase change memory (PCM), a spin-transfer torque RAM (STT-RAM), another memory chip or box, and / or a combination thereof.

[0125] Storage device 1430 may include software services, servers, services, etc., which, when the code defining such software is executed by processor 1410, causes the system to perform a function. In some aspects, a hardware service that performs a particular function may include a software component stored in a computer-readable medium that is combined with the necessary hardware components (such as processor 1410, connection 1405, output device 1435, etc.) to perform the function. The term "computer-readable medium" includes, but is not limited to, portable or non-portable storage devices, optical storage devices, and various other media capable of storing, containing or carrying instructions and / or data. Computer-readable media may include non-transitory media in which data may be stored and does not include carrier waves and / or transient electronic signals propagated wirelessly or via a wired connection. Examples of non-transitory media may include, but are not limited to, disks or tapes, optical storage media (such as CDs or DVDs), flash memory, memory, or memory devices. The computer readable medium may have stored thereon codes and / or machine executable instructions, which may represent a process, function, subprogram, program, routine, subroutine, module, software package, category, or any combination of instructions, data structures, or program statements. A code segment may be coupled to another code segment or hardware circuit by transmitting and / or receiving information, data, independent variables, parameters, or memory contents. Information, independent variables, parameters, data, etc. may be transmitted, forwarded, or sent via any suitable means including memory sharing, message passing, token passing, network transmission, etc.

[0126] In some cases, a computing device or apparatus may include various components, such as one or more input devices, one or more output devices, one or more processors, one or more microprocessors, one or more microcomputers, one or more cameras, one or more sensors, and / or other components configured to perform the steps of the processes described herein. In some examples, a computing device may include a display, one or more network interfaces configured to transmit and / or receive data, any combination thereof, and / or other components. The one or more network interfaces may be configured to transmit and / or receive wired and / or wireless data, including data in accordance with 3G, 4G, 5G, and / or other cellular standards, data in accordance with the Wi-Fi (802.11x) standard, data in accordance with the Bluetooth standard, and / or other wireless communication protocols. TM Standard data, data according to IP standards and / or other types of data.

[0127] The components of the computing device may be implemented in circuits. For example, the components may include and / or may be implemented using electronic circuits or other electronic hardware, which may include one or more programmable electronic circuits (e.g., microprocessors, GPUs, DSPs, CPUs, and / or other suitable electronic circuits), and / or may include and / or may be implemented using computer software, firmware, or any combination thereof to perform the various operations described herein.

[0128] In some aspects, computer readable storage devices, media, and memories may include cables or wireless signals containing bit streams, etc. However, when referred to, non-transitory computer readable storage media explicitly excludes media such as energy, carrier signals, electromagnetic waves, and signals themselves.

[0129] Specific details are provided in the above description to provide a thorough understanding of the aspects and examples provided herein. However, it will be appreciated by those skilled in the art that these aspects can be practiced without these specific details. For clarity, in some cases, the present technology may be presented as including separate functional blocks, which include functional blocks comprising devices, device components, steps or routines in methods embodied in software or a combination of hardware and software. In addition to those components shown in the accompanying drawings and / or described herein, additional components may be used. For example, circuits, systems, networks, processes, and other components may be shown as components in block diagram form, so as not to obscure these aspects with unnecessary details. In other instances, known circuits, processes, algorithms, structures, and techniques may be shown without unnecessary details, so as to avoid obscuring these aspects.

[0130] Various aspects may be described above as a process or method depicted as a flow chart, flow diagram, data flow diagram, structure diagram, or block diagram. Although a flow chart may describe an operation as a sequential process, many operations may be performed in parallel or simultaneously. In addition, the order of the operations may be rearranged. A process terminates when its operations are completed, but may have additional steps not included in the figure. A process may correspond to a method, function, program, subroutine, subprogram, etc. When a process corresponds to a function, its termination may correspond to the function returning to a calling function or a main function.

[0131] The processes and methods according to the above examples can be implemented using computer executable instructions stored in or otherwise available from computer readable media. Such instructions may include, for example, instructions and data that enable or otherwise configure a general-purpose computer, a special-purpose computer, or a processing device to perform a specific function or group of functions. Parts of the computer resources used may be accessed over a network. Computer executable instructions may be, for example, binary files, intermediate format instructions such as assembly language, firmware, source code, etc. Examples of computer readable media that can be used to store instructions, information used, and / or information created during the method according to the described examples include disks or optical disks, flash memory, USB devices provided with non-volatile memory, networked storage devices, etc.

[0132] Devices implementing the processes and methods disclosed herein may include hardware, software, firmware, middleware, microcode, hardware description languages, or any combination thereof, and may be implemented in any of a variety of form factors. When implemented in software, firmware, middleware, or microcode, program code or code segments (e.g., computer program products) for performing the necessary tasks may be stored in a computer-readable or machine-readable medium. (One or more) processors may perform the necessary tasks. Typical examples of form factors include laptop computers, smart phones, mobile phones, tablet devices or other small form factor personal computers, personal digital assistants, rack-mounted devices, stand-alone devices, etc. The functions described herein may also be embodied in peripheral devices or add-on cards. As another example, such functions may also be implemented between different chips or different processes executed in a single device on a circuit board.

[0133] The instructions, the media for carrying such instructions, the computing resources for executing them, and other structures for supporting such computing resources are example means for providing the functions described in this disclosure.

[0134] In the foregoing description, various aspects of the present application are described with reference to their specific aspects, but those skilled in the art will recognize that the present application is not limited thereto. Therefore, although the illustrative aspects of the present application have been described in detail herein, it should be understood that the inventive concept can be implemented and adopted differently in other ways, and the appended claims are intended to be interpreted as including such variations, except for being limited by the prior art. The various features and aspects of the above-mentioned application can be used individually or in combination. In addition, without departing from the broader spirit and scope of this specification, various aspects can be utilized in any number of environments and applications outside the environment and application described herein. Therefore, the description and the accompanying drawings are considered to be illustrative rather than restrictive. For the purpose of illustration, the method is described in a specific order. It should be understood that, in alternative aspects, the method can be performed in an order different from the described order.

[0135] One of ordinary skill will understand that the less than ("<") and greater than (">") symbols or terms used herein may be replaced by less than or equal to ("≤") and greater than or equal to ("≥") symbols, respectively, without departing from the scope of the present specification.

[0136] Where a component is described as being “configured to” perform certain operations, such configuration may be achieved, for example, by designing electronic circuits or other hardware to perform those operations, by programming programmable electronic circuits (e.g., a microprocessor or other suitable electronic circuits) to perform those operations, or any combination thereof.

[0137] The phrase "coupled to" means that any component is directly or indirectly physically connected to another component, and / or any component is in direct or indirect communication with another component (e.g., connected to another component via a wired or wireless connection, and / or other suitable communication interface).

[0138] Claim language or other language reciting "at least one of" a set and / or "one or more of" a set indicates that one member of a set or multiple members of a set (in any combination) satisfy the claim. For example, claim language reciting "at least one of A and B" or "at least one of A or B" means A, B, or A and B. In another example, claim language reciting "at least one of A, B, and C" or "at least one of A, B, or C" means A, B, C, or A and B, or A and C, or B and C, or A and B and C. The language "at least one of" a set and / or "one or more of" a set does not limit the set to the items listed in the set. For example, claim language reciting "at least one of A and B" or "at least one of A or B" may mean A, B, or A and B, and may additionally include items not listed in the set of A and B.

[0139] The various illustrative logic blocks, modules, circuits, and algorithmic steps described in conjunction with the aspects disclosed herein may be implemented as electronic hardware, computer software, firmware, or a combination thereof. To clearly illustrate this interchangeability of hardware and software, various illustrative components, blocks, modules, circuits, and steps have been generally described above with respect to their functionality. Whether such functions are implemented as hardware or software depends on the specific application and the design constraints imposed on the entire system. The technician may implement the described functions in different ways for each specific application, but such implementation decisions should not be interpreted as causing a departure from the scope of the present application.

[0140] The technology described herein can also be implemented in electronic hardware, computer software, firmware, or any combination thereof. The technology can be implemented in any of a variety of devices, such as a general-purpose computer, a wireless communication device mobile phone, or an integrated circuit device, which has multiple uses for applications including wireless communication devices and other devices. Any features described as modules or components can be implemented together in an integrated logic device, or separately implemented as a discrete but interoperable logic device. If implemented in software, the technology can be implemented at least in part by a computer-readable data storage medium including a program code, and the program code contains instructions for executing one or more of the methods described above when executed. The computer-readable data storage medium can form part of a computer program product, which can include packaging materials. The computer-readable medium can include a memory or data storage medium, such as a RAM, such as a synchronous dynamic random access memory (SDRAM), a ROM, a non-volatile random access memory (NVRAM), an EEPROM, a flash memory, a magnetic or optical data storage medium, and the like. Additionally or alternatively, the technology may be implemented at least in part by a computer-readable communication medium that carries or communicates program code in the form of instructions or data structures and that can be accessed, read, and / or executed by a computer, such as a propagated signal or wave.

[0141] The program code may be executed by a processor, which may include one or more processors, such as one or more DSPs, general-purpose microprocessors, application-specific integrated circuits (ASICs), field-programmable logic arrays (FPGAs), or other equivalent integrated or discrete logic circuits. The processor may be configured to perform any of the techniques described in this disclosure. A general-purpose processor may be a microprocessor; however, in an alternative, the processor may be any conventional processor, controller, microcontroller, or state machine. The processor may also be implemented as a combination of computing devices, for example, a combination of a DSP and a microprocessor, a plurality of microprocessors, a combination of one or more microprocessors with a DSP core, or any other such configuration. Therefore, the term "processor" used herein may refer to any of the foregoing structures, any combination of the foregoing structures, or any other structure or device suitable for implementation in the technology described herein.

[0142] Illustrative aspects of the disclosure include:

[0143] Aspect 1. A method for processing one or more images, comprising: determining ROIs in the image, wherein a first ROI is associated with a first object; determining one or more image characteristics of the first ROI; and determining whether to perform an upsampling process on image data in the first ROI based on the one or more image characteristics of the first ROI.

[0144] Aspect 2. The method of aspect 1, wherein the one or more image characteristics include at least one of a size of the first ROI or a distance of the first ROI from a focal point associated with the image.

[0145] Aspect 3. The method according to Aspect 2 further includes: determining that the size of the first ROI is smaller than a size threshold; and determining to perform an upsampling process on the image data in the ROI based on the size of the first ROI being smaller than the size threshold.

[0146] Aspect 4. The method according to any one of Aspects 2 or 3 further includes: determining that the distance between the first ROI and the focal point is greater than a threshold distance; and based on the distance between the first ROI and the focal point being greater than the threshold distance, determining to perform an upsampling process on the image data in the first ROI.

[0147] Aspect 5. The method according to any one of Aspects 1 to 4 further includes: determining a second ROI in the image, wherein the second ROI is associated with a second object; determining one or more image characteristics of the second ROI; and determining not to perform an upsampling process on the image data in the second ROI based on the one or more image characteristics of the second ROI.

[0148] Clause 6. The method according to clause 5, wherein the first ROI and the second ROI are associated with a common object type.

[0149] Aspect 7. The method according to aspect 5, wherein the common object type includes a facial region of a person.

[0150] Aspect 8. The method according to any one of Aspects 1 to 7 further includes: detecting key points associated with a first ROI, wherein the first ROI includes a face of a person; and transforming the portion of the image based on aligning the key points associated with the first ROI.

[0151] Aspect 9. The method according to Aspect 7 further includes: obtaining an output image from an ML model based on the transformed partial image, the ML model being trained to increase resolution and enhance the face at least in part by inputting the first ROI.

[0152] Aspect 10. The method according to aspect 9 further comprises: superimposing the output image on the upsampled version of the image.

[0153] Aspect 11. The method according to any one of Aspects 1 to 10 further includes: detecting a second ROI in the image; determining a sharpness difference between the first ROI and the second ROI; adjusting the focus of the lens of the image sensor to increase the sharpness of the first ROI and reduce the sharpness of the second ROI for the additional image; and obtaining the additional image based on adjusting the focus of the lens.

[0154] Aspect 12. The method according to any one of Aspects 1 to 11 further includes: adjusting the size of a bounding box associated with the first ROI; and determining that the resized bounding box crops the skin information based on a boundary area of ​​the resized bounding box; and modifying the resized bounding box to include an area outside the resized bounding box corresponding to the skin information.

[0155] Aspect 13. An apparatus comprising at least one memory and at least one processor coupled to the at least one memory, and configured to: determine a first ROI in an image, wherein the first ROI is associated with a first object; determine one or more image characteristics of the first ROI; and determine whether to perform an upsampling process on image data in the first ROI based on the one or more image characteristics of the first ROI.

[0156] Clause 14. The apparatus of clause 13, wherein the one or more image characteristics include at least one of a size of the first ROI or a distance of the first ROI from a focal point associated with the image.

[0157] Aspect 15. An apparatus according to Aspect 14, wherein at least one processor is configured to: determine that the size of the first ROI is less than a size threshold; and determine to perform an upsampling process on the image data in the ROI based on the size of the first ROI being less than the size threshold.

[0158] Aspect 15. An apparatus according to any one of Aspects 14 or 15, wherein at least one processor is configured to: determine that a distance between the first ROI and the focal point is greater than a threshold distance; and determine to perform an upsampling process on the image data in the first ROI based on the distance between the first ROI and the focal point being greater than the threshold distance.

[0159] Aspect 17. An apparatus according to any one of Aspects 13 to 16, wherein at least one processor is configured to: determine a second ROI in the image, wherein the second ROI is associated with a second object; determine one or more image characteristics of the second ROI; and based on the one or more image characteristics of the second ROI, determine not to perform an upsampling process on the image data in the second ROI.

[0160] Clause 18. The apparatus according to clause 17, wherein the first ROI and the second ROI are associated with a common object type.

[0161] Clause 19. The apparatus according to clause 28, wherein the common object type comprises a facial region of a person.

[0162] Aspect 20. An apparatus according to any one of Aspects 13 to 19, wherein at least one processor is configured to: detect key points associated with a first ROI, wherein the first ROI includes a face of a person; and transform the portion of the image based on aligning the key points associated with the first ROI.

[0163] Aspect 21. An apparatus according to Aspect 20, wherein at least one processor is configured to: obtain an output image from an ML model based on the transformed partial image, the ML model being trained to increase resolution and enhance the face at least in part by inputting the first ROI.

[0164] Clause 22. The apparatus according to Clause 21, wherein at least one processor is configured to: superimpose the output image on an upsampled version of the image.

[0165] Aspect 23. An apparatus according to any one of Aspects 13 to 22, wherein at least one processor is configured to: detect a second ROI in the image; determine a sharpness difference between the first ROI and the second ROI; adjust the focus of the lens of the image sensor to increase the sharpness of the first ROI and reduce the sharpness of the second ROI for the additional image; and obtain an additional image based on adjusting the focus of the lens.

[0166] Aspect 24. An apparatus according to any one of Aspects 13 to 23, wherein at least one processor is configured to: resize a bounding box associated with the first ROI; determine that the resized bounding box crops the skin information based on a bounding area of ​​the resized bounding box; and modify the resized bounding box to include an area outside the resized bounding box corresponding to the skin information.

[0167] Aspect 25. A non-transitory computer readable medium comprising instructions which, when executed by one or more processors, cause the one or more processors to perform the operations according to any one of Aspects 1 to 12.

[0168] Aspect 26. An apparatus comprising means for performing the operations according to any one of aspects 1 to 12.

Claims

1. A method of processing one or more images, include: determining a first region of interest (ROI) in the image, wherein the first ROI is associated with a first object; determining one or more image characteristics of the first ROI; and Based on the one or more image characteristics of the first ROI, it is determined to perform an upsampling process on the image data in the first ROI.

2. The method according to claim 1, in, The one or more image characteristics include at least one of a size of the first ROI or a distance of the first ROI from a focal point associated with an image.

3. The method according to claim 2, further comprising: include: Determining that a size of the first ROI is less than a size threshold; as well as Based on the fact that the size of the first ROI is smaller than the size threshold, it is determined to perform an upsampling process on the image data in the ROI.

4. The method according to claim 2, further comprising: include: Determining that the distance between the first ROI and the focus is greater than a threshold distance; as well as An upsampling process is performed on the image data in the first ROI based on the distance between the first ROI and the focus being greater than the threshold distance.

5. The method according to claim 1, further comprising: include: determining a second ROI in the image, wherein the second ROI is associated with a second object; determining one or more image characteristics of the second ROI; and Based on one or more image characteristics of the second ROI, it is determined not to perform an upsampling process on the image data in the second ROI.

6. The method according to claim 5, in, The first ROI and the second ROI are associated with a common object type.

7. The method according to claim 6, in, The common object types include facial regions of people.

8. The method according to claim 1, further comprising: include: detecting key points associated with the first ROI, wherein the first ROI includes a face of a person; and The portion of the image is transformed based on aligning keypoints associated with the first ROI.

9. The method according to claim 8, further comprising: include: Based on transforming the portion of the image, an output image is obtained from a machine learning (ML) model, the ML model being trained to increase resolution and enhance faces at least in part by inputting the first ROI.

10. The method according to claim 9, further comprising: include: The output image is superimposed on the upsampled version of the image.

11. The method according to claim 1, further comprising: include: Detecting a second ROI in the image; determining a sharpness difference between the first ROI and the second ROI; adjusting a focus of a lens of an image sensor to increase sharpness of the first ROI and decrease sharpness of the second ROI for additional images; as well as Based on adjusting the focus of the lens, the additional image is obtained.

12. The method according to claim 1, further comprising: include: adjusting the size of a bounding box associated with the first ROI; determining a resized bounding box to crop the skin information based on a bounding area of ​​the resized bounding box; as well as The resized bounding box is modified to include an area outside the resized bounding box corresponding to the skin information.

13. An apparatus for processing one or more images, include: at least one memory; as well as at least one processor, the at least one memory coupled to the at least one memory and configured to: Determining a first region of interest ROI in the image, wherein the first ROI is associated with a first object; determining one or more image characteristics of the first ROI; and Based on one or more image characteristics of the first ROI, it is determined whether to perform an upsampling process on the image data in the first ROI.

14. The device according to claim 13, in, The one or more image characteristics include at least one of a size of the first ROI or a distance of the first ROI from a focal point associated with an image.

15. The device according to claim 14, in, The at least one processor is configured to: determining that a size of the first ROI is less than a size threshold; and Based on the fact that the size of the first ROI is smaller than the size threshold, it is determined to perform an upsampling process on the image data in the ROI.

16. The device according to claim 14, in, The at least one processor is configured to: Determining that the distance between the first ROI and the focus is greater than a threshold distance; and Based on the fact that the distance between the first ROI and the focus is greater than the threshold distance, it is determined to perform an upsampling process on the image data in the first ROI.

17. The device according to claim 13, in, The at least one processor is configured to: determining a second ROI in the image, wherein the second ROI is associated with a second object; determining one or more image characteristics of the second ROI; and Based on one or more image characteristics of the second ROI, it is determined not to perform an upsampling process on the image data in the second ROI.

18. The device according to claim 17, in, The first ROI and the second ROI are associated with a common object type.

19. The device according to claim 18, in, The common object types include facial regions of people.

20. The device according to claim 13, in, The at least one processor is configured to: detecting key points associated with the first ROI, wherein the first ROI includes a face of a person; and The portion of the image is transformed based on aligning keypoints associated with the first ROI.

21. The device according to claim 20, in, The at least one processor is configured to: Based on transforming the portion of the image, an output image is obtained from a machine learning (ML) model, the ML model being trained to increase resolution and enhance faces at least in part by inputting the first ROI.

22. The device according to claim 21, in, The at least one processor is configured to: The output image is superimposed on the upsampled version of the image.

23. The device according to claim 13, in, The at least one processor is configured to: Detecting a second ROI in the image; determining a sharpness difference between the first ROI and the second ROI; adjusting a focus of a lens of an image sensor to increase sharpness of the first ROI and decrease sharpness of the second ROI for additional images; as well as Based on adjusting the focus of the lens, the additional image is obtained.

24. The device according to claim 13, in, The at least one processor is configured to: adjusting the size of a bounding box associated with the first ROI; determining a resized bounding box to crop the skin information based on a bounding area of ​​the resized bounding box; and The resized bounding box is modified to include an area outside the resized bounding box corresponding to the skin information.

25. A non-transitory computer readable medium having stored thereon instructions which, when executed by at least one processor, cause the at least one processor to: Determine the first region of interest (ROI) in the image, in, The first ROI is associated with a first object; determining one or more image characteristics of the first ROI; and Based on one or more image characteristics of the first ROI, it is determined whether to perform an upsampling process on the image data in the first ROI.

26. The non-transitory computer readable medium of claim 25, in, The one or more image characteristics include at least one of a size of the first ROI or a distance of the first ROI from a focal point associated with an image.

27. The non-transitory computer readable medium of claim 26, in, When executed by at least one processor, the instructions cause the at least one processor to: determining that a size of the first ROI is less than a size threshold; and Based on the fact that the size of the first ROI is smaller than the size threshold, it is determined to perform an upsampling process on the image data in the ROI.

28. The non-transitory computer readable medium of claim 26, in, When executed by at least one processor, the instructions cause the at least one processor to: Determining that the distance between the first ROI and the focus is greater than a threshold distance; and Based on the fact that the distance between the first ROI and the focus is greater than the threshold distance, it is determined to perform an upsampling process on the image data in the first ROI.

29. The non-transitory computer readable medium of claim 25, in, When executed by at least one processor, the instructions cause the at least one processor to: determining a second ROI in the image, wherein the second ROI is associated with a second object; determining one or more image characteristics of the second ROI; and Based on one or more image characteristics of the second ROI, it is determined not to perform an upsampling process on the image data in the second ROI.

30. The non-transitory computer readable medium of claim 25, in, When executed by at least one processor, the instructions cause the at least one processor to: detecting key points associated with the first ROI, wherein the first ROI includes a face of a person; and The portion of the image is transformed based on aligning keypoints associated with the first ROI.