Image Processing Based on Object Category Classification

By applying region-specific ISP settings based on object category classification, the system optimizes image processing for diverse scenes, enhancing image quality by improving clarity and reducing noise in specific regions.

JP7713513B2Active Publication Date: 2025-07-25QUALCOMM INC
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
JP2023510415
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Priority Date
2021-01-26
Filing Date
2021-07-20
Publication Date
2025-07-25
Estimated Expiration
2041-07-20

AI Technical Summary

Technical Problem

Conventional image capture devices tune their image signal processors (ISPs) uniformly across all images, leading to suboptimal performance for diverse scenes, as they are designed to function smoothly across various types of scenes rather than optimizing for specific object categories within an image.

Method used

The system applies different ISP setting values to different image regions based on object category classification, adjusting parameters like sharpness and noise reduction based on the detected objects and their confidence levels, optimizing image processing for each region.

Benefits of technology

This approach enhances image quality by improving clarity in specific regions, such as human hair and skin, by applying tailored ISP settings, resulting in sharper textures and reduced noise, thus producing higher-quality images.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007713513000001
    Figure 0007713513000001
  • Figure 0007713513000002
    Figure 0007713513000002
  • Figure 0007713513000003
    Figure 0007713513000003
Patent Text Reader

Abstract

Examples are described for applying different settings for image capture to different portions of image data. For example, an image sensor can capture image data of a scene and send the image data to an image signal processor (ISP) and a classification engine for processing. The classification engine can determine that a first object image region depicts a first object category and a second object image region depicts a second object category. Different confidence regions of the image data can identify different levels of confidence in the classification. The ISP can generate an image by applying different settings to different portions of the image data. The different portions of the image data can be identified based on the object image regions and the confidence regions.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001]

[0001] This application relates to image capture and image processing. More particularly, this application relates to systems and methods for automatically guiding the image processing of a photograph based on the categorization of objects within a scene to be photographed.

Background Art

[0002]

[0002] An image capture device uses an image sensor having an array of photodiodes to capture an image by a first light from a scene. An image signal processor (ISP) then processes the raw image data captured by the photodiodes of the image sensor to store and make an image that can be viewed by a user. How the scene is depicted in the image depends in part on capture settings that control how much light is received by the image sensor, such as an exposure time setting value and an aperture size setting value. How the scene is depicted in the image also depends on how the ISP is tuned to process the photodiode data captured by the image sensor into an image.

[0003]

[0003] Conventionally, the ISP of an image capture device is tuned only once during manufacturing. Tuning of the ISP affects how every image is processed in the image capture device and affects every pixel of every image. A user typically expects the image capture device to capture high-quality images regardless of what scene is being photographed. To avoid situations where the image capture device cannot properly photograph some types of scenes, the ISP tuning is generally selected to function smoothly and without difficulty for as many types of scenes as possible. However, for this reason, conventional ISP tuning is generally not optimal for photographing all types of scenes.

Summary of the Invention

[0004] Systems and techniques for determining and applying different ISP setting values for different image regions are described herein. In some examples, an image capture / processing device can use different ISP setting values for different image regions to process raw image data captured by an image sensor. In some cases, a classification engine can classify the raw image data into different object image regions based on the detection of different types of objects within different image regions in the raw image data. By applying different ISP setting values to different regions within an image, the ISP can be optimized for the type of object depicted within the image. In one exemplary example, the ISP can use an ISP setting value that improves the sharpness of a region of an image depicting human hair, which can improve the clarity of the texture of the hair. Within the same image, the ISP can use different ISP setting values that reduce sharpness and improve noise reduction in a region of the image depicting human skin, which can result in a processed image depicting smoother skin. Different confidence regions of the image data can identify different confidence levels in the classification. The setting values can be further modified based on the confidence level. The strength of certain ISP parameters, such as noise reduction, sharpness, color saturation, or tone mapping, can be adjusted from the default value of a pixel based on the category of the object depicted at the pixel and the confidence level of that category classification.For example, an increase or decrease from a default value associated with a particular category of an object may be suppressed when the confidence level of that category classification is low, or may be expanded when the confidence level of that category classification is high.

[0005]

[0005] In one example, an apparatus for data encoding is provided. The apparatus includes a memory and one or more processors coupled to the memory (e.g., implemented in a circuit). The one or more processors are configured to receive image data captured by an image sensor, determine that a first object image region in the image data depicts a first category of object among a plurality of categories of object, determine that a second object image region in the image data depicts a second category of object among the plurality of categories of object, identify a plurality of confidence levels corresponding to a plurality of confidence image regions in the image data, wherein each confidence level of the plurality of confidence levels identifies the confidence that the corresponding confidence image region of the plurality of confidence image regions depicts one of the plurality of categories of object, generate an image based on the image data using an image capture process, including applying different settings of the image capture process to different portions of the image data, and the different portions of the image data are identified based on the first object image region, the second object image region, and the plurality of confidence image regions, and are capable of performing these operations.

[0006]

[0006] In another example, a method of image processing is provided. The method includes receiving image data captured by an image sensor. The method includes determining that a first object image region in the image data depicts a first object category among a plurality of object categories. The method includes determining that a second object image region in the image data depicts a second object category among the plurality of object categories. The method includes identifying corresponding plurality of confidence levels among a plurality of confidence image regions of the image data, wherein each confidence level of the plurality of confidence levels identifies a confidence that a corresponding confidence image region among the plurality of confidence image regions depicts one of the plurality of object categories. The method includes generating an image based on the image data using an image capture process, including by applying different setting values of the image capture process to different portions of the image data, wherein the different portions of the image data are identified based on the first object image region, the second object image region, and the plurality of confidence image regions.

[0007]

[0007] In another example, when executed by one or more processors, the one or more processors are caused to receive image data captured by an image sensor, determine that a first object image region in the image data depicts a first object category out of a plurality of object categories, determine that a second object image region in the image data depicts a second object category out of the plurality of object categories, identify a plurality of confidence levels corresponding to a plurality of confidence image regions of the image data, where each confidence level of the plurality of confidence levels identifies the confidence that the corresponding confidence image region of the plurality of confidence image regions depicts one of the plurality of object categories, generate an image based on the image data using an image capture process, including by applying different setting values of the image capture process to different portions of the image data, where the different portions of the image data are identified based on the first object image region, the second object image region, and the plurality of confidence image regions, and a non-transitory computer readable storage medium storing instructions for causing the foregoing is provided.

[0008]

[0008] In another example, an apparatus for image processing is provided. The apparatus includes means for receiving image data captured by an image sensor. The apparatus includes means for determining that a first object image region in the image data depicts a first object category among a plurality of object categories. The apparatus includes means for determining that a second object image region in the image data depicts a second object category among a plurality of object categories. The apparatus includes means for identifying a plurality of confidence levels corresponding to a plurality of confidence image regions of the image data, wherein each confidence level of the plurality of confidence levels identifies a confidence that the corresponding confidence image region of the plurality of confidence image regions depicts one of the plurality of object categories. The apparatus includes means for generating an image based on the image data using an image capture process, including by applying different setting values of the image capture process to different portions of the image data, wherein the different portions of the image data are identified based on the first object image region, the second object image region, and the plurality of confidence image regions.

[0009]

[0009] In some embodiments, the methods, apparatuses, and computer-readable media described above further comprise generating one or more modifiers, the one or more modifiers identifying at least one of a first deviation from a default setting of an image capture process for a first object image region and a second deviation from a default setting of an image capture process for a second object image region, wherein different settings of the image capture process are based on the one or more modifiers. In some embodiments, the methods, apparatuses, and computer-readable media described above further comprise adjusting the one or more modifiers, including blending the one or more modifiers with a blending update value based on a plurality of confidence levels corresponding to a plurality of confidence image regions, wherein blending the one or more modifiers with the blending update value adjusts at least one of the first deviation and the second deviation in at least one area of the image data.

[0010]

[0010] In some aspects, the methods, apparatuses, and computer-readable media described above further comprise generating a category map that divides image data into a plurality of object image regions including a first object image region and a second object image region, wherein each object image region of the plurality of object image regions corresponds to one of a plurality of object categories, identifying that a first object category corresponds to a first set value of an image capture process, and identifying that a second object category corresponds to a second set value of the image capture process. In some aspects, the methods, apparatuses, and computer-readable media described above further comprise generating a confidence map that divides image data into a plurality of confidence image regions corresponding to a plurality of confidence levels, wherein different portions of the image data are identified based on the category map and the confidence map.

[0011]

[0011] In some embodiments, the image capture process includes processing image data using an image signal processor (ISP) of one or more processors, where different setting values of the image capture process are different tuning setting values of the ISP. In some embodiments, the different tuning settings of the ISP include different intensities to which an ISP tuning parameter is applied during the processing of the image data using the ISP, where the ISP tuning parameter is one of noise reduction, sharpening, color saturation, color mapping, color processing, and tone mapping. In some embodiments, the different setting values include settings related to at least one of lens position, flash, focus, exposure, white balance, aperture size, shutter speed, ISO, analog gain, digital gain, denoising, sharpening, tone mapping, color saturation, demosaicking, color space conversion, shading, edge enhancement, high dynamic range (HDR) image combining, special effect, artificial noise addition, edge-directed upscaling, upscaling, downscaling, and electronic image stabilization.In some aspects, the methods, apparatuses, and computer-readable media described above further comprise processing the image data, including at least one of demosaicking the image data and converting the image data from a first color space to a second color space.

[0012]

[0012] In some aspects, the methods, apparatuses, and computer-readable media described above further comprise receiving a user input related to at least one of a first object image region and a second object image region, wherein at least one of different setting values is defined based on the user input and corresponds to one of the first object image region and the second object image region. In some aspects, applying different setting values of an image capture process to different portions of the image data includes applying different setting values of the image capture process to different portions of the image data using an image signal processor (ISP). In some aspects, identifying the first object image region and the second object image region includes identifying the first object image region and the second object image region using a classification engine that is at least partially located on an integrated circuit chip. In some aspects, the methods, apparatuses, and computer-readable media described above further comprise displaying an image on a display.

[0013]

[0013] In some embodiments, the device comprises a camera, a mobile device (e.g., a mobile phone or so-called "smartphone" or other mobile device), a wireless communication device, a wearable device, an extended reality device (e.g., a virtual reality (VR) device, an augmented reality (AR) device, or a mixed reality (MR) device), a personal computer, a laptop computer, a server computer, or other device. In some embodiments, one or more processors include an image signal processor (ISP). In some embodiments, the device includes one or more cameras for capturing one or more images. In some embodiments, the device includes an image sensor for capturing image data. In some embodiments, the device further includes a display for displaying images, one or more notifications (e.g., related to the processing of images), and / or other displayable data.

[0014]

[0014] The summary of the present invention is not intended to identify key features or essential features of the claimed subject matter, nor is it intended to be used alone to determine the scope of the claimed subject matter. The subject matter should be understood by reference to the entire specification of this patent, any or all of the drawings, and the appropriate portions of each claim.

[0015]

[0015] The above will become more apparent when considered in conjunction with other features and embodiments, with reference to the following specification, claims, and accompanying drawings.

[0016]

[0016] Exemplary embodiments of the present application will be described in detail below with reference to the following figures.

Brief Description of the Drawings

[0017]

Figure 1

Figure 2

[0018] Conceptual diagram showing image processing using a category map and a confidence map.

Figure 3

[0019] Conceptual diagram showing an image signal processor (ISP) pipeline for image processing based on object category classification.

Figure 4

[0020] Conceptual diagram showing a map decoder pipeline.

Figure 5A

[0021] Conceptual diagram showing the application of a modifier in an ISP module applied as a multiplier.

Figure 5B

[0022] Conceptual diagram showing the application of a modifier in an ISP module applied as an offset.

Figure 5C

[0023] Conceptual diagram showing the application of a modifier in an ISP module applied using parameter-based logic.

Figure 6

[0024] Conceptual diagram showing visual image artifacts caused by anomalies in the segmentation of an image into image regions during the generation of a category map.

Figure 7

[0025] Conceptual diagram showing a pipeline of a smooth transition map processor.

Figure 8

[0026] Conceptual diagram showing the smoothing of a modifier corresponding to an image region using a smooth transition map processor.

Figure 9

[0027] Diagram showing the pipeline of a category map upscaler (CMUS).

Figure 10

[0028] A diagram showing a comparison between a category map upscaled using nearest neighbor upscaling and the same category map upscaled using nearest neighbor upscaling modified by spatial weight filtering applied using a category map upscaler (CMUS).

Figure 11

[0029] A conceptual diagram showing exemplary resolutions of image data corresponding to a category map during downscaling and upscaling operations.

Figure 12A

[0030] A flowchart showing an image processing technique.

Figure 12B

[0031] A flowchart showing an image processing technique.

Figure 13

[0032] A flowchart showing a transition smoothing technique.

Figure 14

[0033] A flowchart showing an image upscaling technique.

Figure 15

[0034] A diagram showing an example of a system for implementing some aspects of the present technology.

Best Mode for Carrying Out the Invention

[0018]

[0035] Some aspects and embodiments of the present disclosure are provided below. As will be apparent to those skilled in the art, some of these aspects and embodiments may be applied independently, and some of them may be applied in combination. In the following description, for the purpose of explanation, specific details are set forth in order to provide a complete understanding of the embodiments of the present application. However, it will be apparent that various embodiments may be practiced without these specific details. The figures and the description are not intended to be limiting.

[0019]

[0036] The following description merely provides exemplary embodiments and is not intended to limit the scope, applicability, or configuration of the present disclosure. Rather, the following description of the exemplary embodiments will provide those skilled in the art with an explanation that enables the implementation of the exemplary embodiments. It should be understood that various changes can be made to the functions and configurations of the elements without departing from the spirit and scope of the present application as set forth in the appended claims.

[0020]

[0037] An image capture device (e.g., a camera) is a device that uses an image sensor to receive light and capture an image frame such as a still image or a video frame. The terms "image", "image frame", and "frame" are used interchangeably herein. An image capture device typically includes at least one lens that receives light from a scene and bends the light towards the image sensor of the image capture device. The light received by the lens passes through an aperture controlled by one or more control mechanisms and is received by the image sensor. The one or more control mechanisms can control exposure, focus, and / or zoom based on information from the image sensor and / or information from an image processor (e.g., a host process or application process and / or an image signal processor). In some examples, the one or more control mechanisms include a motor or other control mechanism that moves the lens of the image capture device to a target lens position.

[0021]

[0038] As described in more detail below, systems and techniques are described herein for determining and applying different setting values for image capture processing of different image regions of image data from an image sensor. In some examples, an image capture / processing device can process image data captured by an image sensor using different setting values for different image regions. The image data may be raw image data or data that has been partially processed by an image signal processor (ISP) or other components. For example, the raw image data or partially processed image data can be processed by the ISP using demosaicing, color space conversion, and / or other processing operations described herein.

[0022]

[0039] In some cases, the classification engine can partition the image data into different image regions based on the detection of different types of objects within different image regions of the image data. By applying different setting values to different regions within the image data, an image is generated in which the capture and / or processing of the image data within the image is optimized for each of the types of objects depicted within the image. In some examples, the setting values can be associated with some of the ISP tuning parameters of the ISP. In one exemplary example, the ISP can process the image data using an ISP setting value that improves the sharpness of the region of the image data depicting human hair, and this ISP setting value can improve the clarity of the texture of the hair in the processed image. The ISP can use different ISP setting values that reduce sharpness and improve noise reduction in different regions of the image data depicting human skin. The different ISP setting values result in a processed image depicting smooth skin, but can also depict sharp and textured hair. In some examples, the setting values can be applied to image capture setting values such as focus, exposure time, aperture size, ISO, flash, any combination thereof, and / or other image capture setting values described herein. In some examples, the setting values can be applied to post-processing setting values that are applied after the ISP has already converted the image data from the raw image data from the image sensor into an image. The post-processing setting values can include setting values such as brightness, contrast, saturation, tone level, histogram, any combination setting values thereof, and / or other processing setting values described herein.

[0023]

[0040] Figure 1 is a block diagram showing the architecture of an image capture / processing system 100. The image capture / processing system 100 includes various components used to capture and process an image of a scene (e.g., an image of scene 110). The image capture / processing system 100 can capture an independent image (or photograph) and / or can capture a video including a plurality of images (or video frames) of a particular sequence. The lens 115 of system 100 faces the scene 110 and receives light from the scene 110. The lens 115 bends the light towards the image sensor 130. The light received by the lens 115 passes through an aperture controlled by one or more control mechanisms 120 and is received by the image sensor 130.

[0024]

[0041] One or more control mechanisms 120 may control exposure, focus, and / or zoom based on information from the image sensor 130 and / or based on information from the image processor 150. One or more control mechanisms 120 may include a plurality of mechanisms and components. For example, the control mechanism 120 may include one or more exposure control mechanisms 125A, one or more focus control mechanisms 125B, and / or one or more zoom control mechanisms 125C. One or more control mechanisms 120 may also include additional control mechanisms other than those shown, such as control mechanisms that control analog gain, flash, HDR, depth of field, and / or other image capture characteristics.

[0025]

[0042] The focus control mechanism 125B of the control mechanism 120 can obtain a focus setting value. In some examples, the focus control mechanism 125B stores the focus setting value in a memory register. Based on the focus setting value, the focus control mechanism 125B can adjust the position of the lens 115 relative to the position of the image sensor 130. For example, based on the focus setting value, the focus control mechanism 125B can move the lens 115 closer to or farther from the image sensor 130 by operating a motor or a servo (or other lens mechanism), thereby adjusting the focus. In some cases, the system 100 may include additional lenses, such as one or more microlenses over each photodiode of the image sensor 130, and each additional lens bends the light received from the lens 115 toward the corresponding photodiode before the light reaches the photodiode. The focus setting value can be determined via contrast detection autofocus (CDAF), phase detection autofocus (PDAF), hybrid autofocus (HAF), or some combination thereof. The focus setting value can be determined using the control mechanism 120, the image sensor 130, and / or the image processor 150. The focus setting value can be referred to as an image capture setting value and / or an image processing setting value.

[0026]

[0043] The exposure control mechanism 125A of the control mechanism 120 can obtain an exposure setting value. In some cases, the exposure control mechanism 125A stores the exposure setting value in a memory register. Based on this exposure setting value, the exposure control mechanism 125A can control the size of the aperture (e.g., aperture size or f-stop), the duration for which the aperture is open (e.g., exposure time or shutter speed), the sensitivity of the image sensor 130 (e.g., ISO speed or film speed), the analog gain applied by the image sensor 130, or any combination thereof. The exposure setting value can be referred to as an image capture setting value and / or an image processing setting value.

[0027]

[0044] The zoom control mechanism 125C of the control mechanism 120 can acquire a zoom setting value. In some examples, the zoom control mechanism 125C stores the zoom setting value in a memory register. Based on the zoom setting value, the zoom control mechanism 125C can control the focal length of a lens assembly (lens assembly) including the lenses 115 and one or more additional lenses. For example, the zoom control mechanism 125C can control the focal length of the lens assembly by operating one or more motors or servos (or other lens mechanisms) to move one or more of the lenses relative to each other. The zoom setting value may be referred to as an image capture setting value and / or an image processing setting value. In some examples, the lens assembly may include an afocal zoom lens or a variable focal length zoom lens. In some examples, the lens assembly may include a focusing lens (which may be the lens 115 in some cases) that first receives light from the scene 110, and then the light passes through an afocal zoom system between the focusing lens (e.g., the lens 115) and the image sensor 130 before the light reaches the image sensor 130. The afocal zoom system may include two positive (e.g., converging, convex) lenses of equal or similar focal lengths (e.g., within a threshold difference of each other), which may have a negative (e.g., diverging, concave) lens in between in some cases. In some cases, the zoom control mechanism 125C moves one or more of the lenses within the afocal zoom system, such as one or both of the negative and positive lenses.

[0028]

[0045] The image sensor 130 includes one or more arrays of photodiodes or other photosensitive elements. Each photodiode measures the amount of light that ultimately corresponds to a particular pixel in the image generated by the image sensor 130. In some cases, different photodiodes may be covered by different color filters, and thus may measure light that corresponds to the color of the filter covering the photodiode. For example, a Bayer color filter includes a red color filter, a blue color filter, and a green color filter, and each pixel of the image is generated based on red light data from at least one photodiode covered by the red color filter, blue light data from at least one photodiode covered by the blue color filter, and green light data from at least one photodiode covered by the green color filter. Other types of color filters may use yellow, magenta, and / or cyan (also called "emerald") color filters instead of, or in addition to, red, blue, and / or green color filters. Some image sensors (e.g., image sensor 130) may have no color filter, and instead may use different photodiodes across the pixel array (which may be vertically stacked in some cases). Different photodiodes across the pixel array have different spectral sensitivity curves and can thus respond to light of different wavelengths. Monochrome image sensors may also lack a color filter and thus may lack color depth.

[0029]

[0046] In some cases, the image sensor 130 may alternatively or additionally include an opaque mask and / or a reflective mask that blocks light from reaching a certain photodiode or a part of a certain photodiode at a certain time and / or from a certain angle, for use in phase detection autofocus (PDAF). The image sensor 130 may also include an analog gain amplifier for amplifying the analog signal output by the photodiode, and / or an analog-to-digital converter (ADC) for converting the analog signal output of the photodiode (and / or amplified by the analog gain amplifier) into a digital signal. In some cases, some of the components or functions described with respect to one or more of the control mechanisms 120 may instead or in addition be included in the image sensor 130. The image sensor 130 can be a charge-coupled device (CCD) sensor, an electron-multiplying CCD (EMCCD) sensor, an active pixel sensor (APS), a complementary metal-oxide-semiconductor (CMOS), an N-type metal-oxide-semiconductor (NMOS), a hybrid CCD / CMOS sensor (e.g., sCMOS), or some other combination thereof.

[0030]

[0047] The image processor 150 may include one or more processors, such as one or more image signal processors (ISPs) (including the ISP 154), one or more host processors (including the host processor 152), and / or one or more of any other type of processor 1110 described with respect to the computing device 1100. The host processor 152 can be a digital signal processor (DSP) and / or other types of processors. In some implementations, the image processor 150 is a single integrated circuit or chip (e.g., called a system-on-chip or SoC) that includes the host processor 152 and the ISP 154. In some cases, the chip can also include one or more input / output ports (e.g., input / output (I / O) port 156), a central processing unit (CPU), a graphics processing unit (GPU), a broadband modem (e.g., 3G, 4G, or LTE (registered trademark), 5G, etc.), memory, connectivity components (e.g., Bluetooth (registered trademark), global positioning system (GPS), etc.), any combination thereof, and / or other components. The I / O port 156 can include any suitable input / output port or interface according to one or more protocols or specifications, such as an inter-integrated circuit 2 (I2C) interface, an inter-integrated circuit 3 (I3C) interface, a serial peripheral interface (SPI) interface, a serial general-purpose input / output (GPIO) interface, a mobile industry processor interface (MIPI::Mobile Industry Processor Interface) (such as a MIPI CSI-2 physical (PHY) layer port or interface), an advanced high-performance bus (AHB) bus, any combination thereof, and / or other input / output ports. In one exemplary example, the host processor 152 can communicate with the image sensor 130 using an I2C port, and the ISP 154 can communicate with the image sensor 130 using a MIPI port.

[0031]

[0048] The image processor 150 may perform several tasks such as demosaicing, color space conversion, image frame downsampling, pixel interpolation, automatic exposure (AE) control, automatic gain control (AGC), CDAF, PDAF, automatic white balance, merging of image frames to form an HDR image, image recognition, object recognition, feature recognition, receiving of inputs, management of outputs, management of memory, or some combination thereof. The image processor 150 may store the image frames and / or the processed images in a random access memory (RAM) 140 / 1020, a read only memory (ROM) 145 / 1025, a cache, a memory unit, another storage device, or some combination thereof.

[0032]

[0049] A variety of input / output (I / O) devices 160 can be connected to the image processor 150. The I / O device 160 can include a display screen, a keyboard, a keypad, a touch screen, a trackpad, a touch-sensitive surface, a printer, any other output device 1035, any other input device 1045, or some combination thereof. In some cases, captions may be input into the image processing device 105B through the physical keyboard or keypad of the I / O device 160, or through the virtual keyboard or keypad of the touch screen of the I / O device 160. The I / O 160 may include one or more ports, jacks, or other connectors that enable a wired connection between the system 100 and one or more peripheral devices, through which the system 100 can receive data from and / or transmit data to one or more peripheral devices. The I / O 160 may include one or more wireless transceivers that enable a wireless connection between the system 100 and one or more peripheral devices, through which the system 100 can receive data from and / or transmit data to one or more peripheral devices. The peripheral devices may include any of the types of I / O devices 160 described above, and when coupled to a port, jack, wireless transceiver, or other wired and / or wireless connector, they can themselves be considered I / O devices 160.

[0033]

[0050] In some cases, the image capture / processing system 100 can be a single device. In some cases, the image capture / processing system 100 can be two or more separate devices, including an image capture device 105A (e.g., a camera) and an image processing device 105B (e.g., a computing device coupled to the camera). In some implementations, the image capture device 105A and the image processing device 105B can be coupled to each other, for example, via one or more wires, cables, or other electrical connectors and / or wirelessly via one or more wireless transceivers. In some implementations, the image capture device 105A and the image processing device 105B may be separated from each other.

[0034]

[0051] As shown in FIG. 1, the vertical dashed line divides the image capture / processing system 100 of FIG. 1 into two parts, each representing an image capture device 105A and an image processing device 105B. The image capture device 105A includes a lens 115, a control mechanism 120, and an image sensor 130. The image processing device 105B includes an image processor 150 (including an ISP 154 and a host processor 152), a RAM 140, a ROM 145, and an I / O 160. In some cases, some components shown within the image capture device 105A, such as the ISP 154 and / or the host processor 152, can be included in the image capture device 105A.

[0035]

[0052] The image capture / processing system 100 can include an electronic device such as a mobile or fixed telephone handset (e.g., smartphone, cellular phone, etc.), desktop computer, laptop or notebook computer, tablet computer, set-top box, television, camera, display device, digital media player, video game console, video streaming device, Internet Protocol (IP) camera, or any other suitable electronic device. In some examples, the image capture / processing system 100 can include one or more wireless transceivers for wireless communication such as cellular network communication, 802.11 Wi-Fi® communication, wireless local area network (WLAN) communication, or some combination thereof. In some implementations, the image capture device 105A and the image processing device 105B can be different devices. For example, the image capture device 105A can include a camera device, and the image processing device 105B can include a computing device such as a mobile handset, desktop computer, or other computing device.

[0036]

[0053] Although the image capture / processing system 100 is shown as including several components, those skilled in the art will understand that the image capture / processing system 100 can include more components than those shown in FIG. 1. The components of the image capture / processing system 100 can include software, hardware, or one or more combinations of software and hardware. For example, in some implementations, the components of the image capture / processing system 100 can include one or more programmable electronic circuits (e.g., microprocessors, GPUs, DSPs, CPUs, and / or other suitable electronic circuits), can include electronic circuits or other electronic hardware, and / or may be implemented using them, and / or can include computer software, firmware, or any combination thereof, and / or may be implemented using them to perform the various operations described herein. The software and / or firmware can be stored on a computer-readable storage medium and can include one or more instructions executable by one or more processors of an electronic device implementing the image capture / processing system 100.

[0037]

[0054] The ISP is tuned by selecting setting values for some ISP tuning parameters. The setting values of the ISP tuning parameters can be referred to as ISP setting values, ISP tuning setting values, ISP tuning parameter setting values, tuning setting values, tuning parameter setting values, or some combination thereof. The ISP processes an image using the setting values selected for the ISP tuning parameters. Tuning the ISP is a computationally expensive process, and thus, the ISP has conventionally been tuned only once during manufacturing using fixed tuning techniques. The setting values of the ISP tuning parameters are not conventionally changed after manufacturing and are thus globally used across each pixel of every image processed by the ISP. To avoid situations where an image capture device cannot properly capture some types of scenes, the tuning of the ISP is generally selected to function smoothly without difficulty for as many types of scenes as possible. However, for this reason, conventional ISP tuning is generally not optimal for capturing any type of scene. As a result, conventional ISP tuning leaves the conventional ISP as a jack of all trades but potentially a master of none.

[0038]

[0055] Although ISP tuning has a high computational cost, it is possible to generate multiple setting values for several ISP tuning parameters during manufacturing. For example, for an ISP tuning parameter such as sharpness, a high sharpness setting value may correspond to an increased sharpness level, while a low sharpness setting value may correspond to a decreased sharpness level. Different setting values may be useful when the captured image mainly depicts a single type of object, such as a close-up image of a plant, a human face, a vehicle, or food. In the case of an image of a human face, a low sharpness setting value may be selected through the user interface or automatically based on the detection of the face in the preview image to depict smoother facial skin. In the case of an image of a plant, a high sharpness setting value may be selected through the user interface or automatically based on the detection of the plant in the preview image to depict the texture of the plant's leaves and flowers in more detail. However, since most images depict many types of objects, images that depict only one type of object are rare. In the case of an image that depicts multiple types of objects, the use of adjusted setting values may have unwanted effects. For example, if the image depicts both a face and a plant, using a high sharpness setting value may make the facial skin look uneven, while using a low sharpness setting value may make the plant's leaves and flowers look mixed together. To avoid such unwanted effects, such adjusted setting values are likely to be used very sparingly in an ISP that globally applies the tuning setting values to all pixels.

[0039]

[0056] Figure 2 is a conceptual diagram 200 showing image processing using the category map 230 and the confidence map 235. The diagram 200 shows three hardware components of the image capture / processing device 100, namely, the image sensor 205, the ISP 240, and the classification engine. The image sensor 205 can be an example of the image sensor 130 in FIG. 1. The ISP 240 can be an example of the ISP 154 in FIG. 1. The classification engine 220 can be an example of the host processor 152 in FIG. 1, the ISP 154 in FIG. 1, the image processor 150 in FIG. 1, a central processing unit (CPU), a graphics processing unit (GPU), an accelerated processing unit (APU), another type of processor 1510, or some combination thereof.

[0040]

[0057] The image sensor 205 receives light 202 from the scene being photographed by the image capture / processing device 100. The scene being photographed in FIG. 200 is a scene of a child eating food from a plate set on a table. The image sensor 205 captures raw image data 210 based on the light 202 from the scene. The raw image data 210 is a signal from the photodiodes of the photodiode array of the image sensor 205 and, in some cases, is amplified via an analog gain amplifier and, in some cases, is converted from an analog format to a digital format using an analog-to-digital converter (ADC). The raw image data 210 generally includes separate image data from different photodiodes having different color filters. For example, if the image sensor 205 uses a Bayer color filter, the raw image data 210 includes image data corresponding to red-filtered photodiodes, image data corresponding to green-filtered photodiodes, and image data corresponding to blue-filtered photodiodes.

[0041]

[0058] The image sensor 205 transmits a first copy 210 of the raw image data to the ISP 240. The image sensor 205 transmits a second copy 215 of the raw image data to the classification engine 220. Optionally, the second copy 215 of the raw image data can be downscaled at, for example, a ratio of 1:2, a ratio of 1:4, a ratio of 1:8, a ratio of 1:16, a ratio of 1:32, a ratio of 1:64, a ratio of higher than 1:64, or another ratio between any of the ratios listed above. The second copy 215 of the raw image data is shown in FIG. 200 as already downscaled for simplicity, but it should be understood that the second copy 215 of the raw image data can be downscaled by the classification engine 220 when the second copy 215 of the raw image data is received. Optionally, the second copy 215 of the raw image data can be downscaled before the second copy 215 of the raw image data is transmitted to the classification engine 220 and / or before it is received by the classification engine 220. For example, the second copy 215 of the raw image data can be downscaled in the image sensor 205, in the ISP 240, in a downscaler component (not shown) separate from the image sensor 205 and the ISP 240, or in any combination thereof. The downscaling techniques used to downscale the second copy 215 of the raw image data can include nearest neighbor downscaling, bilinear interpolation, bicubic interpolation, Sinc resampling, Lanczos resampling, box sampling, mipmapping, Fourier transform scaling, edge-directed interpolation, high-quality scaling (hqx), or any combination thereof. Optionally, the second copy 215 of the raw image data can be at least partially processed by the ISP 240 and / or one or more additional components before the second copy 215 of the raw image data is transmitted to the classification engine 220 and / or before it is received by the classification engine 220. For example, the second copy 215 of the raw image data can be demosaicked by the ISP 240 before the second copy 215 of the raw image data is received by the classification engine 220.The second copy 215 of the raw image data can be converted from one color space (e.g., RGB color space) to another color space (e.g., YUV color space) before the second copy 215 of the raw image data is received by the classification engine 220. This process can be performed before or after downscaling the second copy 215 of the raw image data.

[0042]

[0059] In some examples, the classification engine 220 receives a downscaled second copy 215 of the raw image data. In some examples, the classification engine 220 receives a second copy 215 of the raw image data and performs downscaling to generate a downscaled second copy 215 of the raw image data. The classification engine 220 can use the downscaled second copy 215 of the raw image data to generate a category map 230 and a confidence map 235. For example, the classification engine 220 can partition the downscaled second copy 215 of the raw image data into different image regions based on the detection of different object categories within different image regions of the downscaled second copy 215 of the raw image data. As an illustrative example, the category map 230 shown in FIG. 200 is shaded with a first (diagonal stripe) shading pattern and includes two image regions corresponding to the face and arm of a child labeled "skin," meaning that the classification engine 220 detects skin in those image regions. Similarly, the category map 230 shown in FIG. 200 is shaded with a second (diagonal stripe) shading pattern and includes an image region corresponding to the hair on the head of the child labeled "hair," meaning that the classification engine 220 detects hair in that image region. In some cases, regions with eyelashes, eyebrows, mustaches, beards, and / or other hair objects can also be identified as hair by the classification engine 220.The other image regions shown in the category map 230 include an image region shaded with a third (diagonal stripe) shading pattern and labeled "shirt" where the classification engine 220 detects a shirt, several image regions shaded with a fourth (diagonal stripe) shading pattern and labeled "food" where the classification engine 220 detects food, two image regions shaded with a fifth (diagonal stripe) shading pattern and labeled "cloth" where the classification engine 220 detects cloth, three image regions shaded with a sixth (diagonal stripe) shading pattern and labeled "metal" where the classification engine 220 detects metal, and two image regions shaded with a seventh (cross - hatch) shading pattern and labeled "undefined" where the classification engine 220 does not know what is depicted. The image regions classified by the classification engine 220 as depicting different object categories can depict different objects, different types of objects, different materials, different substances, different elements, different components, objects with different attributes, or some combination thereof. The different shading patterns within the category map 230 of FIG. 2 can represent different values stored at the corresponding pixel locations within the category map 230, such as different colors, different shades of gray, different numbers, different characters, or different bit sequences. In some cases, an image region determined to depict a particular object category may be referred to as an image region, an object image region, an image object region, a category image region, an image category region, a category region, an object category region, an object category image region, an image object category region, or some combination thereof. For example, an image region determined to depict a first object category may be referred to as a first object image region, an image region determined to depict a second object category may be referred to as a second object image region, and so on.

[0043]

[0060] The confidence map 235 identifies the confidence that the classification engine 220 has in the classification of a given pixel within the category map 230. An area of image data having a particular confidence level may be referred to as a confidence region, a confidence image region, an image region, a region, a portion, or a combination thereof. Pixels shown in white in the confidence map 235 represent high confidence levels, such as a confidence level exceeding a high threshold percentage, such as 90%. Pixels shown in black in the confidence map 235 represent low confidence levels, such as a confidence level below a low threshold percentage, such as 10%. The confidence map 235 also includes six different shades of gray (other than black and white), with each shade of gray representing a confidence level falling within a different confidence range between the high threshold percentage and the low threshold percentage. For example, the lightest shade of gray (still darker than white) may represent a confidence value between 90% and 80%, the next shade of gray, which is one shade darker than the previously listed shade of gray, may represent a confidence value between 80% and 70%, the next shade of gray, which is one shade darker than the previously listed shade of gray, may represent a confidence value between 70% and 60%, and so on. Including black, white, and the six shades of gray between black and white, an exemplary confidence map 235 includes a total of eight shades of gray, which correspond to eight possible confidence levels. In some examples, the confidence level of a particular pixel may be stored as a 3-bit value. The classification engine 220 transmits the category map 230 and the confidence map 235 to the ISP 240. The confidence map 235 may visually appear to have visible banding between different shades of gray, for example, if there is a gradient between the shades. In some examples, the shades of gray may be mapped to confidence levels in the opposite direction as described above, such that black and darker shades of gray represent higher confidence values and white and lighter shades of gray represent lower confidence values.

[0044]

[0061] In some cases, the category map 230 and the confidence map 235 can be a single file, data stream, and / or set of metadata. It should be understood that any description herein of either the category map 230 or the confidence map 235 potentially includes both the category map 230 and the confidence map 235. In one example, the single file can be an image. For example, each pixel of the image can include one or more values corresponding to the category classification and confidence associated with the corresponding pixel of the second copy 215 of the raw image data. In another example, the single file can be a matrix or table, and each cell of the matrix or table stores a value corresponding to the category classification and confidence associated with the corresponding pixel of the second copy 215 of the raw image data. For a given pixel of the second copy 215 of the raw image data, the file stores the value in the corresponding cell or pixel. In some examples, the first plurality of bits in the stored value represent the object category that the classification engine 220 classifies as depicting the pixel. In such examples, the second plurality of bits in the stored value can represent the confidence of the classification engine 220 in classifying the pixel as depicting the object category.

[0045]

[0062] In one exemplary example, the value stored in the file can be 8 bits long, which can be referred to as a byte or octet. The first plurality of bits that identify the object category can be 5 bits out of the 8 - bit value, such as the first or most significant 5 bits of the 8 - bit value. In the case of 5 bits, the first plurality of bits can identify 32 possible object categories. The first plurality of bits can represent the most significant bit (MSB) of the stored value. In the above - mentioned exemplary example, the second plurality of bits that represent the confidence level can be 3 bits out of the 8 - bit value, such as the last or least significant 3 bits of the 8 - bit value. In the case of 3 bits, the second plurality of bits can identify 8 possible confidence values. The second plurality of bits can represent the least significant bit (LSB) of the stored value. In some cases, the first plurality of bits can be later in the value than the second plurality of bits. In some cases, different decompositions in bit - length division are possible. For example, the first plurality of bits and the second plurality of bits can each include 4 bits. The first plurality of bits can include 1 bit, 2 bits, 3 bits, 4 bits, 5 bits, 6 bits, 7 bits, or 8 bits. The second plurality of bits can include 1 bit, 2 bits, 3 bits, 4 bits, 5 bits, 6 bits, 7 bits, or 8 bits. In some examples, the value stored in the file can be shorter or longer than 8 bits.

[0046]

[0063] In some examples, as described above, the confidence map 235 and the category map 230 are a single file that stores two separate values for each pixel of the second copy 215 of the raw image data. One of the values represents the object category and the other value represents the confidence level. In some examples, the confidence map 235 and the category map 230 are separate files, where one file stores the value representing the object category and the other file stores the value representing the confidence level.

[0047]

[0064] In some cases, the classification engine 220 upsamples the category map 230 and the confidence map 235 from the size of the downscaled second copy 215 of the raw image data to the size of the first copy 215 of the raw image data and / or the size of the processed image 250. This upscaling process is also called nearest neighbor (NN) upscaling or special category map upscaling (CMUS), which is shown and described with respect to at least FIGS. 9, 10, and 11, and is modified by spatial weight filtering. In some cases, the classification engine 220 blends or merges the category map 230 and the confidence map 235 into a single file before sending the single file to the ISP 240.

[0048]

[0065] The ISP 240 receives the first copy 210 of the raw image data from the image sensor 205 and receives the category map 230 and the confidence map 235 from the classification engine 220. In some cases, the ISP 240 can perform some early processing tasks on the first copy 210 of the raw image data while the classification engine 220 is generating the category map 230 and the confidence map 235 and / or before the ISP 240 receives the category map 230 and the confidence map 235 from the classification engine 220. These early image processing tasks can include, for example, demosaicing, color space conversion (e.g., from RGB to YUV), pixel interpolation, and / or downsampling. However, in other cases, the ISP 240 can delay some or all of these early image processing tasks until the ISP 240 receives the category map 230 and the confidence map 235 from the classification engine 220.

[0049]

[0066] When the ISP240 receives the category map 230 and the confidence map 235 from the classification engine 220, the ISP240 uses the category map 230 and the confidence map 235 from the classification engine 220 to process the image. The ISP240 includes a plurality of modules that control the application of different ISP tuning parameters that the ISP240 can set to different values. In one example, the ISP tuning parameters include noise reduction (NR), sharpening, tone mapping (TM), and color saturation (CS). In some cases, the ISP tuning parameters may include gamma, gain, luminance, shading, edge enhancement, color correction (CC), color mapping (CM) (e.g., based on a 2D look-up table and / or a 3D look-up table), color shift, color enhancement, high dynamic range (HDR) image composition, special effect processing (e.g., background replacement, blur effect), artificial noise (e.g., grain) adder, demosaicing, edge-directed upscaling, other processing parameters described herein, or combinations thereof. The different setting values of each ISP module can include a default setting value (also referred to as a default ISP tuning setting value), one or more setting values that increase the application of the ISP tuning parameter relative to the default setting value, and one or more setting values that decrease the application of the ISP tuning parameter relative to the default. For example, for the noise reduction (NR) ISP tuning parameter, the available setting values can include a default level of noise reduction, one or more increased levels of noise reduction that perform more noise reduction than the default level, and one or more decreased levels of noise reduction that perform less noise reduction than the default level. In some cases, one or more of the different ISP tuning parameters can include sub-parameters. The setting values of such ISP tuning parameters can include values or modifiers for one or more of these sub-parameters.For example, the NR ISP tuning parameters can include a luma NR strength, a chroma NR strength, and sub-parameters including a temporal filter (e.g., for noise reduction of a video or sequence). The set value of NR can include modifiers for the luma NR strength sub-parameter, the chroma NR strength sub-parameter, and / or the temporal filter sub-parameter. In some examples, the color saturation (CS) module 350 can control color correction (CC) and / or color mapping (CM) and / or color shift and / or color enhancement instead of or in addition to color saturation (CS). The color saturation (CS:color saturation) module 350 can be referred to as a module for any of the above parameters (e.g., a CC module, a CM module, a color shift module, a color enhancement module). The color saturation (CS) module 350 can sometimes be referred to as a color processing module. In some examples, the ISP pipeline 305 can include separate color correction (CC) and / or color mapping (CM) and / or color shift and / or color enhancement modules. In some examples, noise reduction (NR) can include spatial noise reduction, temporal noise reduction, or both. The NR module 320 is shown as a single module in FIG. 3, but in some examples, the ISP pipeline 305 can include a separate spatial noise reduction module and a separate temporal noise reduction module.

[0050]

[0067] When the ISP240 processes the first copy 210 of the raw image data, the ISP240 processes each image region of the first copy 210 of the raw image data differently based on which object category is depicted in that image region according to the category map 230. In particular, if the first image region is identified by the category map 230 as depicting skin, the ISP240 processes that first image region according to the set value corresponding to skin. The set value corresponding to skin may be stored as a specific correction value (corresponding to skin) with respect to the default intensity level regarding how strongly a specific parameter is applied. If the second image region is identified by the category map 230 as depicting hair, the ISP240 processes that second image region using the set value corresponding to hair. The set value corresponding to hair may be stored as a specific deviation (corresponding to hair) with respect to the default intensity level regarding how strongly a specific parameter is applied. The ISP240 may identify the set value to be used from a lookup table, database, or other data structure that maps the object category identifier within the category map 230 to the corresponding set value.

[0051]

[0068] In some cases, the ISP 240 can process different pixels within the image region differently based on the reliability levels associated with each pixel in the reliability map 235. To do so, the ISP 240 can use the combination of the category map 230 and the reliability map 235. For example, the map decoder component of the ISP 240 shown in and described with respect to FIG. 3 can generate a modifier using the category map 230 and the reliability map 235 as shown in and described with respect to FIG. 4. Similarly, in some cases, the smooth transition map processor 365 shown in and described with respect to FIG. 3 can generate a modifier using the category map 230 and the reliability map 235 as shown in and described with respect to FIG. 7. The ISP 240 can apply specific ISP tuning parameters such as noise reduction, sharpening, tone mapping, and / or color saturation to the image data based on the modifier. In particular, the setting values at which the ISP 240 applies specific ISP tuning parameters to a given pixel of the image data depend on the application of the modifier as shown in and described with respect to FIGS. 4, 5A, 5B, and 5C. The modifier effectively controls the intensity or weight at which the ISP 240 applies specific ISP tuning parameters to a given pixel of the image data based on the modifier.

[0052]

[0069] Finally, ISP240 generates a processed image 250 by processing different image regions of the first copy 210 of the raw image data using different modules set to different setting values for each image region based on the category map 230 and / or based on the reliability map 235. As described above, the setting value corresponding to the object image region can be stored as a specific deviation (corresponding to the object category) with respect to the default intensity level of how strongly a particular parameter is applied. Although not shown in FIG. 200, ISP240 can send the processed image 250 to a storage device, for example, using I / O 156 and / or I / O 160. The storage device can include an image buffer, random access memory (RAM) 140 / 1525, read-only memory (ROM) 145 / 1520, cache 1512, storage device 1530, secure digital (SD) card, mini SD card, micro SD card, smart card, integrated circuit (IC) memory card, compact disk, portable storage medium, hard drive (HDD), solid state drive (SSD), flash memory drive, non-transitory computer-readable storage medium, any other type of memory or storage device described herein, or any combination thereof. Although not shown in FIG. 200, ISP240 can send the processed image 250 to a display buffer and / or to a display screen or projector so that the image is rendered and displayed on the display screen and / or using a projector.

[0053]

[0070] In some cases, either or both of the first copy 210 of the raw image data and the second copy 215 of the raw image data may simply be referred to as raw image data or image data. As shown in FIG. 2, the image sensor 205 can send the first copy 210 of the raw image data and the second copy 215 of the raw image data to the ISP 240 and / or the classification engine 220. Alternatively, the image sensor 205 can simply send a single copy of the raw image data to the receiving component, which can be the ISP 240 and / or the classification engine 220 and / or another image processing component not shown in FIG. 2. This receiving component can generate one or more copies of the raw image data, and one or more of them are sent by the receiving component to the ISP 240 and / or the classification engine 220 as the first copy 210 of the raw image data and / or the second copy 215 of the raw image data. The receiving component can use and / or send the original copy 210 of the raw image data received from the image sensor as the first copy 210 of the raw image data and / or the second copy 215 of the raw image data.

[0054]

[0071] In some cases, different image capture setting values, including, for example, setting values for focus, exposure time, aperture size, ISO, flash, combinations of any of them, and / or other image capture setting values described herein, may also be generated for different image regions of the image data. In some examples, one or more of these image capture setting values may be determined in the ISP240. In some examples, one or more of these image capture setting values may be returned to the image sensor 205 for application to the image data. In some cases, different image frames are captured by the image sensor 205 using different image capture setting values such that different image regions are taken from the image frames captured using different image capture setting values and then merged with each other by the ISP240, the host processor 152, the image processor 150, another image processing component, or some combination thereof. In some examples, different post-processing setting values, including, for example, setting values for brightness, contrast, saturation, tone level, histogram, combinations of any of them, and / or other processing setting values described herein, may also be generated for different image regions of the image. In some cases, applying setting values in the ISP240 can enable better control over the resulting processed image (e.g., as compared to post-processing setting values) and better quality in the applied processing effects. This is because the ISP240 receives raw image data from the image sensor 205 as its input, while post-processing is typically applied to the image 250 already generated by the ISP240.

[0055]

[0072] In some cases, at least one of different setting values for different object categories, such as ISP tuning setting values, image capture setting values, and / or post-processing setting values, can be manually set by a user using a user interface. For example, the user interface can receive an input from the user specifying a setting value indicating that an enhanced sharpening setting value should always be applied to an image area depicting text. Similarly, the user interface can receive an input from the user specifying a setting value indicating that a reduced sharpening setting value and an increased noise reduction setting value should always be applied to an image area depicting a face so that the skin of the face appears smoother. In some cases, at least one of different setting values for different object categories can be automatically set by ISP240, a classification engine 220, a host processor 152, an image processor 150, an application (e.g., an application used for post-processing of images), another image processing component, or some combination thereof. In some cases, at least one of different setting values for different object categories can be automatically set based on a manually set setting value by automatically determining a setting value that is offset from the manually set setting value based on a modifier, such as a multiplier, an offset, or a logic-based modifier. The deviation from the modifier can be determined in advance or automatically based on how one object category is determined to be different from another object category with respect to a certain visual characteristic, such as texture or color. For example, ISP240 may determine that similar or identical setting values for an image area depicting text should be applied to an image area depicting line art. ISP240 may determine that the setting value applied to an image area depicting skin with sparse hair should be approximately intermediate between the setting value applied to an image area depicting skin and the setting value applied to an image area depicting longer hair.The ISP240 may determine that the setting values applied to the image area depicting the stack wall should be the same as those applied to the image area depicting the brick wall, but with a 10% increased noise reduction.

[0056]

[0073] FIG. 3 is a conceptual diagram 300 showing an ISP pipeline 305 for image processing based on object categorization. The ISP pipeline 305 shows the operations performed by the components of the ISP240. The operations and components of the ISP pipeline 305 are laid out in an exemplary configuration and order, like a flowchart.

[0057]

[0074] The input to the ISP pipeline 305, and thus to the ISP240, is shown on the left side of the ISP pipeline 305. The input to the ISP pipeline 305 includes a category map 230, a confidence map 235, and a first copy 210 of the image data. The first copy 210 of the image data may be in a color filter domain (e.g., Bayer domain), RGB domain, YUV domain, or another color domain described herein. Demosaicking and color domain conversion are not shown in FIG. 300, but it should be understood that demosaicking and / or color domain conversion may be performed by the ISP240 before, after, and / or between any two of the operations shown in FIG. 200. The category map 230 and the confidence map 235 are shown as being received twice by different elements of the ISP pipeline 305. However, it should be understood that the ISP240 may receive the category map 230 and the confidence map 235 once and internally distribute the category map 230 and the confidence map 235 to all appropriate components and elements of the ISP240.

[0058]

[0075] The ISP pipeline 305 receives the category map 230 and the confidence map 235 from the classification engine 220 and passes them to a plurality of map decoders 325, 335, 345, and 355, each corresponding to a different module. Before passing the category map 230 and the confidence map 235 to the map decoders 325, 335, 345, and 355, the ISP pipeline 305 can upscale the category map 230 and the confidence map 235 using an upscaler 310, for example, using a nearest neighbor (NN) upscaling and / or a special category map upscaling (CMUS). The upscaler 310 can upscale the category map 230 and the confidence map 235 such that the dimensions of the category map 230 and the confidence map 235 match the dimensions of the first copy 210 of the raw image data and / or the dimensions of the processed image 250. In some cases, at least a portion of the upscaling described with respect to the upscaler 310 can be performed in the classification engine 220 before the ISP 240 receives the category map 230 and the confidence map 235. In some cases, the category map 230 and the confidence map 235 can be upscaled once in the classification engine 220 and at another time in the upscaler 310 of the ISP 240.

[0059]

[0076] To upscale the category map 230 and the confidence map 235, whether or not the ISP pipeline 305 uses the upscaler 310, the ISP 240 receives the category map 230 and the confidence map 235 and passes them to the map decoder 325 corresponding to the noise reduction (NR) module 320. Based on the category map 230 and the confidence map 235, the map decoder 325 generates one or more modifiers 327. The NR module 320 can use one or more modifiers 237 to determine the NR setting values to be applied to different pixels of the first copy 210 of the raw image data. The NR module 320 processes the first copy 210 of the raw image data based on the modifier 327 to generate NR-processed image data. The NR module 320 can send the NR-processed image data to the sharpening module 330. In some examples, NR includes spatial noise reduction, temporal noise reduction, or both. Although the NR module 320 is shown as a single module in FIG. 3, in some examples, the ISP pipeline 305 can include a separate spatial noise reduction module and a separate temporal noise reduction module.

[0060]

[0077] The map decoder 325 passes the category map 230 and the confidence map 235 to the map decoder 335 corresponding to the sharpening module 330. Based on the category map 230 and the confidence map 235, the map decoder 335 generates one or more modifiers 337. The sharpening module 330 can use one or more modifiers 337 to determine the setting values for sharpening to be applied to different pixels of the NR-processed image data from the NR module 320. The sharpening module 330 processes the NR-processed image data based on the modifier 337 to generate sharpened image data, and sends the sharpened image data to the tone mapping (TM) module 340.

[0061]

[0078] The map decoder 335 passes the category map 230 and the confidence map 235 to a map decoder 345 corresponding to the TM module 340. Based on the category map 230 and the confidence map 235, the map decoder 345 generates one or more modifiers 347A, and these modifiers are used to determine the TM setting values that the TM module 340 applies to different pixels of the sharpened image data from the sharpening module 330. The TM module 340 generates TM-processed image data by processing the sharpened image data based on the modifiers 347A, and transmits the sharpened image data to a color saturation (CS) module 350.

[0062]

[0079] The map decoder 345 passes the category map 230 and the confidence map 235 to a map decoder 355 corresponding to the CS module 350. Optionally, a delay 315 is applied between the map decoder 345 and the map decoder 355. Based on the category map 230 and the confidence map 235, the map decoder 355 generates one or more modifiers 357A, which are used to determine the CS setting values that the CS module 350 applies to different pixels of the TM-processed image data from the TM module 340. The CS module 350 generates CS-processed image data by processing the TM-processed image data based on the modifiers 357A. In some examples, the delay 315 can function to synchronize the receipt of the modifiers 357A from the map decoder 355 in the CS module 350 with the receipt of the TM-processed image data in the CS module 350. Similar delays can be inserted between any two elements of the image capture / processing device 100 (including the ISP 240) to help synchronize the transmission and / or receipt of other signals within the image capture / processing device 100. Optionally, the CS-processed image data is then output by the ISP 240 as the processed image 250. Optionally, the ISP 240 performs one or more additional image processing operations on the CS-processed image data to generate the processed image 250. These one or more additional image processing operations can include, for example, downscaling, upscaling, gamma adjustment, gain adjustment, another image processing operation described herein, or some combination thereof. Optionally, the image processing using one of the ISP tuning parameter modules 320, 330, 340, or 350 can be skipped. In some examples, the map decoders 325, 335, 345, or 355 corresponding to the skipped ISP tuning parameter modules can be skipped.If one or both of these are skipped, a delay device similar to delay device 315 may be added in place of the skipped module (ISP parameter module and / or corresponding map decoder) to ensure that the processing element remains synchronized while moving forward. In some cases, map decoders 325, 335, 345, and / or 355 may be able to internally track the timing, detect when a module is skipped or removed, generate a modifier, and / or dynamically adjust the timing sent to the corresponding ISP tuning parameter modules 320, 330, 340, or 350.

[0063]

[0080] Delay device 315 receives the category map 230 and the confidence map 235 from map decoder 345 and transmits the category map 230 and the confidence map 235 to the next map decoder 355 without generating any modifiers using the category map 230 and / or the confidence map 235. It is a module that is unaware of the category map 230 and the confidence map 235. In some cases, other components that are unaware of the category map 230 and / or the confidence map 235 may be included within the ISP pipeline 305 (not shown). In some cases, delay device 315 may be removed. In some cases, one or more delay modules similar to delay device 315 may be inserted between any two other components of the ISP pipeline 305, ISP 240, any map decoder, smooth transition map processor (STMP) 365, classification engine 220, image capture / processing device 100, computing system 1500, any component of any of these modules, any other component or module or device described herein, or a combination thereof.

[0064]

[0081] In some cases, ISP240 can pass the category map 230 and the confidence map 235 to a smooth transition map processor (STMP) 365. ISP240 uses STMP365 to generate modifiers and can pass those modifiers to at least some of the map decoders 325, 335, 345, and 355 instead of or in addition to those map decoders, and to at least some of the modules 320, 330, 340, and 355. STMP365 is shown in at least FIGS. 6, 7, and 8 and can be used to provide a smooth transition between them as described with respect to those figures. In some examples, downscaler 360 downsizes the category map 230 and the confidence map 235 before STMP365 receives them. In conceptual diagram 300, downscaler 360 is shown as a component that is not part of ISP240, and thus ISP240 can receive the category map 230 and the confidence map 235 from classification engine 220 and the downscaled versions of the category map 230 and the confidence map 235 from downscaler 360. Downscaler 360 can receive the confidence map 235 from classification engine 220. In some examples, downscaler 360 may be part of ISP240 such that ISP240 routes the category map 230 and the confidence map 235 to downscaler 360 itself before receiving the category map 230 and the confidence map 235 and passing the downscaled category map 230 and the confidence map 235 to STMP365.

[0065]

[0082] In FIG. 300, STMP365 is shown to generate an alternative modifier 347B for the TM module 340 and pass the alternative modifier 347B that the TM module 340 uses when processing sharpened image data to the TM module 340. STMP365 is shown to generate an alternative modifier 357B for the CS module 350 and pass the alternative modifier 357B that the CS module 350 uses when processing TM-processed image data to the CS module 350. Although not shown in FIG. 300, STMP365 may also generate modifiers for the NR module 320 and / or the sharpening module 325 and pass the generated modifiers to those modules.

[0066]

[0083] FIG. 300 shows the category map 230 and the confidence map 235 that are sequentially passed from one to the next among the map decoders 325, 335, 345, and 355. However, alternatively, the ISP 240 may pass copies of the category map 230 and the confidence map 235 in parallel to two or more of the map decoders 325, 335, 345, and 355. In this way, the map decoders 325, 335, 345, and 355 may generate the modifiers 327, 337, 347A, and 357A in parallel, potentially increasing the image processing efficiency.

[0067]

[0084] Although STMP365 and upscaler 310 are shown as components of ISP pipeline 305 and ISP240, at least one of these may, in some examples, be separate from ISP pipeline 305 and / or ISP240. For example, at least one of STMP365 and / or upscaler 310 may be part of classification engine 220, another component of image capture / processing device 100, another component of computing system 1500, or some combination thereof. Downscaler 360 is shown as a component separate from ISP pipeline 305 and ISP240, but in some examples, downscaler 360 may be part of ISP pipeline 305 and / or ISP240.

[0068]

[0085] For purposes of illustration, different ISP parameter modules are shown and described in a particular order. In alternative embodiments, it should be understood that the ISP parameter modules may be configured in an order different from the order described, and / or that the operations performed by the ISP parameter modules may be performed in an order different from the order described. For example, TM module 340 may be located before sharpening module 330 and / or before NR module 320.

[0069]

[0086] FIG. 4 is a conceptual diagram 400 showing the pipeline of map decoder 325. Conceptual diagram 400 includes operations performed by components of map decoder 325 of ISP240. Map decoder 325 corresponds to NR module 320, and FIG. 400 shows the generation of modifier 455 of modifiers 327 and the transmission of modifier 455 to NR module 320. The operations and components of map decoder 325 are laid out in an exemplary configuration and order, as in a flowchart.

[0070]

[0087] The map decoder 325 receives the category map 230 and the confidence map 235 that are generated by the classification engine 220 and received by the ISP 240. In some examples, the map decoder 325 can include a delay line buffer 410. The delay line buffer 410 can delay the transmission of the category map 230 and the confidence map 235 from the map decoder 325 to a map decoder 335 corresponding to a sharpening module 330 in the next line of the ISP pipeline 305. The delay line buffer 410 can also delay the transmission of a modifier 327 (including a modifier 455) to the NR module 320 such that the timing of the reception of the modifier 327 by the NR module 320 is synchronized with the timing of the reception of the first copy 210 of the raw image data by the NR module 320.

[0071]

[0088] As shown in FIG. 4, the generator 430 of the category-based modifier 465 obtains the category map 230 (e.g., from one of the buffers of the delay line buffer 410 or directly). An exemplary category map 230 is shown at the bottom of the conceptual diagram 400, which labels each of several differently colored image regions with different numerical values, each corresponding to one of several object categories. In particular, the category map 230 of the conceptual diagram 400 is shaded in a first shading pattern and includes a first image region labeled with "0" representing a human as an object category. The category map 230 of FIG. 4 is all shaded in a second shading pattern and includes several second image regions depicting trees and grass labeled with "1" representing plants as object categories. The third image region in the category map 230 is shaded in a third shading pattern and is labeled with "2" representing empty as an object category. The fourth image region in the category map 230 is shaded in a fourth shading pattern and is labeled with "6" representing an asphalt road as an object category. Finally, the fifth image region in the category map 230 is shaded in a fifth shading pattern and is labeled with "9" representing a vehicle as an object category. The different shading patterns in the category map 230 of FIG. 4 can represent different values stored at corresponding pixel locations in the category map 230, such as different colors, different shades of gray, different numbers, different characters, or different bit sequences.

[0072]

[0089] To generate the category-based modifier 465, the generator 430 cross-references the object categories in different image regions of the category map 230 with respect to the data structure 480. The data structure 480 can be, for example, a lookup table, a database, a dictionary, a list, an array, an array list, a different data structure capable of storing the relationship between values, or some combination thereof. The data structure 480 stores a predetermined set value suitable for each of the object categories and the ISP tuning parameters of the problem. Since the map decoder 325 corresponds to the NR module 320, the data structure 480 stores a predetermined set value suitable for each of the object categories and NR. Different predetermined set values can essentially represent different intensities of applying NR. In some examples, the different predetermined set values are expressed as values on an absolute scale. In some examples, the different predetermined set values are expressed as values relative to other values, for example, values relative to the values in the default settings.

[0073]

[0090] An exemplary category-based modifier 465 generated by the generator 430 is shown at the bottom of the conceptual diagram 400. The exemplary category-based modifier 465 is shown in grayscale, with lighter shades representing predetermined set values associated with higher levels of NR and darker shades representing predetermined set values associated with lower levels of NR. Based on the category map 230, a total of four different gray shades are used. Thus, both the image regions categorized within the category map 230 as depicting the sky (2) and the vehicle (9) should be processed using a high level of NR, the image region categorized within the category map 230 as depicting a human (0) should be processed using an intermediate level of NR, the image region categorized within the category map 230 as depicting an asphalt road (6) should be processed using a low level of NR, and the image region categorized within the category map 230 as depicting a plant (1) should be processed using the lowest level of NR.

[0074]

[0091] Within the map decoder 325, the reliability map 235 is sent to a generator 435 for generating blending update values for the category-based modifier 465. An exemplary reliability map 235 is shown at the bottom of the conceptual diagram 400. The reliability map 235 is shown in eight shades of gray as previously described, with lighter shades representing higher levels or reliability and darker shades representing lower levels or reliability. The reliability map 235 of FIG. 4 may visually appear to have visible banding between different gray shades, for example, when there is a gradient between the shades. The reliability map 235 indicates that the classification engine 220 has generated a category map 230 that generally has a high level of reliability and that most of the portions with lower reliability levels are around the edges between different image regions representing different object categories. The generator 435 can identify appropriate adjustments to predetermined set values within the category-based modifier 465 based on different reliability levels. These adjustments may also be specified in the data structure 480.

[0075]

[0092] In the category-confidence blending operation 440, the category-based modifier 465 generated by the generator 430 is blended with a blending update value for the category-based modifier 465 generated by the generator 435. The category-confidence blending operation 440 generates a category-confidence blended modifier 470. The modifier value for a particular pixel can be reduced from the amount determined in the category-based modifier 465 based on the confidence level of that pixel in the confidence map 235. For example, a pixel having the maximum confidence level within the confidence map 235 can hold its modifier value from the category-based modifier 465. On the other hand, the modifier value of a pixel having a low confidence level within the confidence map 235 can be decreased from its modifier value from the category-based modifier 465, thus reducing the intensity of the effect applied by the ISP tuning parameter module at that pixel. For example, as shown in the conceptual diagram 400, when the ISP tuning parameter module is the NR module 320, a weaker NR effect is applied to pixels that the classification engine 220 has categorized with low confidence than to pixels that the classification engine 220 has categorized with high confidence. In some examples, the blending of the category-based modifier 465 is performed using a no-operation value. In some examples, the category-confidence blending operation 440 can be referred to as a category-confidence adjustment operation.

[0076]

[0093] An exemplary category confidence blend modifier 470 is shown at the bottom of the conceptual diagram 400. The category confidence blend modifier 470 is shown in grayscale. Similar to the exemplary category-based modifier 465, the lighter gray shading within the category confidence blend modifier 470 represents set values associated with higher levels of NR, while the darker shading represents set values associated with lower levels of NR. The blending update value based on the confidence value is blended with a predetermined set value (e.g., the modifier value in the category-based modifier 465), thus adjusting the predetermined set value, and the category confidence blend modifier 470 may have portions corresponding to set values that do not match the predetermined set value associated with any particular object category. Blending can be performed, for example, via addition, subtraction, or multiplication of the blending update value for the category-based modifier 465 and the set value in the category-based modifier 465. This more finely tuned control over the ISP tuning parameter module allows the ISP 240 to apply its ISP tuning parameters more weakly, for example, to regions with a lower confidence of being category classified. This reduces the risk of inappropriate set values being applied to portions of the image due to incorrect category classification of parts of the image. In these parts of the image that are most likely to be misclassified, i.e., parts with a low confidence of category classification, the set values are weakened or, in some cases, adjusted so that the ISP tuning parameters are applied more conservatively. The category confidence blend modifier 470 in FIG. 4 may visually appear to have visible banding between different gray shades, for example, if there is a gradient between the shades. The banding within the category confidence blend modifier 470 can be inherited from the banding within the confidence map 235 in FIG. 4, from separate shades corresponding to separate image regions of the category-based modifier 465, or from a combination thereof.

[0077]

[0094] Next, the category confidence blend modifier 470 passes through the low-pass filter 445 to generate a filtered modifier 475. An example of the filtered modifier 475 is shown at the bottom of the conceptual diagram 400. The filtered modifier 475 is similar to the category confidence blend modifier 470, but the transitions between different set values are smoothed. For example, the boundaries between image regions within the category confidence blend modifier 470 may include banding resulting from different levels of confidence from the confidence map 235, different image regions within the category-based modifier 465, or combinations thereof. In the filtered modifier 475, the banding is smoothed with a median value, resulting in a blurring effect. This enables the transitions between different set values applied using a module (here, NR) to be smoother than those using the category confidence blend modifier 470. The filtered modifier 475 is upscaled using the upscaler 450 to generate the final modifier 455. The upscaler 450 can perform this upscaling using, for example, nearest neighbor (NN) upscaling, bilinear interpolation, bicubic interpolation, Sinc resampling, Lanczos resampling, box sampling, mipmapping, Fourier transform scaling, edge-directed interpolation, high-quality scaling (hqx), or some combination thereof. The map decoder 325 sends the final modifier 455 to the NR module logic 405 of the NR module 320. The final modifier 455 can be one of a set of one or more modifiers 327. The NR module logic 405 of the NR module 320 applies NR to each pixel of the first copy 210 of the raw image data at an intensity based on the final modifier 455. The intensity can range from a minimum intensity within a predetermined range represented by the darkest gray in an exemplary filtered modifier 475 to a maximum intensity within a predetermined range represented by white in an exemplary filtered modifier 475.In some examples, the low-pass filter 445 may include a Gaussian blur filter. In some examples, the low-pass filter 445 may be supplemented by or replaced with another type of filter or blur effect, such as an average filter, box blur, lens blur, radial blur, motion blur, shape blur, smart blur, surface blur, or a combination thereof.

[0078]

[0095] As described above, FIG. 400 shows the generation of the modifier 327 by the map decoder 325 corresponding to the NR module 320. Substantially the same process can be used by the map decoder 335 to generate a modifier 337 for the sharpening module 330, by the map decoder 345 to generate a modifier 347 for the TM module 340, and by the map decoder 355 to generate a modifier 357 for the CS module 350. The main difference between the other map decoders 335, 345, and 355 is that different data structures 480 that store predetermined setting values of the modules corresponding to these map decoders are used. Alternatively, the data structure 480 may store the predetermined setting values of a plurality of modules in different columns of a table, for example, in which case the same data structure 480 may be used but queried in different columns.

[0079]

[0096] Figure 5A is a conceptual diagram 510 showing the application of modifier 545A in the ISP module, and modifier 545A is applied as a multiplier. An internal signal 540A is received. The internal signal 540A represents a default setting value or a default intensity for the module to apply ISP tuning parameters to a part of the image data. A modifier 545A (such as one to which a low-pass filter (LPF) 445 is applied) is received. The modifier 545A identifies a value for each pixel of the image data. The filtered modifier 475 and / or the final modifier 455 are examples of the modifier 545A. Similar to the filtered modifier 475, the modifier 545A can be stored as an image having different shades of gray, i.e., different luminance values, at each pixel. Alternatively, the modifier 545A can be stored as a matrix or table or other data structure having cells corresponding to every pixel of the image data. For a given pixel of the image data, the module obtains the internal signal 540A, i.e., the default setting value for applying the ISP tuning parameters, and multiplies that internal signal 540A by the value in the modifier 545A corresponding to that pixel in the image data.

[0080]

[0097] For example, the internal signal 540A can indicate that the default setting value or default intensity for applying a particular ISP parameter is 3. The modifier 545A may include a value 1.6 corresponding to a given pixel of the image data, which means that the ISP tuning parameter is 1.6 times the default setting value indicated by the internal signal 540 at that pixel of the image data, i.e., 3 * ×1.6 = 4.8 is applied. The modifier 545A may include a value 0.8 corresponding to different pixels of the image data, which means that the ISP tuning parameter is 0.8 times the default setting value indicated by the internal signal 540A at that pixel of the image data, i.e., 3 *It means that it is applied at 0.8 = 2.4. The modified internal signal 550A is the result of this multiplication and, thus, is the intensity at which the module finally applies the ISP tuning parameters to a given portion of the image data. The modified internal signal 550A can, in some cases, be expressed as a decimal value or a fraction. Alternatively, the modified internal signal 550A can be rounded to the nearest integer, or the floor or ceiling function can be applied, respectively, to round the modified internal signal 550A to the nearest integer smaller than or larger than the multiplication result.

[0081]

[0098] In some cases, the modified internal signal 550A generated by multiplying the default setting value from the internal signal 540A by the modifier 545A can be equivalent to one of a set of predetermined setting values corresponding to different object categories and / or confidence levels. In some cases, the modified internal signal 550A can be between two setting values of the set of predetermined setting values or outside the range represented by the set of predetermined setting values.

[0082]

[0099] FIG. 5B is a conceptual diagram 520 showing the application of the modifier 545B in the ISP module, where the modifier 545B is applied as an offset. The conceptual diagram 520 includes the internal signal 540B and the modifier 545B. The internal signal 540B may be the same as the internal signal 540A. The modifier 545B may be the same as the modifier 545A. However, in the conceptual diagram 520, the value of the modifier 545B for a given pixel is added to the value of the internal signal 540B to generate the modified internal signal 550B.

[0083]

[0100] For example, the internal signal 540B may indicate that the default setting value or default strength for which a particular ISP parameter should be applied is 3. The modifier 545B may include a value 1.6 corresponding to a given pixel of the image data, which means that the ISP tuning parameter is applied as the addition of the default setting value indicated by the internal signal 540B and 1.6 at that pixel of the image data, i.e., 3 + 1.6 = 4.6. The modifier 545 may include a value -0.8 corresponding to a different pixel of the image data, which means that the ISP tuning parameter is applied as the addition of the default setting value indicated by the internal signal 540B and -0.8 at that pixel of the image data, i.e., 3 - 0.8 = 2.2. The modified internal signal 550B is the result of this addition and is thus the strength at which the module finally applies the ISP tuning parameter to a given portion of the image data.

[0084]

[0101] In some cases, the modified internal signal 550B generated by adding the modifier 545B to the default setting value from the internal signal 540B may be equivalent to one of a set of predetermined setting values corresponding to different object categories and / or confidence levels. In some cases, the modified internal signal 550B may be between two setting values of the set of predetermined setting values or outside the range represented by the set of predetermined setting values.

[0085]

[0102] FIG. 5C is a conceptual diagram 530 showing the application of modifier 545C in the ISP module, and modifier 535C is applied using logic 555 based on parameter 560. The conceptual diagram 530 includes an internal signal 540C and a modifier 545C. The internal signal 540C may be the same as the internal signal 540A and / or the internal signal 540B. The modifier 545C may be the same as the modifier 545A and / or the modifier 545B. However, in the conceptual diagram 530, the value of the modifier 545C for a given pixel represents a change from one predetermined set value to another predetermined set value.

[0086]

[0103] For example, the internal signal 540C may indicate that the default setting value or default strength to which a particular ISP parameter should be applied is 3. The modifier 545C may include a value 2 corresponding to a given pixel of the image data, which means that the ISP tuning parameter is applied to that pixel of the image data at an intensity selected from a list of predetermined setting values by selecting a second predetermined setting value greater than and consecutive from a list of predetermined setting values. If the list of predetermined setting values includes the set {1.5, 3, 4, 6, 8}, for example, 6 is two values higher than the default setting value (3) in the list, so the modified internal signal 550C is 6. Similarly, if the modifier 545C has a value of -1 corresponding to a different pixel of the image data and the same list of predetermined setting values is used, 1.5 is one value lower than the default setting value (3) in the list, so the modified internal signal 550C is 1.5. The list of predetermined setting values may be obtained from the data structure 480, may each correspond to a different object category, and may sometimes be referred to as the parameter 560. In some examples, each predetermined setting value in the list of predetermined setting values may represent a different object category and / or a different confidence level. For example, if the object category and confidence level are represented within 8 bits in the category map 230 and the confidence map 235, the list of predetermined setting values can include 256 different predetermined setting values. In some cases, the list of predetermined setting values can include a predetermined intermediate setting value that is between two other predetermined setting values corresponding to a particular object category and confidence level. The use of such a predetermined intermediate setting value can help produce a smoother transition with less banding in the processed image 250 and can assist in enabling the smoothing generated by the low-pass filter 445. In some examples, the logic 555 can include a multiplication 510, an offset 520, and / or a combination of other arithmetic operations, instead of or in addition to the operations described above.In some examples, the logic 555 can blend the modifier 545C with data from the parameter 560 instead of or in addition to the operations described above. In some examples, the logic 555 can include conditional programming (e.g., if...else), loops, and / or other programming logic instead of or in addition to the operations described above. The determination of the modified internal signal 550C described above may be referred to as the application of the logic 555 based on the parameter 560, the application of a logic engine that determines the modified internal signal 550C using the logic 555 based on the parameter 560, or a combination thereof.

[0087]

[0104] FIG. 6 is a conceptual diagram showing visual image artifacts caused by anomalies in the segmentation of an image into image regions during the generation of a category map. In particular, FIG. 6 includes a first image 610, a second image 620, and a third image 630. The first image 610 is an image of two buildings and trees with a blue sky as the background. The first image 610 of FIG. 6 represents raw image data that has not yet been processed by the ISP 240 based on object categories.

[0088]

[0105] The second image 620 is similar to the first image 610 but includes a white image region marked as an empty image region 625. The empty image region 625 represents a portion of the first image 610 that the classification engine 220 detected as depicting empty. Empty is an object category in this example. Thus, the empty image region 625 is an image region that the classification engine 220 detected as depicting the object category of "empty". The boundaries of the empty image region 625 include artifacts 640 in some areas near the boundaries between the region of the first image 610 depicting empty and the region depicting buildings and trees. These anomalies can be generated as a result of an incomplete detection of object categories, for example, due to similar blue shades appearing in the empty and building windows or due to the complex boundaries of the tree leaves.

[0089]

[0106] The third image 630 includes an empty image region 625 and represents a version of the first image 610 processed by the ISP 240 based on object categories using a category map based on the category segmentation of the second image 620. At the location of the artifact 640 within the empty image region 625 of the second image 620, the third image 630 includes visual image artifacts 645 in the tone and color transitions. For example, in the third image 630, the empty regions near the boundary between the building and the sky and near the boundary between the tree and the sky appear brighter and less saturated than the rest of the empty part. These artifacts 645 that did not fall within the empty image region 625 in the second image 620 due to the artifact 640 are brighter and have a lower saturation. In such situations, a sudden transition from one set value to another set value may generate these types of visual image artifacts, or similar visual image artifacts. A smoother transition between one set value and another set value can reduce the occurrence of such artifacts in similar situations. A smoother transition can be achieved by generating a smoother modifier corresponding to the empty image region 625, for example, by using the technique shown in FIG. 7 to generate a smoother modifier such as that shown in FIG. 8.

[0090]

[0107] FIG. 7 is a conceptual diagram 700 showing the pipeline of a smooth transition map processor (STMP) 365. The STMP 365 generates a smooth transition from one set value to another within the same image. The STMP 365 functions in a similar manner to two map decoders 435 that are combined into one component as described below, providing a modifier 755A to the TM module 340 and a modifier 755B to the CS module 350. However, the category map 230 and the confidence map 235 can be passed to an additional downscaler 360 and a front end (FE) 705 before the modifiers 755A and 755B are generated in the STMP 365. The additional downscaling provided by the downscaler 360 effectively provides a blur effect at the boundaries between different set values when upscaled by the upscalers 750A and 750B. The STMP 365 can also optionally use low-pass filters 745A and 745B that are stronger than the low-pass filter 445, which can further smooth the transition between different set values to reduce banding in the transition. In some examples, the low-pass filters 745A and 745B can include Gaussian blur filters. In some examples, the low-pass filters 745A and 745B can be supplemented or replaced by another type of filter or blur effect, such as an average filter, box blur, lens blur, radial blur, motion blur, shape blur, smart blur, surface blur, or a combination thereof.

[0091]

[0108] Similar to map decoder 435 and STMP365, line buffer 710 routes category map 230 to two generators 730A and 730B of category-based modifiers 765A and 765B, and these generators each function similarly to generator 430 of category-based modifier 465. Line buffer 710 routes reliability map 235 to two generators 735A and 735B of blending update values for category-based modifiers 765A and 765B, and these generators each function similarly to generator 435 of blending update values for category-based modifier 465. Category reliability blending operation 740A blends category-based modifier 765A with the blending update value for category-based modifier 765A, similar to category reliability blending operation 440. The resulting blended modifier is filtered by low-pass filter 745A and upscaled by upscaler 750A to generate a modifier 755A (an example of modifier 347B) with a smooth transition. Modifier 755A with a smooth transition is sent from STMP365 to TM module 340 instead of modifier 347A. Similarly, category reliability blending operation 740B blends category-based modifier 765B with the blending update value for category-based modifier 765B, similar to category reliability blending operation 440. The resulting blended modifier is filtered by low-pass filter 745B and upscaled by upscaler 750B to generate a modifier 755B (an example of modifier 357B) with a smooth transition. Modifier 755B with a smooth transition is sent from STMP365 to CS module 350 as an example of modifier 347B.

[0092]

[0109] The TM module 340 processes the sharpened image data 770A based on the modifier 755A with a smooth transition to generate the TM-processed image data 770B, and the TM-processed image data 770B is transmitted to the CS module 350. The CS module 350 processes the TM-processed image data 770B based on the modifier 755B with a smooth transition to generate the CS-processed image data 770C, and the CS-processed image data 770C can be used as the processed image 250 or transmitted to another component within the ISP 240 for further processing to generate the processed image 250. Although not shown in FIG. 700, the STMP 365 can also generate modifiers for the NR module 320 and / or the sharpening module 325 and pass the generated modifiers to those modules. In some examples, the TM module 340 can use both the STMP-based modifier 347B / 755A and the map decoder-based modifier 347A in parallel to enjoy both low-resolution processing correction and high-resolution processing correction. In some examples, the CS module 350 can use both the STMP-based modifier 357B / 755B and the map decoder-based modifier 357A in parallel to enjoy both low-resolution processing correction and high-resolution processing correction.

[0093]

[0110] FIG. 8 is a conceptual diagram showing the smoothing of modifiers corresponding to image regions using a smooth transition map processor. In particular, FIG. 8 shows four versions of the modifier generated based on the empty image region 625 of FIG. 6 using different scalings by the STMP 365. The first modifier 810 is generated using a 1:1 scaling, which means that the downscaler 360 lacks or does not perform the downscale of the category map 230 and / or the confidence map 235. As a result, the boundary between the region corresponding to the empty image region 625 within the first modifier 810 and the other regions within the first modifier 810 is sharp.

[0094]

[0111] The second modifier 820 is generated using a 1:4 scaling, which means that the downscaler 360 downsizes the category map 230 and / or the confidence map 235 to 1 / 4 of their original size. As a result, the boundary between the region in the second modifier 820 corresponding to the empty image region 625 and the other regions in the second modifier 820 is blurrier than the same boundary in the first modifier 810. Similarly, a 1:16 scaling is used by the downscaler 360 to generate the third modifier 830, and thus the category map 230 and / or the confidence map 235 are downsized to 1 / 16 of their original size. Therefore, the boundary in the third modifier 830 is blurrier than the boundary in the second modifier 820. Finally, a 1:64 scaling is used by the downscaler 360 to generate the fourth modifier 840, and thus the category map 230 and / or the confidence map 235 are downsized to 1 / 64 of their original size. Therefore, the boundary in the fourth modifier 840 is blurrier than the boundary in the third modifier 830. Higher levels of downscaling than those shown in FIG. 8, such as 1:256 scaling, are possible. Downscaling levels between any of the previously described downscaling levels, such as 1:3 scaling, 1:6 scaling, 1:10 scaling, 1:32 scaling, 1:50 scaling, 1:100 scaling, or 1:128 scaling, are also possible.

[0095]

[0112] FIG. 9 is a diagram 900 showing the pipeline of a Category Map Upscaler (CMUS) 905. Since all of the values within the category map 230 represent a particular object category, the category map 230 generally cannot be upscaled using an interpolation-based upscaling such as bilinear interpolation or bicubic interpolation. Interpolation can create intermediate values that may reference unintended or non-existent object categories. For example, a category map may exist with a pixel having a value of 2 adjacent to a pixel having a value of 4. The value 2 may, for example, represent an empty object category, while the value 4 may represent a plant object category. Upscaling using interpolation may create a pixel having a value of 3 between the pixel having a value of 2 and the pixel having a value of 4. The value 3 may represent yet another category of objects, such as cloth, that may correspond to an ISP tuning setting value that is completely different from empty or plant, and may result in visual artifacts when used for image processing based on object categorization. Alternatively, the value 3 may not represent any known category of objects at all, which may result in errors or visual artifacts when used for image processing based on object categorization.

[0096]

[0113] One method that can be used to upscale a category map without problems resulting from interpolation-based upscaling is nearest neighbor (NN) upscaling. NN upscaling does not generate intermediate values. However, NN upscaling can generate sharp and blocky edges. Sometimes, as a result of NN upscaling, thin, curved objects depicted in an image, such as a person's eyebrows, shadows, or straps of clothing, may appear particularly blocky and inaccurate. Category map upscaling (CMUS), also called NN with spatial weight filtering, more accurately upscales the category map without introducing problems related to interpolation, such as intermediate values. The improvement in upscaling is particularly prominent at the boundaries between image regions and in narrow image regions.

[0097]

[0114] The pipeline of CMUS905 in the conceptual diagram 900 uses the input of the category map 230 for the filter and upscaling operations 920, as well as the filter sizing operation 910. The filter sizing operation 910 can adaptively select one of a set of filter sizes for upscaling a given pixel, such as a filter having a size of 2×2 pixels, 4×2 pixels, 2×4 pixels, or 4×4 pixels. To preserve finer details, smaller filters are used for narrow or small image regions within the category map 230. The neighborhood weights 915 based on the confidence map 235 are provided to the filter and upscaling operations 920. Two examples using 2x (×2) upscaling 970 and 4x (×4) upscaling 975 are shown, where circular dots (pixels) are interpolated from square dots (adjacent pixels) based on the confidence of the square dot category, the distance from the square dot to the circular dot, or a combination thereof.

[0098]

[0115] The accumulated weight per category is calculated in operation 925. An example of this calculation is provided in example 965. Adjacent pixels having a lower reliability level within the reliability map 235 may have a lower weight in the neighborhood weight 915 than adjacent pixels having a higher reliability level within the reliability map 235, and thus may contribute less to the sum of weights. Adjacent pixels that are farther from the upscaled pixels may also have a lower weight than adjacent pixels that are closer to the upscaled pixels, and thus may contribute less to the sum of weights. In other words, adjacent pixels that are closer to the upscaled pixels can have a higher weight than adjacent pixels that are farther from the upscaled pixels. In operation 930, for the upscaled pixels within the upscaled category map 950, the category having the maximum weight is used. In some examples, the upscaler 310 can perform upscaling using NN upscaling, CMUS upscaling, or some combination thereof. In some examples, one or more of the upscalers 450, 750A, and 750B can perform upscaling using NN upscaling, CMUS upscaling, another upscaling technique described herein, or some combination thereof.

[0099]

[0116] FIG. 10 is a diagram showing a comparison between a category map upscaled using nearest neighbor upscaling and the same category map upscaled using nearest neighbor upscaling modified by spatial weight filtering applied using a category map upscaler (CMUS) 905. The category maps in FIG. 10 are based on the same images, namely, a photograph of a man, a photograph of a woman, a desktop with three pens placed thereon, and an image of the corner of a tablet device placed on the desk. Both category maps include various image regions depicted using different shades of gray. The first category map 1010 in FIG. 10 is upscaled using NN upscaling. The boundaries of the various image regions within the first category map 1010 are extremely blocky and jagged due to the use of NN upscaling.

[0100]

[0117] The second category map 1020 in FIG. 10 is the same category map as the first category map 1010, but is upscaled using a category map upscaler (CMUS) instead of normal nearest neighbor upscaling. The category map upscaler (CMUS) is sometimes referred to as NN upscaling modified by spatial weight filtering. As a result, the boundaries between different image regions are, where appropriate, more rounded and not blocky overall. The improvement in upscaling fidelity is particularly noticeable in narrow image regions such as the image region representing the strap of the woman's clothing in the photograph.

[0101]

[0118] FIG. 11 is a conceptual diagram 1100 showing an exemplary resolution of image data corresponding to a category map during downscale and upscale operations. In conceptual diagram 1100, raw image data 1105 captured by image sensor 205 has a 4K resolution and an electronic image stabilization (EIS) margin, resulting in a resolution of 4800×2700. The raw image data is downscaled to a resolution of 848×480 in operation 1110. This is the downscale shown in the second copy 215 of the raw image data in FIG. 2, and the downscale can be performed using a downscaler. This downscaler can be part of image sensor 205 in classification engine 220, in ISP 240, or in another component not shown in FIG. 2. Classification engine 220 partitions the downscaled image data from operation 1110 into image regions in operation 1115 to generate a category map 1120 that also has a resolution of 848×480. In the example of conceptual diagram 1100, category map 1120 is upscaled once to a resolution of 1200×675 using NN upscale operation 1125. This once-upscaled category map is upscaled again to a resolution of 1920×1080 using CMUS upscale operation 1130. In some examples, upscaler 310 can perform one or both of upscale operations 1125 and / or 1130. In some examples, the upscaler within classification engine 220 can perform one or both of upscale operations 1125 and / or 1130. In some examples, the upscaler within ISP 240 can perform one or both of upscale operations 1125 and / or 1130.

[0102]

[0119] FIG. 12A is a flowchart 1200 showing an image processing technique. The image processing technique shown by the flowchart 1200 can be executed by a device. The device can be an image capture / processing device 100, an image capture device 105A, an image processing device 105B, a classification engine 220, an ISP 240, an image sensor 205, one or more network servers of a cloud service, a computing system 1500, or some combination thereof.

[0103]

[0120] In operation 1205, as part of the image processing technique, the device receives image data captured by the image sensor 205. In some cases, the device may include a connector coupled to the image sensor 205 and the image data can be received using the connector. The connector can include a port, a jack, a wire, input / output (IO) pins, conductive traces on a printed circuit board (PCB), any other type of connector described herein, or some combination thereof. In some cases, the device may include the image sensor 205.

[0104]

[0121] In some examples, the image data can be raw image data. In some examples, the device can demosaic the image data. In one exemplary example, the device can demosaic the image data after receiving the image data in operation 1205 but before at least one of the other operations 1210-1235. In some examples, the device can convert the image data from a first color space to a second color space. In one exemplary example, the device can convert the image data from the first color space to the second color space after receiving the image data in operation 1205 but before at least one of the other operations 1210-1235. In some examples, the second color space is the YUV color space. In some examples, the second color space is the RGB color space. In some examples, the first color space is the RGB color space. In some examples, the first color space is the Bayer color space, or another color space associated with one or more color filters on the image sensor 205.

[0105]

[0122] In operation 1210, as part of the image processing technique, the device determines that the first object image region in the image data depicts the first object category among a plurality of object categories. In operation 1215, as part of the image processing technique, the device determines that the second object image region in the image data depicts the second object category among a plurality of object categories.

[0106]

[0123] In operation 1220, as part of the image processing technique, the device generates the category map 230 by partitioning the image data into a plurality of object image regions including a first object image region and a second object image region. Each of the plurality of regions corresponds to one of a plurality of object categories (e.g., the first region corresponds to a first object, the second region corresponds to a second object, etc.). In some embodiments, the device also generates a downscaled copy of the image data by downscaling the image data. Generating the category map based on the image data includes generating the category map based on the downscaled copy of the image data.

[0107]

[0124] Although not shown in the flow diagram 1200, operation 1220 can also include generating a confidence map 235 based on the image data. The confidence map 235 identifies a plurality of confidence levels corresponding to a plurality of portions of the image data. Each confidence level of the plurality of confidence levels identifies the confidence when determining that the corresponding portion of the plurality of portions depicts one of the plurality of object categories. In one example, the confidence map 235 and the category map 230 are a single file that stores a single value for each pixel. The first plurality of bits in that value represent the object category that classifies the pixel as depicted by the classification engine 220. The second plurality of bits in that value represent the confidence of the classification engine 220 when classifying the pixel as depicting the object category. In another example, the confidence map 235 and the category map 230 are a single file that stores two values for each pixel, where one value represents the object category and the other value represents the confidence.

[0108]

[0125] In some embodiments, the device also upscales the category map. In some examples, upscaling the category map can include upscaling the category map to a size that matches at least one of the size of the image data and the size of the image. In some examples, the category map can be upscaled using nearest neighbor upscaling or nearest neighbor upscaling modified (e.g., modified by spatial weight filtering, which can also be referred to herein as category map upscaling (CMUS)). Upscaling the category map using nearest neighbor upscaling modified by spatial weight filtering can include identifying a first filter size corresponding to a first object image region and a second filter size corresponding to a second object image region. The first filter size is smaller than the second filter size. Upscaling the category map using nearest neighbor upscaling modified by spatial weight filtering can include upscaling a first pixel within the first object image region based on the first filter size and one or more weights associated with one or more confidence values from a confidence map corresponding to one or more pixels adjacent to the first pixel. Upscaling the category map using nearest neighbor upscaling modified by spatial weight filtering can include upscaling a second pixel within the second object image region based on the second filter size and one or more weights associated with one or more confidence values from a confidence map corresponding to one or more pixels adjacent to the second pixel.

[0109]

[0126] In operation 1225, as part of the image processing technique, the device identifies that a first object category corresponds to a first tuning setting value of an image signal processor (ISP). In operation 1230, the image processing technique includes identifying that a second object category corresponds to a second tuning setting value of the ISP. The first tuning setting value and the second tuning setting value can include indicators of different intensities to which noise reduction (NR) ISP tuning parameters are applied during the processing of image data. The first tuning setting value and the second tuning setting value can include indicators of different intensities to which sharpening ISP tuning parameters are applied during the processing of image data. The first tuning setting value and the second tuning setting value can include indicators of different intensities to which color saturation (CS) ISP tuning parameters are applied during the processing of image data. The first tuning setting value and the second tuning setting value can include indicators of different intensities to which tone mapping (TM) ISP tuning parameters are applied during the processing of image data. The first tuning setting value and the second tuning setting value can include indicators of different intensities to which gamma ISP tuning parameters are applied during the processing of image data. The first tuning setting value and the second tuning setting value can include indicators of different intensities to which different ISP tuning parameters are applied during the processing of image data. Different ISP tuning parameters can include, for example, gain, luminance, shading, edge enhancement, high dynamic range (HDR) image composition, special effect processing (e.g., background replacement, blur effect), artificial noise adder, demosaic, edge-directed upscaling, other processing parameters described herein, or combinations thereof.

[0110]

[0127] In some cases, as described above, the first setting value, the second setting value, and / or the first and second tuning setting values are defined based on user input related to the first object image region and the second object image region. In some cases, at least one of the first setting value and the second setting value is automatically determined, as described above.

[0111]

[0128] In operation 1235, as part of the image processing technique, the device generates an image by processing image data using an ISP tuned based on a category map. For example, the ISP can process a first object image region in the image data using a first tuning setting value. The ISP can process a second object image region in the image data using a second tuning setting value. In some cases, generating an image includes processing raw image data using an ISP tuned based on a category map and a confidence map.

[0112]

[0129] In some aspects, the device also generates one or more modifiers based on a category map. The one or more modifiers identify at least one of a first deviation or a second deviation. The first deviation is a deviation from a default setting value applied by the ISP within a first object image region during processing of the image data. The second deviation is a deviation from a default ISP tuning setting value applied by the ISP within a second object image region during processing of the image data. In some aspects, the ISP identifies at least one of the first deviation or the second deviation by performing an arithmetic function of the one or more modifiers and the default tuning setting values. The arithmetic function can include at least one of a multiplication function, an addition function, a subtraction function, a division function, or some combination thereof. The multiplication function can multiply the default tuning setting values by the one or more modifiers, as shown, for example, in FIG. 5A. The addition function can add the one or more modifiers to the default tuning setting values, as shown, for example, in FIG. 5C. The subtraction function can subtract the one or more modifiers from the default tuning setting values, or vice versa. The division function can divide the default tuning setting values by the one or more modifiers, or vice versa. In some aspects, the ISP identifies at least one of the first deviation and the second deviation based on increments within a list of predetermined possible setting values including the default setting values, and these increments are based on the modifiers, as shown, for example, in FIG. 5C. In some aspects, the device can downscale the category map before generating one or more modifiers based on the category map.

[0113]

[0130] The device can generate one or more blended modifiers by blending one or more modifiers with information corresponding to the confidence map. The device can generate one or more filtered modifiers by filtering one or more blended modifiers using a low-pass filter. The device can generate one or more upscaled modifiers by upscaling one or more filtered modifiers. Processing the image data using an ISP tuned based on the category map, such as operation 1235, can include processing the image data using one or more modifiers, one or more blended modifiers, one or more filtered modifiers, one or more upscaled modifiers, or some combination thereof.

[0114]

[0131] The image processing techniques illustrated in flowchart 1200 can also include any operations illustrated and described with respect to, or in, any of flowcharts 1250, 1300, and / or 1400.

[0115]

[0132] FIG. 12B is a flowchart 1250 showing an image processing technique. The image processing technique shown by flowchart 1250 can be executed by a device. The device can be an image capture / processing device 100, an image capture device 105A, an image processing device 105B, a classification engine 220, an ISP 240, an image sensor 205, one or more network servers of a cloud service, a computing system 1500, or some combination thereof.

[0116]

[0133] In operation 1255, as part of the image processing technique, the device receives image data captured by the image sensor 205. Operation 1205 of flowchart 1200 can be an example of operation 1255 of flowchart 1250.

[0117]

[0134] In operation 1250, as part of the image processing technique, the device determines that a first object image region in the image data depicts a first object category among a plurality of object categories. Operation 1210 of flowchart 1200 can be an example of operation 1260 of flowchart 1250.

[0118]

[0135] In operation 1265, as part of the image processing technique, the device determines that a second object image region in the image data depicts a second object category among a plurality of object categories. Operation 1215 of flowchart 1200 can be an example of operation 1265 of flowchart 1250.

[0119]

[0136] In operation 1270, as part of the image processing technique, the device identifies a plurality of confidence levels corresponding to a plurality of confidence image regions in the image data, where each confidence level of the plurality of confidence levels identifies the confidence that the corresponding confidence image region among the plurality of confidence image regions depicts one of a plurality of object categories. Operation 1220 of flowchart 1200 can include operation 1270 of flowchart 1250.

[0120]

[0137] In operation 1275, as part of the image processing technique, the device generates an image based on image data using an image capture process, including by applying different setting values of the image capture process to different portions of the image data, where the different portions of the image data are identified based on a first object image region, a second object image region, and a plurality of reliable image regions. Operation 1275 of flow diagram 1250 may, in some examples, include at least a subset of at least one of operations 1220, 1225, 1230, and / or 1235 of flow diagram 1200. For example, some of the different portions of the image data may be different portions of the first object image region having different levels of reliability from each other. Some of the different portions of the image data may be different portions of the second object image region having different levels of reliability from each other. Some of the different portions of the image data may be outside the first object image region and / or the second object image region.

[0121]

[0138] In some examples, the image capture process includes generating one or more modifiers. The one or more modifiers can identify a first deviation from a default setting value of the image capture process for a first object image region, a second deviation from the default setting value of the image capture process for a second object image region, or both. Different setting values of the image capture process can be based on the one or more modifiers. The default setting value can be a default intensity to which a particular parameter (e.g., an ISP parameter) is applied, and each deviation corresponding to each modifier can represent weakening or strengthening of that default intensity. In some examples, as part of an image processing technique, the device adjusts the one or more modifiers. Adjusting the one or more modifiers can include blending the one or more modifiers with a blending update value (e.g., the blending update value generated by generator 435 of FIG. 4) based on a plurality of confidence levels corresponding to a plurality of reliable image regions. Blending the one or more modifiers with the blending update value can adjust at least one of the first deviation and the second deviation in at least one area of the image data. The modification can further weaken or strengthen the intensity to which a particular parameter (e.g., an ISP parameter) is applied.

[0122]

[0139] In some examples, as part of an image processing technique, the device generates a category map that divides the image data into a plurality of object image regions including a first object image region and a second object image region. Each object image region of the plurality of object image regions corresponds to one of a plurality of object categories. The device can identify that the first object category corresponds to a first set value of the image capture process. The device can identify that the second object category corresponds to a second set value of the image capture process. In some examples, as part of an image processing technique, the device generates a reliability map that divides the image data into a plurality of reliability image regions corresponding to a plurality of reliability levels. Different portions of the image data can be identified (e.g., by the device) based on the category map and the reliability map.

[0123]

[0140] In some examples, the image capture process can include processing of the image data of operation 1235. In some examples, the first set value of the image capture process can be the first tuning set value described with respect to operations 1225 and 1235. In some examples, the second set value of the image capture process can be the second tuning set value described with respect to operations 1230 and 1235.

[0124]

[0141] In some examples, the image capture process includes processing the image data using an Image Signal Processor (ISP). Different setting values of the image capture process can be different tuning setting values of the ISP. In some examples, the different tuning setting values of the ISP include different intensities at which ISP tuning parameters are applied during the processing of the image data using the ISP. The ISP tuning parameters can be, for example, one of noise reduction, sharpening, color saturation, color mapping, color processing, and tone mapping. In some examples, the different setting values include setting values related to at least one of lens position, flash, focus, exposure, white balance, aperture size, shutter speed, ISO, analog gain, digital gain, noise removal, sharpening, tone mapping, color saturation, demosaicing, color space conversion, shading, edge enhancement, high dynamic range (HDR) image composition, special effects, artificial noise addition, edge-directed upscaling, upscaling, downscaling, electronic image stabilization, or a combination thereof. In some examples, the device processes the image data. Processing the image data can include demosaicing the image data and / or converting the image data from a first color space to a second color space (e.g., between a Bayer color space, an RGB color space, and / or a YUV color space).

[0125]

[0142] In some examples, as part of an image processing technique, the device receives user input related to at least one of a first object image region and a second object image region. At least one of different set values can be defined based on the user input and can correspond to either the first object image region or the second object image region. In some examples, applying different set values of the image capture process to different portions of the image data includes applying different set values of the image capture process to different portions of the image data using an image signal processor (ISP). In some examples, identifying the first object image region and the second object image region includes using a classification engine that is at least partially located on an integrated circuit (IC) chip, such as an application specific integrated circuit (ASIC) chip, to identify the first object image region and the second object image region. In some examples, as part of an image processing technique, the device displays an image on a display.

[0126]

[0143] The image processing technique shown in flow diagram 1250 may also include any operations illustrated or described with respect to any of flow diagrams 1200, 1300, and / or 1400.

[0127]

[0144] FIG. 13 is a flow diagram 1300 showing a transition smoothing technique. The image processing technique shown by flow diagram 1300 can be executed by a device. The device can be an image capture / processing device 100, an image capture device 105A, an image processing device 105B, a classification engine 220, an ISP 240, an image sensor 205, one or more network servers of a cloud service, a computing system 1500, or some combination thereof.

[0128]

[0145] In operation 1305, the transition smoothing technique includes receiving a category map and a confidence map. In operation 1310, the transition smoothing technique includes downscaling the category map. In some examples, operation 1310 may be skipped so that the category map is not downscaled.

[0129]

[0146] In operation 1315, the transition smoothing technique includes generating one or more modifiers based on the category map. For example, the one or more modifiers identify at least one of a first deviation from a default setting value applied by the ISP in a first image region during processing of the image data and a second deviation from a default setting value applied by the ISP in a second image region during processing of the image data. The internal signal 540 shown in FIGS. 5A, 5B, and 5C may represent an example of a default setting value. The category-based modifiers 465, 765A, and 765B may represent examples of one or more modifiers based on the category map. The generators 430, 730A, and 730B may perform operation 1315.

[0130]

[0147] In operation 1320, the transition smoothing technique includes generating one or more blended modifiers by blending one or more modifiers with information corresponding to the confidence map. The category confidence blending operations 440, 740A, and 740B (for example, using the confidence as a blending factor by blending the modifier with a no-operation equivalent modifier value according to the confidence, etc.) may represent examples of operation 1320. The blending update values for the category-based modifiers 465, 765A, and 765B generated by the generators 435, 735A, and 735B may represent examples of information corresponding to the confidence map.

[0131]

[0148] In operation 1325, the transition smoothing technique includes generating one or more filtered modifiers by filtering one or more blended modifiers using a low-pass filter (LPF). LPFs 445, 745A, and 745B may represent examples of the LPF for operation 1325.

[0132]

[0149] In operation 1330, the transition smoothing technique includes generating one or more upscaled modifiers by upscaling one or more filtered modifiers. Upscalers 450, 750A, and 750B may perform operation 1330.

[0133]

[0150] In operation 1335, the transition smoothing technique includes processing image data using an ISP tuned based on one or more upscaled modifiers. In one example, operation 1335 may be performed by the module logic of an ISP tuning parameter module such as the NR module logic 405 of the NR module 320.

[0134]

[0151] In some cases, one or more of operations 1305 - 1335 of flowchart 1300 may be performed by a device that performs one or more of operations 1205 - 1235 of flowchart 1200. In some cases, the transition smoothing technique of FIG. 13 may be part of the image processing technique of FIG. 12A. The image processing technique of FIG. 12A may represent at least a part of the operations of the classification engine 220 and / or the ISP 240. The transition smoothing technique of FIG. 13 may represent at least a part of the operations of the STMP 365 and / or the downscaler 360.

[0135]

[0152] The transition smoothing technique shown in flowchart 1300 may also include any operation illustrated and described in or with respect to any of flowcharts 1200, 1250, and / or 1400.

[0136]

[0153] FIG. 14 is a flowchart 1400 showing an image upscaling technique. The image processing technique shown by flowchart 1400 can be executed by a device. The device can be an image capture / processing device 100, an image capture device 105A, an image processing device 105B, a classification engine 220, an ISP 240, an image sensor 205, one or more network servers of a cloud service, a computing system 1500, or some combination thereof.

[0137]

[0154] In operation 1405, the image upscaling technique includes receiving a category map 230 and a confidence map 235. Optionally, the category map 230 and the confidence map 235 can be a single file having both category information and confidence information for each pixel, as previously described.

[0138]

[0155] In operation 1410, the image upscaling technique includes identifying a first image region and a second image region of the category map 230, where the first image region is narrower than the second image region.

[0139]

[0156] In operation 1415, the image upscaling technique includes identifying a first filter size corresponding to the first image region and a second filter size corresponding to the second image region, where the first filter size is smaller than the second filter size.

[0140]

[0157] In operation 1420, the image upscaling technique includes upscaling a first pixel in the first image region based on the first filter size and one or more weights associated with one or more confidence values from the confidence map 235 corresponding to one or more pixels adjacent to the first pixel. The confidence level may sometimes be referred to as the confidence level or confidence.

[0141]

[0158] In operation 1425, the image upscaling technique includes upscaling a second pixel in a second image region based on a second filter size and one or more weights associated with one or more confidence values from a confidence map corresponding to one or more pixels adjacent to the second pixel.

[0142]

[0159] The image upscaling technique illustrated in flowchart 1400 may also include any operation illustrated and described in or with respect to any of flowcharts 1200, 1250, and / or 1300.

[0143]

[0160] In some cases, one or more of operations 1405 - 1435 of flowchart 1400 may be performed by a device that performs one or more of operations 1205 - 1235 of flowchart 1200. In some cases, the image upscaling technique of FIG. 14 may be part of the image processing technique of FIG. 12A. The image processing technique of FIG. 12A may represent at least a portion of the operations of classification engine 220 and / or ISP 240. The image upscaling technique of FIG. 14 may represent at least a portion of the operations of category map upscaler (CMUS) 905 that may be used in upscaler 310.

[0144]

[0161] In some cases, at least a subset of the techniques shown by flowcharts 1200, 1250, 1300, and 1400 may be remotely executed by one or more network servers of a cloud service. In some examples, the processes described herein (e.g., the processes including operations 200, 300, 400, 700, 900, 1100, 1200, 1250, 1300, 1400, and / or other processes described herein) may be executed by a computing device or apparatus. In one example, processes 200, 300, 400, 700, 900, 1100, 1200, 1250, 1300, and / or 1400 may be executed by the image capture device 105A of FIG. 1. In another example, the process including operations 200, 300, 400, 700, 900, 1100, 1200, 1250, 1300, and / or 1400 may be executed by the image processing device 105B of FIG. 1. The process including operations 200, 300, 400, 700, 900, 1100, 1200, 1250, 1300, and / or 1400 may also be executed by the image capture / processing system 100 of FIG. 1. The process including operations 200, 300, 400, 700, 900, 1100, 1200, 1250, 1300, and / or 1400 may be executed by a computing device having the architecture of the computing system 1500 shown in FIG. 15. The computing device can include a mobile device (e.g., a mobile phone), a desktop computing device, a tablet computing device, a wearable device (e.g., a VR headset, an AR headset, AR glasses, a network-connected watch or smartwatch, or other wearable device), a server computer, an autonomous vehicle or the computing device of an autonomous vehicle, a robotic device, a television, and / or any other suitable device including any device having resource capabilities for executing the processes described herein, including the processes including operations 200, 300, 400, 700, 900, 1100, 1200, 1250, 1300, and / or 1400.In some cases, a computing device or apparatus may include various components such as one or more input devices, one or more output devices, one or more processors, one or more microprocessors, one or more microcomputers, one or more cameras, one or more sensors, and / or other components configured to perform the steps of the processes described herein. In some examples, the computing device may include a display, a network interface configured to communicate and / or receive data, any combination thereof, and / or other components. The network interface may be configured to communicate and / or receive Internet Protocol (IP)-based data or other types of data.

[0145]

[0162] The components of the computing device may be implemented within a circuit. For example, the components can include an electronic circuit or other electronic hardware that can include one or more programmable electronic circuits (such as a microprocessor, a graphics processing unit (GPU), a digital signal processor (DSP), a central processing unit (CPU), and / or other suitable electronic circuits), and / or can be implemented using them, and / or can include computer software, firmware, or any combination thereof, and / or can be implemented using them to perform the various operations described herein.

[0146]

[0163] The processes shown by the conceptual and flow diagrams 200, 300, 400, 700, 900, 1100, 1200, 1250, 1300, 1400 are organized as logical flow diagrams, and their operations represent a sequence of operations that can be implemented in hardware, computer instructions, or a combination thereof. In the context of computer instructions, an operation represents computer-executable instructions stored on one or more computer-readable storage media that, when executed by one or more processors, perform the recited operations. Generally, computer-executable instructions include routines, programs, objects, components, data structures, etc. that perform a particular function or implement a particular data type. The order in which the operations are described is not intended to be construed as limiting, and any number of the described operations may be combined in any order and / or in parallel to implement the process.

[0147]

[0164] In addition, the processes shown by the conceptual and flow diagrams 200, 300, 400, 700, 900, 1100, 1200, 1250, 1300, 1400 and / or other processes described herein may be executed under the control of one or more computer systems configured with executable instructions and may be implemented in hardware, or a combination thereof, as code (e.g., executable instructions, one or more computer programs, or one or more applications) executed collectively on one or more processors. As described above, the code may be stored on a computer-readable storage medium or a machine-readable storage medium, for example, in the form of a computer program comprising a plurality of instructions executable by one or more processors. The computer-readable storage medium or the machine-readable storage medium may be non-transitory.

[0148]

[0165] FIG. 15 is a diagram showing an example of a system for implementing some aspects of the present technology. In particular, FIG. 15 shows an example of a computing system 1500 that can be any of the computing devices, remote computing systems, cameras, or components of the system that make up an internal computing system, where any of those components communicate with each other using connection 1505. Connection 1505 can be a physical connection using a bus or a direct connection to processor 1510 in a chipset architecture or the like. Connection 1505 can also be a virtual connection, a network connection, or a logical connection.

[0149]

[0166] In some embodiments, computing system 1500 is a distributed system in which the functions described in this disclosure can be distributed across one data center, multiple data centers, within a peer network, etc. In some embodiments, one or more of the system components described represent many such components, each of which performs some or all of the functions described for the component. In some embodiments, the components can be physical devices or virtual devices.

[0150]

[0167] The exemplary system 1500 includes at least one processing unit (CPU or processor) 1510 and a connection 1505 that couples various system components, including system memory 1515 such as read only memory (ROM) 1520 and random access memory (RAM) 1525, to the processor 1510. Computing system 1500 can include a cache 1512 of high-speed memory that is directly connected to, very close to, or integrated as part of the processor 1510.

[0151]

[0168] The processor 1510 can include any general-purpose processor, hardware services or software services such as services 1532, 1534, and 1536 stored in the storage device 1530 and configured to control the processor 1510, and a dedicated processor in which software instructions are incorporated into an actual processor design. The processor 1510 may essentially be a fully self - contained computing system including multiple cores or processors, buses, memory controllers, caches, etc. The multi - core processor may be symmetric or asymmetric.

[0152]

[0169] To enable user interaction, computing system 1500 includes an input device 1545 that can represent any number of input mechanisms such as a microphone for voice, a touch-sensitive screen for gesture or graphical input, a keyboard, a mouse, motion input, voice, etc. Computing system 1500 can also include an output device 1535 that can be one or more of several output mechanisms. In some cases, a multimodal system can be enabled to provide multiple types of input and output for a user to communicate with computing system 1500. Computing system 1500 can include a communication interface 1540, which can generally orchestrate and manage user input and system output.The communication interface may use a wired transceiver and / or a wireless transceiver, including those that utilize an audio jack / plug, a microphone jack / plug, a Universal Serial Bus (USB) port / plug, an Apple® Lightning® port / plug, an Ethernet® port / plug, an optical fiber port / plug, a proprietary wired port / plug, BLUETOOTH® wireless signal transfer, BLUETOOTH Low Energy (BLE) wireless signal transfer, iBeacon® wireless signal transfer, Radio Frequency Identification (RFID) wireless signal transfer, Near Field Communication (NFC) wireless signal transfer, Dedicated Short Range Communication (DSRC) wireless signal transfer, 802.11 Wi-Fi wireless signal transfer, Wireless Local Area Network (WLAN) signal transfer, Visible Light Communication (VLC), Worldwide Interoperability for Microwave Access (WiMAX®), infrared (IR) communication wireless signal transfer, Public Switched Telephone Network (PSTN) signal transfer, Integrated Services Digital Network (ISDN) signal transfer, 3G / 4G / 5G / LTE cellular data network wireless signal transfer, ad hoc network signal transfer, radio wave signal transfer, microwave signal transfer, infrared signal transfer, visible light signal transfer, ultraviolet light signal transfer, wireless signal transfer along the electromagnetic spectrum, or some combination thereof, to perform or facilitate the reception and / or transmission of wired communication or wireless communication. The communication interface 1540 may also include one or more GNSS system receivers or transceivers that are used to determine the location of the computing system 1500 based on the reception of one or more signals from one or more satellites associated with one or more Global Navigation Satellite Systems (GNSS). GNSS systems include, but are not limited to, the United States-based Global Positioning System (GPS), the Russia-based Global Navigation Satellite System (GLONASS), the China-based BeiDou Navigation Satellite System (BDS), and the Europe-based Galileo GNSS.There is no restriction on operating on any specific hardware configuration. Thus, the basic features here can be easily replaced with improved hardware configurations or firmware configurations when they are developed.

[0153]

[0170] The storage device 1530 can be a non-volatile memory device and / or a non-transitory memory device and / or a computer-readable memory device, such as a magnetic cassette, a flash memory card, a solid-state memory device, a digital versatile disk, a cartridge, a floppy (registered trademark) disk, a flexible disk, a hard disk, a magnetic tape, a magnetic strip / stripe, any other magnetic storage medium, flash memory, memristor memory, any other solid-state memory, a compact disk read-only memory (CD-ROM) optical disk, a rewritable compact disk (CD) optical disk, a digital video disk (DVD) optical disk, a Blu-ray (registered trademark) disk (BDD) optical disk, a holographic optical disk, another optical medium, a Secure Digital (SD) card, a micro Secure Digital (microSD) card, a Memory Stick (registered trademark) card, a smart card chip, an EMV chip, a subscriber identity module (SIM) card, a mini / micro / nano / pico SIM card, another integrated circuit (IC) chip / card, a random access memory (RAM), a static RAM (SRAM), a dynamic RAM (DRAM), a read-only memory (ROM), a programmable read-only memory (PROM), an erasable programmable read-only memory (EPROM), an electrically erasable programmable read-only memory (EEPROM (registered trademark)), a flash EPROM (FLASHEPROM), a cache memory (L1 / L2 / L3 / L4 / L5 / L#), a resistive random access memory (RRAM (registered trademark) / ReRAM), a phase change memory (PCM), a spin transfer torque RAM (STT-RAM), another memory chip or cartridge, and / or a combination thereof, and can be a hard disk or other type of computer-readable medium capable of storing data accessible by a computer.

[0154]

[0171] The memory device 1530 can include software services, servers, services, etc., which cause the system to perform functions when the code defining such software is executed by the processor 1510. In some embodiments, a hardware service that performs a particular function can include software components stored on a computer-readable medium together with the necessary hardware components such as the processor 1510, the connection portion 1505, the output device 1535, etc. for performing that function.

[0155]

[0172] As used herein, the term "computer-readable medium" includes, but is not limited to, portable or non-portable storage devices, optical storage devices, and various other media capable of storing, containing, or carrying instructions and / or data. The computer-readable medium can include non-transitory media that can store data and do not include carrier waves and / or transient electronic signals that propagate wirelessly or via a wired connection. Examples of non-transitory media can include, but are not limited to, magnetic disks or tapes, optical storage media such as compact discs (CDs) or digital versatile discs (DVDs), flash memory, memory, or memory devices. The computer-readable medium can store code and / or machine-executable instructions that can represent any combination of procedures, functions, subprograms, programs, routines, subroutines, modules, software packages, classes, or instructions, data structures, or program statements. A code segment can be coupled to another code segment or hardware circuit by passing and / or receiving information, data, arguments, parameters, or memory contents. Information, arguments, parameters, data, etc. can be passed, forwarded, or transmitted using any suitable means including memory sharing, message passing, token passing, network transmission, etc.

[0156]

[0173] In some embodiments, a computer-readable storage device, medium, and memory can include a cable or wireless signal including a bitstream and the like. However, when stated, non-transitory computer-readable storage media expressly exclude media such as energy, carrier signals, electromagnetic waves, and signals themselves.

[0157]

[0174] To provide a complete understanding of the embodiments and examples provided herein, specific details are given in the above description. However, one of ordinary skill in the art will understand that the embodiments can be practiced without these specific details. For the sake of clarity of explanation, in some instances, the present technology may be presented as including individual functional blocks including devices, device components, steps or routines in methods implemented in software, or functional blocks comprising combinations of hardware and software. Additional components other than those shown in the figures and / or described herein may be used. For example, circuits, systems, networks, processes, and other components may be shown as components in block diagram form so as not to obscure the embodiments with unnecessary detail. In other instances, well-known circuits, processes, algorithms, structures, and techniques may be shown without unnecessary detail so as not to obscure the embodiments.

[0158]

[0175] Individual embodiments may be described above as a process or method depicted as a flowchart, flow diagram, data flow diagram, structure diagram, or block diagram. A flowchart may describe operations as a sequential process, but many of the operations may be performed in parallel or simultaneously. Further, the order of the operations may be rearranged. When the operations thereof are completed, the process ends, but may have additional steps not included in the figure. A process may correspond to a method, function, procedure, subroutine, subprogram, etc. When the process corresponds to a function, its end may correspond to the return value of the function to a calling function or main function.

[0159]

[0176] The processes and methods according to the examples described above can be implemented using computer-executable instructions that are stored or otherwise available from a computer-readable medium. Such instructions can include, for example, instructions and data that cause a general-purpose computer, a special-purpose computer, or a processing device to perform a certain function or group of functions, or otherwise configure a general-purpose computer, a special-purpose computer, or a processing device to perform a certain function or group of functions. Portions of the computer resources used can be made accessible over a network. The computer-executable instructions can be, for example, binary, intermediate format instructions such as assembly language, firmware, source code, and the like. Examples of computer-readable media that can be used to store the instructions, the information used, and / or the information created in the methods according to the examples described include magnetic or optical disks, flash memory, USB devices with non-volatile memory, networked storage devices, and the like.

[0160]

[0177] Devices implementing the processes and methods according to these disclosures can include hardware, software, firmware, middleware, microcode, hardware description languages, or any combination thereof, and can take any of various form factors. When implemented in software, firmware, middleware, or microcode, the program code or code segments (e.g., a computer program product) for performing the required tasks can be stored in a computer-readable medium or a machine-readable medium. A processor can perform the required tasks. Typical examples of form factors include laptop computers, smartphones, mobile phones, tablet devices or other small form factor personal computers, personal digital assistants, rack-mounted devices, stand-alone devices, etc. The functionality described herein can also be embodied in peripheral devices or add-in cards. Such functionality can also be implemented on a circuit board, as a further example, between different chips or different processes running on a single device.

[0161]

[0178] Instructions, a medium for transmitting such instructions, computing resources for executing them, and other structures for supporting such computing resources are exemplary means for providing the functionality described in this disclosure.

[0162]

[0179] In the above description, although the aspects of the present application have been described with reference to specific embodiments thereof, those skilled in the art will recognize that the present application is not limited thereto. Thus, although exemplary embodiments of the present application are described in detail herein, except where limited by the prior art, the inventive concept may, in some cases, be variously implemented and adopted, and it should be understood that the appended claims are to be construed to include such variations. The various features and aspects of the application examples described above may be used individually or together. Further, the embodiments may be utilized in any number of environments and application examples other than those described herein and the environments and application examples, without departing from the broader spirit and scope of this specification. Accordingly, this specification and the drawings should be considered as exemplary rather than restrictive. For purposes of illustration, the methods are described in a particular order. It should be understood that in alternative embodiments, the methods may be performed in an order different from that described.

[0163]

[0180] It will be understood by those skilled in the art that the symbols or terms "less than" ("<") and "greater than" (">") used herein may be replaced, without departing from the scope of this specification, by the symbols "less than or equal to" ("≦") and "greater than or equal to" ("≧"), respectively.

[0164]

[0181] When a component is described as being "configured to" perform a certain operation, such a configuration can be achieved, for example, by designing an electronic circuit or other hardware to perform that operation, by programming a programmable electronic circuit (e.g., a microprocessor, or other suitable electronic circuit) to perform that operation, or by any combination thereof.

[0165]

[0182] The term "coupled" refers to any component that is physically connected to another component, either directly or indirectly, and / or that communicates with another component, either directly or indirectly (e.g., is connected to another component via a wired or wireless connection and / or via another suitable communication interface).

[0166]

[0183] The language of a claim or other language that recites "at least one" of a set and / or "one or more" of a set indicates that one member of the set or multiple members of the set (in any combination) satisfy the claim. For example, the language of a claim that recites "at least one of A and B" means A, B, or A and B. In another example, the language of a claim that recites "at least one of A, B, and C" means A, B, C, or A and B, or A and C, or B and C, or A and B and C. The language "of the set of at least one" and / or "one or more" of a set does not limit the set to the items recited in the set. For example, the language of a claim that recites "at least one of A and B" can mean A, B, or A and B, and can further include items not recited in the set of A and B.

[0167]

[0184] With respect to the embodiments disclosed in this specification, the various illustrative logical blocks, modules, circuits, and algorithm steps described may be implemented as electronic hardware, computer software, firmware, or any combination thereof. To clearly illustrate this interchangeability of hardware and software, various illustrative components, blocks, modules, circuits, and steps have been described above generally in terms of their functionality. Whether such functionality is implemented as hardware or software depends upon the particular application and design constraints imposed on the overall system. Skilled artisans may implement the described functionality in varying ways for each particular application, but such implementation decisions should not be interpreted as causing a departure from the scope of the present application.

[0168]

[0185] The techniques described herein may also be implemented in electronic hardware, computer software, firmware, or any combination thereof. Such techniques may be implemented in any of a variety of devices, including integrated circuit devices having multiple applications, such as general-purpose computers, wireless communication device handsets, or wireless communication device handsets and other devices. Any features described as modules or components may be implemented together in an integrated logic device or separately as discrete but interoperable logic devices. When implemented in software, the techniques may be at least partially realized by a computer-readable data storage medium comprising program code that includes instructions, which when executed, perform one or more of the methods described above. The computer-readable data storage medium may form part of a computer program product that may include packaging material. The computer-readable medium may comprise a memory or data storage media, such as random access memory (RAM), such as synchronous dynamic random access memory (SDRAM), read only memory (ROM), non-volatile random access memory (NVRAM), electrically erasable programmable read only memory (EEPROM), flash memory, magnetic or optical data storage media. The techniques may additionally or alternatively be at least partially realized by a computer-readable communication medium that conveys or communicates program code in the form of instructions or data structures and that can be accessed, read, and / or executed by a computer, such as a propagated signal or wave.

[0169]

[0186] The program code can be executed by a processor that can include one or more processors such as one or more digital signal processors (DSPs), general-purpose microprocessors, application-specific integrated circuits (ASICs), field programmable logic arrays (FPGAs), or other equivalent integrated logic circuits or discrete logic circuits. Such a processor can be configured to execute any of the techniques described in this disclosure. The general-purpose processor can be a microprocessor, but alternatively, the processor can be any conventional processor, controller, microcontroller, or state machine. The processor can also be implemented as a combination of computing devices, such as a combination of a DSP and a microprocessor, multiple microprocessors, one or more microprocessors cooperating with a DSP core, or any other such configuration. Thus, the term "processor" as used herein can refer to any of the above structures, any combination of the above structures, or any other structure or device suitable for implementation of the techniques described herein. Additionally, in some aspects, the functions described herein can be provided within dedicated software modules or hardware modules configured for encoding and decoding, or incorporated into a combined video encoder decoder (CODEC).

[0170]

[0187] Exemplary aspects of the present disclosure include the following.

[0171]

[0188] Aspect 1: A method for processing video data. The method includes receiving image data captured by an image sensor, determining that a first image region in the image data depicts a first object category among a plurality of object categories, determining that a second image region in the image data depicts a second object category among the plurality of object categories, and generating an image based on the image data using an image capture process by applying a first setting value of the image capture process to the first image region and a second setting value of the image capture process to the second image region.

[0172]

[0189] Aspect 2: The method according to aspect 1, further comprising generating a category map by dividing the image data into a plurality of image regions including a first image region and a second image region, wherein each image region of the plurality of image regions corresponds to one of a plurality of object categories, identifying that the first object category corresponds to a first setting value of the image capture process, and identifying that the second object category corresponds to a second setting value of the image capture process.

[0173]

[0190] Aspect 3: A method further comprising generating a downscaled copy of the image data by downscaling the image data, wherein generating a category map based on the image data includes generating a category map based on the downscaled copy of the image data, the method according to aspect 1 or 2.

[0174]

[0191] Aspect 4: A method further comprising generating a confidence map based on image data, the confidence map identifying a plurality of confidence levels corresponding to a plurality of portions of the image data, wherein each confidence level of the plurality of confidence levels identifies the confidence that the corresponding portion of the plurality of portions depicts one of a plurality of object categories, and generating the image includes processing the image data using an image signal processor (ISP) tuned based on the category map and the confidence map, the method according to any one of Aspects 1 to 3.

[0175]

[0192] Aspect 5: The method according to any one of Aspects 1 to 4, further comprising upscaling the category map.

[0176]

[0193] Aspect 6: Upscaling the category map includes upscaling the category map to a size that matches at least one of the sizes of the image data and the image, the method according to any one of Aspects 1 to 5.

[0177]

[0194] Aspect 7: Upscaling the category map is performed using nearest neighbor upscaling, the method according to any one of Aspects 1 to 6.

[0178]

[0195] Aspect 8: Upscaling the category map is performed using nearest neighbor upscaling modified by spatial weight filtering, the method according to any one of Aspects 1 to 7.

[0179]

[0196] Aspect 9: Upscaling the category map using nearest neighbor upscaling modified by spatial weight filtering involves identifying a first filter size corresponding to a first image region and a second filter size corresponding to a second image region, where the first filter size is smaller than the second filter size, and upscaling a first pixel in the first image region based on the first filter size and one or more weights associated with one or more confidence values from a confidence map corresponding to one or more pixels adjacent to the first pixel, and upscaling a second pixel in the second image region based on the second filter size and one or more weights associated with one or more confidence values from a confidence map corresponding to one or more pixels adjacent to the second pixel, the method according to any one of Aspects 1 to 8.

[0180]

[0197] Aspect 10: The image capture process includes processing image data using an image signal processor (ISP), where a first set value of the image capture process is a first tuning set value of the ISP, and a second set value of the image capture process is a second tuning set value of the ISP, the method according to any one of Aspects 1 to 9.

[0181]

[0198] Aspect 11: The first tuning set value and the second tuning set value include different intensities to which noise reduction ISP tuning parameters are applied during the processing of the image data, the method according to any one of Aspects 1 to 10.

[0182]

[0199] Aspect 12: The first tuning set value and the second tuning set value include different intensities to which sharpening ISP tuning parameters are applied during the processing of the image data, the method according to any one of Aspects 1 to 11.

[0183]

[0200] Aspect 13: The method according to any one of Aspects 1 to 12, wherein the first tuning setting value and the second tuning setting value include different intensities to which the color saturation ISP tuning parameter is applied during the processing of the image data.

[0184]

[0201] Aspect 14: The method according to any one of Aspects 1 to 13, wherein the first tuning setting value and the second tuning setting value include different intensities to which the tone mapping ISP tuning parameter is applied during the processing of the image data.

[0185]

[0202] Aspect 15: The method according to any one of Aspects 1 to 14, wherein the first tuning setting value and the second tuning setting value include different intensities to which the gamma ISP tuning parameter is applied during the processing of the image data.

[0186]

[0203] Aspect 16: The ISP (Image Signal Processor) tuning setting value sets at least one value related to at least one of the ISP noise reduction module, the ISP sharpening module, the ISP tone mapping module, the ISP color saturation module, the ISP gamma module, the ISP blurring module, the ISP demosaicing module, the ISP color space conversion module, the ISP gain module, the ISP luminance module, the ISP shading module, the ISP edge enhancement module, the ISP high dynamic range (HDR) image synthesis module, the ISP special effect processing module, the ISP artificial noise (e.g., particles) adder module, the ISP edge-directed upscaling module, the ISP autofocus module, the ISP automatic exposure module, the ISP automatic white balance module, the ISP aperture size control module, the ISP shutter speed control module, the ISP ISO control module, the ISP lens position module, the ISP electronic image stabilization module, and the ISP flash control module, according to the method described in any one of Aspects 1 to 15.

[0187]

[0204] Aspect 17: A method further comprising generating one or more modifiers, wherein the one or more modifiers identify at least one of a first deviation from default tuning setting values applied by an ISP in a first image region during processing of image data and a second deviation from default tuning setting values applied by the ISP in a second image region during processing of the image data, the method according to any one of Aspects 1 to 16.

[0188]

[0205] Aspect 18: The method according to any one of Aspects 1 to 17, wherein the one or more modifiers identify at least one of the first deviation and the second deviation by multiplying the one or more modifiers by default ISP tuning setting values.

[0189]

[0206] Aspect 19: The method according to any one of Aspects 1 to 18, wherein the one or more modifiers identify at least one of the first deviation and the second deviation by adding the one or more modifiers to default ISP tuning setting values.

[0190]

[0207] Aspect 20: The method according to any one of Aspects 1 to 19, wherein the one or more modifiers identify at least one of the first deviation and the second deviation based on increments within a list of predetermined possible ISP tuning settings including the default ISP tuning setting values, the increments being based on the modifiers.

[0191]

[0208] Aspect 21: A method further comprising generating a category map by partitioning image data into a plurality of image regions including a first image region and a second image region, wherein each image region of the plurality of image regions corresponds to one of a plurality of object categories, and the one or more modifiers are generated based at least on the category map, the method according to any one of Aspects 1 to 20.

[0192]

[0209] Aspect 22: The method according to any one of Aspects 1 to 21, further comprising downscaling the category map before generating one or more modifiers based on the category map.

[0193]

[0210] Aspect 23: A method further comprising generating one or more blend modifiers by blending one or more modifiers with information corresponding to a reliability map that identifies a plurality of reliability levels corresponding to a plurality of portions of the image data, wherein each reliability level of the plurality of reliability levels identifies the reliability when determining that the corresponding portion of the plurality of portions depicts one of a plurality of object categories. The method according to any one of Aspects 1 to 22.

[0194]

[0211] Aspect 24: The method according to any one of Aspects 1 to 23, further comprising generating one or more filtered modifiers by filtering one or more blend modifiers using a low-pass filter.

[0195]

[0212] Aspect 25: The method according to any one of Aspects 1 to 24, further comprising generating one or more upscaled modifiers by upscaling one or more filtered modifiers.

[0196]

[0213] Aspect 26: Processing the image data using an ISP includes processing the image data using at least one of one or more modifiers, one or more blend modifiers, one or more filtered modifiers, and one or more upscaled modifiers. The method according to any one of Aspects 1 to 25.

[0197]

[0214] Aspect 27: A method according to any one of Aspects 1 to 26, wherein at least one of a first set value of an image capture process and a second set value of the image capture process is a tuning set value related to at least one of a lens position, a flash, a focus, an exposure, a white balance, an aperture size, a shutter speed, an ISO, an analog gain, a digital gain, noise removal, sharpening, tone mapping, color saturation, demosaicing, color space conversion, shading, edge enhancement, high dynamic range (HDR) image composition, special effects, grain addition, artificial noise addition, edge-directed upscaling, upscaling, downscaling, and electronic image stabilization.

[0198]

[0215] Aspect 28: A method according to any one of Aspects 1 to 27, wherein the image data is raw image data.

[0199]

[0216] Aspect 29: A method according to any one of Aspects 1 to 28, further comprising demosaicing the image data.

[0200]

[0217] Aspect 30: A method according to any one of Aspects 1 to 29, further comprising converting the image data from a first color space to a second color space.

[0201]

[0218] Aspect 31: A method according to any one of Aspects 1 to 30, wherein the second color space is a YUV color space.

[0202]

[0219] Aspect 32: An apparatus for image processing, comprising one or more memory units for storing instructions and one or more processors for executing the instructions, wherein execution of the instructions by the one or more processors causes the one or more processors to execute the method according to any one of Aspects 1 to 31.

[0203]

[0220] Aspect 33: An apparatus according to Aspect 32, which is a mobile device.

[0204]

[0221] Aspect 34: The apparatus according to aspect 32 or 33, which is a wireless communication device.

[0205]

[0222] Aspect 35: The apparatus according to any one of aspects 32 to 34, which is a camera including at least an image sensor and one or more processors.

[0206]

[0223] Aspect 36: The apparatus according to any one of aspects 32 to 35, wherein the one or more processors include an Image Signal Processor (ISP).

[0207]

[0224] Aspect 37: The apparatus according to any one of aspects 32 to 36, wherein the one or more processors include a classification engine.

[0208]

[0225] Aspect 38: The apparatus according to any one of aspects 32 to 37, including a display configured to display an image.

[0209]

[0226] Aspect 39: A non-transitory computer-readable storage medium embodying a program executable by a processor to perform an image processing method comprising the method according to any one of aspects 1 to 31.

[0210]

[0227] Aspect 40: An apparatus for image processing, comprising means for performing the method according to any one of aspects 1 to 31.

[0211]

[0228] Aspect 41: An apparatus for image processing, the apparatus comprising a memory and one or more processors coupled to the memory, the one or more processors being configured to receive image data captured by an image sensor, determine that a first object image region in the image data depicts a first object category among a plurality of object categories, determine that a second object image region in the image data depicts a second object category among the plurality of object categories, identify a plurality of confidence levels corresponding to a plurality of reliable image regions of the image data, wherein each confidence level of the plurality of confidence levels identifies a confidence that a corresponding reliable image region of the plurality of reliable image regions depicts one of the plurality of object categories, generate an image based on the image data using an image capture process, including by applying different setting values of the image capture process to different portions of the image data, the different portions of the image data being identified based on the first object image region, the second object image region, and the plurality of reliable image regions.

[0212]

[0229] Aspect 42: The one or more processors are configured to generate one or more modifiers, the one or more modifiers identifying at least one of a first deviation from a default setting value of the image capture process for the first object image region and a second deviation from the default setting value of the image capture process for the second object image region, wherein different setting values of the image capture process are based on the one or more modifiers, the apparatus according to aspect 41.

[0213]

[0230] Aspect 43: One or more processors are configured to adjust one or more modifiers, including blending the one or more modifiers with a blending update value based on a plurality of confidence levels corresponding to a plurality of reliable image regions, wherein blending the one or more modifiers with the blending update value adjusts at least one of a first deviation and a second deviation in at least one area of the image data, the apparatus according to aspect 41 or 42.

[0214]

[0231] Aspect 44: One or more processors are configured to generate a category map that divides the image data into a plurality of object image regions including a first object image region and a second object image region, wherein each object image region of the plurality of object image regions corresponds to one of a plurality of object categories, identifying that the first object category corresponds to a first setting value of the image capture process, and identifying that the second object category corresponds to a second setting value of the image capture process, the apparatus according to any one of aspects 41 to 43.

[0215]

[0232] Aspect 45: One or more processors are configured to generate a confidence map that divides the image data into a plurality of reliable image regions corresponding to a plurality of confidence levels, and different portions of the image data are identified based on the category map and the confidence map, the apparatus according to any one of aspects 41 to 44.

[0216]

[0233] Aspect 46: The image capture process includes processing the image data using an image signal processor (ISP) among one or more processors, wherein different setting values of the image capture process are different tuning setting values of the ISP, the apparatus according to any one of aspects 41 to 45.

[0217]

[0234] Aspect 47: Different tuning setting values of the ISP include different intensities to which ISP tuning parameters are applied during the processing of image data using the ISP, where the ISP tuning parameter is one of noise reduction, sharpening, color saturation, color mapping, color processing, and tone mapping, for the apparatus according to any one of Aspects 41 to 46.

[0218]

[0235] Aspect 48: Different tuning setting values of the ISP include different intensities to which ISP tuning parameters are applied during the processing of image data using the ISP, where the ISP tuning parameter is one of noise reduction, sharpening, color saturation, color mapping, color processing, and tone mapping, for the method according to any one of Aspects 41 to 47.

[0219]

[0236] Aspect 49: Different setting values include setting values related to at least one of lens position, flash, focus, exposure, white balance, aperture size, shutter speed, ISO, analog gain, digital gain, noise removal, sharpening, tone mapping, color saturation, demosaicing, color space conversion, shading, edge enhancement, high dynamic range (HDR) image composition, special effects, artificial noise addition, edge-directed upscaling, upscaling, downscaling, and electronic image stabilization, for the apparatus according to any one of Aspects 41 to 48.

[0220]

[0237] Aspect 50: One or more processors are configured to process image data, including at least one of demosaicing the image data and converting the image data from a first color space to a second color space, for the apparatus according to any one of Aspects 41 to 49.

[0221]

[0238] Aspect 51: One or more processors are configured to receive user input related to at least one of a first object image region and a second object image region, where at least one of different setting values is defined based on the user input and corresponds to one of the first object image region and the second object image region, the apparatus according to any one of Aspects 41 to 50.

[0222]

[0239] Aspect 52: One or more processors include an image signal processor (ISP) that applies different setting values of an image capture process to different portions of image data, the apparatus according to any one of Aspects 41 to 51.

[0223]

[0240] Aspect 53: One or more processors include a classification engine that identifies at least a first object image region and a second object image region, where the classification engine is at least partially located on an integrated circuit chip, the apparatus according to any one of Aspects 41 to 52.

[0224]

[0241] Aspect 54: The apparatus according to any one of Aspects 41 to 53, which is one of a mobile device, a wireless communication device, and a camera.

[0225]

[0242] Aspect 55: The apparatus according to any one of Aspects 41 to 54, further comprising an image sensor.

[0226]

[0243] Aspect 56: The apparatus according to any one of Aspects 41 to 55, further comprising a display for displaying an image.

[0227]

[0244] Aspect 57: A method of image processing, the method comprising receiving image data captured by an image sensor; determining that a first object image region in the image data depicts a first object category among a plurality of object categories; determining that a second object image region in the image data depicts a second object category among the plurality of object categories; identifying a plurality of confidence levels corresponding to a plurality of reliable image regions of the image data, wherein each confidence level of the plurality of confidence levels identifies the confidence that the corresponding reliable image region among the plurality of reliable image regions depicts one of the plurality of object categories; generating an image based on the image data using an image capture process, including by applying different setting values of the image capture process to different portions of the image data, wherein the different portions of the image data are identified based on the first object image region, the second object image region, and the plurality of reliable image regions.

[0228]

[0245] Aspect 58: A method further comprising generating one or more modifiers, the one or more modifiers identifying at least one of a first deviation from a default setting value of an image capture process for the first object image region and a second deviation from a default setting value of the image capture process for the second object image region, wherein the different setting values of the image capture process are based on the one or more modifiers, the method according to aspect 57.

[0229]

[0246] Aspect 59: A method further comprising adjusting one or more modifiers, including blending one or more modifiers with a blending update value based on a plurality of confidence levels corresponding to a plurality of reliable image regions, wherein blending one or more modifiers with the blending update value adjusts at least one of a first deviation and a second deviation in at least one area of the image data, the method according to aspect 57 or 58.

[0230]

[0247] Aspect 60: Generating a category map that divides the image data into a plurality of object image regions including a first object image region and a second object image region, wherein each object image region of the plurality of object image regions corresponds to one of a plurality of object categories, identifying that the first object category corresponds to a first set value of the image capture process, and identifying that the second object category corresponds to a second set value of the image capture process, the method according to any one of aspects 57 to 59.

[0231]

[0248] Aspect 61: A method further comprising generating a confidence map that divides the image data into a plurality of reliable image regions corresponding to a plurality of confidence levels, wherein different portions of the image data are identified based on the category map and the confidence map, the method according to any one of aspects 57 to 60.

[0232]

[0249] Aspect 62: The image capture process includes processing the image data using an image signal processor (ISP) among one or more processors, wherein different set values of the image capture process are different tuning set values of the ISP, the method according to any one of aspects 57 to 61.

[0233]

[0250] Aspect 63: The method according to any one of Aspects 57 to 62, wherein different setting values include setting values related to at least one of lens position, flash, focus, exposure, white balance, aperture size, shutter speed, ISO, analog gain, digital gain, noise reduction, sharpening, tone mapping, color saturation, demosaicing, color space conversion, shading, edge enhancement, high dynamic range (HDR) image composition, special effects, particle addition, artificial noise addition, edge-directed upscaling, upscaling, downscaling, and electronic image stabilization.

[0234]

[0251] Aspect 64: The method according to any one of Aspects 57 to 63, further comprising processing the image data, including at least one of demosaicing the image data and converting the image data from a first color space to a second color space.

[0235]

[0252] Aspect 65: A method further comprising receiving a user input related to at least one of a first object image region and a second object image region, wherein at least one of the different setting values is defined based on the user input and corresponds to one of the first object image region and the second object image region, the method according to any one of Aspects 57 to 64.

[0236]

[0253] Aspect 66: Applying different setting values of an image capture process to different parts of the image data includes applying different setting values of the image capture process to different parts of the image data using an image signal processor (ISP), the method according to any one of Aspects 57 to 65.

[0237]

[0254] Aspect 67: Identifying a first object image region and a second object image region includes identifying the first object image region and the second object image region using a classification engine that is at least partially located on an integrated circuit chip, the method according to any one of Aspects 57 to 66.

[0238]

[0255] Aspect 68: The method according to any one of Aspects 57 to 67, further comprising displaying an image on a display.

[0239]

[0256] Aspect 69: A method for image processing, comprising receiving image data captured by an image sensor, determining that a first object image region in the image data depicts a first object category among a plurality of object categories, determining that a second object image region in the image data depicts a second object category among the plurality of object categories, identifying a plurality of confidence levels corresponding to a plurality of reliable image regions of the image data, wherein each confidence level of the plurality of confidence levels identifies the confidence that the corresponding reliable image region among the plurality of reliable image regions depicts one of the plurality of object categories, generating an image based on the image data using an image capture process, including by applying different setting values of the image capture process to different portions of the image data, and the different portions of the image data are identified based on the first object image region, the second object image region, and the plurality of reliable image regions. A non-transitory computer-readable storage medium embodying a program executable by a processor to perform the method. The invention described in the claims of the present application at the time of filing is appended below. [C1] An apparatus for image processing, comprising: a memory; one or more processors coupled to the memory, wherein the one or more processors are configured to: receive image data captured by an image sensor; determine that a first object image region in the image data depicts a first object category among a plurality of object categories; determine that a second object image region in the image data depicts a second object category among the plurality of object categories; identify a plurality of confidence levels corresponding to a plurality of reliable image regions of the image data, wherein each confidence level of the plurality of confidence levels identifies a confidence that a corresponding reliable image region of the plurality of reliable image regions depicts one of the plurality of object categories; generate an image based on the image data using the image capture process, at least in part, by applying different setting values of the image capture process to different portions of the image data, wherein the different portions of the image data are identified based on the first object image region, the second object image region, and the plurality of reliable image regions; An apparatus configured to perform the above. [C2] The one or more processors are configured to: generate one or more modifiers, the one or more modifiers identifying at least one of a first deviation from a default setting value of the image capture process for the first object image region and a second deviation from the default setting value of the image capture process for the second object image region, wherein the different setting values of the image capture process are based on the one or more modifiers. The apparatus according to C1. [C3] The one or more processors are: The apparatus according to C2, wherein adjusting the one or more modifiers is configured to include blending the one or more modifiers with a blending update value based on the plurality of confidence levels corresponding to the plurality of reliable image regions, and blending the one or more modifiers with the blending update value adjusts at least one of the first deviation and the second deviation in at least one area of the image data. [C4] The one or more processors are configured to generate a category map that divides the image data into a plurality of object image regions including the first object image region and the second object image region, wherein each object image region of the plurality of object image regions corresponds to one of the plurality of object categories. identifying that the first object category corresponds to a first set value of the image capture process; identifying that the second object category corresponds to a second set value of the image capture process The apparatus according to C1, which is configured to perform. [C5] The one or more processors are configured to generate a confidence map that divides the image data into the plurality of reliable image regions corresponding to the plurality of confidence levels, and different portions of the image data are identified based on the category map and the confidence map. The apparatus according to C4. [C6] The image capture process includes processing the image data using an image signal processor (ISP) among the one or more processors, wherein the different set values of the image capture process are different tuning set values of the ISP. The apparatus according to C1. [C7] The different tuning set values of the ISP include different intensities at which ISP tuning parameters are applied during the processing of the image data using the ISP, wherein the ISP tuning parameter is one of noise reduction, sharpening, color saturation, color mapping, color processing, and tone mapping. The apparatus according to C6. [C8] The device according to C1, wherein the different set values include set values related to at least one of lens position, flash, focus, exposure, white balance, aperture size, shutter speed, ISO, analog gain, digital gain, noise removal, sharpening, tone mapping, color saturation, demosaicing, color space conversion, shading, edge enhancement, high dynamic range (HDR) image synthesis, special effects, artificial noise addition, edge-directed upscaling, upscaling, downscaling, and electronic image stabilization. [C9] The one or more processors are The device according to C1, configured to process the image data, including at least one of demosaicing the image data and converting the image data from a first color space to a second color space. [C10] The one or more processors are The device according to C1, configured to receive a user input related to at least one of the first object image region and the second object image region, wherein at least one of the different set values is defined based on the user input and corresponds to one of the first object image region and the second object image region. [C11] The device according to C1, wherein the one or more processors include an image signal processor (ISP) that applies the different set values of the image capture process to different portions of the image data. [C12] The device according to C1, wherein the one or more processors include a classification engine that identifies at least the first object image region and the second object image region, and wherein the classification engine is at least partially located on an integrated circuit chip. [C13] The device according to C1, which is one of a mobile device, a wireless communication device, and a camera. [C14] The device according to C1, further comprising the image sensor. [C15] The device according to C1, further comprising a display for displaying the image. [C16] An image processing method, comprising: Receiving image data captured by an image sensor; Determining that a first object image region in the image data depicts a first object category among a plurality of object categories; determining that a second object image region in the image data depicts a second object category among a plurality of object categories; identifying a plurality of confidence levels corresponding to a plurality of reliable image regions of the image data, wherein each confidence level of the plurality of confidence levels identifies a confidence that a corresponding reliable image region of the plurality of reliable image regions depicts one of the plurality of object categories; generating an image based on the image data using the image capture process, including by applying different setting values of the image capture process to different portions of the image data, wherein the different portions of the image data are identified based on the first object image region, the second object image region, and the plurality of reliable image regions; A method comprising: [C17] A method further comprising generating one or more modifiers, wherein the one or more modifiers identify at least one of a first deviation from a default setting value of the image capture process for the first object image region and a second deviation from the default setting value of the image capture process for the second object image region, wherein the different setting values of the image capture process are based on the one or more modifiers, the method according to C16. [C18] A method further comprising adjusting the one or more modifiers, including blending the one or more modifiers with a blending update value based on the plurality of confidence levels corresponding to the plurality of reliable image regions, wherein blending the one or more modifiers with the blending update value adjusts at least one of the first deviation and the second deviation in at least one area of the image data, the method according to C17. [C19] generating a category map that divides the image data into a plurality of object image regions including the first object image region and the second object image region, wherein each object image region of the plurality of object image regions corresponds to one of the plurality of object categories; Identifying that the first object category corresponds to a first set value of the image capture process; Identifying that the second object category corresponds to a second set value of the image capture process The method according to C16, further comprising. [C20] A method further comprising generating a confidence map that divides the image data into the plurality of confidence image regions corresponding to the plurality of confidence levels, wherein different portions of the image data are identified based on the category map and the confidence map, the method according to C19. [C21] The method according to C16, wherein the image capture process includes processing the image data using an image signal processor (ISP) among the one or more processors, wherein the different set values of the image capture process are different tuning set values of the ISP. [C22] The different tuning set values of the ISP include different intensities to which ISP tuning parameters are applied during processing of the image data using the ISP, wherein the ISP tuning parameter is one of noise reduction, sharpening, color saturation, color mapping, color processing, and tone mapping, the method according to C21. [C23] The different set values include set values related to at least one of lens position, flash, focus, exposure, white balance, aperture size, shutter speed, ISO, analog gain, digital gain, noise removal, sharpening, tone mapping, color saturation, demosaicing, color space conversion, shading, edge enhancement, high dynamic range (HDR) image synthesis, special effects, artificial noise addition, edge-directed upscaling, upscaling, downscaling, and electronic image stabilization, the method according to C16. [C24] The method according to C16, further comprising processing the image data, including at least one of demosaicing the image data and converting the image data from a first color space to a second color space. [C25] A method further comprising receiving user input related to at least one of the first object image region and the second object image region, wherein at least one of the different setting values is defined based on the user input and corresponds to one of the first object image region and the second object image region, the method according to C16. [C26] Applying the different setting values of the image capture process to the different portions of the image data includes applying the different setting values of the image capture process to the different portions of the image data using an image signal processor (ISP), the method according to C16. [C27] Identifying the first object image region and the second object image region includes identifying the first object image region and the second object image region using a classification engine that is at least partially located on an integrated circuit chip, the method according to C16. [C28] The method according to C16, further comprising displaying the image on a display. [C29] A non-transitory computer-readable storage medium embodying a program executable by a processor to perform a method of image processing, the method comprising: Receiving image data captured by an image sensor; Determining that a first object image region in the image data depicts a first object category among a plurality of object categories; Determining that a second object image region in the image data depicts a second object category among a plurality of object categories; Identifying a plurality of confidence levels corresponding to a plurality of reliable image regions of the image data, wherein each confidence level of the plurality of confidence levels identifies the confidence that the corresponding reliable image region of the plurality of reliable image regions depicts one of the plurality of object categories. Generating an image based on the image data using the image capture process, including by applying different setting values of the image capture process to different portions of the image data, wherein the different portions of the image data are identified based on the first object image region, the second object image region, and the plurality of reliable image regions. A non-transitory computer-readable storage medium comprising.

Claims

Claim 1 An apparatus for image processing, comprising a memory, and one or more processors coupled to the memory, wherein the one or more processors are configured to receive image data captured by an image sensor, determine that a first object image region in the image data depicts a first object category among a plurality of object categories, determine that a second object image region in the image data depicts a second object category among the plurality of object categories, generate a category map that divides the image data into a plurality of object image regions including the first object image region and the second object image region, wherein each object image region of the plurality of object image regions corresponds to one of the plurality of object categories, generate a confidence map that divides the image data into a plurality of reliable image regions, and identify a plurality of confidence levels corresponding to the plurality of reliable image regions of the image data, wherein each confidence level of the plurality of confidence levels identifies a confidence associated with a category classification in the category map that the corresponding reliable image region of the plurality of reliable image regions depicts one of the plurality of object categories, generate a plurality of modifiers based on the category map and the confidence map, wherein the plurality of modifiers identify a first deviation from a default setting value of an image signal processor (ISP) tuning parameter for the first object image region and a second deviation from the default setting value of the ISP tuning parameter for the second object image region Generating an image based on the image data using an image capture process by applying different set values of the ISP tuning parameters to at least partially different portions of the image data, wherein the different set values of the ISP tuning parameters are based on the plurality of modifiers, and the different portions of the image data are identified based on the first object image region, the second object image region, and the plurality of reliable image regions. An apparatus configured to perform the above. **Claim 2** The one or more processors are configured to adjust the one or more modifiers, including blending the one or more modifiers with a blending update value based on the plurality of reliability levels corresponding to the plurality of reliable image regions, wherein blending the one or more modifiers with the blending update value adjusts at least one of the first deviation and the second deviation in at least one area of the image data. The apparatus according to claim 1. **Claim 3** The image capture process includes processing the image data using an image signal processor (ISP) among the one or more processors, wherein the different set values of the image capture process are different tuning set values of the ISP. The apparatus according to claim 1. **Claim 4** The different tuning set values of the ISP include different intensities at which ISP tuning parameters are applied during processing of the image data using the ISP, wherein the ISP tuning parameter is one of noise reduction, sharpening, color saturation, color mapping, color processing, and tone mapping. The apparatus according to claim 3. **Claim 5** The apparatus according to claim 1, wherein the different set values include set values related to at least one of lens position, flash, focus, exposure, white balance, aperture size, shutter speed, ISO, analog gain, digital gain, noise reduction, sharpening, tone mapping, color saturation, demosaicing, color space conversion, shading, edge enhancement, high dynamic range (HDR) image synthesis, special effects, artificial noise addition, edge-directed upscaling, upscaling, downscaling, and electronic image stabilization.

6. The one or more processors are configured to process the image data, including at least one of demosaicing the image data and converting the image data from a first color space to a second color space, the apparatus according to claim 1.

7. The one or more processors are configured to receive a user input related to at least one of the first object image region and the second object image region, wherein at least one of the different set values is defined based on the user input and corresponds to one of the first object image region and the second object image region, the apparatus according to claim 1.

8. The apparatus according to claim 1, wherein the one or more processors include an image signal processor (ISP) that applies the different set values of the image capture process to different portions of the image data.

9. The apparatus according to claim 1, wherein the one or more processors include a classification engine that identifies at least the first object image region and the second object image region, and wherein the classification engine is at least partially located on an integrated circuit chip.

10. The apparatus according to claim 1, which is one of a mobile device, a wireless communication device, and a camera.

11. The apparatus according to claim 1, further comprising the image sensor.

12. The apparatus according to claim 1, further comprising a display for displaying the image.

13. An image processing method, comprising: receiving image data captured by an image sensor, Determining that the first object image region in the image data depicts a first object category among a plurality of object categories; Determining that the second object image region in the image data depicts a second object category among a plurality of object categories; Generating a category map that divides the image data into a plurality of object image regions including the first object image region and the second object image region, wherein each object image region of the plurality of object image regions corresponds to one of the plurality of object categories; Generating a reliability map that divides the image data into a plurality of reliable image regions, and identifying a plurality of reliability levels corresponding to the plurality of reliable image regions of the image data, wherein each reliability level of the plurality of reliability levels identifies a reliability associated with a category classification in the category map, wherein the corresponding reliable image region of the plurality of reliable image regions depicts one of the plurality of object categories; Generating a plurality of modifiers based on the category map and the reliability map, wherein the plurality of modifiers identify a first deviation from a default setting value of an image signal processor (ISP) tuning parameter for the first object image region and a second deviation from the default setting value of the ISP tuning parameter for the second object image region; Generating an image based on the image data using the ISP tuning parameters, including applying different setting values of the ISP tuning parameters to different portions of the image data, wherein the different setting values of the ISP tuning parameters are based on the plurality of modifiers, and the different portions of the image data are identified based on the first object image region, the second object image region, and the plurality of reliable image regions; A method comprising. Claim 14 The method further comprises adjusting the one or more modifiers, including blending the one or more modifiers with a blending update value based on the plurality of confidence levels corresponding to the plurality of reliable image regions, wherein blending the one or more modifiers with the blending update value adjusts at least one of the first deviation and the second deviation in at least one area of the image data. The method according to claim 13.

15. A non-transitory computer-readable storage medium embodying a program executable by a processor for performing a method of image processing, the method comprising: Receiving image data captured by an image sensor; Determining that a first object image region in the image data depicts a first object category among a plurality of object categories; Determining that a second object image region in the image data depicts a second object category among a plurality of object categories; Generating a category map that divides the image data into a plurality of object image regions including the first object image region and the second object image region, wherein each object image region of the plurality of object image regions corresponds to one of the plurality of object categories; Generating a confidence map that divides the image data into a plurality of reliable image regions and identifying a plurality of confidence levels corresponding to the plurality of reliable image regions of the image data, wherein each confidence level of the plurality of confidence levels identifies a confidence level associated with a category classification in the category map in which the corresponding reliable image region of the plurality of reliable image regions depicts one of the plurality of object categories; Generating a plurality of modifiers based on the category map and the confidence map, wherein the plurality of modifiers identify a first deviation from a default setting value of an image signal processor (ISP) tuning parameter for the first object image region and a second deviation from the default setting value of the ISP tuning parameter for the second object image region. Generating an image based on the image data using the ISP tuning parameter, including applying different setting values of the ISP tuning parameter to different portions of the image data, wherein the different setting values of the ISP tuning parameter are based on the plurality of modifiers, and the different portions of the image data are identified based on the first object image region, the second object image region, and the plurality of confidence image regions. A non-transitory computer-readable storage medium comprising the above.

Citation Information

Patent Citations

  • Image processing method and device

    JP2004062604A

  • Image processing apparatus, image processing method, and program

    JP2015069361A

  • Digital image processing system and method for emphasizing a main subject of an image

    US20060098889A1

  • Class-based image enhancement system

    US20080317358A1

  • Electronic device for adjusting image including multiple objects and control method thereof

    US20200053293A1