Content based image processing

The image processor uses content-based coefficients to sharpen image segments, addressing inefficient resource use and improving image quality by applying differential processing.

JP2025111543AActive Publication Date: 2025-07-30APPLE INC
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
JP2025067362
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Priority Date
2020-04-13
Filing Date
2025-04-16
Publication Date
2025-07-30
Estimated Expiration
2041-03-19

AI Technical Summary

Technical Problem

Existing image processing pipelines apply uniform adjustment parameters across entire images, adversely affecting the appearance of segments with different content, leading to inefficient resource consumption and increased power usage.

Method used

An image processor with a content image processing circuit that determines content coefficients based on texture and chroma values, using a content map generated by a neural processor, to apply differential sharpening to image segments.

Benefits of technology

Enhances image quality by selectively sharpening segments based on content, reducing resource consumption and power usage while maintaining image integrity.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2025111543000001_ABST
    Figure 2025111543000001_ABST
Patent Text Reader

Abstract

To provide a device and a method for performing sharpening operation on segments of image data based upon content related to the segments of an image.SOLUTION: There is provided a content based sharpening method performed by a content image processing circuit that receives luminance values of an image and a content map, wherein the content map identifies categories of content in segments of the image, and a content factor associated with a pixel is determined based on one or more of the identified categories of content. The content factor may also be based on a texture and / or chroma values. A texture value indicates a likelihood of a category of content and is based on detected edges in the image. A chroma value indicates a likelihood of a category of content and is based on color information of the image. The method also receives the content factor and applies it to a version of the luminance value of the pixel to generate a sharpened version of the luminance value.SELECTED DRAWING: Figure 12
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to a circuit for processing images, and more particularly to differentially sharpening segments of an image based on the content within the image.

Background Art

[0002] Image data captured by an image sensor or received from other data sources is often processed in an image processing pipeline before further processing or consumption. For example, RAW image data may be corrected, filtered, or modified before being provided to subsequent components such as a video encoder. Various components, unit stages, or modules may be used to correct or enhance the captured image data.

[0003] Such an image processing pipeline can be configured to conveniently perform corrections or enhancements to the captured image data without consuming other system resources. Many image processing algorithms can be executed by running a software program on a Central Processing Unit (CPU), but the execution of such a program on the CPU significantly consumes the CPU's bandwidth and other peripheral device resources and also increases power consumption. Thus, the image processing pipeline is often implemented as a dedicated hardware component separate from the CPU for executing one or more image processing algorithms.

[0004] The image processing pipeline often includes a sharpening process or a smoothing process. These processes are implemented using one or more adjustment parameters that are applied uniformly across the entire image. For this reason, one or more segments of the content within the image may be smoothed or sharpened in a way that adversely affects the appearance of the final image.

[0005] Similarly, other image processing processes such as tone mapping, white balance, and noise reduction are typically performed using one or more adjustment parameters that are applied uniformly across the entire image. As a result, one or more segments of the content within the image can be processed in a way that adversely affects the appearance of the final image. SUMMARY OF THE INVENTION

[0006] Some embodiments relate to an image processor that includes a content image processing circuit that determines content coefficients using a content map that identifies categories of content within segments of an image. The content image processing circuit includes a content coefficient circuit and a content modification circuit coupled to the content coefficient circuit. The content coefficient circuit determines content coefficients associated with pixels of the image according to the identified categories of content within the image and at least one of a texture value of a pixel within the image or a chroma value of a pixel within the image. The texture value indicates a likelihood of a category of content associated with the pixel based on the texture within the image. The chroma value indicates a likelihood of a category of content associated with the pixel based on the color information of the image. The content modification circuit receives the content coefficients from the content coefficient circuit. The content modification circuit generates a sharpened version of the luminance pixel value of the pixel by applying at least the content coefficients to a version of the luminance pixel value of the pixel.

[0007] In some embodiments, the content map is generated by a neural processor circuit that performs a machine learning operation on a version of the image to generate the content map.

[0008] In some embodiments, the content map is downscaled with respect to the image. The content coefficient circuit may determine the content coefficient by upsampling the content map. The content map may be upsampled by (1) obtaining content coefficients associated with lattice points in the content map and (2) interpolating content coefficients associated with lattice points surrounding a pixel when the content map is enlarged until it matches the size of the image.

[0009] In some embodiments, the content coefficients are weighted according to likelihood values. The likelihood values are based on one of an identified category of content, a texture value, and a chroma value in the image.

[0010] In some embodiments, when the luminance version of the image is divided into first information and second information including a frequency component lower than the frequency component of the first component, the luminance pixel value is included in the first information of the image.

[0011] In some embodiments, the image processor includes a bilateral filter coupled to the content image processing circuit. The bilateral filter generates a version of the luminance pixel value.

[0012] In some embodiments, the content modification circuit applies the content coefficient to the version of the luminance pixel value by multiplying the content coefficient by the version of the luminance pixel value when the content coefficient exceeds a threshold. The content modification circuit applies the content coefficient to the version of the luminance pixel value by blending the version of the luminance pixel value based on the content coefficient in response to the content coefficient being less than the threshold.

[0013] In some embodiments, the content map is a heat map indicating the amount of sharpening applied to the pixels of the image corresponding to the lattice points in the content map.

Brief Description of the Drawings

[0014]

Figure 1

Figure 2

Figure 3

Figure 4

Figure 5

Figure 6A

Figure 6B

Figure 7

Figure 8

Figure 9

Figure 10

Figure 11

Figure 12

[0015] For the purpose of illustration only, various non-limiting embodiments are shown in the figures and will be described in the detailed description.

Embodiments for Carrying Out the Invention

[0016] Reference will now be made in detail to embodiments shown in the accompanying drawings. In the following detailed description, numerous specific details are set forth in order to provide a thorough understanding of the various embodiments described. However, the embodiments to be described can be practiced without these specific details. In other instances, well-known methods, procedures, components, circuits, and networks have not been described in detail so as not to unnecessarily obscure aspects of the embodiments.

[0017] Embodiments of the present disclosure relate to sharpening segments of an image based on content within a segment indicated by a content map. A content coefficient for a pixel or segment of the image is determined based on one or more of the identified categories of content associated with the pixel or segment. The content coefficient can also be adjusted based on a set of texture values and / or a set of chroma values. The texture values indicate the likelihood of one of the identified categories of content and are based on the texture within the image. The chroma values indicate the likelihood of one of the identified categories of content and are based on the color information of the image. The content coefficient is applied to the pixel or segment to generate a sharpened version of the luminance value.

[0018] Exemplary Electronic Device Embodiments of electronic devices, user interfaces for such devices, and processes related to the use of such devices are described. In some embodiments, the device is a portable communication device such as a mobile phone that also includes other functions such as a personal digital assistant (PDA) function and / or a music player function. Exemplary embodiments of portable multifunctional devices include, but are not limited to, the iPhone®, iPod Touch®, Apple Watch®, and iPad® devices from Apple Inc. of Cupertino, California. Other portable electronic devices such as wearable, laptop, or tablet computers may optionally be used. In some embodiments, the device is not a portable communication device but a desktop computer or other computing device that is not designed for portable use. In some embodiments, the disclosed electronic device may include a touch-sensitive surface (e.g., a touch screen display and / or a touch pad). The exemplary electronic device (e.g., device 100) described below in relation to FIG. 1 may include a touch-sensitive surface for receiving user input. The electronic device may also include one or more other physical user interface devices such as a physical keyboard, a mouse, and / or a joystick.

[0019] FIG. 1 is a schematic diagram of an electronic device 100 according to an embodiment. The device 100 may include one or more physical buttons, such as a "home" button or a menu button 104. The menu button 104 is used, for example, to navigate to any application within a set of applications executed on the device 100. In some embodiments, the menu button 104 includes a fingerprint sensor that identifies a fingerprint on the menu button 104. The fingerprint sensor may be used to determine whether a finger on the menu button 104 has a fingerprint that matches a fingerprint stored to unlock the device 100. Alternatively, in some embodiments, the menu button 104 is implemented as a soft key within a graphical user interface (GUI) displayed on the touch screen.

[0020] In some embodiments, device 100 includes a touch screen 150, a menu button 104, a push button 106 for turning the device on / off and locking the device, volume adjustment buttons 108, a Subscriber Identity Module (SIM) card slot 110, a headset jack 112, and a docking / charging external port 124. The push button 106 can be used to turn the device on / off by pressing the button and maintaining the pressed state for a predetermined time interval, lock the device by pressing the button and releasing the button before the predetermined time interval has elapsed, and / or unlock the device or initiate an unlock process. In an alternative embodiment, device 100 also accepts verbal input via a microphone 113 to activate or deactivate some functions. Device 100 includes various components including, but not limited to, a memory (which can include one or more computer-readable storage media), a memory controller, one or more central processing units (CPUs), a peripheral device interface, an RF circuit, an audio circuit, a speaker 111, a microphone 113, an Input / Output (I / O) subsystem, and other input or control devices. Device 100 can include one or more image sensors 164, one or more proximity sensors 166, and one or more accelerometers 168. Device 100 can include two or more types of image sensors 164. Each type may include two or more image sensors 164. For example, one type of image sensor 164 can be a camera, and another type of image sensor 164 can be an infrared sensor that can be used for face recognition. Additionally or alternatively, the image sensors 164 can be associated with different lens configurations. For example, device 100 may include rear image sensors, one having a wide-angle lens and the other having a telephoto lens. Device 100 can include components not shown in FIG. 1, such as an ambient light sensor, a dot projector, and a projection illuminator.

[0021] Device 100 is merely one example of an electronic device, and device 100 may have more or fewer components than those listed above, and some of the components may be combined into one component or have another configuration or arrangement. The various components of device 100 listed above may be embodied in hardware, software, firmware, or a combination thereof, including one or more signal processing circuits and / or application specific integrated circuits (ASICs). While the components in FIG. 1 are generally shown as being located on the same side as touchscreen 150, one or more components may be located on the opposite side of device 100. For example, the front of device 100 may include an infrared image sensor 164 for facial recognition and another image sensor 164 as the device's front camera. The back of device 100 may also include two additional image sensors 164 as the device's rear cameras.

[0022] FIG. 2 is a block diagram illustrating components of device 100, according to one embodiment. Device 100 may perform various operations, including image processing. For this and other purposes, device 100 may include, among other components, an image sensor 202, a system-on-a-chip (SOC) component 204, system memory 230, persistent storage (e.g., flash memory) 228, an orientation sensor 234, and a display 216. The components shown in FIG. 2 are merely exemplary. For example, device 100 may include other components (such as a speaker or microphone) not shown in FIG. 2. Additionally, some components (such as orientation sensor 234) may be omitted from device 100.

[0023] The image sensor 202 is a component for capturing image data. Each of the image sensors 202 is a component for capturing image data and can be embodied as, for example, a Complementary Metal-Oxide-Semiconductor (CMOS) active pixel sensor, a camera, a video camera, or other device. The image sensor 202 generates RAW image data that is transmitted to the SOC component 204 for further processing. In some embodiments, the image data processed by the SOC component 204 is displayed on the display 216, stored in the system memory 230, the persistent storage 228, or transmitted to a remote computing device via a network connection. The RAW image data generated by the image sensor 202 can be in a Bayer Color Filter Array (CFA) pattern (hereinafter also referred to as a "Bayer pattern"). The image sensor 202 may also include optical and mechanical components that assist the image sensing components (e.g., pixels) in capturing an image. The optical and mechanical components may include an aperture, a lens system, and an actuator that controls the focal length of the image sensor 202.

[0024] The motion sensor 234 is a component or set of components for sensing the motion of the device 100. The motion sensor 234 can generate a sensor signal indicating the orientation and / or acceleration of the device 100. The sensor signal is transmitted to the SOC component 204 for various operations such as turning on the device 100 or rotating an image displayed on the display 216.

[0025] The display 216 is a component for displaying an image generated by the SOC component 204. The display 216 may include, for example, a Liquid Crystal Display (LCD) device or an Organic Light Emitting Diode (OLED) device. Based on the data received from the SOC component 204, the display 116 can display various images such as a menu, selected operation parameters, an image captured by the image sensor 202 and processed by the SOC component 204, and / or other information received from a user interface (not shown) of the device 100.

[0026] The system memory 230 is a component for storing instructions to be executed by the SOC component 204 and data processed by the SOC component 204. The system memory 230 may be embodied as any type of memory including, for example, Dynamic Random Access Memory (DRAM), Synchronous DRAM (SDRAM), Double Data Rate (DDR, DDR2, DDR3, etc.) RAMBUS DRAM (RDRAM), Static RAM (SRAM), or a combination thereof. In some embodiments, the system memory 230 may store pixel data or other image data, or statistical data in various formats.

[0027] The persistent storage 228 is a component for storing data non-volatilely. The persistent storage 228 retains data even without power. The persistent storage 228 can be embodied as a read-only memory (ROM), flash memory, or other non-volatile random access memory device. The persistent storage 228 stores the operating system of the device 100 and various software applications. The persistent storage 228 can also store one or more machine learning models such as support vector machines (SVMs) like regression models, random forest models, kernel SVMs, convolutional neural networks (CNNs), recurrent neural networks (RNNs), autoencoders, and long short-term memory (LSTM) artificial neural networks (ANNs). The machine learning model can be an independent model that operates with the neural processor circuit 218 and various software applications or sensors of the device 100. The machine learning model can also be part of a software application. The machine learning model can perform various tasks such as face recognition, image classification, object classification, concept classification, and information classification, speech recognition, machine translation, voice recognition, voice command recognition, text recognition, text and context analysis, other natural language processing, prediction, and recommendation.

[0028] The various machine learning models stored in device 100 may be fully trained, not trained, or partially trained so as to enable device 100 to enhance or continue to train the machine learning model when device 100 is used. The operations of the machine learning model include various calculations used when training the model and determining results at runtime using the model. For example, in some cases, device 100 captures a user's face image and uses the image to continuously improve the machine learning model used to lock or unlock device 100.

[0029] The SOC component 204 is embodied as one or more integrated circuit (IC) chips and executes various data processing processes. The SOC component 204 can include, among other sub-components, an Image Signal Processor (ISP) 206, a Central Processor Unit (CPU) 208, a network interface 210, a motion sensor interface 212, a display controller 214, a Graphics Processor Unit (GPU) 220, a memory controller 222, a video encoder 224, a storage controller 226, a neural processor circuit 218, and a bus 232 that connects these sub-components. The SOC component 204 may include more or fewer sub-components than those shown in FIG. 2.

[0030] The ISP206 is a circuit that executes various stages of an image processing pipeline. In some embodiments, the ISP206 may receive RAW image data from the image sensor 202 and process the RAW image data into a form usable by other sub-components of the SOC component 204 or components of the device <100>. The ISP206 may execute various image operation operations, such as image conversion operations, horizontal and vertical scaling, color space conversion, and / or image stabilization conversion, as will be described in detail below with reference to Figure 3.

[0031] The CPU 208 may be implemented using any suitable instruction set architecture and may be configured to execute instructions defined by that instruction set architecture. The CPU 208 may be a general-purpose processor or an embedded processor that uses any of various ISAs, such as the x86, PowerPC, SPARC, RISC, ARM, or MIPS instruction set architectures (ISAs), or any other suitable ISA. Although a single CPU is shown in Figure 2, the SOC component 204 may include multiple CPUs. In a multi-processor system, although not necessarily so, each of the CPUs may typically implement the same ISA in common.

[0032] The Graphics Processing Unit (GPU) 220 is a graphics processing circuit for executing operations on graphic data. For example, the GPU 220 may render an object (e.g., one containing pixel data for an entire frame) that is to be displayed in a frame buffer. The GPU 220 may include one or more graphics processors that may execute graphic software or hardware acceleration of specific graphic operations to execute some or all of the graphic operations.

[0033] The neural processor circuit 218 is a programmable circuit that performs machine learning operations on the input data of the neural processor circuit 218. The machine learning operations may include different calculations for training a machine learning model and for performing estimations or predictions based on the trained machine learning model. The neural processor circuit 218 is a circuit that performs various machine learning operations based on calculations including multiplication, addition, and accumulation. Such calculations can be arranged to perform various types of tensor multiplications such as tensor products and convolutions of input data and kernel data, for example. The neural processor circuit 218 is a configurable circuit that efficiently executes these operations at high speed while freeing the CPU 208 from resource-intensive operations related to neural network operations. The neural processor circuit 218 can receive input data from other sources such as the sensor interface 212, the image signal processor 206, the persistent storage 228, the system memory 230, or the network interface 210 or GPU 220. The output of the neural processor circuit 218 can be provided to various components of the device 100, such as the image signal processor 206, the system memory 230, or the CPU 208, for various operations.

[0034] The network interface 210 is a sub-component that enables the exchange of data between the device 100 and other devices (e.g., carrier devices or agent devices) via one or more networks. For example, video or other image data can be received from other devices via the network interface 210 and stored in the system memory 230 for subsequent processing and display (e.g., via a backend interface to the image signal processor 206 as described later with respect to FIG. 3). The network can include, but is not limited to, a Local Area Network (LAN) (e.g., Ethernet or a corporate network) and a Wide Area Network (WAN). The image data received via the network interface 210 can be subjected to an image processing process by the ISP 206.

[0035] The motion sensor interface 212 is a circuit for interfacing with the motion sensor 234. The motion sensor interface 212 receives sensor information from the motion sensor 234 and processes this sensor information to determine the orientation or motion of the device 100.

[0036] The display controller 214 is a circuit for transmitting the image data to be displayed on the display 216. The display controller 214 receives the image data from the ISP 206, the CPU 208, the graphics processor, or the system memory 230 and processes the image data into a format suitable for display on the display 216.

[0037] The memory controller 222 is a circuit for communicating with the system memory 230. The memory controller 222 can read data from the system memory 230 by the ISP 206, CPU 208, GPU 220, or other sub-components of the SOC component 204 and have the data processed by other sub-components. The memory controller 222 can also write data received from various sub-components of the SOC component 204 to the system memory 230.

[0038] The video encoder 224 is hardware, software, firmware, or a combination thereof for encoding video data into a format suitable for storage in the persistent storage 128 or for passing data to the network interface 210 for transmission to another device via the network.

[0039] In some embodiments, one or more sub-components of the SOC component 204 or some functionality of these sub-components can be executed by software components running on the neural processor circuit 218, ISP 206, CPU 208, or GPU 220. Such software components can be stored in another device that communicates with the device 100 via the system memory 230, persistent storage 228, or network interface 210.

[0040] Image data or video data can flow through various data paths within the SOC component 204. In one example, RAW image data is generated by the image sensor 202, processed by the ISP 206, and then transmitted to the system memory 230 via the bus 232 and the memory controller 222. After the image data is stored in the system memory 230, the image data can be accessed via the bus 232 by the video encoder 224 for encoding or by the display 116 for display.

[0041] In another embodiment, the image data is received from a source other than the image sensor 202. For example, video data can be streamed, downloaded, or communicated to the SOC component 204 via a wired or wireless network. The image data can be received via the network interface 210 and written to the system memory 230 via the memory controller 222. Next, the image data can be retrieved from the system memory 230 by the ISP 206 and processed through one or more image processing pipeline stages, as will be described in detail below with reference to FIG. 3. Next, the image data can be returned to the system memory 230 or sent to the storage controller 226 for storage in the video encoder 224, the display controller 214 (for display on the display 216), or the persistent storage 228.

[0042] Exemplary Image Signal Processing Pipeline FIG. 3 is a block diagram showing an image processing pipeline implemented using ISP206 according to one embodiment. In the embodiment of FIG. 3, ISP206 is coupled to an image sensor system 201 that includes one or more image sensors 202A-202N (hereinafter collectively referred to as "image sensor 202" or individually also referred to as "image sensor 202") for receiving RAW image data. The image sensor system 201 may include one or more subsystems for individually controlling the image sensors 202. In some cases, each image sensor 202 can operate independently, and in other cases, the image sensors 202 can share some components. For example, in one embodiment, two or more image sensors 202 may share the same circuit board that controls the mechanical components of the image sensors (e.g., an actuator that changes the focal length of each image sensor). The image sensing components of the image sensor 202 may include different types of image sensing components that can provide RAW image data to ISP206 in different forms. For example, in one embodiment, the image sensing components may include a plurality of focus pixels used for autofocus and a plurality of image pixels used for capturing an image. In another embodiment, the image sensing pixels may be used for both the purpose of autofocus and the purpose of image capture.

[0043] ISP206 implements an image processing pipeline that may include a set of stages for processing image information from creation, capture, or reception to output. ISP206 includes, among other components, a sensor interface 302, a central control module 320, a front-end pipeline stage 330, a back-end pipeline stage 340, an image statistics module 304, a scaler 322, a back-end interface 342, an output interface 316, and autofocus circuits 350A-350N (hereinafter collectively referred to as "autofocus circuits 350" or individually also referred to as "autofocus circuits 350"). ISP206 can include other components not shown in FIG. 3 or can omit one or more of the components shown in FIG. 3.

[0044] In one or more embodiments, different components of ISP206 process image data at different rates. In the embodiment of FIG. 3, front-end pipeline stages 330 (e.g., RAW processing stage 306 and resampling processing stage 308) can process image data at an initial rate. Thus, various different techniques, adjustments, corrections, or other processing operations are performed at the initial rate by these front-end pipeline stages 330. For example, if the front-end pipeline stage 330 processes two pixels per clock cycle, operations of the RAW processing stage 306 (e.g., black level compensation, highlight recovery, and defective pixel correction) can process two pixels of the image data simultaneously. In contrast, one or more back-end pipeline stages 340 can process image data at a different rate that is less than the initial data rate. For example, in the embodiment of FIG. 3, the back-end pipeline stages 340 (e.g., noise processing stage 310, color processing stage 312, and output rescaling module 314) can be processed at a reduced rate (e.g., one pixel per clock cycle).

[0045] The RAW image data captured by the image sensor 202 can be sent to different components of the ISP206 in different ways. In one embodiment, the RAW image data corresponding to the focus pixels can be sent to the autofocus circuit 350, while the RAW image data corresponding to the image pixels can be sent to the sensor interface 302. In another embodiment, the RAW image data corresponding to both types of pixels can be sent simultaneously to both the autofocus circuit 350 and the sensor interface 302.

[0046] The autofocus circuit 350 may include a hardware circuit that analyzes RAW image data to determine the appropriate focal length of each image sensor 202. In one embodiment, the RAW image data may include data transmitted from image sensing pixels specialized for image focusing. In another embodiment, RAW image data from the image capture pixels may also be used for the purpose of autofocus. The autofocus circuit 350 may perform various image processing operations to generate data for determining the appropriate focal length. The image processing operations may include cropping, binning, image compensation, and scaling to generate data used for the purpose of autofocus. The autofocus data generated by the autofocus circuit 350 may be fed back to the image sensor system 201 to control the focal length of the image sensor 202. For example, the image sensor 202 may include a control circuit that analyzes the autofocus data to determine a command signal transmitted to an actuator associated with the lens system of the image sensor to change the focal length of the image sensor. The data generated by the autofocus circuit 350 may also be transmitted to other components of the ISP 206 for other image processing purposes. For example, a portion of the data may be transmitted to the image statistics module 304 to determine information regarding autoexposure.

[0047] The autofocus circuit 350 may be a separate individual circuit distinct from other components such as the image statistics module 304, the sensor interface 302, the front-end module 330, and the back-end module 340. Thereby, the ISP 206 can perform autofocus analysis independently from other image processing pipelines. For example, the ISP 206 can analyze the RAW image data from the image sensor 202A to adjust the focal length of the image sensor 202A using the autofocus circuit 350A while simultaneously performing downstream image processing on the image data from the image sensor 202B. In one embodiment, the number of autofocus circuits 350 may correspond to the number of image sensors 202. In other words, each image sensor 202 may have a corresponding autofocus circuit dedicated to autofocus of the image sensor 202. The device 100 can perform autofocus for different image sensors 202 even when one or more of the image sensors 202 are not actively being used. This enables seamless transition between two image sensors 102 when the device 100 switches from one image sensor 202 to another. For example, in one embodiment, the device 100 may include a wide-angle camera and a telephoto camera as a dual-back camera system for photo and image processing. The device 100 can display an image captured by one of the dual cameras and can switch between the two cameras sometimes. Since two or more autofocus circuits 350 can continuously provide autofocus data to the image sensor system 201, the displayed image can seamlessly transition from the image data captured by one image sensor 202 to the image data captured by another image sensor without waiting for the second image sensor 202 to adjust its focal length.

[0048] RAW image data captured by different image sensors 202 can also be sent to the sensor interface 302. The sensor interface 302 receives RAW image data from the image sensor 202 and processes the RAW image data into image data that can be processed by other stages in the pipeline. The sensor interface 302 can perform various preprocessing operations such as image cropping, binning, or scaling to reduce the image data size. In some embodiments, pixels are sent from the image sensor 202 to the sensor interface 302 in raster order (e.g., horizontally, row by row). Subsequent processes in the pipeline can also be executed in raster order, and the results can also be output in raster order. Although only one image sensor and one sensor interface 302 are shown in FIG. 3, if two or more image sensors are provided in the device 100, a corresponding number of sensor interfaces can be provided in the ISP 206 to process the RAW image data from each image sensor.

[0049] The front-end pipeline stage 330 processes the image data in the RAW color domain or the full-color domain. The front-end pipeline stage 330 can include, but is not limited to, the RAW processing stage 306 and the resampling processing stage 308. The RAW image data can be, for example, in the Bayer RAW format. In the Bayer RAW image format, pixel data having specific values for specific colors (not for all colors) is given to each pixel. In an image capture sensor, the image data is typically provided in a Bayer pattern. The RAW processing stage 306 can process the image data in the Bayer RAW format.

[0050] The operations performed by the RAW processing stage 306 include, but are not limited to, sensor linearization, black level compensation, fixed pattern noise reduction, defective pixel correction, RAW noise filtering, lens shading correction, white balance gain, and highlight recovery. Sensor linearization refers to mapping non-linear image data into a linear space for other processing. Black level compensation refers to independently providing digital gain, offset, and clip for each color component (e.g., Gr, R, B, Gb) of the image data. Fixed pattern noise reduction refers to removing offset-fixed pattern noise and gain-fixed pattern noise by subtracting a dark frame from the input image and multiplying different gains to the pixels. Defective pixel correction refers to detecting defective pixels and then replacing the defective pixel values. RAW noise filtering refers to reducing the noise of the image data by averaging adjacent pixels with similar luminance. Highlight recovery refers to estimating pixel values for pixels clipped (or nearly clipped) from other channels. Lens shading correction refers to applying per-pixel gain to compensate for the reduction in light intensity that is approximately proportional to the distance from the optical center of the lens. White balance gain refers to independently providing digital gain, offset, and clip for white balance for all color components (e.g., Gr, R, B, Gb in the Bayer format). The components of the ISP206 can convert RAW image data into full-color domain image data, and thus, the RAW processing stage 306 can process full-color domain image data in addition to or instead of RAW image data.

[0051] The resampling processing stage 308 performs various operations to convert, resample, or scale the image data received from the RAW processing stage 306. The operations performed by the resampling processing stage 308 may include, but are not limited to, a demosaicing operation, a per-pixel color correction operation, a gamma mapping operation, a color space conversion, and a downscaling or sub-band splitting. The demosaicing operation refers to converting or interpolating the missing color samples from the RAW image data (e.g., into a Bayer pattern) to output the image data into the full-color domain. The demosaicing operation may include directional low-pass filtering on the interpolated samples for obtaining full-color pixels. The per-pixel color correction operation refers to a process of performing color correction for each pixel using information regarding the relative noise standard deviation of each color channel to correct the color without amplifying the noise within the image data. Gamma mapping refers to converting the image data from the input image data values to the output data values to perform gamma correction. For the purpose of gamma mapping, a look-up table (or another structure that indexes pixel values to other values) (e.g., separate look-up tables for the color components of R, G, and B) for different color components or color channels of each pixel can be used. Color space conversion refers to converting the color space of the input image data to a different format. In one embodiment, the resampling processing stage 308 converts the RGG format to the YCbCr format for further processing. In one embodiment, the resampling processing stage 308 converts the RBD format to the RGB format for further processing.

[0052] The central control module 320 can control and coordinate the operations of other components within the ISP206. The central control module 320 monitors various operation parameters (such as clock cycles, memory wait times, quality of service, and logging status information) to control the start and stop of other components of the ISP206, updates or manages the control parameters of other components of the ISP206, and performs operations including but not limited to interfacing with the sensor interface 302. For example, the central control module 320 can update the programmable parameters of other components within the ISP206 while other components are in an idle state. After updating the programmable parameters, the central control module 320 can put these components of the ISP206 into an execution state to execute one or more operations or tasks. The central control module 320 can also command other components of the ISP206 to store image data (e.g., by writing to the system memory 230 in FIG. 2) before, during, or after the resampling processing stage 308. In this way, in addition to or instead of processing the image data output from the resampling processing stage 308 through the backend pipeline stage 340, full-resolution image data in the RAW color domain format or the full-color domain format can be stored.

[0053] The image statistics module 304 performs various operations to collect statistical information associated with the image data. The operations for collecting statistical information may include, but are not limited to, sensor linearization, replacement of patterned defective pixels, subsampling of RAW image data, detection and replacement of non-patterned defective pixels, black level compensation, lens shading correction, and inverse black level compensation. After performing one or more of such operations, 3A statistics (statistical information such as Auto White Balance (AWB), Auto Exposure (AE), histogram (e.g., 2D color or components), and any other image data information) can be collected or tracked. In some embodiments, the value of a particular pixel or an area of pixel values may be excluded from the collection of particular statistical data if the preceding operations identify clipped pixels. Although only one statistical module 304 is shown in FIG. 3, a plurality of image statistics modules may be included in the ISP 206. For example, each image sensor 202 may correspond to an individual image statistics unit 304. In such embodiments, each statistical module can be programmed by the central control module 320 to collect different information for the same or different image data.

[0054] The scaler 322 receives the image data and generates a downscaled version of the image. Thus, the scaler 322 can provide a reduced-resolution image to various components such as the neural processor circuit 218. Although the scaler 322 is coupled to the RAW processing stage 306 in FIG. 3, the scaler 322 may be coupled to receive the input image from other components of the image signal processor 206.

[0055] The back-end interface 342 receives image data from an image source other than the image sensor 102 and transfers it to other components of the ISP 206 for processing. For example, the image data can be received via a network connection and stored in the system memory 230. The back-end interface 342 retrieves the image data stored in the system memory 230 and provides it to the back-end pipeline stage 340 for processing. One of the many operations performed by the back-end interface 342 is to convert the retrieved image data into a format that can be utilized by the back-end processing stage 340. For example, the back-end interface 342 can convert image data formatted in RGB, YCbCr 4:2:0, or YCbCr 4:2:2 to the YCbCr 4:4:4 color format.

[0056] The back-end pipeline stage 340 processes the image data according to a specific full-color format (e.g., YCbCr 4:4:4 or RGB). In some embodiments, the components of the back-end pipeline stage 340 can convert the image data into a specific full-color format before further processing. The back-end pipeline stage 340 can include, among other stages, a noise processing stage 310 and a color processing stage 312. The back-end pipeline stage 340 can include other stages not shown in FIG. 3.

[0057] The noise processing stage 310 performs various operations to reduce noise in the image data. The operations performed by the noise processing stage 310 include, but are not limited to, color space conversion, gamma / de-gamma mapping, temporal filtering, noise filtering, luma sharpening, and chroma noise reduction. Color space conversion can convert the image data from one color space format to another (e.g., convert the RGB format to the YCbCr format). Gamma / de-gamma operations convert the image data from input image data values to output data values to perform gamma correction or inverse gamma correction. Temporal filtering filters noise using previously filtered image frames to reduce noise. For example, combine the pixel values of the current image frame with those of the previous image frame. Noise filtering can include, for example, spatial noise filtering. Luma sharpening can sharpen the luma value of the pixel data, while chroma suppression can attenuate chroma to gray (e.g., no color). In some embodiments, luma sharpening and chroma suppression can be performed simultaneously with spatial noise filtering. The degree of aggressiveness of the noise filtering may be determined differently for different regions of the image. Spatial noise filtering may be included as part of the temporal loop that performs the temporal filtering. For example, the previous image frame may be processed by the temporal filter and the spatial noise filter before being stored as a reference frame for the next image frame to be processed. In other embodiments, spatial noise filtering may not be included as part of the temporal loop for temporal filtering (e.g., the spatial noise filter may be applied to the image frame after it has been stored as the reference image frame, and thus the reference frame is not spatially filtered).

[0058] The color processing stage 312 can perform various operations associated with adjusting color information in the image data. The operations performed in the color processing stage 312 include, but are not limited to, local tone mapping, gain / offset / clip, color correction, 3D color lookup, gamma conversion, and color space conversion. Local tone mapping refers to spatially varying the local tone curve to provide additional control when rendering an image. For example, bilinear interpolation may be performed so that a flatly varying tone curve is generated across the image (which can be programmed by the central control module 320). In some embodiments, local tone mapping may also apply a spatially varying and luminance-varying color correction matrix that can be used, for example, to darken the blue in shadows in the image while making the sky bluer. Digital gain / offset / clip may be provided for each color channel or color component of the image data. Color correction can apply a color correction transformation matrix to the image data. 3D color lookup may utilize a 3D array of color component output values (e.g., R, G, B) to perform extended tone mapping, color space conversion, and other color conversions. Gamma conversion can be performed, for example, by mapping input image data values to output data values to perform gamma correction, tone mapping, or histogram matching. Color space conversion may be performed to convert the image data from one color space to another (e.g., from RGB to YCbCr). Other processing techniques may also be performed as part of the color processing stage 312 to perform other special image effects, including black and white conversion, sepia tone conversion, negative conversion, or solarization conversion.

[0059] The output resizing module 314 can resample, transform, and correct distortions on the fly while the ISP 206 is processing image data. The output resizing module 314 can calculate the fractional input coordinates of each pixel and use these fractional coordinates to interpolate the output pixels through a polyphase resampling filter. The fractional input coordinates can be generated from various possible transformations of the output coordinates, such as resizing or cropping of the image (e.g., through simple horizontal and vertical scaling transformations), rotation and shearing of the image (e.g., through non-separable matrix transformations), perspective warping (e.g., through additional depth transformations), and per-pixel perspective division applied separately to strips caused by changes in the image sensor during capture of the image data (e.g., due to a rolling shutter), as well as geometric distortion correction (e.g., by calculating the radial distance from the optical center to index an interpolated radial gain table and applying a radial perturbation to the coordinates caused by the radial distortion of the lens).

[0060] The output resizing module 314 can apply a conversion to the image data when the image data is processed by the output resizing module 314. The output resizing module 314 can include a horizontal scaling component and a vertical scaling component. The vertical portion of the design can implement a series of image data line buffers to hold the "support" required by the vertical filter. Since the ISP 206 can be a streaming device, only the lines of the image data within a sliding window consisting of finite-length lines can be available for the filter. When a line is discarded to make room for incoming lines, that line can become unavailable. The output resizing module 314 can statistically monitor the calculated input Y coordinates of the previous lines and use that to calculate an optimal set of lines to hold within the vertical support window. For each subsequent line, the output resizing module can automatically generate an estimate regarding the center of the vertical support window. In some embodiments, the output resizing module 314 can implement a table of piecewise perspective transforms encoded as a Digital Difference Analyzer (DDA) stepper that performs a pixel-by-pixel perspective transform between the input image data and the output image data to correct artifacts and motion caused by the movement of the sensor during the capture of the image frame. Output resizing can provide the image data to various other components of the device 100 via the output interface 316, as described above with respect to FIGS. 1 and 2.

[0061] In various embodiments, the functionality of components 302-350 may be executed in an order different from the order implied by the order of these functional units within the image processing pipeline shown in FIG. 3, or may be executed by functional components different from those shown in FIG. 3. Further, the various components as described in FIG. 3 can be embodied in various combinations of hardware, firmware, or software.

[0062] Exemplary pipeline related to a multiple band noise reduction circuit FIG. 4 is a block diagram showing a portion of an image processing pipeline including a multiple band noise reduction (MBNR) circuit 420 according to one embodiment. In the embodiment of FIG. 4, the MBNR circuit 420 is part of a resampling processing stage 308 that includes, among other components, a scaler 410 and a sub-band splitter circuit 430. The resampling processing stage 308 recursively performs scaling, noise reduction, and sub-band splitting.

[0063] As a result of the recursive processing, the resampling processing stage 308 outputs a series of high-frequency component image data HF(N) and low-frequency component image data LF(N) derived from the original input image 402. Here, N represents the level of downsampling performed on the original input image 402. For example, HF(0) and LF(0) represent high-frequency component image data and low-frequency component image data respectively, divided from the original input image 402, and HF(1) and LF(1) represent high-frequency component image data and low-frequency component image data respectively, divided from the first downscaled version of the input image 402.

[0064] The MBNR circuit 420 is a circuit that performs noise reduction on the multiple bands of the input image 402. The input image 402 is first passed to the MBNR circuit 420 via the multiplexer 414 for noise reduction. The noise reduction version 422 of the original input image 402 is generated by the MBNR circuit 420 and supplied to the sub-band splitter 430. The sub-band splitter 430 divides the noise reduction version 422 of the original input image 402 into high-frequency component image data HF(0) and low-frequency component image data LF(0). The high-frequency component image data HF(0) is passed to the sub-band processing pipeline 448 and then to the sub-band merger 352. In contrast, the low-frequency component image LF(0) passes through the demultiplexer 440 and is fed back to the resampling processing stage 308 for downscaling by the scaler 410.

[0065] The scaler 410 generates a downscaled version 412 of the low-frequency component image LF(0) supplied to the scaler 410 and passes it to the MBNR circuit 420 via the multiplexer 414 for noise reduction. The MBNR circuit 420 performs noise reduction to generate a noise reduction version 432 of the downscaled image 412, sends it to the sub-band splitter 430, and re-divides the processed low-frequency image data LF(0) into high-frequency component image data HF(1) and low-frequency component image data LF(1). The high-frequency component image data HF(1) is sent to the sub-band processing pipeline 448 and then to the sub-band merger 352, while the low-frequency component image data LF(1) is fed back to the scaler 410 again to repeat the process within the resampling processing stage 308. The process of generating the high-frequency component image data HF(N) and the low-frequency component image data LF(N) is repeated until the final level of band splitting by the sub-band splitter 430 is performed. When the final level of band splitting is reached, the low-frequency component image data LF(N) is passed to the sub-band processing pipeline 448 and the sub-band merger 352 via the demultiplexer 440 and the multiplexer 446.

[0066] As described above, the MBNR circuit 420 performs noise reduction on the input image 402 and the downscaled low-frequency version of the input image 402. This enables the MBNR circuit 420 to perform noise reduction on multiple bands of the original input image 402. However, it should be noted that without sub-band splitting and scaling, the noise reduction on the input image 402 that can be performed by the MBNR circuit 420 is only a single pass.

[0067] The sub-band merger 352 merges the processed high-frequency component image data HF(N)' and the processed low-frequency component image data LF(N)' to generate the processed LF(N - 1)'. The processed LF(N - 1)' is fed back to the sub-band merger 352 via the demultiplexer 450 and the multiplexer 446 to merge with the processed HF(N - 1)' in order to generate the processed LF(N - 2)'. The process of merging the processed high-frequency component image data and the processed low-frequency component data is repeated until the sub-band merger 352 generates a processed version 454 of the input image output via the demultiplexer 450.

[0068] The first contrast enhancement stage 450 and the second contrast enhancement stage 452 perform a sharpening operation or a smoothing operation on segments of the image data based on the content associated with the segments of the image. The first contrast enhancement stage 450 is a component of the sub-band processing pipeline 448 and performs a sharpening operation on the downscaled high-frequency component image data HF(N) with respect to the input image 402. On the other hand, the second contrast enhancement stage 452 performs sharpening on the output of the sub-band merger 352, which can be full-resolution image data having the same spatial size as the input image 402. The first and second contrast enhancement stages 450, 452 are further described with reference to FIG. 7.

[0069] Exemplary pipeline related to the content map FIG. 5 is a block diagram showing the provision of a content map 504 (also referred to herein as a "segmentation map") to the image signal processor 206 by the neural processor circuit 218 according to one embodiment. The image signal processor 206 provides an input image 502 to the neural processor circuit 218. The input image 502 may be the same or different compared to the original input image 402. Based on the input image 502, the neural processor circuit 218 generates a content map 504 and provides the content map 504 to the image signal processor 206. For example, the content map 504 is sent to the contrast enhancement stages 450, 452.

[0070] As previously described with reference to FIG. 2, the neural processor circuit 218 can be machine-learned. Thus, the neural processor circuit 218 can determine the content map 504 by performing one or more machine learning operations on the input image 502. In some embodiments, the input image 502 is a reduced-resolution image compared to the original input image 402 (e.g., provided by the scaler 322). By reducing the resolution of the image, the processing time for generating the content map 504 can be shortened. However, since the resolution of the content map 504 is typically the same or similar to the resolution of the input image 502, the content map 504 can be downscaled with respect to the original input image 402. An exemplary input image 602 is provided in FIG. 6A, and an exemplary content map 604 is provided in FIG. 6B. The content map 604 identifies the grass 506 and the person 508 within the image 602.

[0071] Each segment of the content can be associated with one or more predetermined categories of the content. The content map 504 can be associated with a grid having a plurality of grid points used during the upscaling process. The content map 504 may be the same size as the full-scale image. In some embodiments, to facilitate various processes, the number of grid points is less than the number of pixels in the full-scale image to be sharpened. In one or more embodiments, each grid point of the map can be associated with a category of content (also referred to herein as a "content category") and can be used to determine the content coefficients of nearby pixels in the full-scale image, as will be described in detail below with reference to FIG. 11.

[0072] Examples of content categories in the content map 504 have different desired content coefficients. Different content categories may include skin, leaves, grass, and sky. For other categories (e.g., skin), it is generally desirable to sharpen specific categories (e.g., leaves and grass). In one or more embodiments, the neural processor circuit 218 is trained using various machine learning algorithms to classify different segments of the input image 502, identify the content categories of the content in the input image 502, and then this is used by the image signal processor 206 to sharpen different segments of the input image 502 with different content coefficients indicating the degree of sharpening to be applied to the different segments, as will be described in detail below with reference to FIG. 8.

[0073] In other embodiments, the neural processor circuit 218 is trained using various machine learning algorithms to generate a heat map as the content map 504. The heat map directly indicates the desired degree of sharpening in different segments of the input image 502, rather than indicating the content categories associated with the different segments.

[0074] Example of a contrast enhancement stage circuit FIG. 7 is a block diagram showing components of a first contrast enhancement stage circuit 450 according to an embodiment. The first contrast enhancement stage circuit 450 performs content-based image sharpening and smoothing on luminance information Y to generate a sharpened version Y' of the luminance information. The luminance information Y refers to an image including only the luminance component of the input image 402, and the sharpened luminance information Y' refers to an image including only the luminance component of the output image. The first contrast enhancement stage circuit 450 may include, among other components, an image sharpener 702, content image processing 704, and an addition circuit 706. The second contrast enhancement stage circuit 452 has substantially the same structure as the first contrast enhancement stage circuit 450, except that the luminance image is not downscaled with respect to the full input image 402, and thus, for the sake of brevity, its detailed description is omitted herein.

[0075] The image sharpener 702 is a circuit that performs contrast enhancement (e.g., sharpening) on the luminance information Y and generates an output delta Y. The delta Y represents a mask of Y. For example, the delta Y is the result of an unsharp masking process. In one or more embodiments, the image sharpener 702 is implemented as a bilateral filter or a high-pass frequency filter that performs processing on the luminance information Y. Thus, for example, the delta Y can be a high-frequency component of the image. The delta Y is further adjusted by components downstream of the first contrast enhancement stage 450.

[0076] The content image processing 704 is a circuit that adjusts the delta Y based on the content category identified by the content map and the likelihood of such classification. The content image processing 704 receives the luminance information Y and the content map 504 and generates an adjusted delta Y' that is increased or decreased with respect to the delta Y according to the desired degree of sharpening based on the content category, as will be further described with respect to FIG. 8.

[0077] In some embodiments, the addition circuit 706 adds the adjusted delta Y’ from the content image processing 704 to the luminance information Y to generate sharpened luminance information Y’. In some embodiments, the addition circuit 706 adds the adjusted delta Y’ to the low-frequency component of the luminance information Y (for example, here, the low-frequency component = Y - delta Y). For some pixels, the adjusted delta Y’ is positive, whereby the addition in the addition circuit 706 results in sharpening of the relevant segment of the image. For pixels having a negative adjusted delta Y’, the addition circuit 706 performs a blurring operation such as alpha blurring. As will be further described below, delta Y can be alpha blurred with the low-frequency component of the luminance information Y. This can result in the blurring being limited to the low-frequency component. This can prevent image artifacts that may occur when delta Y’ includes large negative values.

[0078] <Example of content image processing circuit> FIG. 8 is a block diagram showing components of the content image processing 704 according to an embodiment. As described above with reference to FIG. 7, the content image processing 704 performs a sharpening operation and a smoothing operation based on the content in the image. The content image processing 704 may include, among other components, a texture circuit 802, a chroma circuit 804, a content coefficient circuit 806, and a content correction circuit 810.

[0079] The content coefficient circuit 806 determines a content coefficient for pixels in the input image. The content coefficient can be determined for each pixel in the input image. The content coefficient is based on one or more values in the content map and indicates the amount of sharpening to be applied to the pixel. The content coefficient can also be based on one or more texture values from the texture circuit 802 and / or chroma values from the chroma circuit 804.

[0080] As described above, the content categories of the content map can be associated with content coefficients. For example, a content coefficient is predetermined for each content category. When the content map has the same resolution as the image, the content coefficient for a pixel can be obtained by referring to the information at the corresponding position within the content map. When the content map is downscaled compared to the image, the content map can be upscaled to match the size of the input image so that the content coefficient can be interpolated from nearby pixels within the upscaled version of the content map. A grid having a plurality of grid points may be overlaid on the input image, and the information associated with the grid points can be used to determine the information of the pixels within the full image by interpolation. For example, when the pixel position of the full image does not coincide with the grid points within the content map (e.g., the pixel is located between a set of grid points), the content coefficient of the pixel can be determined by interpolating the content coefficients of the grid points closest to (e.g., surrounding) the pixel. The upsampling of the content map for determining the content coefficients of the pixels between the grid points is further described with respect to FIG. 11.

[0081] In some embodiments, the content coefficients are weighted according to likelihood values. For example, the content coefficient Q of a pixel is determined by multiplying an initial content coefficient Q0 by a likelihood value. Q = (Q0) * (likelihood value) (1) Here, the initial content coefficient Q0 is the content coefficient associated with a specific content category. The likelihood value can be based on the content category of the pixel, the texture value and / or chroma value of the pixel, as will be described below with reference to the texture circuit 802 and the chroma circuit 804. In some embodiments, the likelihood value is determined by a likelihood model such as the following. (likelihood value) = C1 + C2 * (texture value) + C3 * (chroma value) (2) Here, C1, C2, and C3 are predetermined constants (e.g., adjustment parameters). The predetermined constants may have values based on the content categories within the content map. In some embodiments, the likelihood model is a polynomial function with respect to texture values and chroma values. The likelihood model represents a model regarding the classification accuracy and can be determined empirically or by a machine learning process.

[0082] The texture circuit 802 is a circuit that determines a texture value representing a likelihood corrected based on texture information for the content category identified by the content map 504. In one or more embodiments, the texture value is determined by applying one or more edge detection operations to the input image. For example, an edge detection method such as a Sobel filter or a high - pass frequency filter is applied to the luminance input image to obtain edge values at the pixel positions of the input image. After the edge values are determined, the texture values for the lattice points can be determined by applying the edge values to the texture model. The texture circuit 802 may store a plurality of different texture models corresponding to different content categories. Examples of texture models include a leaf texture model, an empty texture model, a grass texture model, and a skin texture model. An exemplary leaf texture model will be described below with respect to FIG. 9.

[0083] The chroma circuit 804 is a circuit that determines a chroma value representing a likelihood corrected based on chroma information for the content category identified by the content map 504. The chroma value is based on the color information of the image (e.g., Cb value and Cr value). The chroma circuit 804 may store different chroma models for different content categories. Examples of chroma models include a leaf chroma model, an empty chroma model, a grass chroma model, and a skin chroma model. The chroma model may be determined manually or by machine learning techniques. An exemplary empty chroma model will be described below with respect to FIG. 10.

[0084] The content correction circuit 810 receives content coefficients from the content coefficient circuit 806 and delta Y values from the image sharpener 702. The content correction circuit 810 applies the content coefficients to the delta Y values to generate delta Y' values. For example, when the content coefficient of a pixel exceeds a predetermined threshold (e.g., 0), the content correction circuit 810 performs a sharpening operation such as multiplying the content coefficient of the pixel by the delta Y value of the pixel. When the content coefficient of a pixel is below the predetermined threshold, the content correction circuit 810 can perform a smoothing operation by blending delta Y' based on the content coefficient. For example, alpha blending is performed according to the following formula. Y’=(1 - alpha) * Y + alpha * (Y - delta Y) (3) And Y’ = Y + delta Y’ (4) Therefore delta Y’ = -(alpha) * (delta Y) (5). Here, alpha = |Q| * As a scale, alpha is a value between 0 and 1. The scale is a predetermined positive constant. |Q| * Note that when |Q| is large enough such that scale > 1, alpha is clipped to 1.

[0085] FIG. 9 is a plot showing a texture model according to an embodiment. When a pixel (or a lattice point if a lattice is used) is classified as a "leaf" by the content map, for example, the texture value can be determined by applying the edge value of the pixel (or the lattice point if a lattice is used) to the model of FIG. 9. The x-axis represents the input edge value, and the Y-axis represents the output texture value. The range of the texture value is set to 0 to 1. When the edge value of the pixel exceeds the high threshold, the texture value is set to 1, and when the edge value is below the low threshold, the texture value is set to 0. The texture value increases linearly from 0 to 1 with respect to the edge value between the low threshold and the high threshold. The values of the low threshold and the high threshold can be determined empirically. For example, the threshold is set according to the edge value typical for a leaf. Therefore, the fact that the edge value exceeds the high threshold may indicate that the pixel or the lattice point is in a region having a high texture corresponding to a leaf. Similarly, the fact that the edge value is below the low threshold may indicate that the pixel or the lattice point is in a region having a flat texture not corresponding to a leaf. For example, according to Equation 2, a high texture value indicates that the likelihood that the content of the pixel is a leaf is high (assuming that C2 is positive).

[0086] For different categories, the corresponding texture models can be represented by different texture parameters (e.g., different low thresholds, different high thresholds, and / or different gradients, and / or inversion of 1 and 0). For the "grass" category, the low threshold and the high threshold may be higher than the thresholds of the "leaf" category. Some categories may be more likely to be correct when the texture is flat than when the texture is complex. In such categories, the 1 and 0 of the texture values can be inverted. For example, the texture model of "sky" can have a value of 1 when the edge value is lower (e.g., below the low threshold), and can have a value of 0 when the edge value is higher (e.g., above the high threshold). That is, pixels in the region of the input image where the texture is flat are likely to indicate "sky". In some embodiments, instead of inverting the 1 and 0 values, a predetermined constant in Equation 2 changes based on the content category. For example, for categories such as "skin" or "sky", C2 can be negative such that a high texture value indicates a low likelihood that the pixel content is "skin" or "sky".

[0087] FIG. 10 is a plot showing a chroma model according to an embodiment. The chroma model of FIG. 10 may represent a "blank" model. When a pixel or lattice point is classified as "blank" by the content map, the chroma circuit 804 determines whether the combination of the Cb value and the Cr value for the pixel or lattice point falls within one of the ranges of areas 1010, 1020, and 1030. The x-axis represents the input Cb value, and the Y-axis represents the input Cr value. The ellipse is located at the upper right corner of the plot diagram representing the "blank" color range. When the Cb value / Cr value of a pixel or lattice point is within the inner ellipse in area 1010, the chroma value is 1, which may indicate that the pixel or lattice point has a high likelihood of corresponding to "blank" (e.g., according to Equation 2 where C3 is positive). When the Cb value / Cr value is outside the outer ellipse in area 1030, the chroma value is 0, which indicates that the pixel or lattice point has a low possibility of corresponding to "blank" (e.g., according to Equation 2 where C3 is positive). When the Cb value / Cr value is between the inner ellipse and the outer ellipse in area 1020, the chroma value can be between 0 and 1 (e.g., the chroma value increases as the distance from the edge of the inner ellipse increases). The position and size of the ellipse can be determined empirically or by a statistical model. In some embodiments, instead of an ellipse, other shapes such as a square, triangle, or circle can be used.

[0088] For different categories, the corresponding chroma models can be represented by different chroma parameters (e.g., the center, radius, angle, inclination, and ratio of the outer / inner ellipse of the ellipse). For example, the chroma model of "leaf" is usually green, so the chroma parameters of the "leaf" category cover the Cb value / Cr value corresponding to green, while the chroma parameters of the "blank" category cover the Cb value / Cr value corresponding to blue. In another example, the chroma parameters are generally selected according to colors not associated with the category. For example, "leaf" generally does not contain blue. Therefore, the chroma model can be configured to output a low value when the Cb value / Cr value corresponds to blue.

[0089] Example of interpolation using lattice points FIG. 11 is a diagram showing a method of upsampling a content map according to an embodiment. When the content map has a lower resolution than the input image, the pixels in the image and the pixels in the content map do not correspond one-to-one. Therefore, upscaling of the content map can be performed. One way to perform such upscaling is by using a grid having a plurality of grid points superimposed on the input image 1102. The grid points may be sparser than the pixels in the input image 1102. In such a case, the content coefficients of the input image 1102 with a higher resolution can be determined using the content coefficients of the grid points.

[0090] Taking an example where the grid superimposed on the input image 1102 has grid points 1 to 4 and the input image 1102 includes a pixel 1104 located between the grid points 1 to 4, the content coefficient Q(pixel) for the pixel 1104 can be determined by performing bilinear interpolation with respect to the content coefficients Q(1) to Q(4) associated with the grid points 1 to 4 in consideration of the spatial distance from the grid points to the pixel 1104.

[0091] In one or more embodiments, the content coefficients Q(1) to Q(4) for determining Q(pixel) of pixel 1104 are determined using the texture parameters and chroma parameters of pixel 1104 of input image 1102. The categories of grid points 1 to 4 indicated by the content map are used, but the likelihood values for these categories (described above with reference to Equation (2)) are determined using the edge value and Cb value / Cr value of pixel 1104, rather than the edge value and Cb value / Cr value of the grid point. For example, if grid point 1 is classified as "skin", the texture of pixel 1104 is applied to a texture model having texture parameters corresponding to "skin", and the Cb value / Cr value of pixel 1104 is applied to a chroma model having chroma parameters corresponding to "skin" in order to obtain Q(1) according to Equation (1). Similarly, if grid point 2 is classified as "leaf", the edge value of pixel 1104 is applied to a texture model having texture parameters corresponding to "leaf", and the Cb value / Cr value of pixel 1104 is applied to a chroma model having chroma parameters corresponding to "leaf" in order to obtain Q(2) according to Equation (1). After repeating the same process for grid points 3 and 4, Q(pixel) is obtained by bilinear interpolation.

[0092] In other embodiments, the content coefficients Q(1) to Q(4) are obtained by using the texture value and Cb value / Cr value of the grid point instead of the texture value and Cb value / Cr value of the pixel.

[0093] Exemplary Method of Image Sharpening Based on Content FIG. 12 is a flowchart showing a method of sharpening one or more pixels of an image based on the content within a segment of the image according to one embodiment. The steps of the method may be performed in a different order, and the method may include different, additional, or fewer steps.

[0094] Receive the luminance pixel values of the input image (1202). Receive a content map (1204). The content map identifies the category of content within segments of the image. The content map can be generated by a neural processor circuit that performs at least one machine learning operation on a version of the image to generate the content map. The categories identified by the content map may include skin, leaves, grass, or sky. In some embodiments, the content map is a heat map indicating the amount of sharpening to be applied to the pixels of the image.

[0095] Determine a content coefficient associated with a pixel in the image (1206). The content coefficient is determined according to at least one of the identified content category in the input image and the texture value of the pixel in the image or the chroma value of the pixel in the image. The texture value indicates the likelihood of the content category associated with the pixel based on the texture in the image. The chroma value indicates the likelihood of the content category associated with the pixel based on the color information of the image.

[0096] In some embodiments, the content map is downscaled with respect to the input image. In these embodiments, the content coefficient can be determined by upsampling the content map. The content map can be upsampled by obtaining the content coefficients of the grid points overlaid on the content map and then interpolating the content coefficients of the grid points to obtain the content coefficients of the pixels in the input image.

[0097] By applying at least a content coefficient to a version of the luminance pixel value of a pixel, a sharpened version of the luminance pixel value of the pixel is generated (1208). The version of the luminance pixel value can be generated by a bilateral filter or a high-pass filter. In some embodiments, the content coefficient is applied to the version of the luminance pixel value by multiplying the content coefficient by the version of the luminance pixel value in response to the content coefficient exceeding a threshold. If the content coefficient is below the threshold, the content correction circuit applies a version of the luminance pixel value multiplied by a negative parameter.

[0098] The teachings described herein relate to generating a content coefficient for each pixel in an image. The content coefficient is based on the content in the image identified via a content map. Although the teachings herein are described in the context of image sharpening, this is for convenience. The teachings described herein can also be applied to other image processing processes such as noise reduction, tone mapping, and white balance processes. For example, for noise reduction, the content coefficient can be applied (e.g., multiplied) to the noise standard deviation.

[0099] Although specific embodiments and applications have been illustrated and described, the present invention is not limited to the exact structures and components disclosed herein, and various modifications, changes, and variations that will be apparent to those skilled in the art may be made to the construction, operation, and details of the methods and apparatuses disclosed herein without departing from the spirit and scope of the disclosure.

Claims

1. An apparatus for image processing comprising a content image processing circuit configured to receive a luminance pixel value of an image and a content map identifying a category of content within a segment of the image, wherein the content image processing circuit is a content coefficient circuit configured to determine a content coefficient associated with a pixel of the image according to the identified category of the content in the image and at least one of a texture value of a pixel in the image or a chroma value of a pixel in the image, the texture value indicating a likelihood of a category of content associated with the pixel based on a texture within the image, and the chroma value indicating a likelihood of a category of content associated with the pixel based on color information of the image, the content coefficient circuit is a content correction circuit coupled to the content coefficient circuit for receiving the content coefficient and configured to generate a sharpened version of the luminance pixel value of the pixel by applying at least the content coefficient to a version of the luminance pixel value of the pixel including An apparatus for image processing.

2. The apparatus according to claim 1, wherein the content map is generated by a neural processor circuit configured to perform a machine learning operation on a version of the image to generate the content map.

3. The apparatus according to claim 1, wherein the content map is downscaled with respect to the image.

4. The apparatus according to claim 3, wherein the content coefficient circuit is configured to determine the content coefficient by upsampling the content map.

5. The content map when the content map is enlarged until it matches the size of the image, obtaining a content coefficient associated with a lattice point surrounding the pixel, and interpolating the content coefficient to obtain the content coefficient associated with the pixel is upsampled by the method according to claim 4.

6. The apparatus according to claim 1, wherein the content coefficient is weighted according to a likelihood value, and the likelihood value is based on one of the identified category of the content, texture value, and chroma value within the image.

7. The apparatus according to claim 1, wherein when the luminance pixel value is divided into first information of the image and second information including a frequency component lower than the frequency component of the first information, the luminance pixel value is included in the first information of the image.

8. The apparatus according to claim 1, further comprising a bilateral filter coupled to the content image processing circuit, wherein the bilateral filter is configured to generate the version of the luminance pixel value.

9. The apparatus according to claim 1, wherein the content correction circuit is configured to apply the content coefficient to the version of the luminance pixel value by (i) multiplying the content coefficient by the version of the luminance pixel value in response to the content coefficient exceeding a threshold, and (ii) blending the version luminance pixel value based on the content coefficient in response to the content coefficient being less than the threshold.

10. The apparatus according to claim 1, wherein the content map is a heat map indicating an amount of sharpness to be applied to a pixel of the image corresponding to a grid point in the content map.

11. Receiving a luminance pixel value of an image; Receiving a content map for identifying a category of content within a segment of the image; Determining a content coefficient associated with a pixel in the image according to at least one of the identified category of content in the image and a texture value of a pixel in the image or a chroma value of a pixel in the image, wherein the texture value indicates a likelihood of a category of content associated with the pixel based on a texture in the image, and the chroma value indicates a likelihood of a category of content associated with the pixel based on color information of the image, and generating a sharpened version of the luminance pixel value of the pixel by applying at least the content coefficient to a version of the luminance pixel value of the pixel. A method comprising.

12. The method according to claim 11, further comprising generating the content map by a neural processor circuit, wherein the neural processor circuit performs a machine learning operation on a version of the image to generate the content map.

13. The method according to claim 11, wherein the content map is downscaled with respect to the image.

14. The method according to claim 13, wherein determining the content coefficient includes upsampling the content map.

15. Upsampling the content map includes: when the content map is enlarged until it matches the size of the image, obtaining content coefficients associated with grid points surrounding the pixel; interpolating the content coefficients to obtain the content coefficients associated with the pixel; The method according to claim 14, comprising:

16. A neural processor circuit configured to generate a content map that identifies categories of content within segments of an image by executing a machine learning algorithm on the image; A content image processing circuit coupled to the neural processor circuit, the content image processing circuit being configured to receive the luminance pixel values of the image and the content map, the content image processing circuit comprising: A content coefficient circuit configured to determine content coefficients associated with pixels of the image according to at least one of the identified categories of content in the image and a texture value of a pixel in the image or a chroma value of a pixel in the image, the texture value indicating a likelihood of a category of content associated with the pixel based on a texture in the image, and the chroma value indicating a likelihood of a category of content associated with the pixel based on color information of the image; A content correction circuit coupled to the content coefficient circuit for receiving the content coefficients and configured to generate a sharpened version of the luminance pixel values of the pixels by applying at least the content coefficients to a version of the luminance pixel values of the pixels; Including An electronic device.

17. The electronic device according to claim 16, wherein the content map is downscaled with respect to the image.

18. The electronic device according to claim 16, wherein the content coefficient circuit is configured to determine the content coefficients by upsampling the content map.

19. The content map is When the content map is enlarged until it matches the size of the image, obtaining content coefficients associated with lattice points surrounding the pixel; Interpolating the content coefficients to obtain the content coefficients associated with the pixel; The electronic device according to claim 18, which is upsampled thereby.

20. The electronic device according to claim 16, further comprising a bilateral filter coupled to the content image processing circuit, the bilateral filter being configured to generate the version of the luminance pixel value.

Citation Information

Patent Citations

  • Method for sharpening digital image without amplifying noise

    JP2003256831A

  • Image processing apparatus, imaging apparatus, control method, and program

    JP2015032966A

  • Content based image processing

    JP2024037722A

  • Method and system for selectively applying enhancement to an image

    US20030108250A1