Machine learning models for mapping between image sensor color patterns

A machine learning model trained on simulated and empirical samples addresses the challenge of remosaicing and demosaicing nonstandard Bayer CFAs, enhancing image quality by mapping pixel values to standard patterns and reducing distortions.

WO2026072055A1PCT designated stage Publication Date: 2026-04-02GOOGLE LLC
View PDF 1 Cites 0 Cited by

Patent Information

Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Filing Date
2024-09-30
Publication Date
2026-04-02

AI Technical Summary

Technical Problem

Nonstandard Bayer color filter arrays (CFAs) in cameras are difficult to accurately remosaic and/or demosaic, leading to image quality loss due to visual distortions.

Method used

A machine learning model, such as a convolutional neural network, is trained to map pixel values from a nonstandard CFA to a standard Bayer CFA or demosaiced RGB pattern using simulated and empirical training samples, enhancing image quality by jointly performing remosaicing/demosaicing and image enhancement.

Benefits of technology

The model effectively remosaices or demosaices images from nonstandard CFAs to standard Bayer or RGB patterns, improving image quality by reducing visual artifacts and enhancing image clarity.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 00000038_0000
    Figure 00000038_0000
  • Figure 00000039_0000
    Figure 00000039_0000
  • Figure 00000040_0000
    Figure 00000040_0000
Patent Text Reader

Abstract

A method includes obtaining an input image that has been captured by an image sensor using a first color filter array (CFA) having a first color pattern. The method also includes generating an output image by remosaicing or demosaicing the input image using a machine learning model configured to map pixel values determined using the first CFA to a second color pattern that (i) differs from the first color pattern of the first CFA and (ii) is associated with a second CFA. The machine learning model may have been trained using a simulated training sample. The method further includes outputting the output image.
Need to check novelty before this filing date? Find Prior Art

Description

Machine Learning Models for Mapping Between Image Sensor Color PatternsBACKGROUND

[0001] A camera may include an image sensor and a color filter array (CFA). Example CFAs include the standard Bayer CFA and nonstandard Bayer CFAs such as Quad Bayer, Nona Bayer, and QxQ Bayer, among others. While the nonstandard Bayer CFAs may provide various advantages, they may be difficult to remosaic and / or demosaic, which may sometimes contribute to a loss of image quality.SUMMARY

[0002] A machine learning model may be trained to map pixel values of an image generated using a first color filter array (CFA) that has a first color pattern to a second color pattern, which maybe associated with a second CFA different from the first CFA. For example, the machine learning model may be configured to remosaic an input image from a nonstandard Bayer CFA (e.g., Quad Bayer, Nona Bayer, QxQ Bayer, etc.) to the standard Bayer CFA, or to demosaic the input image from the nonstandard Bayer CFA to a demosaiced RGB pattern. The machine learning model may be trained using simulated training samples, each of which includes a simulated training input image that represents a corresponding training scene through the first CFA and a simulated ground-truth image that represents the corresponding training scene through the second CFA. The machine learning model may also be trained using empirical training samples, each of which includes (i) an empirical training input image representing a reference image as displayed on a display and captured using a camera having the first CFA and (ii) an empirical ground-truth image having the second color pattern and generated by fusing image data of the reference image with image data of the empirical training input image.

[0003] In a first example embodiment, a method includes obtaining an input image that has been captured by an image sensor using a first CFA having a first color pattern. The method also includes generating an output image by remosaicing or demosaicing the input image using a machine learning model configured to map pixel values determined using the first CFA to a second color pattern that (i) differs from the first color pattern of the first CFA and (ii) is associated with a second CFA. The machine learning model may have been trained using a simulated training sample. The method also includes outputting the output image.

[0004] In a second example embodiment, a system may include a processor and a non- transitory computer-readable medium having stored thereon instructions that, when executedby the processor, cause the processor to perform operations in accordance with the first example embodiment.

[0005] In a third example embodiment, a system may include a processor configured to perform operations in accordance with the first example embodiment.

[0006] In a fourth example embodiment, a non-transitory computer-readable medium may have stored thereon instructions that, when executed by a computing device, cause the computing device to perform operations in accordance with the first example embodiment.

[0007] In a fifth example embodiment, a computer program product for carrying out the operations in accordance with the first example embodiment.

[0008] In a sixth example embodiment, a system may include various means for carrying out each of the operations of the first example embodiment.

[0009] These, as well as other embodiments, aspects, advantages, and alternatives, will become apparent to those of ordinary skill in the art by reading the following detailed description, with reference where appropriate to the accompanying drawings. Further, this summary and other descriptions and figures provided herein are intended to illustrate embodiments by way of example only and, as such, that numerous variations are possible. For instance, structural elements and process steps can be rearranged, combined, distributed, eliminated, or otherwise changed, while remaining within the scope of the embodiments as claimed.BRIEF DESCRIPTION OF THE DRAWINGS

[0010] Figure 1 illustrates a computing device, in accordance with examples described herein.

[0011] Figure 2 illustrates a computing system, in accordance with examples described herein.

[0012] Figures 3A, 3B, 3C, and 3D illustrate color filter arrays, in accordance with examples described herein.

[0013] Figure 4 illustrates an image processing system, in accordance with examples described herein.

[0014] Figure 5 illustrates a training system, in accordance with examples described herein.

[0015] Figure 6 illustrates a simulation system, in accordance with examples described herein.

[0016] Figure 7 illustrates a training sample capture system, in accordance with examples described herein.

[0017] Figure 8 is a flow chart, in accordance with examples described herein.

[0018] Figure 9 is a flow chart, in accordance with examples described herein.DETAILED DESCRIPTION

[0019] Example methods, devices, and systems are described herein. It should be understood that the words “example” and “exemplary” are used herein to mean “serving as an example, instance, or illustration.” Any embodiment or feature described herein as being an “example,” “exemplary,” and / or “illustrative” is not necessarily to be construed as preferred or advantageous over other embodiments or features unless stated as such. Thus, other embodiments can be utilized and other changes can be made without departing from the scope of the subject matter presented herein.

[0020] Accordingly, the example embodiments described herein are not meant to be limiting. It will be readily understood that the aspects of the present disclosure, as generally described herein, and illustrated in the figures, can be arranged, substituted, combined, separated, and designed in a wide variety of different configurations.

[0021] Further, unless context suggests otherwise, the features illustrated in each of the figures may be used in combination with one another. Thus, the figures should be generally viewed as component aspects of one or more overall embodiments, with the understanding that not all illustrated features are necessary for each embodiment.

[0022] Additionally, any enumeration of elements, blocks, or steps in this specification or the claims is for purposes of clarity. Thus, such enumeration should not be interpreted to require or imply that these elements, blocks, or steps adhere to a particular arrangement or are carried out in a particular order. Unless otherwise noted, figures are not drawn to scale.I. Overview

[0023] Cameras may use various nonstandard-Bayer color filter arrays (CFAs) as part of an effort to improve the quality of images captured thereby. Examples of nonstandard-Bayer CFAs include the Quad Bayer CFA, the Nona Bayer CFA, and the QxQ Bayer CFA, among other possibilities. However, images captured using nonstandard-Bayer patterns may be difficult to accurately remosaic (e.g., to the standard Bayer color pattern) and / or demosaic (e.g., to the standard red-green-blue (RGB) color pattern), and may thus include various undesirable visual distortions once remosaiced or demosaiced.

[0024] A machine learning (ML) model (e.g., a convolutional neural network) may be trained to perform the remosaicing and / or demosaicing operations. For example, the ML modelmay be configured to process an input image that has been captured by an image sensor using a first CFA having a first color pattern and, based thereon, generate an output image representing a remosaiced or demosaiced version (that has a second color pattern) of the input image. The ML model may be trained to map pixel values determined using the first CFA, which has the first color pattern, to the second color pattern that differs from the first color pattern. The second color pattern may be associated with a second CFA, such as the standard Bayer CFA (in which case the input image is remosaiced) or a three-channel monochrome RGB CFA (in which case the input image is demosaiced). Unlike predetermined rule-based demosaicing / remosaicing algorithms that do not rely on ML and / or training, the ML model may be able to vary the manner in which pixel values are combined across different parts of the input image (e.g., based on various visual attributes quantified by the ML model) to improve and / or optimize the quality of the output image.

[0025] One challenge in training such an ML model is the generation of sufficient training samples for adequately training the ML model with respect to a varied range of input images, especially as the size of the ML model increases. For example, it may be difficult to capture two images of the same scene using, respectively, a nonstandard-Bayer CFA camera and a standard Bayer CFA camera, since the scene and / or camera positioning may change between capture of the two images. Accordingly, provided herein are various techniques for generating simulated training samples and empirical training samples to be used in training the ML model to perform image remosaicing and / or demosaicing.

[0026] Specifically, each simulated training sample may include a corresponding simulated training input image and a corresponding simulated ground-truth image. The simulated training input image may be generated by simulating a training image sensor capturing a training scene using the first CFA, and the simulated ground-truth image may be generated by simulating the training image sensor capturing the training scene using the second CFA. In some implementations, the training scene may be determined based on a reference image (e.g., captured by a high-quality DSLR camera) having a resolution that exceeds a resolution of the training image sensor, thus providing a high-resolution and / or high-quality input for the simulation. In some cases, the training scene may be represented as a three- dimensional model, and the simulation may be configured to represent how light reflected from this three-dimensional model is captured by the training image sensor. The simulation may be performed using, for example, the ISET software provided by the Stanford Center for Image Systems Engineering.

[0027] For example, the simulation may be used to determine, for the training scene, a red monochrome image corresponding to a red monochrome CFA, a green monochrome image corresponding to a green monochrome CFA, and a blue monochrome image corresponding to a blue monochrome CFA. The red monochrome image, the green monochrome image, and the blue monochrome image may be combined to form a demosaiced RGB image (e.g., a 16-bit raw RGB image). When the second color pattern is a demosaiced RGB color pattern, the demosaiced RGB image may be used as the simulated ground-truth image.

[0028] In some cases, the simulated training input image may be determined by selecting, from the demosaiced RGB image and for each respective pixel of the simulated training input image, a corresponding color value according to the color of the respective pixel in the first CFA (thus keeping 1 of the 3 values for each respective pixel). In other cases, the simulated training input image may be determined by separately determining the simulated training input image using a separate simulation of the first CFA. Since both the simulated training input image and the simulated ground-truth image are simulated based on the same training scene, both the simulated training input image and the simulated ground-truth image may represent the training scene from the same viewpoint and using the same training image sensor but with different CFAs.

[0029] The ML model may also be configured to enhance the input image by removing at least part of an image degradation present in the input image. For example, the ML model may be configured to perform the image enhancement and the remosaicing / demosaicing jointly, simultaneously, and / or in parallel, thus generating an output image that is both (i) enhanced and (ii) remosaiced or demosaiced. To train the ML model to perform image enhancement, the simulated training input image may be generated to include an image degradation (e.g., contrast degradation, blur degradation, and / or noise degradation), while the simulated ground-truth image may be degradation-free. Rather than adding the image degradation to the simulated training input image directly (i.e., on top of the first color pattern), which may result in various undesirable color artifacts, the image degradation may be added to the demosaiced RGB image, which may then be sampled to determine the simulated training input image.

[0030] In some implementations, rather than training the ML model using the fullresolution output of the simulation of the training image sensor, the simulated training input image may be determined by cropping the output of the simulation. The crop of the fullresolution output may be taken at a consistent part of the first color pattern, such that a given part of each simulated training input image always starts at the same part of the first colorpatern. That is, the crop of the full-resolution output may be taken such that, across a plurality of training samples, a particular pixel position of the simulated training input image is aligned with a predetermined portion of the first color patern. For example, starting at the top left pixel of the simulated training input image, the four pixels of the top row of simulated training input image may be aligned with a G, G, B, B pixel sequence of the Quad Bayer CFA, and the four pixels of the leftmost column of simulated training input image may be aligned with a G, G, R, R pixel sequence of the Quad Bayer CFA. By consistently structuring the crops in this manner, the ML model need not explicitly learn to identify the starting position of each crop, thus simplifying and / or accelerating the training process.

[0031] The ML model may also be trained using empirical training samples. For example, the ML model may be pre-trained using the simulated training samples and fine-tuned using the empirical training samples. Each empirical training sample may include a corresponding empirical training input image and a corresponding empirical ground-truth image, each of which represents a second training scene as may be depicted in another reference image. This reference image may be displayed on a display having at least a threshold resolution (e.g., an 8K display), and the empirical training input image may be captured using a camera with the first CFA to represent the training scene as displayed on the display.

[0032] The empirical ground-truth image may be generated by fusing image data from the reference image with image data from a demosaiced (e.g., using existing demosaicing algorithms) version of the empirical training input image. For example, the reference image may be aligned with the empirical training input image and / or a merged image determined based on an image burst that includes the empirical training input image. Once aligned, the image data for these two images may be combined to provide an enhanced representation of the second training scene using the second color pattern. Such image fusion may correct at least some image degradations introduced by demosaicing of the empirical training input image (or the image burst), while preserving at least some visual properties introduced by the training image sensor.II. Example Computing Devices and Systems

[0033] Figure 1 illustrates an example computing device 100. Computing device 100 is shown in the form factor of a mobile phone. However, computing device 100 may be alternatively implemented as a laptop computer, a tablet computer, and / or a wearable computing device, among other possibilities. Computing device 100 may include various elements, such as body 102, display 106, and butons 108 and 110. Computing device 100 mayfurther include one or more cameras, such as front-facing camera 104 and rear-facing camera 112.

[0034] Front-facing camera 104 may be positioned on a side of body 102 typically facing a user while in operation (e.g., on the same side as display 106). Rear-facing camera 112 may be positioned on a side of body 102 opposite front-facing camera 104. Referring to the cameras as front and rear facing is arbitrary, and computing device 100 may include multiple cameras positioned on various sides of body 102.

[0035] Display 106 could represent a cathode ray tube (CRT) display, a light emitting diode (LED) display, a liquid crystal (LCD) display, a plasma display, an organic light emitting diode (OLED) display, or any other type of display known in the art. In some examples, display 106 may display a digital representation of the current image being captured by front-facing camera 104 and / or rear-facing camera 112, an image that could be captured by one or more of these cameras, an image that was recently captured by one or more of these cameras, and / or a modified version of one or more of these images. Thus, display 106 may serve as a viewfinder for the cameras. Display 106 may also support touchscreen functions that may be able to adjust the settings and / or configuration of one or more aspects of computing device 100.

[0036] Front-facing camera 104 may include an image sensor and associated optical elements such as lenses. Front-facing camera 104 may offer zoom capabilities or could have a fixed focal length. In other examples, interchangeable lenses could be used with front-facing camera 104. Front-facing camera 104 may have a variable mechanical aperture and a mechanical and / or electronic shutter. Front-facing camera 104 also could be configured to capture still images, video images, or both. Further, front-facing camera 104 could represent, for example, a monoscopic, stereoscopic, or multiscopic camera. Rear-facing camera 112 may be similarly or differently arranged. Additionally, one or more of front-facing camera 104 and / or rear-facing camera 112 may be an array of one or more cameras. For example, computing device 100 may include multiple instances of front-facing camera 104 and / or multiple instances of rear-facing camera 112, each of which may have different optical properties to allow for image capture under various environmental contexts.

[0037] Computing device 100 could be configured to use display 106 and front-facing camera 104 and / or rear-facing camera 112 to capture images of a target object. The captured images could be a plurality of still images or a video stream. The image capture could be triggered by activating button 108, pressing a softkey on display 106, or by some other mechanism. Depending upon the implementation, the images could be captured automatically at a specific time interval, for example, upon pressing button 108, upon appropriate lightingconditions of the target object, upon moving computing device 100 a predetermined distance, or according to a predetermined capture schedule.

[0038] Figure 2 is a simplified block diagram showing some of the components of an example computing system 200. By way of example and without limitation, computing system 200 may be a cellular mobile telephone (e.g., a smartphone), a computer (such as a desktop, notebook, tablet, server, or handheld computer), a home automation component, a digital video recorder (DVR), a digital television, a remote control, a wearable computing device, a gaming console, a robotic device, a vehicle, or some other type of device. Computing system 200 may represent, for example, aspects of computing device 100.

[0039] As shown in Figure 2, computing system 200 may include communication interface 202, user interface 204, processor 206, data storage 208, and camera components 224, all of which may be communicatively linked together by a system bus, network, or other connection mechanism 210. Computing system 200 may be equipped with at least some image capture and / or image processing capabilities. It should be understood that computing system 200 may represent a physical image processing system, a particular physical hardware platform on which an image sensing and / or processing application operates in software, or other combinations of hardware and software that are configured to carry out image capture and / or processing functions.

[0040] Communication interface 202 may allow computing system 200 to communicate, using analog or digital modulation, with other devices, access networks, and / or transport networks. Thus, communication interface 202 may facilitate circuit-switched and / or packet-switched communication, such as plain old telephone service (POTS) communication and / or Internet protocol (IP) or other packetized communication. For instance, communication interface 202 may include a chipset and antenna arranged for wireless communication with a radio access network or an access point. Also, communication interface 202 may take the form of or include a wireline interface, such as an Ethernet, Universal Serial Bus (USB), or High- Definition Multimedia Interface (HDMI) port, among other possibilities. Communication interface 202 may also take the form of or include a wireless interface, such as a Wi-Fi, BLUETOOTH®, global positioning system (GPS), or wide-area wireless interface (e.g., WiMAX or 3GPP Long-Term Evolution (LTE)), among other possibilities. However, other forms of physical layer interfaces and other types of standard or proprietary communication protocols may be used over communication interface 202. Furthermore, communication interface 202 may comprise multiple physical communication interfaces (e.g., a Wi-Fi interface, a BLUETOOTH® interface, and a wide-area wireless interface).

[0041] User interface 204 may function to allow computing system 200 to interact with a human or non-human user, such as to receive input from a user and to provide output to the user. Thus, user interface 204 may include input components such as a keypad, keyboard, touch-sensitive panel, computer mouse, trackball, joystick, microphone, and so on. User interface 204 may also include one or more output components such as a display screen, which, for example, may be combined with a touch-sensitive panel. The display screen may be based on CRT, LCD, LED, and / or OLED technologies, or other technologies now known or later developed. User interface 204 may also be configured to generate audible output(s), via a speaker, speaker jack, audio output port, audio output device, earphones, and / or other similar devices. User interface 204 may also be configured to receive and / or capture audible utterance(s), noise(s), and / or signal(s) by way of a microphone and / or other similar devices.

[0042] In some examples, user interface 204 may include a display that serves as a viewfinder for still camera and / or video camera functions supported by computing system 200. Additionally, user interface 204 may include one or more buttons, switches, knobs, and / or dials that facilitate the configuration and focusing of a camera function and the capturing of images. It may be possible that some or all of these buttons, switches, knobs, and / or dials are implemented by way of a touch-sensitive panel.

[0043] Processor 206 may comprise one or more general purpose processors - e.g., microprocessors - and / or one or more special purpose processors - e.g., digital signal processors (DSPs), graphics processing units (GPUs), floating point units (FPUs), network processors, application-specific integrated circuits (ASICs), and / or tensor processing units (TPUs). In some instances, special purpose processors may be capable of image processing, image alignment, and merging images, among other possibilities. Data storage 208 may include one or more volatile and / or non-volatile storage components, such as magnetic, optical, flash, or organic storage, and may be integrated in whole or in part with processor 206. Data storage 208 may include removable and / or non-removable components.

[0044] Processor 206 may be capable of executing program instructions 218 (e.g., compiled or non-compiled program logic and / or machine code) stored in data storage 208 to carry out the various functions described herein. Therefore, data storage 208 may include a non-transitory computer-readable medium, having stored thereon program instructions that, upon execution by computing system 200, cause computing system 200 to carry out any of the methods, processes, or operations disclosed in this specification and / or the accompanying drawings. The execution of program instructions 218 by processor 206 may result in processor 206 using data 212.

[0045] By way of example, program instructions 218 may include an operating system 222 (e.g., an operating system kernel, device driver(s), and / or other modules) and one or more application programs 220 (e.g., camera functions, address book, email, web browsing, social networking, audio-to-text functions, text translation functions, and / or gaming applications) installed on computing system 200. Similarly, data 212 may include operating system data 216 and application data 214. Operating system data 216 may be accessible primarily to operating system 222, and application data 214 may be accessible primarily to one or more of application programs 220. Application data 214 may be arranged in a file system that is visible to or hidden from a user of computing system 200.

[0046] Application programs 220 may communicate with operating system 222 through one or more application programming interfaces (APIs). These APIs may facilitate, for instance, application programs 220 reading and / or writing application data 214, transmitting or receiving information via communication interface 202, receiving and / or displaying information on user interface 204, and so on.

[0047] In some cases, application programs 220 may be referred to as “apps” for short. Additionally, application programs 220 may be downloadable to computing system 200 through one or more online application stores or application markets. However, application programs can also be installed on computing system 200 in other ways, such as via a web browser or through a physical interface (e.g., a USB port) on computing system 200.

[0048] Camera components 224 may include, but are not limited to, an aperture, shutter, recording surface (e.g., photographic film and / or an image sensor), lens, shutter button, infrared projectors, and / or visible-light projectors. Camera components 224 may include components configured for capturing of images in the visible-light spectrum (e.g., electromagnetic radiation having a wavelength of 380 - 700 nanometers) and / or components configured for capturing of images in the infrared light spectrum (e.g., electromagnetic radiation having a wavelength of 701 nanometers - 1 millimeter), among other possibilities. Camera components 224 may be controlled at least in part by software executed by processor 206.

[0049] In further examples, computing system 200 may include and / or be associated with one or more remote camera(s) 230. Remote camera(s) 230 may be controlled by computing system 200. For instance, computing system 200 may transmit control signals to remote camera(s) 230 through a wireless or wired connection. Such signals may be transmitted as part of an ambient computing environment. In such examples, inputs received at the computing system 200 (for instance, physical movements of a wearable device) may bemapped to movements or other functions of remote camera(s) 230. Images captured by remote camera(s) 230 may be transmitted to computing system 200 for further processing. Such images may be treated as images captured by cameras physically located on computing system 200.III. Example Color Filter Arrays and Color Patterns

[0050] Figures 3 A, 3B, 3C, and 3D illustrate example CFAs, each of which is structured according to a corresponding color pattern. A given color pattern may specify at least part of a repeating array / arrangement of colors that a group of two or more pixels of an image or image sensor are configured to represent. Specifically, Figure 3A illustrates CFA 300, which represents an example of the standard Bayer CFA, while each of Figures 3B, 3C, and 3D illustrates an example of a nonstandard-Bayer CFA.

[0051] Turning to Figure 3A, CFA 300 includes block 302 (which may alternatively be referred to as an array or arrangement) of 4 pixels, as shown at the intersections of rows R1 and R2 and columns Cl and C2. Block 302 includes two green pixels (G) at (Rl, Cl) and (R2, C2), a red pixel (R) at (R2, Cl), and a blue pixel (B) at (Rl, C2). Block 302 may be repeated across the area of an image sensor, as shown by the ellipses. In some cases, the placement of colored pixels within block 302 may be modified.

[0052] Figure 3B illustrates CFA 310, which represents an example of the Quad Bayer CFA. CFA 310 includes block 312 of 16 pixels, as shown at the intersection of (i) rows Rl through R4 and (ii) columns Cl through C4. Block 312 includes eight green pixels (G) at (Rl- R2, C1-C2) (i.e., at the intersections of rows Rl through R2 and columns Cl through C2) and (R3-R4, C3-C4), four red pixels (R) at (R3-R4, C1-C2), and four blue pixels (B) at (R1-R2, C3-C4). Block 312 may be repeated across the area of an image sensor, as shown by the ellipses. In some cases, the placement of the groups of same-colored pixels within block 312 may be modified (e.g., blue may replace green, green may replace red, and red may replace blue).

[0053] Figure 3C illustrates CFA 320, which represents an example of the Nona Bayer CFA. CFA 320 includes block 322 of 36 pixels, as shown at the intersection of (i) rows Rl through R6 and (ii) columns Cl through C6. Block 322 includes eighteen green pixels (G) at (R1-R3, C1-C3) and (R4-R6, C4-C6), nine red pixels (R) at (R4-R6, C1-C3), and nine blue pixels (B) at (R1-R3, C4-C6). Block 322 may be repeated across the area of an image sensor, as shown by the ellipses. In some cases, the placement of the groups of same-colored pixels within block 322 may be modified.

[0054] Figure 3D illustrates CFA 330, which represents an example of the QxQ Bayer CFA. CFA 330 includes block 332 of 64 pixels, as shown at the intersection of (i) rows R1 through R8 and (ii) columns Cl through C8. Block 332 includes thirty two green pixels (G) at (R1-R4, C1-C4) and (R5-R8, C5-C8), sixteen red pixels (R) at (R5-R8, C1-C4), and sixteen blue pixels (B) at (R1-R4, C5-C8). Block 332 may be repeated across the area of an image sensor, as shown by the ellipses. In some cases, the placement of the groups of same-colored pixels within block 322 may be modified.

[0055] Each of CFAs 310, 320, and 330 may be more difficult to demosaic than CFA 300, at least in part due to the larger average separation between differently colored pixels therein. Thus, demosaicing CFAs 310, 320, and / or 330 using rule-based algorithms may result in undesirable visual artifacts and / or loss of image quality.IV. Example Image Processing System

[0056] Figure 4 illustrates an example image processing system 410 that includes machine learning model 402. Image processing system 410 may be implemented using hardware, software, or a combination thereof. Image processing system 410 may be implemented using computing device 100, computing system 200, a client device (e.g., a mobile device), a server device, and / or a combination thereof. In some cases, image processing system 410 may be implemented using the same device as the camera that captured input image 400.

[0057] Machine learning model 402 may be configured to determine output image 404 based on input image 400. Input image 400 may have been captured by an image sensor that uses a first CFA having a first color pattern. For example, input image 400 may have been captured using CFA 310, CFA 320, or CFA 330, each of which may be difficult and / or impractical to accurately remosaic and / or demosaic using, for example, predefined rule-based algorithms.

[0058] Machine learning model 402 may be configured to generate output image 404 by remosaicing or demosaicing input image 400. Specifically, machine learning model 402 may be configured to map pixel values that have been determined using the first CFA (and are thus arranged according to the first color pattern) to a second color pattern that differs from the first color pattern of the first CFA. In some cases, the second color pattern may correspond to a second CFA that differs from the first CFA.

[0059] For example, machine learning model 402 may be configured to remosaic CFA 310, CFA 320, and / or CFA 330 to the standard Bayer pattern, as represented by CFA 300. Thus, output image 404 may represent a remosaiced version of input image 400. Othercomponents of image processing system 410 may be configured to further process output image 404 to, for example, demosaic output image 404 to a demosaiced RGB color pattern. The demosaiced RBG pattern may include, for each respective pixel thereof, a corresponding red pixel value, a corresponding green pixel value, and a corresponding blue pixel value.

[0060] As another example, machine learning model 402 may be configured to demosaic CFA 310, CFA 320, and / or CFA 330 to the demosaiced RGB color pattern. Thus, output image 404 may represent a demosaiced version of input image 400. Output image 404 may thus be an RGB output image in which each pixel is associated with a corresponding red pixel value, a corresponding green pixel value, and a corresponding blue pixel value. Whether machine learning model 402 is trained to perform remosaicing or demosaicing may depend on other image processing components of image processing system 410 (e.g., availability of an algorithm for demosaicing the standard Bayer pattern) and / or may be a design choice.

[0061] Machine learning model 402 may include an artificial neural network, such as a convolutional neural network. For example, machine learning model 402 may be fully convolutional. The convolutional neural network may include and / or be based on the UNet architecture. For example, the convolutional neural network may include a portion of the UNet architecture that corresponds to a largest scale / resolution of the UNet architecture, but may omit (i) downscaled / downsampled scales / resolutions and / or (ii) space-to-depth transformations. Omitting each of the downscaled / downsampled scales / resolutions and the space-to-depth transformations from the UNet architecture may reduce and / or prevent the introduction of checkerboard artifacts into output image 404 by machine learning model 402.

[0062] In some implementations, machine learning model 402 may also be configured to enhance input image 400 by removing therefrom at least some image degradations present therein. For example, machine learning model 402 may be configured to improve a contrast of input image 400, reduce an extent of blurring present in input image 400, and / or remove noise present in input image 400, among other possibilities. Thus, output image 404 may be enhanced relative to input image 400 in that it may lack at least some of the visual degradations present in input image 400.

[0063] Machine learning model 402 may be configured to perform the enhancement of input image 400 concurrently with, in parallel with, and / or jointly with performing the remosaicing or demosaicing of input image 400. Thus, machine learning model 402 may be configured to generate output image 404 by both (i) mapping pixel values of input image 400 from the first color pattern to the second color pattern and (ii) adjusting the pixel values of input image 400 to reduce the extent of visual degradations in input image 400. Having machinelearning model 402 perform both of these operations may allow for simplification of image processing system 410 be allowing other components thereof to be simplified or entirely omitted, thus improving the image quality of output image 404 while reducing the amount of computational resources used by image processing system 410.

[0064] In some implementations, image processing system 410 may include and / or be used with other image processing components and / or operations. For example, image processing system 410 may include and / or be used with (e.g., output image 404 may be further processed by) high dynamic range (HDR) operation(s), upsampling operation(s), super resolution operation(s), sharpening operation(s), and / or color adjustment operations), among other possibilities.V. Example Training System

[0065] Figure 5 illustrates an example training system 500 for training of machine learning model 402. Training system 500 may include machine learning model 402, loss function(s) 522, and model parameter adjuster 526. Training system 500 may be implemented using software, hardware, or a combination thereof. Training system 500 may be implemented using computing device 100, computing system 200, a client device (e.g., a mobile device), a server device, and / or a combination thereof. Training system 500 may be configured to generate a trained version of machine learning model 402 based on simulated training sample 502 and / or empirical training sample 512.

[0066] Simulated training sample 502 may include simulated training input image 504 and simulated ground-truth image 506, each of which may represent a first training scene. Simulated training input image 504 may include a representation of the first training scene as generated by simulating a training image sensor capturing the first training scene using the first CFA having the first color pattern (e.g., the nonstandard-Bayer color pattern). Simulated ground-truth image 506 may include a representation of the first training scene as generated by simulating the training image sensor capturing the first training scene using the second CFA having the second color pattern (e.g., the standard Bayer color pattern, or the demosaiced RGB color pattern).

[0067] Simulated training sample 502 may be generated using simulation system 600, as shown and discussed in connection with Figure 6. In implementations where machine learning model 402 is being trained to perform image enhancement jointly with image remosaicing and / or demosaicing, simulated training input image 504 may include one or more image degradations (which may be simulated or added after the simulation) that are absent from simulated ground-truth image 506. Thus, simulated ground-truth image 506 mayrepresent an enhanced version of simulated training input image 504. For example, simulated training input image 504 may include contrast degradation(s), blur degradation(s), and / or noise degradation(s), among other possibilities, each of which may be wholly or partially absent from simulated ground-truth image 506. Thus, simulated training sample 502 may indicate both (i) how to remap the representation of the first training scene from the first color pattern to the second color pattern and (ii) how to remove image degradations present in simulated training input image 504.

[0068] Empirical training sample 512 may include empirical training input image 514 and empirical ground-truth image 516, each of which may represent a second training scene. Empirical training input image 514 may include a representation of the second training scene as captured by the training image sensor using the first CFA having the first color pattern. Specifically, empirical training input image 514 may be captured by the training image sensor and may represent a reference image of the second training scene as displayed on a display. Empirical ground-truth image 516 may represent the second training scene using the second CFA having the second color pattern. Specifically, empirical ground-truth image 516 may be generated by fusing image data from the reference image with image data from a demosaiced version of empirical training input image 514. Empirical training sample 512 may thus represent an actual capture of a real-world scene (e.g., rather than a synthetic scene).

[0069] Empirical training sample 512 may be generated using training sample capture system 700, as shown in and discussed in connection with Figure 7. In implementations where machine learning model 402 is being trained to perform image enhancement jointly with image remosaicing and / or demosaicing, empirical training input image 514 may include one or more image degradations (which may be inherently added by the training image sensor during image capture) that are absent from empirical ground-truth image 516. Thus, empirical ground-truth image 516 may represent an enhanced version of empirical training input image 514. For example, empirical training input image 514 may include contrast degradation(s), blur degradation(s), and / or noise degradation(s), among other possibilities, each of which may be wholly or partially absent from empirical ground-truth image 516. Thus, empirical training sample 512 may indicate both (i) how to remap the representation of the second training scene from the first color pattern to the second color pattern and (ii) how to remove image degradations present in empirical training input image 514.

[0070] Machine learning model 402 may be trained using a plurality of instances of simulated training sample 502, which may provide simulated representations of different training scenes, and / or a plurality of instances of empirical training sample 512, which mayprovide actual (i.e., non-simulated) representations of different training scenes. Specifically, each of simulated training input image 504 and empirical training input image 514 may be used to generate a corresponding instance of training output image 520. Thus, each of simulated training input image 504 and empirical training input image 514 may be analogous to input image 400, but may be processed at training time rather than at inference time. Training on instances of simulated training sample 502 may allow machine learning model 402 to learn from image pairs that are accurately spatially aligned as a result of being simulated, while training on instances of empirical training samples 512 may allow machine learning model 402 to learn from image pairs that represent characteristics of real-world scenes and / or physical imaging hardware. Thus, training machine learning model 402 using both types of training samples may improve the quality and accuracy of the resulting model and / or allow it to accurately process a wide range of scenes.

[0071] In some implementations, simulated training sample 502 and / or empirical training sample 512 may include and / or be based on respective cropped portions of corresponding full-resolution images. For example, rather than training machine learning model 402 using full-resolution 6144 pixel by 8160 pixel images, machine learning model 402 may instead be trained using 256 pixel by 256 pixel images. Training in this manner may reduce usage of processing and memory resources and reduce training time, without affecting the ability of machine learning model 402 to be applied to images of varying resolutions at inference time. For example, machine learning model 402 may include a convolutional neural network, thus allowing machine learning model 402 to process input images of varying resolutions.

[0072] However, taking random crops from the full-resolution images may result in different cropped portions starting at different positions relative to the nonstandard-Bayer pattern. In the example of CFA 310, taking the cropped image with the top left pixel thereof located at (Rl, Cl) of CFA 310 will result in a different pattern of colors across the cropped image than, for example, taking the cropped image with the top left pixel thereof located at (R4, C2) of CFA 310. Since the training input images processed by machine learning model 402 do not explicitly indicate the color associated with each pixel thereof (i.e., the input image is a single-channel image, with different portions thereof captured using different color filters), cropped portions that are inconsistently aligned with a given CFA may make it more difficult to train machine learning model 402. Specifically, when training on inconsistently cropped images, machine learning model 402 may learn both (i) how to map from the first color pattern to the second color pattern and (ii) how to figure out the portion of the first CFA at which agiven training input image starts. At inference time, input images may all start at a consistent and / or predetermined portion of the first CFA (e.g., as determined by the CFA used by a given camera), and thus the machine learning model 402 need not be configured to handle different starting positions relative to the CFA (i.e., part of the skills gained in training are wasted).

[0073] Accordingly, simulated training sample 502 and / or empirical training sample 512 may include and / or be based on cropped portions of full-resolution images, where each cropped portion is aligned with a consistent and / or predetermined portion of the first color pattern of the first CFA. Thus, for each respective training sample, a particular pixel position of a corresponding training input image of the respective training sample may be aligned with a consistent portion of the first color pattern. In the example of CFA 310, the top left 4 by 4 block of pixels of simulated training input image 504 and / or empirical training input image 514 may be consistently aligned with the color pattern shown in block 312 of CFA 310. The top left pixel of pixels of simulated training input image 504 and / or empirical training input image 514 might not always start at (Rl, Cl) ofCFA 310, and may instead start at another coordinate of CFA 310 that includes and is surrounded by the same pattern of colors as (Rl, Cl). Thus, for example, the top left 4 by 4 block of pixels may consistently correspond to (i) G, G, B, B left-to-right across the top row thereof and (ii) G, G, R, R, top-to-bottom across the leftmost column thereof. Simulated ground-truth image 506 may be cropped at the same coordinates as simulated training input image 504, and empirical ground-truth image 516 may be cropped at the same coordinates as empirical training input image 514.

[0074] In some implementations, the predetermined portion of the first color pattern at which the cropped portions start may be based on (e.g., may be the same as) the starting pattern of colors of the CFA of the camera from which input images are to be obtained and processed by machine learning model 402 at inference time. Thus, the cropped portions used at training time may have the same starting pattern of colors as input image 400, thereby allowing machine learning model 402 to avoid explicitly learning how to handle different starting positions relative to the first CFA.

[0075] Over the course of training, machine learning model 402 may get progressively better at remapping training input images from the first color pattern to the second color pattern and / or removing image degradations from the training input images. Thus, training output image 520 may be analogous to output image 404, but may be generated at training time rather than at inference time. Performance of machine learning model 402 during training may be quantified and improved using loss function(s) 522, which may be configured to comparetraining output image 520 to the corresponding training input image on which training output image 520 is based.

[0076] Specifically, loss function(s) 522 may be configured to determine loss value 524 based on training output image 520 and a corresponding ground-truth image. When training output image 520 is determined based on simulated training input image 504, the corresponding ground-truth image is simulated ground-truth image 506. When training output image 520 is determined based on empirical training input image 514, the corresponding ground-truth image is empirical ground-truth image 516.

[0077] Loss function(s) 522 may be configured to determine loss value 524 based on a difference (e.g., a mean square error, a mean absolute error, etc.) between training output image 520 and the corresponding ground-truth image (and / or between respective gradient images based on gradients of training output image 520 and the corresponding ground-truth image). Loss function(s) 522 may additionally or alternatively be configured to compare a latent representation of training output image 520 to a latent representation of the corresponding ground-truth image using a perceptual loss function. In implementations that utilize diffusion models, loss function(s) 522 may additionally or alternatively be configured to compare training noise added by a forward diffusion process to noise detected by machine learning model 402 (e.g., when machine learning model 402 and / or aspects thereof are configured to predict the noise rather than explicitly predict a denoised image).

[0078] Model parameter adjuster 526 may be configured to determine updated model parameters 528 based on loss value 524 and / or loss function(s) 522. Specifically, updated model parameters 528 may be selected such that, during a subsequent iteration of processing of the training input image (e.g., simulated training input image 504 and / or empirical training input image 514), training output image 520 more closely matches the corresponding groundtruth image (e.g., simulated ground-truth image 506 and / or empirical ground-truth image 516). Updated model parameters 528 may include one or more updated parameters of any trainable component of machine learning model 402.

[0079] Model parameter adjuster 526 may be configured to determine updated model parameters 528 by, for example, determining a gradient of loss function(s) 522. Based on this gradient and loss value 524, model parameter adjuster 526 may be configured to select (e.g., using gradient descent) updated model parameters 528 that are expected to reduce loss value 524, and thus improve a performance of machine learning model 402. After applying updated model parameters 528 to machine learning model 402, the operations discussed above may be repeated to compute another instance of loss value 524 and, based thereon, another instance ofupdated model parameters 528 may be determined and applied to machine learning model 402 to further improve the performance thereof. Such training of machine learning model 402 may be repeated until, for example, loss value 524 is reduced to below a target loss value.VI. Example Simulation System

[0080] Figure 6 illustrates an example simulation system 600. Simulation system 600 may include CFA simulator 618, degradation calculator 628, pixel selector 634, and pixel selector 638. Simulation system 600 may be configured to generate simulated training input image 504 and simulated ground-truth image 506 based on training scene 602. Simulation system 600 may include, represent, and / or be based on, for example, the ISET software provided by the Stanford Center for Image Systems Engineering, which is publicly available from the “ / ISET” github repository. Simulation system 600 may be implemented using software, hardware, or a combination thereof (e.g., using computing device 100, computing system 200, a client device, a server device, and / or a combination thereof).

[0081] Training scene 602 may include a two-dimensional and / or a three-dimensional representation of one or more objects, environments, and / or visual features. In some cases, training scene 602 may include and / or be used by simulation system 600 to generate a representation of light reflected from the one or more objects, environments, and / or visual features. In some cases, training scene 602 may be a synthetic scene (e.g., based on a computer- aided design (CAD) model), rather than a real-world scene. In other cases, training scene may represent and / or approximate a real-world scene.

[0082] For example, training scene 602 may be based on reference image 640 that depicts a real-world scene. Reference image 640 may have a resolution that is equal to or greater than a target resolution of simulated training input image 504, simulated ground-truth image 506, and / or the training image sensor being simulated by simulation system 600. Thus, reference image 640 may include sufficient and / or excess image data for determining pixel values of simulated training input image 504 and / or simulated ground-truth image 506. Reference image 640 may be captured using, for example, a digital single-reflex (DSLR) camera, and may thus be of a higher visual quality than the same image captured using the training sensor.

[0083] Simulation system 600 may provide a user interface for specifying training scene 602, reference image 640, properties of the CFAs to be simulated, properties of the training image sensor, and / or properties of any optical components utilized by the training image sensor, among other possibilities.

[0084] CFA simulator 618 may be configured to generate red monochrome image 622 by simulating the training image sensor capturing training scene 602 through red monochrome CFA 612. CFA simulator 618 may also be configured to generate blue monochrome image 624 by simulating the training image sensor capturing training scene 602 through blue monochrome CFA 614. CFA simulator 618 may further be configured to generate green monochrome image 626 by simulating the training image sensor capturing training scene 602 through green monochrome CFA 616. In some cases, the value of each pixel of red monochrome image 622, blue monochrome image 624, and green monochrome image 626 (“monochrome images 622- 626”) may be represented using 16 bits. Monochrome images 622-626 may collectively (e.g., when stacked together) form a demosaiced RGB image. CFA simulator 618 may determine monochrome images 622-626 by, for example, simulating light reflected from the training scene and the corresponding irradiance of each pixel of the training image sensor.

[0085] Degradation calculator 628 may be configured to generate degraded RGB image 630 based on monochrome images 622-626. For example, degradation calculator may be configured to modify a contrast of monochrome images 622-626, blur one or more portions of monochrome images 622-626, add noise to monochrome images 622-626, and / or otherwise degrade the visual quality of monochrome images 622-626. Degradation calculator 628 may thus introduce visual artifacts that are expected and / or likely to be present in real-world images captured by the training sensor (e.g., due to camera imperfections and / or properties of the scene).

[0086] Notably, degradation calculator 628 may add degradations to monochrome images 622-626 in the RGB domain before application of pixel selector 634, rather than in the domain of the first CFA (i.e., in the domain of the first color pattern) after application of pixel selector 634. Adding the degradations in this manner may prevent machine learning model 402 from introducing color artifacts into output images at inference time. Specifically, since simulated training input image 504 is a single-channel image corresponding to the first CFA, adding the degradations directly to simulated training input image 504 (i.e., after application of pixel selector 634) would cause visual information from different colors of the CFA to be mixed in physically inconsistent and / or implausible ways. For example, a blur kernel applied to each of monochrome images 622-626 individually blurs together image data of the same color, while the blur kernel applied to simulated training input image 504 might blur together image data of different colors.

[0087] Pixel selector 634 may be configured to determine simulated training input image 504 by sampling degraded RGB image 630 in accordance with input color pattern 632.Input color pattern 632 may represent the color pattern of the nonstandard-Bayer CFA that machine learning model 402 is being trained to demosaic or remosaic. For example, input color pattern 632 may represent Quad Bayer CFA 310, Nona Bayer CFA 320, or QxQ Bayer CFA 330. Pixel selector 634 may be configured to select, for each respective pixel of simulated training input image 504 and from degraded RGB image 630, a corresponding red, green, or blue pixel value in accordance with the color indicated by input color pattern 632 for the respective pixel.

[0088] As one example, when input color pattern 632 corresponds to CFA 310, pixel selector 634 may be configured to select (i) for the pixels of simulated training input image 504 corresponding to (R1-R2, C1-C2) and (R3-R4, C3-C4) of CFA 310, the green values of the spatially-corresponding pixels in green monochrome image 626 (while discarding the red and blue values, respectively, of spatially-corresponding pixels in red monochrome image 622 and blue monochrome image 624), (ii) for the pixels of simulated training input image 504 corresponding to (R1-R2, C3-C4) of CFA 310, the blue values of the spatially-corresponding pixels in blue monochrome image 624 (while discarding the corresponding red and green values), and (iii) for the pixels of simulated training input image 504 corresponding to (R3-R4, C1-C2) of CFA 310, the red values of the spatially-corresponding pixels in red monochrome image 622 (while discarding the corresponding blue and green values).

[0089] As another example, when input color pattern 632 corresponds to CFA 320, pixel selector 634 may be configured to select (i) for the pixels of simulated training input image 504 corresponding to (R1-R3, C1-C3) and (R4-R6, C4-C6) of CFA 320, the green value of the spatially-corresponding pixel in green monochrome image 626 (while discarding the corresponding blue and red values), (ii) for the pixel of simulated training input image 504 corresponding to (R1-R3, C4-C6) of CFA 320, the blue value of the spatially-corresponding pixel in blue monochrome image 624 (while discarding the corresponding red and green values), and (iii) for the pixel of simulated training input image 504 corresponding to (R4-R6, C1-C3) of CFA 320, the red value of the spatially-corresponding pixel in red monochrome image 622 (while discarding the corresponding blue and green values).

[0090] Pixel selector 638 may perform operations similar to those of pixel selector 634. Specifically, pixel selector 638 may be configured to determine simulated ground-truth image 506 by sampling monochrome images 622-626, as indicated by lines 642, 644, and 646, respectively, in accordance with output color pattern 636. Output color pattern 636 may represent the color pattern of the standard Bayer CFA when machine learning model 402 is being trained to remosaic image data, and may represent the demosaiced RGB color patternwhen machine learning model 402 is being trained to demosaic image data. Pixel selector 638 may be configured to select, for each respective pixel of simulated ground-truth image 506 and from monochrome images 622-626, one or more corresponding red, green, and / or blue pixel values in accordance with the color indicated by output color pattern 636 for the respective pixel.

[0091] For example, when output color pattern 636 corresponds to CFA 300 (i.e., the standard Bayer CFA), pixel selector 638 may be configured to select (i) for the pixels of simulated ground-truth image 506 corresponding to (Rl, Cl) and (R2, C2) of CFA 300, the green value of the spatially-corresponding pixel in green monochrome image 626 (while discarding the red and blue values, respectively, of spatially-corresponding pixels in red monochrome image 622 and blue monochrome image 624), (ii) for the pixel of simulated ground-truth image 506 corresponding to (Rl , C2) of CFA 300, the blue value of the spatially- corresponding pixel in blue monochrome image 624 (while discarding the corresponding red and green values), and (iii) for the pixel of simulated ground-truth image 506 corresponding to (R2, Cl) of CFA 300, the red value of the spatially-corresponding pixel in red monochrome image 622 (while discarding the corresponding blue and green values).

[0092] As another example, when output color pattern 636 corresponds to the demosaiced RGB color pattern, pixel selector 638 may be configured to stack monochrome images 622-626 together, thus combining three one-channel images into one three-channel image.

[0093] In some implementations, when simulated ground-truth image 506 is represented using the standard Bayer pattern, rather than sampling simulated ground-truth image 506 from monochrome images 622-626, CFA simulator 618 may instead be configured to simulate the standard Bayer pattern directly. That is, CFA simulator 618 may be configured to generate simulated ground-truth image 506 by simulating the training image sensor capturing training scene 602 through the standard Bayer CFA. Similarly, in implementations wherein degradation calculator 5628 is omitted, CFA simulator 618 may be configured to simulate the standard Bayer pattern directly instead of simulating each of CFAs 612, 614, and 616.VII. Example Training Sample Capture System

[0094] Figure 7 illustrates example training sample capture system 700. Training sample capture system 700 includes display 704, camera 706, burst processor 710, image preprocessor 714, and fusion model 718. Training sample capture system 700 may be configured to generate empirical training input image 514 and empirical ground-truth image 516 based onreference image 702. Reference image 702 may be analogous to reference image 640, and may represent a training scene different from that represented by reference image 640. In some cases, the same reference image may be used by both simulation system 600 and training sample capture system 700. Training sample capture system 700 may be implemented using software, hardware, or a combination thereof (e.g., using computing device 100, computing system 200, a client device, a server device, and / or a combination thereof).

[0095] Display 704 may be configured to display reference image 702. Camera 706 may be configured to capture one or more images of reference image 702 as displayed by display 704. That is, camera 706 may photograph a display surface of display 704 while display 704 generates a visual representation of reference image 702. Display 704 may include, for example, a television, a computer monitor, and / or a projector, among other possibilities. Camera 706 may include an image sensor and the first CFA (e.g., CFA 310, 320, or 330). The image sensor of camera 706 may be the same as or similar to the image sensor simulated by simulation system 600 and / or the image sensor used to generate input image 400 at inference time. Since this image sensor is used to generate training samples, the image sensor may also be referred to as a training image sensor.

[0096] Display 704 may have at least a threshold display resolution (e.g., 4k, 8k, 16k), which may be predetermined. In some cases, the threshold display resolution may represent at least a first threshold fraction (e.g., 50%, 55%, 66%, etc.) of a resolution of the image sensor of camera 706. For example, when the image sensor of camera 706 is configured to generate images having a resolution of 50 megapixels, an 8k resolution display 704 may be used. Thus, reference image 702 may be displayed using approximately 33 million pixels, which approximately corresponds to 66% of a resolution of the image sensor. In some implementations, reference image 702 may have at least a threshold image resolution, which may represent at least a second threshold fraction (e.g., 66%, 100%, 150%, 200%, etc.) of the resolution of the image sensor of camera 706. Thus, in some cases, the resolution of reference image 702 may exceed a resolution of the image sensor of camera 706.

[0097] Camera 706 may be configured to generate image burst 708 based on reference image 702 as displayed on display 704. Image burst 708 may include a plurality of mosaiced images captured using the first CFA having the first color pattern. For example, image burst 708 may include five successively-captured images of reference image 702 as displayed on display 704. Alternatively, in some implementations (e.g., where camera 706 is not configured to capture image bursts), camera 706 may instead be configured to generate a single image of reference image 702 as displayed on display 704.

[0098] Burst processor 710 may be configured to determine, based on image burst 708, (i) demosaiced RGB image 712 and (ii) empirical training input image 514. In some implementations, burst processor 710 may be configured to determine empirical training input image 514 by aligning and merging two or more images of image burst 708. Thus, burst processor 710 may combine image data from multiple images of image burst 708 but without demosaicing the resulting image, thus generating empirical training input image 514 that is expressed using the first color pattern. In other implementations, burst processor 710 may be configured to select empirical training input image 514 from image burst 708. For example, empirical training input image 514 may correspond to a base image frame of image burst 708 relative to which other images of image burst 708 are merged together. Empirical training input image 514 may include image degradations introduced by camera 706 and / or properties of the training scene of reference image 702.

[0099] Burst processor 710 may be configured to determine demosaiced RGB image 712 by aligning, merging, and demosaicing the image burst 708. Burst processor 710 may perform demosaicing using, for example, a rule-based algorithm that does not rely on training of a machine learning model. In some cases, burst processor 710 may introduce one or more visual artifacts into demosaiced RGB image 712 as part of the demosaicing operation. That is, due to larger groupings of pixels of the same color in nonstandard Bayer CFAs, it may be difficult to determine a rule-based algorithm that accurately demosaics a wide variety of images. These visual artifacts may be the same as or similar to the visual artifacts that machine learning model 402 is trained to avoid introducing into images demosaiced or remosaiced thereby. At least some of these visual artifacts may be removed from demosaiced RGB image 712 using fusion model 718, as discussed below.

[0100] In some implementations, demosaiced RGB image 712 may be a linear RGB image in which each pixel value is linearly proportional to the irradiance measured by the corresponding pixel. In some implementations, burst processor 710 may be configured to apply min-max normalization as part of the determination of demosaiced RGB image 712.

[0101] Image pre-processor 714 may be configured to determine aligned reference image 716 by aligning reference image 702 with demosaiced RGB image 712. Specifically, image pre-processor 714 may be configured to transform (e.g., translate, rotate, scale, warp, etc.) reference image 702 such that a plurality of features in reference image 702 are spatially aligned in pixel space with corresponding features in demosaiced RGB image 712. Accordingly, both aligned reference image 716 and demosaiced RGB image 712 may represent the second training scene using exactly or substantially the same point of view and / or field ofview. In some implementations, image pre-processor 714 may also be configured to adjust the brightness of reference image 702 and / or demosaiced RGB image 712 to more closely match one another. In some implementations, image pre-processor 714 may additionally be configured to resize reference image 702 to a resolution of demosaiced RGB image 712.

[0102] Fusion model 718 may be configured to combine the image data of demosaiced RGB image 712 with image data of aligned reference image 716. Fusion model 718 may include and / or be based on the techniques and / or models discussed in a paper titled “Efficient Hybrid Zoom using Camera Fusion on Mobile Phones,” authored by Wu et al., and published as arXiv:2401.01461, which is incorporated herein by reference. Specifically, demosaiced RGB image 712 may be used as the source image (W) and aligned reference image 716 may be treated as the reference image (T). Thus, detail present in reference image 702 may allow fusion model 718 to compensate for any detail loss due to photographing reference image as displayed by display 704, especially when the resolution of reference image 702 is equal to or exceeds the resolution of camera 706 and / or the resolution of display 704.

[0103] Since empirical training input image 514 and demosaiced RGB image 712 are determined using camera 706, each of these images may provide a representative depiction of the training scene in reference image 702 as captured using camera 706. That is, demosaiced RGB image 712 may include various visual characteristics that may be present at inference time in input images captured using camera 706, and may thus be representative of the actual distribution of samples that may be encountered at inference time. However, due to demosaiced RGB image 712 being based on (i) a capture of the training scene as represented on display 704 and (ii) processing of image burst 708 using demosaicing algorithm(s) that may introduce visual artifacts, at least some portions of demosaiced RGB image 712 may include visual distortions and / or artifacts, may lack sufficient image data, and / or may otherwise suffer from reduced image quality. Since aligned reference image 716 includes the relatively higher quality image data of reference image 702, the fusion of demosaiced RGB image 712 and aligned reference image 716 may correct at least some undesirable visual inaccuracies in demosaiced RGB image 712 while preserving at least some of the desirable visual characteristics present in images captured using camera 706.

[0104] In cases where the second color pattern is the demosaiced RGB color pattern, the output of fusion model 718 may form empirical ground-truth image 516. In cases where the second color pattern corresponds to the standard Bayer CFA, the output of fusion model 718 may be processed by mosaic calculator 720 to determine empirical ground-truth image 516. In either case, empirical ground-truth image 516 may provide, using the second colorpatern, a relatively high quality representation of reference image 702 and may include various visual characteristics introduced by camera 706 as part of the image capture process. In some implementations, training sample capture system 700 may be configured to generate empirical ground-truth image 516 by combining the Y channel of the image output by fusion model 718 with the UV channels of demosaiced RGB image 712.VIII. Additional Example Operations

[0105] Figure 8 illustrates a flow chart of operations related to demosaicing and / or remosaicing an input image using a machine learning model. The operations may be carried out by computing device 100, computing system 200, image processing system 410, and / or training system 500, among other possibilities. The embodiments of Figure 8 maybe simplified by the removal of any one or more of the features shown therein. Further, these embodiments may be combined with features, aspects, and / or implementations of any of the previous figures or otherwise described herein.

[0106] Block 800 may involve obtaining an input image that has been captured by an image sensor using a first CFA having a first color patern.

[0107] Block 802 may involve generating an output image by remosaicing or demosaicing the input image using a machine learning model configured to map pixel values determined using the first CFA to a second color patern that (i) differs from the first color patern of the first CFA and (ii) is associated with a second CFA. The machine learning model may have been trained using a simulated training sample.

[0108] Block 804 may involve outputing the output image.

[0109] In some examples, the first CFA may include a nonstandard-Bayer CFA.

[0110] In some examples, the nonstandard-Bayer CFA may include a Quad Bayer CFA.

[0111] In some examples, the nonstandard-Bayer CFA may include a Nona Bayer CFA.

[0112] In some examples, the nonstandard-Bayer CFA may include a QxQ Bayer CFA.

[0113] In some examples, the second CFA may include a Bayer CFA. The output image may be generated by remosaicing the input image from the first color patern to the second color patern of the Bayer CFA.

[0114] In some examples, the second color patern may include a demosaiced red- green-blue (RGB) color patern. The output image may be generated by demosaicing the input image from the first color patern to the demosaiced RGB color patern. Each pixel of thedemosaiced RGB color patern may be associated with a corresponding red pixel value, a corresponding green pixel value, and a corresponding blue pixel value.

[0115] In some examples, the simulated training sample may include (i) a simulated training input image generated by simulating a training image sensor capturing a training scene using the first CFA and (ii) a simulated ground-truth image generated by simulating the training image sensor capturing the training scene using a second CFA having the second color patern.

[0116] In some examples, the simulated training input image may include an image degradation added as part of simulating the training image sensor capturing the training scene using the first CFA. The simulated ground-truth image may represent the training scene without the image degradation. Generating the output image may include enhancing the input image using the machine learning model by removing at least part of the image degradation present in the input image.

[0117] In some examples, the image degradation may include one or more of a contrast degradation, a blur degradation, or a noise degradation.

[0118] In some examples, the second CFA may include a red monochrome CFA, a green monochrome CFA, and a blue monochrome CFA. Simulating the training image sensor capturing the training scene using the second CFA may include determining a red monochrome image representing the training scene through the red monochrome CFA, determining a green monochrome image representing the training scene through the green monochrome CFA, determining a blue monochrome image representing the training scene through the blue monochrome CFA, and determining an RGB image by combining the red monochrome image, the green monochrome image, and the blue monochrome image. Each pixel of the RGB image may be associated with a corresponding red pixel value, a corresponding green pixel value, and a corresponding blue pixel value.

[0119] In some examples, simulating the training image sensor capturing the training scene using the first CFA may include selecting, for each respective pixel of the simulated training input image, a corresponding color value from the RGB image based on a color associated with the respective pixel in the first CFA.

[0120] In some examples, simulating the training image sensor capturing the training scene using the first CFA may include applying one or more image degradations to the RGB image before selecting, for each respective pixel of the simulated training input image, the corresponding color value from the RGB image. The simulated ground-truth image may correspond to the RGB image without the one or more image degradations.

[0121] In some examples, the machine learning model may include a convolutional neural network.

[0122] In some examples, the simulated training input image may have been generated by cropping a simulated output of the training image sensor in alignment with a predetermined portion of the first color pattern such that, for each respective simulated training sample of a plurality of simulated training samples, a particular pixel position of a corresponding simulated training input image of the respective simulated training sample is aligned with a consistent portion of the first color pattern.

[0123] In some examples, the machine learning model may have been trained using a training process that includes obtaining the simulated training sample, generating a training output image by processing the simulated training input image using the machine learning model, determining a loss value based on the simulated ground-truth image and the training output image, and adjusting one or more parameters of the machine learning model based on the loss value.

[0124] In some examples, the machine learning model may have been trained using an empirical training sample that includes an empirical training input image and an empirical ground-truth image. The empirical training input image may have been captured by the training image sensor using the first CFA, and may represent a reference image of a second training scene displayed on a display having at least a threshold resolution. The empirical ground-truth image may be generated by fusing (i) image data from the reference image with (ii) image data from a demosaiced version of the empirical training input image. The empirical ground-truth image may represent the second training scene using the second color pattern.

[0125] In some examples, the empirical ground-truth image may have been generated by determining the demosaiced version of the empirical training input image by aligning, merging, and demosaicing an image burst that has been captured by the training image sensor using the first CFA and represents the reference image of the second training scene displayed on the display. The reference image may be aligned to match a field of view of the demosaiced version of the empirical training input image. Image data from the reference image as aligned may be fused, using an image fusion model, with image data from the demosaiced version of the empirical training input image.

[0126] In some examples, the machine learning model may be pre-trained using the simulated training sample and fine-tuned using the empirical training sample.

[0127] In some examples, the training process may further include obtaining the empirical training sample, generating a second training output image by processing theempirical training input image using the machine learning model, determining a second loss value based on the empirical ground-truth image and the second training output image, and adjusting the one or more parameters of the machine learning model based on the second loss value.

[0128] Figure 9 illustrates a flow chart of operations related to training a machine learning model to perform image demosaicing and / or remosaicing. The operations may be carried out by computing device 100, computing system 200, image processing system 410, training system 500, simulation system 600, and / or training sample capture system 700, among other possibilities. The embodiments of Figure 9 may be simplified by the removal of any one or more of the features shown therein. Further, these embodiments may be combined with features, aspects, and / or implementations of any of the previous figures or otherwise described herein.

[0129] Block 900 may involve obtaining a simulated training sample that includes (i) a simulated training input image generated by simulating a training image sensor capturing a training scene using a first CFA having a first color pattern and (ii) a simulated ground-truth image generated by simulating the training image sensor capturing the training scene using a second CFA having a second color pattern that differs from the first color pattern of the first CFA.

[0130] Block 902 may involve generating a training output image by processing the simulated training input image using a machine learning model.

[0131] Block 904 may involve determining a loss value based on the simulated groundtruth image and the training output image.

[0132] Block 906 may involve adjusting one or more parameters of the machine learning model based on the loss value such that the machine learning model is trained to remosaic or demosaic the simulated training input image by mapping pixel values determined using the first CFA to the second color pattern.

[0133] Block 908 may involve outputting the machine learning model as trained.

[0134] In some examples, the first CFA may include a nonstandard-Bayer CFA.

[0135] In some examples, the nonstandard-Bayer CFA may include a Quad Bayer CFA.

[0136] In some examples, the nonstandard-Bayer CFA may include a Nona Bayer CFA.

[0137] In some examples, the nonstandard-Bayer CFA may include a QxQ Bayer CFA.

[0138] In some examples, the second CFA may include a Bayer CFA. Thus, the machine learning model may be trained to remosaicing input images from the first color pattern to the second color pattern of the Bayer CFA.

[0139] In some examples, the second color pattern may include a demosaiced red- green-blue (RGB) color pattern. Each pixel of the demosaiced RGB color pattern may be associated with a corresponding red pixel value, a corresponding green pixel value, and a corresponding blue pixel value.

[0140] In some examples, obtaining the simulated training sample may include determining the training sample by executing a simulation based on the training scene. Executing the simulation may include adding an image degradation to an RGB image determined by simulating the training image sensor capturing the training scene using the first CFA. The simulated ground-truth image may represent the training scene without the image degradation. The machine learning model as trained may be configured to enhance an input image by removing at least part of the image degradation present in the input image.

[0141] In some examples, the image degradation may include one or more of a contrast degradation, a blur degradation, or a noise degradation.

[0142] In some examples, the second CFA may include a red monochrome CFA, a green monochrome CFA, and a blue monochrome CFA. Simulating the training image sensor capturing the training scene using the second CFA may include determining a red monochrome image representing the training scene through the red monochrome CFA, determining a green monochrome image representing the training scene through the green monochrome CFA, determining a blue monochrome image representing the training scene through the blue monochrome CFA, and determining an RGB image by combining the red monochrome image, the green monochrome image, and the blue monochrome image. Each pixel of the RGB image may be associated with a corresponding red pixel value, a corresponding green pixel value, and a corresponding blue pixel value.

[0143] In some examples, simulating the training image sensor capturing the training scene using the first CFA may include selecting, for each respective pixel of the simulated training input image, a corresponding color value from the RGB image based on a color associated with the respective pixel in the first CFA.

[0144] In some examples, simulating the training image sensor capturing the training scene using the first CFA may include applying one or more image degradations to the RGB image before selecting, for each respective pixel of the simulated training input image, thecorresponding color value from the RGB image. The simulated ground-truth image may correspond to the RGB image without the one or more image degradations.

[0145] In some examples, the machine learning model may include a convolutional neural network.

[0146] In some examples, executing the simulation may include generating the simulated training input image by cropping a simulated output of the training image sensor in alignment with a predetermined portion of the first color pattern such that, for each respective simulated training sample of a plurality of simulated training samples, a particular pixel position of a corresponding simulated training input image of the respective simulated training sample is aligned with a consistent portion of the first color pattern.

[0147] In some examples, an empirical training sample may be obtained. The empirical training sample may include an empirical training input image and an empirical ground-truth image. The empirical training input image may have been captured by the training image sensor using the first CFA, and may represent a reference image of a second training scene displayed on a display having at least a threshold resolution. The empirical ground-truth image may have been generated by fusing (i) image data from the reference image with (ii) image data from a demosaiced version of the empirical training input image. The empirical ground-truth image may represent the second training scene using the second color pattern. A second training output image may be generated by processing the empirical training input image using the machine learning model. A second loss value may be determined based on the empirical ground-truth image and the second training output image. The one or more parameters of the machine learning model may be adjusted based on the second loss value.

[0148] In some examples, the empirical ground-truth image may be generated by determining the demosaiced version of the empirical training input image by aligning, merging, and demosaicing an image burst that has been captured by the training image sensor using the first CFA and represents the reference image of the second training scene displayed on the display. The reference image may be aligned to match a field of view of the demosaiced version of the empirical training input image. Image data from the reference image as aligned may be fused, using an image fusion model, with image data from the demosaiced version of the empirical training input image.

[0149] In some examples, the machine learning model may be pre-trained using the simulated training sample and fine-tuned using the empirical training sample.IX. Conclusion

[0150] The present disclosure is not to be limited in terms of the particular embodiments described in this application, which are intended as illustrations of various aspects. Many modifications and variations can be made without departing from its scope, as will be apparent to those skilled in the art. Functionally equivalent methods and apparatuses within the scope of the disclosure, in addition to those described herein, will be apparent to those skilled in the art from the foregoing descriptions. Such modifications and variations are intended to fall within the scope of the appended claims.

[0151] The above detailed description describes various features and operations of the disclosed systems, devices, and methods with reference to the accompanying figures. In the figures, similar symbols typically identify similar components, unless context dictates otherwise. The example embodiments described herein and in the figures are not meant to be limiting. Other embodiments can be utilized, and other changes can be made, without departing from the scope of the subject matter presented herein. It will be readily understood that the aspects of the present disclosure, as generally described herein, and illustrated in the figures, can be arranged, substituted, combined, separated, and designed in a wide variety of different configurations.

[0152] With respect to any or all of the message flow diagrams, scenarios, and flow charts in the figures and as discussed herein, each step, block, and / or communication can represent a processing of information and / or a transmission of information in accordance with example embodiments. Alternative embodiments are included within the scope of these example embodiments. In these alternative embodiments, for example, operations described as steps, blocks, transmissions, communications, requests, responses, and / or messages can be executed out of order from that shown or discussed, including substantially concurrently or in reverse order, depending on the functionality involved. Further, more or fewer blocks and / or operations can be used with any of the message flow diagrams, scenarios, and flow charts discussed herein, and these message flow diagrams, scenarios, and flow charts can be combined with one another, in part or in whole.

[0153] A step or block that represents a processing of information may correspond to circuitry that can be configured to perform the specific logical functions of a herein-described method or technique. Alternatively or additionally, a block that represents a processing of information may correspond to a module, a segment, or a portion of program code (including related data). The program code may include one or more instructions executable by a processor for implementing specific logical operations or actions in the method or technique.The program code and / or related data may be stored on any type of computer readable medium such as a storage device including random access memory (RAM), a disk drive, a solid state drive, or another storage medium.

[0154] The computer readable medium may also include non-transitory computer readable media such as computer readable media that store data for short periods of time like register memory, processor cache, and RAM. The computer readable media may also include non-transitory computer readable media that store program code and / or data for longer periods of time. Thus, the computer readable media may include secondary or persistent long term storage, like read only memory (ROM), optical or magnetic disks, solid state drives, compactdisc read only memory (CD-ROM), for example. The computer readable media may also be any other volatile or non-volatile storage systems. A computer readable medium may be considered a computer readable storage medium, for example, or a tangible storage device.

[0155] Moreover, a step or block that represents one or more information transmissions may correspond to information transmissions between software and / or hardware modules in the same physical device. However, other information transmissions may be between software modules and / or hardware modules in different physical devices.

[0156] The particular arrangements shown in the figures should not be viewed as limiting. It should be understood that other embodiments can include more or less of each element shown in a given figure. Further, some of the illustrated elements can be combined or omitted. Yet further, an example embodiment can include elements that are not illustrated in the figures.

[0157] While various aspects and embodiments have been disclosed herein, other aspects and embodiments will be apparent to those skilled in the art. The various aspects and embodiments disclosed herein are for purpose of illustration and are not intended to be limiting, with the true scope being indicated by the following claims.

Claims

CLAIMSWhat is claimed is:

1. A computer-implemented method comprising: obtaining an input image that has been captured by an image sensor using a first color filter array (CFA) having a first color pattern; generating an output image by remosaicing or demosaicing the input image using a machine learning model configured to map pixel values determined using the first CFA to a second color pattern that (i) differs from the first color pattern of the first CFA and (ii) is associated with a second CFA, wherein the machine learning model has been trained using a simulated training sample; and outputting the output image.

2. The computer-implemented method of claim 1, wherein the first CFA comprises a nonstandard-Bayer CFA.

3. The computer-implemented method of claim 2, wherein the nonstandard-Bayer CFA comprises a Quad Bayer CFA.

4. The computer-implemented method of any of claims 1-3, wherein the second CFA comprises a Bayer CFA, and wherein the output image is generated by remosaicing the input image from the first color pattern to the second color pattern of the Bayer CFA.

5. The computer-implemented method of any of claims 1-3, wherein the second color pattern comprises a demosaiced red-green-blue (RGB) color pattern, wherein the output image is generated by demosaicing the input image from the first color pattern to the demosaiced RGB color pattern, and wherein each pixel of the demosaiced RGB color pattern is associated with a corresponding red pixel value, a corresponding green pixel value, and a corresponding blue pixel value.

6. The computer-implemented method of any of claims 1-5, wherein the machine learning model comprises a convolutional neural network.

7. The computer-implemented method of any of claims 1-6, wherein the simulated training sample comprises (i) a simulated training input image generated by simulating atraining image sensor capturing a training scene using the first CFA and (ii) a simulated groundtruth image generated by simulating the training image sensor capturing the training scene using a second CFA having the second color pattern.

8. The computer-implemented method of claim 7, wherein the simulated training input image comprises an image degradation added as part of simulating the training image sensor capturing the training scene using the first CFA, wherein the simulated ground-truth image represents the training scene without the image degradation, and wherein generating the output image comprises enhancing the input image using the machine learning model by removing at least part of the image degradation present in the input image.

9. The computer-implemented method of claim 8, wherein the image degradation comprises one or more of a contrast degradation, a blur degradation, or a noise degradation.

10. The computer-implemented method of any of claims 7-9, wherein the second CFA comprises a red monochrome CFA, a green monochrome CFA, and a blue monochrome CFA, and wherein simulating the training image sensor capturing the training scene using the second CFA comprises: determining a red monochrome image representing the training scene through the red monochrome CFA; determining a green monochrome image representing the training scene through the green monochrome CFA; determining a blue monochrome image representing the training scene through the blue monochrome CFA; and determining an RGB image by combining the red monochrome image, the green monochrome image, and the blue monochrome image, wherein each pixel of the RGB image is associated with a corresponding red pixel value, a corresponding green pixel value, and a corresponding blue pixel value.

11. The computer-implemented method of claim 10, wherein simulating the training image sensor capturing the training scene using the first CFA comprises: selecting, for each respective pixel of the simulated training input image, a corresponding color value from the RGB image based on a color associated with the respective pixel in the first CFA.

12. The computer-implemented method of claim 11, wherein: simulating the training image sensor capturing the training scene using the first CFA comprises applying one or more image degradations to the RGB image before selecting, for each respective pixel of the simulated training input image, the corresponding color value from the RGB image, and the simulated ground-truth image corresponds to the RGB image without the one or more image degradations.

13. The computer-implemented method of any of claims 7-12, wherein the simulated training input image has been generated by cropping a simulated output of the training image sensor in alignment with a predetermined portion of the first color pattern such that, for each respective simulated training sample of a plurality of simulated training samples, a particular pixel position of a corresponding simulated training input image of the respective simulated training sample is aligned with a consistent portion of the first color pattern.

14. The computer-implemented method of any of claims 7-13, wherein the machine learning model has been trained by: obtaining the simulated training sample; generating a training output image by processing the simulated training input image using the machine learning model; determining a loss value based on the simulated ground-truth image and the training output image; and adjusting one or more parameters of the machine learning model based on the loss value.

15. The computer-implemented method of any of claims 7-14, wherein the machine learning model has been trained using an empirical training sample that comprises: an empirical training input image that has been captured by the training image sensor using the first CFA, wherein the empirical training input image represents a reference image of a second training scene displayed on a display having at least a threshold resolution; an empirical ground-truth image generated by fusing (i) image data from the reference image with (ii) image data from a demosaiced version of the empirical training input image,wherein the empirical ground-truth image represents the second training scene using the second color pattern.

16. The computer-implemented method of claim 15, wherein the empirical ground-truth image has been generated by: determining the demosaiced version of the empirical training input image by aligning, merging, and demosaicing an image burst that has been captured by the training image sensor using the first CFA and represents the reference image of the second training scene displayed on the display; aligning the reference image to match a field of view of the demosaiced version of the empirical training input image; and fusing, using an image fusion model, (i) image data from the reference image as aligned with (ii) image data from the demosaiced version of the empirical training input image.

17. The computer-implemented method of any of claims 15-16, wherein the machine learning model has been pre-trained using the simulated training sample and fine-tuned using the empirical training sample.

18. The computer-implemented method of any of claims 15-17, wherein the machine learning model has been trained by: obtaining the empirical training sample; generating a second training output image by processing the empirical training input image using the machine learning model; determining a second loss value based on the empirical ground-truth image and the second training output image; and adjusting the one or more parameters of the machine learning model based on the second loss value.

19. A system comprising a processor configured to perform the method of any of claims 1- 18.

20. A computer program product for carrying out the method of any of claims 1-18.

Citation Information

Patent Citations

  • Image processing methods, image processing apparatus, storage media and terminal devices

    CN113364964B